Screening for technical flaws in multiple-choice items: A generalizability study

Lotte Dyhrberg O'Neill*, Sara Mathilde Radl Mortensen, Cita Nørgaard, Anne Lindebo Holm, Ulla Glenert Friis

*Corresponding author for this work

Research output: Contribution to journalJournal articleResearchpeer-review

91 Downloads (Pure)


Construction errors in multiple-choice items are quite prevalent and constitute threats to test validity of multiple-choice tests. Currently very little research on the usefulness of systematic item screening by local review committees before test administration seem to exist. The aim of this study was therefore to examine validity and feasibility aspects of review committee screening for item flaws. We examined the reliability of item reviewers’ independent judgments of the presence/absence of item flaws with a generalizability study design and found only moderate reliability using five reviewers. Statistical analyses of actual exam scores could be a more efficient way of identifying flaws and improving average item discrimination of tests in local contexts. The question of validity of human judgments of item flaws is important - not just for sufficiently sound quality assurance procedures of tests in local test contexts - but also for the global research on item flaws.
Original languageEnglish
JournalDansk Universitetspædagogisk Tidsskrift
Issue number26
Pages (from-to)51-65
Publication statusPublished - 1. Apr 2019


  • Multiple-choice Tests
  • Higher Education
  • Validity
  • Quality Assurance
  • quality appraisal
  • screening
  • Item flaws


Dive into the research topics of 'Screening for technical flaws in multiple-choice items: A generalizability study'. Together they form a unique fingerprint.

Cite this