Difficulty and discrimination decide an item's value
Two classical item statistics tell you whether an item earns its place. Difficulty p is the fraction answering correctly. Discrimination d is how sharply the item separates stronger from weaker models. Items at the floor or ceiling add noise without information.
Difficulty p
p near 0.0 is a floor item (everyone fails), p near 1.0 a ceiling item (everyone passes). Both provide little signal.
Discrimination d
d = p(top group) − p(bottom group). High d means the item tracks the capability under test.
Aim for a spread
A test that is uniformly hard produces a compressed score distribution and a wide confidence interval.
If the top quartile of models answers 0.9 and the bottom quartile 0.2, d = 0.7 — a strongly discriminating item. If both answer 0.5, d = 0 and the item is uninformative.
- Psychometrics teams routinely cull items with p < 0.1 or p > 0.9 before freezing a suite.
- Adaptive testing picks the next item near the examinee's current ability to maximize information per item.