Every NCLEX prep company says its questions are “NCLEX difficulty.” It is the easiest claim in the industry to make and the hardest to check, because difficulty is not a property you can see by reading an item. A question that looks brutal can be answered correctly by 80% of students, and an innocuous-looking one can split a cohort down the middle.
So rather than assert it, here is our data. 4,600+ published questions, six difficulty tiers, and what students actually score on each — as of August 14, 2026, across 4,127 graded responses.
Why this chart is the whole point
A difficulty label is only meaningful if it predicts performance. Ours does, monotonically: accuracy falls at every step up the ladder, from roughly 72% on items written well below the passing standard to roughly 51% on items written well above it. No inversions, no bunching.
That matters because the adaptive engine is built on top of these labels. When a practice session or a full-length exam decides you are ready for harder items, it is trusting the label to mean something. A bank whose “hard” questions are answered correctly 75% of the time would produce an adaptive test that lies to you — confidently walking you up a ladder whose rungs are all the same height.
Note the tier that sits near a coin flip: items written above the passing standard land in the 50-something range. That is the target zone for a well-calibrated adaptive test, and it is why a good NCLEX feels awful while you are taking it. Feeling like you are drowning is the normal experience of a test aimed correctly.
How an item gets its tier
Difficulty is assigned at authoring time against the cognitive demand of the item, not the obscurity of its content. The things that move an item up the ladder:
- Competing correct-ish options. An item where three answers are defensible nursing actions and one is first is harder than one with three obviously wrong distractors — even though both test the same fact.
- Cues that must be interpreted, not read. “Potassium 6.2” is recall. A rhythm strip plus a potassium plus a new medication, where you have to decide which one explains the other, is analysis.
- Multiple correct components. SATA, Select N, matrix and bowtie items require several simultaneous judgments, and partial credit does not make them easy.
- Distance from the cue to the action. Recognizing hypoglycemia is one step. Recognizing it, choosing the intervention, and then evaluating whether it worked is three, and NGN items chain them deliberately.
What does not move an item up: rare diseases, trivia, and clinician-level detail an entry-level nurse would never act on. A question can be extremely hard and still be a bad NCLEX question. Those get cut in audit.
The bank is weighted where the exam is
Two-thirds of published items sit at or above the passing standard. That is deliberate: the NCLEX decides at the passing line, so the items that carry information about whether you will pass are the ones written near and above it. A bank stuffed with easy recall questions produces comfortable practice scores and no signal.
It is also spread across every format the Next Generation NCLEX uses, not just multiple choice:
- Multiple choice (MCQ) — 2,311 published items
- Select All That Apply (SATA) — 493 published items
- Bowtie (NGN) — 308 published items
- Select N — 281 published items
- Highlight text (NGN) — 207 published items
- Ordered Response (NGN) — 180 published items
- plus matrix, grouping, trend, drag-and-drop cloze, dropdown table, and the rest of the NGN micro-formats
What we are not claiming
Honesty about the limits, since this is a post about evidence:
- These are our students on our items, not a psychometrically equated comparison against the actual NCLEX. Nobody outside the NCSBN has that.
- The response counts are uneven across tiers, because the adaptive engine serves fewer items at the extremes by design. The well-above tier has the thinnest data.
- Accuracy here is all-or-nothing per item — partial credit on multi-part formats is scored separately and is not folded into these percentages.
- Practice accuracy is not a pass probability, and any prep tool that hands you one is guessing.
What the chart does establish is narrower and more useful: our difficulty labels track real student performance, so the adaptive machinery built on them is aiming at something real.
Why publish this at all
Because “true NCLEX difficulty” is unfalsifiable marketing until someone shows the numbers, and we would rather be the ones who did. We will re-run this as the bank and the response volume grow, including if a tier stops behaving.
Go find out where you land against it — a full-length NCLEX-style practice test is free, and unlimited.