Writing assessment items
The rules every item is written and reviewed against, including the flaw catalogue a reviewer checks each one with. Published in full, because a rule nobody can read is a rule nobody can hold us to.
Options on a selected response item. Not three, not five, and never all of the above.
Flaws on the checklist. An item with any one of them goes back to its writer.
Vocabulary items. Knowing a definition is not the skill being certified.
Reviews before an item is piloted, none of them by the person who wrote it.
1. Who this is for
Item writers, and reviewers. It is also public, which is a deliberate choice and worth explaining, because the instinct with assessment material is to keep everything behind a wall.
Publishing the rules does not help a candidate. Knowing that distractors must be plausible does not tell you which option is keyed, and knowing that a stem must be answerable before reading the options does not tell you the answer. What it does is let anybody who thinks an item was unfair check it against the standard it was supposed to meet, which is the only way an appeal about item quality can be argued rather than merely asserted.
What stays closed is the item pool itself, the keys, and the mapping of individual items to tasks. Those are the things that would actually be compromised. See the candidate handbook for what that means for candidates.
2. What an item has to do
Every item assesses one task from the job task analysis, in one domain, and the writer is told which. If an item cannot be traced back to a task an operator said they actually do, it does not belong in the assessment however good it is.
Items are scenario based rather than recall based. The distinction that matters is not whether the item has a story attached: it is whether a candidate who has never done the work but has read about it can answer. If they can, the item is testing reading rather than practice, and it goes back.
There are no vocabulary items. Knowing the definition of operating cadence is not the skill being certified, and an assessment full of definitions is an assessment that rewards preparation products over practice.
3. Part A, the stem
- Ask a complete question. The test is that a competent operator could answer it with the options covered up. If they cannot, the question is in the options rather than in the stem, and it needs rewriting.
- Put the whole situation in the stem. Detail spread across the options means the candidate is assembling the scenario while evaluating it, which measures working memory rather than judgement.
- No unnecessary detail. Every fact in the stem should either be needed to answer or be a plausible distraction that a practitioner would correctly set aside. Filler is neither.
- Avoid negatives. Where the question genuinely has to be negative, the negative word is emphasised, and the item may not also contain negatives in its options.
- If the item asks for the best answer, the stem states the basis. Best for what, on what timescale, against which constraint. Without that, the item is an opinion poll with a key.
4. Part A, the options
- Four options. One is defensibly the best, and the other three are wrong for reasons a competent operator could articulate.
- Every distractor is something a real practitioner might actually choose. An option nobody would pick is not a distractor, it is padding, and it quietly turns a four option item into a three option one.
- Options are homogeneous. Same kind of thing, roughly the same length, same grammatical shape.
- No all of the above, and no none of the above. Both are answerable by partial knowledge.
- Options are ordered logically or alphabetically, never by where the writer feels the key should sit.
5. The flaw catalogue
The technical review is a checklist against this list. It is the least interesting of the three reviews and the one that catches the most, because every flaw here lets a candidate who does not know the answer find it anyway, which means the item is measuring test taking rather than practice.
- Length cue. The key is noticeably longer or more qualified than the distractors, because the writer was busy making it defensible.
- Grammatical cue. The stem ends in "an" and only one option begins with a vowel. Or the stem is plural and one option is singular.
- Word repeat. A distinctive word from the stem appears in the key and nowhere else.
- Absolutes. Always, never, all, none in a distractor. Experienced candidates eliminate these without reading them.
- Convergence. The key shares elements with most of the distractors, so it can be assembled from the overlap without knowing anything.
- Overlap. Two options where one contains the other, so both are right or both are wrong.
- Implausible distractor. An option no practitioner would ever select.
- Vocabulary. The item turns on knowing a term rather than on doing anything with it.
- Unstated basis. The item asks which is best or most appropriate without saying against what.
An item carrying any one of these goes back to its writer with the flaw named. It is not silently fixed by a reviewer, because a writer who is never told keeps making the same item.
6. Part B, the scenarios
Four scenarios, rubric scored, seventy percent of the result. They are where the credential is actually earned, and they are harder to write than selected response items by a wide margin.
- A scenario presents a situation with genuinely competing considerations. If one course of action is obviously correct, there is nothing to assess: the candidate is reporting the obvious rather than reasoning.
- It requires a decision and a justification, not a description. "Explain what is happening here" is a comprehension task. "Decide what you would do first and say what you are trading away" is an operations task.
- It is answerable from the scenario plus general practice. A scenario that needs sector specific knowledge, a particular regulatory frame, or familiarity with a named tool is assessing background rather than skill.
- It has more than one defensible answer. The rubric credits the quality of the reasoning and the honesty about the trade off, not agreement with a preferred conclusion.
- It states the operating context, or asks the candidate to answer against the context they described in their own application. Advice that is right for a thousand person organisation and wrong for a ten person one is not advice.
Each scenario carries a rubric with named dimensions, each scored independently, each with written descriptors for every level. A rubric with a single overall impression score is not a rubric, it is a mark with a form around it.
7. What happens to what you submit
Items are submitted against a specific domain and task. They go through content, technical and sensitivity review, by three different people, none of them you. You get told the outcome and, where an item is rejected or returned, the specific reason.
Accepted items are piloted unscored in a live sitting before they count toward anybody's result, then watched, then retired when their statistics say they have stopped carrying information. The full lifecycle is in how the assessment is built.
Item writing is paid. Committee membership is not. The two are kept separate so that nobody is weighing one role against the other for money, and writers are barred by contract from any preparation product for the credential.
8. What would change these rules
Item statistics. If items written to a rule here fail in a consistent way, the rule is wrong and it changes, with the change and its reason published in the changelog. Nothing here is original: it is the ordinary craft of item writing as practised by credentialing bodies, written down so that it can be pointed at.
Until the standards committee is seated, this document may be amended by NBOP. Once it is seated, these rules sit with the committee under its terms of reference.
No items exist yet
The rules are written first, and the blueprint before that.
Item writing opens once the job task analysis closes and the blueprint is published, because a writer needs to know which task they are writing to. If you hold operational accountability and want to be asked, the analysis is where the pool of writers is drawn from.