How the assessment is built
The rule that turns the job task analysis into an exam blueprint, and the method for setting the pass mark. Both are published here before the analysis has closed, which is the only time publishing them proves anything.
Assessment items written today. The blueprint comes first and the items follow it.
Eligible responses before any weight is published. The floor was set before collection began.
Independent reviews every item passes before it can count toward anybody's result.
Votes the founder holds on the standard setting panel.
1. Why this is published now
A method published after the data has been seen cannot be checked. Anyone reading it has to take on trust that it was not adjusted once the numbers were known, and there is no way for them to establish otherwise, because the only record of the earlier version is in the heads of the people who changed it.
So the rule below is written while the job task analysis is still open and while nobody, including NBOP, knows what the weights will be. If a domain comes back heavier than expected, the blueprint follows it. If a domain comes back light enough to be embarrassing, the blueprint follows that too. That is the entire value of writing this down first, and it is worth nothing at all if the document is quietly revised when the results arrive.
Every change to this page after today appears in the changelog with its date, including changes made before the first sitting. A reader who wants to know whether the method moved after the data landed can check rather than ask.
2. From the analysis to the blueprint
The job task analysis asks practising operators to do two separate things. They rate each domain on criticality and on frequency, measured against fixed intervals rather than adjectives, and they allocate one hundred points across the eight domains according to where the work actually sits. The allocation is what sets the weighting. The two rating scales are what make the allocation interpretable, and they are what a reviewer would use to challenge it.
The rule, stated as arithmetic:
- Take every response marked eligible for round one. Responses excluded by the screening rule are excluded here too, and the count and the reasons are published alongside the weights.
- For each domain, take the mean of the hundred point allocations. Normalise the eight means so they sum to one hundred, because rounding otherwise leaves a published set that does not add up, which is the first thing a reviewer checks.
- Multiply each domain weight by sixty, the number of items in Part A. Round using largest remainder, so the parts sum to exactly sixty rather than to fifty nine or sixty one.
- Apply a floor of three items per domain, taking the items from the domains with the largest remainders. No domain in the standard may be unrepresented in the assessment, and a domain carrying one or two items cannot support the per domain score the handbook promises.
There is deliberately no ceiling. If the analysis says one domain carries a third of the work, the blueprint gives it a third of the items. Capping it would be overriding the evidence with a prior about what a balanced exam should look like, and that prior is exactly what the analysis exists to replace.
Part B works differently because scenarios do not divide cleanly by domain. Four scenarios are written so that, taken together, the domains they require the candidate to reason across approximate the same weighting, and so that every domain appears in at least one. The mapping from each scenario to the domains it touches is published with the blueprint rather than held internally.
3. What the blueprint fixes, and what it does not
It fixes how many items come from each domain, and it fixes that before any item is written. It does not fix which tasks within a domain get asked, because that would make the assessment memorisable from the blueprint, and an assessment you can prepare for by reading its own specification is measuring the wrong thing.
A consequence worth stating plainly: the per domain scores in your report are not all equally precise. A domain assessed with three items supports a much rougher estimate than one assessed with twelve. Every domain score is published with its item count beside it, and scores drawn from fewer than six items are labelled indicative rather than reliable. That is less tidy than a clean set of eight percentages and it is the honest version.
4. Who writes the items
Items are written by practising operators, paid for the work, under contract that assigns the item to NBOP and bars them from any preparation product for the credential. Committee membership is unpaid; item writing is paid, and the two are kept separate so that nobody is choosing between the two roles for money.
Item writers are told which domain and which task from the analysis they are writing to. They are not told the blueprint counts, because a writer who knows a domain is short of items has an incentive to submit weak ones.
Every item passes three reviews, by three different people, none of them the writer:
- Content. Does it assess the task it claims to, in the domain it claims to, and is the keyed answer defensibly the best answer for a competent operator rather than merely the one the writer prefers?
- Technical. Item writing has a known catalogue of flaws: cueing from the stem, implausible distractors, absolutes, all of the above, options that overlap. This review is a checklist against that catalogue, published in full as the item writing rules, and it is the least interesting and most necessary of the three.
- Sensitivity. Does the item advantage or disadvantage anybody for a reason unrelated to operations practice: sector jargon, a national regulatory frame presented as universal, a scenario that assumes a size of organisation or a level of budget authority not required for eligibility?
5. How an item starts counting, and how it stops
A new item is piloted before it counts. It appears in a live sitting, in a position the candidate cannot identify, and it is scored for statistics only. Nobody passes or fails on a pilot item. Only once an item has enough pilot responses to have statistics worth reading does it become eligible to count.
After that it is watched. An item that almost everybody gets right, or almost everybody gets wrong, is carrying no information and is retired. An item that candidates who did well overall get wrong more often than candidates who did badly is measuring something other than what it claims, and it is pulled and rewritten or discarded. Retirement statistics are published in aggregate each cycle.
Where an item is found to be faulty after a sitting has been scored, the item is removed and every affected result is rescored without it. Results move only in the candidate's favour under that correction: nobody who passed is failed by a rescore caused by our own faulty item.
6. Setting the pass mark
The cut score is set by a formal standard setting study before the first sitting, and it is set against a description of the minimally competent operator rather than against a target pass rate. Those are different exercises and only one of them is defensible.
- Part A uses a modified Angoff procedure. Each panellist estimates, item by item, the proportion of minimally competent operators who would answer correctly. Panellists see the item statistics between rounds and may revise. The cut is the mean of the final round.
- Part B uses an extended Angoff on the rubric. Each panellist estimates the score a minimally competent operator would earn on each rubric dimension of each scenario, using the same two round structure.
- The two are combined at the published weighting, thirty percent Part A and seventy percent Part B, which is fixed in the handbook and is not a lever the panel may pull.
The panel is eight to twelve people. A majority must be practising operators who meet the eligibility criteria for the credential. No panellist may have written any item in the pool being judged. Sector, organisation size and geography are spread deliberately rather than left to whoever volunteers. The founder does not sit on the panel and holds no vote in it.
The panel composition, the method, the standard error of the resulting cut, and the cut itself are all published before the first sitting opens. The pass rate is published afterwards, in raw numbers rather than as a percentage on its own, because a percentage without a denominator is how a small cohort is made to look like evidence.
The cut score is not adjusted to control how many people pass. If the first cohort passes at a rate that looks wrong, the response is to examine the items and the panel, publish what was found, and correct the instrument rather than the threshold.
7. After the first sitting
The analysis is not a one off. Practice moves, and a blueprint built on a reading of the work taken five years ago will quietly stop describing the job while continuing to look authoritative. The job task analysis is repeated on a fixed cycle rather than when somebody suspects it is stale, because the point at which it is most obviously needed is the point at which it is least likely to be commissioned.
Between analyses the blueprint does not drift. A domain does not gain items because a committee member argues it has become more important, and does not lose them because items in it are hard to write. Both of those are ordinary pressures and both are the reason the weighting is set by an instrument rather than by a discussion.
Where a repeat analysis moves a weight enough to change the item counts, the change, the size of it and the date it takes effect are published before the first sitting that uses the new blueprint. Nobody sits an assessment built to a specification that was changed after they started preparing.
8. What would make this document change
Being wrong. Specifically: if the largest remainder rule produces a blueprint that a competent reviewer can show misrepresents the analysis, if the floor of three turns out to distort a weighting materially rather than just protecting a score report, or if the standard setting method proves unworkable with a panel of this size.
Any of those would be published as a change with its reasoning, before the sitting it affects. What will not change it is an unwelcome result. A weighting that surprises us is the analysis working, and a method that only survives contact with agreeable data was never a method.
Until the standards committee is seated, this document may be amended by NBOP, and every amendment appears in the changelog. Once the committee is seated, the blueprint rule and the standard setting method sit with the committee under its terms of reference.
None of this is worth anything until the analysis closes
Which is why the analysis is the thing being asked for, and this is not.
The weighting is set by practising operators describing their own work, and the floor is one hundred and twenty eligible responses before a single weight is published. Until then the blueprint above is a rule with nothing to apply itself to.