Assessby the team behind Cogn-IQ

Calibration provisional-2026-08-24

Methods

Everything the Screener does, stated so it can be argued with.

This page is the reason to trust the band, or not to. It says what the items are, what the scale is worth today, and which sittings are thrown away. No number on this page appears without a range.

§ 01

How the items are composed

Every item is a 3 by 3 figural matrix with the bottom-right cell withheld and six choices. Nothing is a bitmap: a cell is a data structure — frame, glyph, rotation, marks, or a 3 by 3 dot pattern — and it is drawn as vector graphics at request time.

There are 10 rule families across four item types: progression, rotation, distribution and set logic. A form takes each family once, then a second instance of six mid-range families, for 16 items. Shapes, rotations and patterns are drawn from a seeded generator, so two takers do not see the same 16 figures and the form cannot be memorised from a forum thread.

Distractors are built from the near-misses of the rule itself: a one-step error, the wrong logical operation, the right shape with the wrong count. Every choice set is checked for duplicates, and rows are rejected and recomposed when the rule produces a degenerate answer — which is what keeps exactly one choice defensible.

§ 02

What an anchor item is, and how many are in

An anchor item is a retired item whose difficulty is already known on the Cogn-IQ scale. Seeding a few into every form is what links a free screener to a calibrated instrument rather than to an assumption.

Stated plainly, because this is the claim the sceptical reader should check: a form is built to carry 4 anchor items, and none are installed yet. Today every item on your form is a generated item with a provisional difficulty, and the band screen says the scale is preliminary for exactly this reason.

An anchor item is never shown in a worked-item explanation and its answer is never published, because a calibrated item that circulates is a calibrated item no longer.

Installing anchors will not, by itself, remove the word preliminary from your band. Anchors link the middle of the scale; the ends still need the table in part 04.

§ 03

How ability is estimated

A three-parameter logistic model, with the lower asymptote fixed at the chance rate for six choices, which is 1 in 6, and discrimination shared at 1.4 for every generated item. An anchor item that arrives with its own fitted discrimination uses it instead. Ability is estimated from the whole response pattern, not by counting correct answers: getting a hard item right and an easy item wrong is not the same evidence as the reverse.

The estimate is the mean of the posterior over ability, taken against a standard normal population prior, and the reported uncertainty is that posterior’s standard deviation. The prior is what keeps a sitting with every item right from producing an unbounded estimate; it is also why such a sitting is reported as a floor rather than a number.

A skipped item is scored as incorrect — the taker saw it and did not answer it. An item that failed to draw is dropped from the sitting entirely and the estimate is computed from the items that remain, so a broken figure costs precision rather than credibility.

§ 04

How the scale is linked, and what preliminary means

Ability is estimated in logits and then linked to a familiar 100 and 15 scale for reporting, clamped to the 60 to 140 window because 16 items cannot resolve the tails.

The status of that link today is preliminary. The linking function is Assess’s own assumption of a standard normal population, not a table fitted to a norming sample, and the item difficulties are the ones the generator declares. The final altitude-to-scale table is supplied separately, under supervision, and until it is in place every band is labelled preliminary on screen and the interval is widened by 10% to pay for the assumption.

A 16-item screener estimate is not a clinical or certified IQ, and Assess will never describe it as one.

§ 05

What the interval means

The band you see is the central 95% of the posterior: the range your ability most plausibly falls in, given this response pattern and the population prior. It is drawn as a shaded block because that is what was measured. The centre of the block is not reported on its own, at any point, in any screen or Report.

The interval also has a floor. However cleanly a taker performs, 16 multiple-choice items carry an irreducible standard error, so the interval is never drawn narrower than the instrument can support. A site that shows you a single number from 16 items is inventing the precision.

§ 06

Which sittings are excluded

Two rules, applied before anything is shown. A sitting with fewer than 12 of the 16 items answered is not scored: below that the estimate is mostly assumption. A sitting whose median response is under 1.5 seconds is not scored either, because a figural item cannot be read that fast.

An excluded sitting sees the withheld-band screen, is offered a fresh form, and is never offered a Report. It is kept as data about response behaviour, which is how these two thresholds get revised.

The same two rules define the company’s own success metric, so Assess cannot count a sitting it refused to score.

§ 07

What is not measured

Fluid reasoning (Gf) only. Verbal reasoning (Gc), quantitative reasoning (Gq) and spatial reasoning (Gv) are unmeasured here, and no combination of 16 figural items can stand in for them. The free battery at Cogn-IQ is where a whole profile comes from.

Start the Screener16 items, untimed, about 15 minutes. No account, no email.

Supervised by Xavier Jouve, Ph.D.

Fluid reasoning is one of four domains. The other three are free at Cogn-IQ.