A “significant” index split happens in about 40% of healthy people.
Statistical significance asks whether a difference is bigger than measurement error. Rarity asks how many ordinary people show it. On the Wechsler indexes those two lines sit a long way apart, and the second one is what a reader actually wants to know.
The two numbers, on the same index
Take Verbal Comprehension on the WAIS-IV, compared against a person’s own average index score.
| Question | Difference needed |
|---|---|
| Bigger than measurement error (p < .05) | 5.55 points |
| Rarer than 95% of the healthy sample | 16.60 points — 1.11 SD |
Three times the gap. A six-point deviation clears the first line and is entirely ordinary against the second. And the historical figure is the one worth carrying around: a p < .05 verbal–performance split was shown by roughly 40% of the WAIS-R standardization sample — people selected to represent the population, not a clinic.
So reporting a significant split as though it were notable is, most of the time, reporting the ordinary. The word does the damage: in every other clinical context “significant” means “matters”, and here it means “probably not zero”.
Then there is the second problem: there are six pairs
Four indexes make six possible comparisons, and each one gets its own chance to look unusual. About a fifth of healthy adults show at least one index difference rarer than 5%, for no reason other than that six comparisons were available.
Which means a single unusual difference, reported on its own, is defensible. Six of them reported as six findings is not — it is the arithmetic of having looked six times, written up as a profile.
What Casebook does with this
Ipsative comparison reports at the rarity line, not the significance line. The threshold is 1.3 SD — the larger of the published index-level rarity figures, since Processing Speed needs 1.28 — so a difference called unusual here is unusual on every index rather than on the most forgiving one.
A smaller difference is not hidden. It is reported as a split worth noting at 1.0 SD, kept separate from the unusual ones, because collapsing the two overstates and dropping the smaller one hides. And when more than one difference clears the line, the page says how many comparisons were available to produce it.
One further refusal: when the person’s own average sits far enough from the population mean, regression alone puts almost everything on one side of it. An anchor 1.7 SD down will produce a long list of measures “above expectation” and none below, which is a fact about regression rather than about the person. The comparison refuses rather than producing that list.
Where these numbers come from
- WAIS-IV index-level significance and rarity tables (Wechsler, technical manual)
- WAIS-R standardization sample verbal–performance discrepancy base rates
- Binder, Iverson & Brooks (2009), on abnormal scores being psychometrically normal
Every reading in Casebook carries its threshold and its source on the page, so the figure you are being asked to accept is visible next to the line it was judged against.
See it on a real profile
The demonstration is the product, holding invented people. Open a case and expand any reading to its arithmetic.