Casebook

We built subtest profile analysis, then took it out.

The retest stability of a subtest profile is around .26. A pattern that does not survive being measured twice in the same person cannot carry an interpretation about that person, however compelling the peaks and troughs look on the page.

What .26 means in practice

Test a person, plot their ten subtest scaled scores, note the high ones and the low ones. Test them again after a decent interval and plot it again. The correlation between those two shapes — the same person, the same instrument, nothing changed — is about .26.

Which is to say the shape is mostly not a property of the person. It is the thing you would get by drawing ten numbers with a bit of signal and a lot of error, and the error is what makes the picture interesting: peaks and troughs are exactly what random variation around a mean produces.

Index scores hold up far better, because they average several subtests and the averaging is where the error goes. That is the whole reason the indexes exist.

Why we removed it rather than caveated it

A caveat next to a chart does not beat the chart. The shape is drawn, the eye reads it, and a sentence underneath saying it may not replicate is read afterwards if at all. Software that draws a compelling picture and then disclaims it has still mostly drawn the picture.

So Casebook does ipsative comparison on composites only. Subtest scores are held, displayed, exported, and counted in the battery — they are real data and you may want them. What does not happen is a reading built on the shape across them.

Three of the first four readings built had to be weakened once the literature was read properly. This one was the only one deleted.

The two that were weakened instead

The ipsative comparison now reports at the rarity line rather than the significance line, because a statistically significant index split occurs in about 40% of healthy people.

And the memory reading names the shape it sees — “recovered by recognition” — rather than asserting the mechanism behind it. “Intact recognition implies a retrieval deficit” is far weaker for one individual than its routine use suggests; across 28 longitudinal studies, cued recall predicted conversion no better than free recall.

Where these numbers come from

  • Published retest stability figures for Wechsler subtest profiles
  • Belleville et al. (2017) — 28 longitudinal studies; cued recall predicted conversion no better than free recall

The corrections are the part worth having. A tool that removed a feature after reading the retest data is telling you something about how the rest of it was built.

See what it does report

More answers