No published norm table describes the people you see.
Norms answer how this person compares to the population. The question a report cannot answer is how they compare to your last three hundred referrals — who came through the same door, for the same reasons, and were given much the same battery.
A worked example you can reproduce
These are the 62 survivors of acute lymphoblastic leukemia in the demonstration, against the other 258 cases in it. The demonstration is 320 invented people — no real practice’s records appear on this site — but the figures below are the ones the software actually produces, and you can open it and get them yourself.
| Domain | Mean | vs everyone else | Hedges’ g |
|---|---|---|---|
| Verbal comprehension | −0.16 SD | +0.10 | +0.19 |
| Fluid reasoning | −0.47 SD | −0.08 | −0.14 |
| Working memory | −0.77 SD | −0.48 | −0.78 |
| Processing speed | −0.83 SD | −0.45 | −0.80 |
| Executive function | −1.00 SD | −0.56 | −0.94 |
Selective rather than global: verbal comprehension sits where the rest of the series sits, while executive function, speed and working memory are half a standard deviation below it. That shape is the finding, and no single report shows it. Each of those 62 reports, read on its own, shows one person who is somewhat slow.
Why the comparison group matters more than it sounds
“Two SD below the mean” means something different in a memory clinic than in a pediatric learning-disability practice, because the referral stream is different and the base rates are different. A profile that is unremarkable among the people you see is worth knowing about, and it is precisely what a population norm cannot tell you.
Casebook also reports how often a given reading fires elsewhere in your own series, with the case codes as links. That is a base rate for your database rather than for somebody else’s sample — and it is the number that tells you whether a finding is routine in your practice.
The confounds are reported, not corrected away
A cohort built this way has real problems, and the design principle is to hand you the numbers to discount a result rather than to adjust it silently:
- Age spread. A cohort drawn across a wide age range is reported with that range, because correcting for it requires a decision about your population that is yours rather than the software’s.
- Effect-size multiplicity. Fifteen domains is fifteen chances. The count is shown beside the effect sizes.
- Retests. A person contributes one profile to a cohort, never an average of their own retests. Where their scores are compared across testings, a rise across an interval under a year is flagged as possibly practice.
- Mixed batteries. Which instruments are carrying a domain is stated, including when it is only one.
Correcting a confound requires a judgment about your own population. Reporting it leaves that judgment where it belongs.
What it takes to get there
A case series you can query is mostly a data-entry problem, which is why import matters more than it sounds: upload a finished report or a vendor export and Casebook reads the score tables, works out which instrument each section belongs to, picks up the informant and the age, and proposes a row per score. Nothing is written until you have looked at it.
Build a cohort in the demonstration
320 invented people across six referral groups with deliberately different profile shapes — because a demonstration where every profile is flat demonstrates nothing.