Notes from a full working session inside Indiana's reading assessment experience — the item families a third grader actually faces — organized into what the walkthrough surfaced, directions worth exploring, and the questions Indiana can ask. The tone here follows the rest of this work: analytical, not adversarial. Every observation is offered as a design question, because every one of them has a design answer.
Taking the test as a test-taker — not reviewing it as a document — surfaces a consistent theme: many items measure something adjacent to reading. Each observation below names that adjacent thing.
A run of early items requires the vocabulary about language — vowel, syllable, prefix, suffix — before a child can even parse what is being asked. A child can read fluently and still stumble on "which word has the same vowel sound," because knowing what a vowel is called is a different skill from hearing one.
"Which word has the same vowel sound as rain?" — the reading is easy; the grammatical framing is the obstacle.Sound-alike answer sets — rain / ran / ray, see / seed — are built so near-neighbors sit next to each other. That measures composure and test-wiseness under pressure at least as much as phonemic awareness. An adult taking these cold feels the trap; a stressed eight-year-old feels it more.
An item hinging on raising the flag on a mailbox assumes a household with a curbside mailbox. Many Indiana children — apartment buildings, city housing, cluster boxes — have never seen one. When an item's difficulty depends on housing type rather than reading skill, it is measuring circumstance.
The question this raises is not rhetorical: what did the bias and sensitivity review say about experience-contingent items, and is it published?Children learn to read with pictures, voices, and familiar stories. The instrument presents dense, unillustrated text — walls of it, in tight paragraphs. The scaffold isn't a crutch to be removed at test time; for early readers it is part of how meaning is constructed. Visual context could be added to many item families without touching the construct.
The long informational passages — the bee passage is the canonical example — are attention and working-memory tests wearing a reading test's clothes. A slow reader pays twice: the passage takes longer, so by the time a question asks what a worker bee does first, the answer is further behind them than it is for a fast reader. Adults would resent this task. We hand it to eight-year-olds with a retention decision attached.
"Which sentence best shows…", "what is the main idea…", and predict-what-the-character-does-next items admit multiple defensible answers — especially for an imaginative child who personifies characters, exactly as we encourage at home. A fable-prediction item is a judgment call about a fictional lion; for a student near the cut, it can also be the difference in a retention decision. The classroom's rational response is formula-teaching: "the main idea is the first or last sentence." That is not comprehension; that is compliance.
Grammar items ("the three dogs barks") and capitalization items test conventions in isolation. All of the wrong answers still communicate. The deeper question for an assessment of reading: which of these sentences did the child fail to understand? None. Something else is being scored.
A listening item buries a permission-slip requirement mid-announcement and scores the child for missing it. In adult life we call that a communication failure by the speaker. The item teaches children that when the main point is missed, the fault is theirs — the inverse of what we teach about clear communication.
The math sections arrive as word problems — situational narratives, explain-your-answer prompts, vocabulary like "quadrilateral." For a child on the reading borderline, the math score partially re-measures reading. And the explain-in-writing prompts penalize the child who could talk through their reasoning perfectly — the verbal pathway exists in the standards but barely in the instrument.
The calculator question is worth asking straight: if a child gets every answer right with a calculator, what exactly did we establish they cannot do — and is it the thing this test is for?None of these require abandoning measurement. They require deciding that the goal is a reading population, not a defensible ranking — and then designing backward from that.
Reading is substantially built at home — the same story requested every night until every word is known. That is reading development. Policy currently treats the home as outside the system. Family-engagement incentives, familiar-story instruments, and opt-in home evidence would align the assessment with where the skill is actually formed.
Some children's oral language runs far ahead of their test performance — including English learners who converse fluently and follow every story. A structured verbal demonstration of understanding is valid evidence of comprehension. Hand-scoring already exists in the system; structured verbal response could too.
The target skills — tricky pairings, decoding, comprehension — can be embedded in a continuous, illustrated narrative a child can engage with, rather than clinical fragments about nothing. Familiar format is not memorization; it is the difference between demonstrating reading and surviving an unfamiliar ritual.
Checkpoints exist; the idea can go further. Low-stakes signals gathered through the year — including a child privately marking "this didn't click" during a lesson, with help arriving later at home, never a spotlight in class — produce a truer picture than a single high-stakes sitting, and produce it in time to act.
One child explains in words, another draws, another makes tally marks. How a child solves reveals how a child thinks — and in aggregate, how a classroom thinks. The instrument currently discards this. Recording solution pathways would serve the student, the teacher, and every conversation the department has about instruction.
High-stakes pressure changes what an instrument measures. A stressed child performs below their skill; a retained child repeats the same instruction while learning mainly that they are "not good at school." If the design goal includes accuracy, the instrument should be engineered to minimize stress artifacts — pacing, framing, layout, spacing, stakes — the way any measurement system controls for interference.
Co-design with teachers and with children, then refine continuously against the actual goal — every child reading — rather than the proxy of a stable score distribution. If the aspiration of one hundred percent is unreasonable, the reasons it is unreasonable are precisely the data the system should be surfacing.
Each of these can be asked in a public meeting, in a procurement conversation, or in a letter — and each has a knowable answer. Grouped by what the answer would inform.
How standardized is the testing experience from school to school — forms, conditions, preparation intensity? If experiences differ, school-to-school score comparisons carry noise the public conversation currently ignores.
For judgment-based item families ("best sentence," "main idea," prediction), what do response-distribution and rater data show about how many answers are defensible? If strong readers split across two answers, the item measures convention, not comprehension.
What did bias and sensitivity review conclude about experience-contingent items — and can those conclusions be published? The mailbox problem generalizes: any item whose difficulty depends on a child's circumstances is measuring the circumstance.
How much of the reading score reflects attention and working-memory load rather than decoding and comprehension — and has passage length and density ever been studied against the construct? The slow-reader double penalty is a measurable artifact, not a hypothesis.
In the math assessment, how much score variance is explained by reading load? A reading-borderline child's math score partially re-measures reading; the state should know the size of that effect.
What would it take to publish a released item form each cycle for public scrutiny? Other states release forms. Public confidence in a $126M instrument should not require taking anyone's word.
Where are the non-passing students in every other measure — attendance, oral language, other subjects, growth trajectories? A retention decision made on one instrument, without the whole-child view the state already possesses, is a data-illiteracy problem at the decision layer.
What are the longitudinal outcomes of retained students in Indiana specifically — retention into the same instruction versus promotion with targeted support? The national literature is contested; Indiana is generating its own evidence right now and should be studying it.
How many students sit within one standard error of the 446 cut each year, and what happens to them? Children inside the measurement-error band could flip outcomes on a different testing day. The count is knowable and should be public — see the briefing.
What did the second-grade opt-in cohort's high pass-through actually measure — skill acquisition or format familiarity? If practice on the format moves scores this much, the format is a larger variable than the conversation admits.
Is the goal a defensible ranking of students, or a population of readers? What would the instrument look like if designed purely for the second? These produce different designs. The current instrument leans heavily toward the first.
What evidence would the state require to admit structured verbal demonstration into the proficiency determination? Naming the evidentiary bar turns "that's not how it's done" into a research agenda.
Where will AI scoring be allowed to operate, where does hand-scoring remain, and what is published about scoring error by item type? The scoring boundary is a policy choice being made inside a vendor relationship; it should be visible.
Where do family-engagement incentives sit in the literacy policy stack? The science of reading points at the home as much as the classroom; the policy stack currently funds only one of them.