Inside the Instrument · observations & questions

Inside the Instrument

Luminary AI Technologies · August 2026 · companion to the working prototype and the analytical briefing

Notes from a full working session inside Indiana's reading assessment experience — the item families a third grader actually faces — organized into what the walkthrough surfaced, directions worth exploring, and the questions Indiana can ask. The tone here follows the rest of this work: analytical, not adversarial. Every observation is offered as a design question, because every one of them has a design answer.

How this was produced. These observations come from sitting the assessment experience end to end using the working prototype reconstructed from the public record — item types, tools, and scoring approach — not the secure operational test. Item examples below are paraphrased from that reconstruction. Where an observation depends on something only IDOE or the vendor can see, it appears as a question in the final section instead of a claim.

01What the walkthrough surfaced

Taking the test as a test-taker — not reviewing it as a document — surfaces a consistent theme: many items measure something adjacent to reading. Each observation below names that adjacent thing.

Metalanguage is not reading

A run of early items requires the vocabulary about language — vowel, syllable, prefix, suffix — before a child can even parse what is being asked. A child can read fluently and still stumble on "which word has the same vowel sound," because knowing what a vowel is called is a different skill from hearing one.

"Which word has the same vowel sound as rain?" — the reading is easy; the grammatical framing is the obstacle.

Distractor sets engineered to be confusable

Sound-alike answer sets — rain / ran / ray, see / seed — are built so near-neighbors sit next to each other. That measures composure and test-wiseness under pressure at least as much as phonemic awareness. An adult taking these cold feels the trap; a stressed eight-year-old feels it more.

Items that assume a life

An item hinging on raising the flag on a mailbox assumes a household with a curbside mailbox. Many Indiana children — apartment buildings, city housing, cluster boxes — have never seen one. When an item's difficulty depends on housing type rather than reading skill, it is measuring circumstance.

The question this raises is not rhetorical: what did the bias and sensitivity review say about experience-contingent items, and is it published?

Text stripped of the scaffolds children learned with

Children learn to read with pictures, voices, and familiar stories. The instrument presents dense, unillustrated text — walls of it, in tight paragraphs. The scaffold isn't a crutch to be removed at test time; for early readers it is part of how meaning is constructed. Visual context could be added to many item families without touching the construct.

The attention tax, and the slow-reader double penalty

The long informational passages — the bee passage is the canonical example — are attention and working-memory tests wearing a reading test's clothes. A slow reader pays twice: the passage takes longer, so by the time a question asks what a worker bee does first, the answer is further behind them than it is for a fast reader. Adults would resent this task. We hand it to eight-year-olds with a retention decision attached.

Subjectivity at the proficiency line

"Which sentence best shows…", "what is the main idea…", and predict-what-the-character-does-next items admit multiple defensible answers — especially for an imaginative child who personifies characters, exactly as we encourage at home. A fable-prediction item is a judgment call about a fictional lion; for a student near the cut, it can also be the difference in a retention decision. The classroom's rational response is formula-teaching: "the main idea is the first or last sentence." That is not comprehension; that is compliance.

Conventions divorced from communication

Grammar items ("the three dogs barks") and capitalization items test conventions in isolation. All of the wrong answers still communicate. The deeper question for an assessment of reading: which of these sentences did the child fail to understand? None. Something else is being scored.

Listening items that blame the listener

A listening item buries a permission-slip requirement mid-announcement and scores the child for missing it. In adult life we call that a communication failure by the speaker. The item teaches children that when the main point is missed, the fault is theirs — the inverse of what we teach about clear communication.

Math that is secretly reading

The math sections arrive as word problems — situational narratives, explain-your-answer prompts, vocabulary like "quadrilateral." For a child on the reading borderline, the math score partially re-measures reading. And the explain-in-writing prompts penalize the child who could talk through their reasoning perfectly — the verbal pathway exists in the standards but barely in the instrument.

The calculator question is worth asking straight: if a child gets every answer right with a calculator, what exactly did we establish they cannot do — and is it the thing this test is for?

02Directions worth exploring

None of these require abandoning measurement. They require deciding that the goal is a reading population, not a defensible ranking — and then designing backward from that.

Bring the home in

Reading is substantially built at home — the same story requested every night until every word is known. That is reading development. Policy currently treats the home as outside the system. Family-engagement incentives, familiar-story instruments, and opt-in home evidence would align the assessment with where the skill is actually formed.

Give speech a place

Some children's oral language runs far ahead of their test performance — including English learners who converse fluently and follow every story. A structured verbal demonstration of understanding is valid evidence of comprehension. Hand-scoring already exists in the system; structured verbal response could too.

Assess inside stories

The target skills — tricky pairings, decoding, comprehension — can be embedded in a continuous, illustrated narrative a child can engage with, rather than clinical fragments about nothing. Familiar format is not memorization; it is the difference between demonstrating reading and surviving an unfamiliar ritual.

Make the signal continuous

Checkpoints exist; the idea can go further. Low-stakes signals gathered through the year — including a child privately marking "this didn't click" during a lesson, with help arriving later at home, never a spotlight in class — produce a truer picture than a single high-stakes sitting, and produce it in time to act.

Capture the pathway, not just the answer

One child explains in words, another draws, another makes tally marks. How a child solves reveals how a child thinks — and in aggregate, how a classroom thinks. The instrument currently discards this. Recording solution pathways would serve the student, the teacher, and every conversation the department has about instruction.

Treat stress as a design constraint

High-stakes pressure changes what an instrument measures. A stressed child performs below their skill; a retained child repeats the same instruction while learning mainly that they are "not good at school." If the design goal includes accuracy, the instrument should be engineered to minimize stress artifacts — pacing, framing, layout, spacing, stakes — the way any measurement system controls for interference.

Design it with the people in it

Co-design with teachers and with children, then refine continuously against the actual goal — every child reading — rather than the proxy of a stable score distribution. If the aspiration of one hundred percent is unreasonable, the reasons it is unreasonable are precisely the data the system should be surfacing.

03Questions Indiana can ask

Each of these can be asked in a public meeting, in a procurement conversation, or in a letter — and each has a knowable answer. Grouped by what the answer would inform.

About the instrument

1

How standardized is the testing experience from school to school — forms, conditions, preparation intensity? If experiences differ, school-to-school score comparisons carry noise the public conversation currently ignores.

2

For judgment-based item families ("best sentence," "main idea," prediction), what do response-distribution and rater data show about how many answers are defensible? If strong readers split across two answers, the item measures convention, not comprehension.

3

What did bias and sensitivity review conclude about experience-contingent items — and can those conclusions be published? The mailbox problem generalizes: any item whose difficulty depends on a child's circumstances is measuring the circumstance.

4

How much of the reading score reflects attention and working-memory load rather than decoding and comprehension — and has passage length and density ever been studied against the construct? The slow-reader double penalty is a measurable artifact, not a hypothesis.

5

In the math assessment, how much score variance is explained by reading load? A reading-borderline child's math score partially re-measures reading; the state should know the size of that effect.

6

What would it take to publish a released item form each cycle for public scrutiny? Other states release forms. Public confidence in a $126M instrument should not require taking anyone's word.

About the data

7

Where are the non-passing students in every other measure — attendance, oral language, other subjects, growth trajectories? A retention decision made on one instrument, without the whole-child view the state already possesses, is a data-illiteracy problem at the decision layer.

8

What are the longitudinal outcomes of retained students in Indiana specifically — retention into the same instruction versus promotion with targeted support? The national literature is contested; Indiana is generating its own evidence right now and should be studying it.

9

How many students sit within one standard error of the 446 cut each year, and what happens to them? Children inside the measurement-error band could flip outcomes on a different testing day. The count is knowable and should be public — see the briefing.

10

What did the second-grade opt-in cohort's high pass-through actually measure — skill acquisition or format familiarity? If practice on the format moves scores this much, the format is a larger variable than the conversation admits.

About the decision frame

11

Is the goal a defensible ranking of students, or a population of readers? What would the instrument look like if designed purely for the second? These produce different designs. The current instrument leans heavily toward the first.

12

What evidence would the state require to admit structured verbal demonstration into the proficiency determination? Naming the evidentiary bar turns "that's not how it's done" into a research agenda.

13

Where will AI scoring be allowed to operate, where does hand-scoring remain, and what is published about scoring error by item type? The scoring boundary is a policy choice being made inside a vendor relationship; it should be visible.

14

Where do family-engagement incentives sit in the literacy policy stack? The science of reading points at the home as much as the classroom; the policy stack currently funds only one of them.

The standing offer. Every observation above is also a buildable feature — several already exist in the working prototype. The point of organizing these notes is not critique; it is that the conversation about what this assessment could be deserves to happen in public, with working examples on the table.
🧪 Working PrototypeThe assessment experience, reconstructed — structure, item types, tools, scoring 📊 Analytical BriefingIREAD-3 & ILEARN — instruments, cut scores, measurement error, the numbers 🗺️ IDOE Systems MapEvery active system, vendor, and data flow in Indiana education