Dossier 1 — Does background knowledge drive reading comprehension?

Historical research appendix; not the current appraisal. This earlier draft retains superseded bounded-returns wording and source metadata. Read the current D1 judgments and the dated source correction. The final 1994 Rosenshine–Meister article has not been read in full.

And does it follow that there is no general, teachable reading-comprehension skill?

DRAFT — MACHINE-WRITTEN, UNREVIEWED BY THE EDITOR. Drafted 2026-08-18 by an LLM (Claude Opus) from SCOPE.md, data/corpus_consolidated.json, data/theses/*, data/evidence_registry/*, and web verification of the external literature. Nothing here has been signed off by a human. The editorial assessment in §10 is supplied in two forms: an empty template that the editor owns, and a clearly-labelled machine provisional read that exists only to be argued with. Where a fact could not be verified against a primary source, it says unverified — that word is load-bearing and should not be smoothed away in editing.

What is checked: every quoted occurrence in §4 is a literal substring of that claim's source_passage in the corpus (38/38 verified by verify_quotes.py; re-run it after any edit). Every occurrence→proposition correspondence in §4 is read off data/theses/mappings.json. Every external citation in §6–§8 carries a verified_by note naming the URL actually opened.


1. Research question and stakes

The question. Does relevant prior knowledge drive reading comprehension — and does it follow that there is no general, teachable reading-comprehension skill?

Why it decides other things. This is the floor of the whole Hirsch project. If background knowledge is the dominant determinant of comprehension then reading tests are knowledge tests, the achievement gap is a knowledge gap, the elementary reading block is misallocated, curriculum content is not interchangeable, and content-neutral "skills" standards are a category error. Dossiers 2 (do knowledge-rich curricula work?), 3 (France) and 4 (what is worth knowing?) all inherit from here.

If the floor is only partly sound — if knowledge matters a great deal and general processes also exist and are partly trainable — then the policy conclusion becomes an allocation argument about marginal returns rather than an existence argument, and it has to be won on much narrower ground than Hirsch fights on.

It also matters because this is where Hirsch is strongest. An atlas that cannot state precisely what he has established here has no claim on a reader's attention for anything else.

What the corpus brings that a reading list does not. Three things, all checkable:

  1. The negative claim ("no general reading skill") and the mechanism claim ("text is incomplete") have different first dates. The mechanism is 1977; the negative claim does not appear as an asserted proposition until 1987 (§12).
  2. The qualifier attached to the decisive experiment — "for that particular text" — is present in 2006 and gone by 2010 (§4, D1-P4).
  3. Across ten books and 892 recorded objections, not one of Hirsch's replies on this question is addressed to a named living critic (§9).

2. Definitions and scope

2.1 Three theses that must not be run together

The central discipline of this dossier. These have different evidential requirements and different truth statuses, and most of the public heat in the reading debates comes from people attacking one while their opponents defend another.

ThesisLogical formWhat would establish it
T1Relevant prior knowledge is a large causal determinant of reading comprehension.Causal, gradedExperimental manipulation of knowledge → comprehension; longitudinal directionality tests
T2At the margin of school time, adding knowledge buys more comprehension than adding strategy instruction.Comparative, causal, policy-relevantHead-to-head equal-time trials; or a well-identified time-substitution study
T3There is no general reading-comprehension skill independent of topic knowledge.Existence claim (negative, universal)Showing that no comprehension variance survives controlling for topic knowledge; or that no content-general process is trainable

Hirsch states T3 in its strongest form — "there is no such thing as a general reading skill independent of specific unstated topic knowledge shared between writer and reader" (sk:ch12_C21, 2023) — while the evidence he marshals establishes T1 and gestures at T2. The job of §4–§7 is to show exactly where the argument crosses from T1 into T3, and on what.

2.2 "General skill" has at least three readings, and Hirsch uses all of them

  1. A content-free process that operates in comprehension (inference-making, coherence monitoring, constraint satisfaction). Denying this is a strong claim about cognitive architecture.
  2. A content-free ability trainable with continuing returns. Denying this is a claim about instructional yield. Brief strategy efficacy is supported, but the continuing-return curve is underdetermined by the reviewed dose evidence.
  3. A measured trait that generalises across topics for a given reader. Hirsch affirms this one when he defends standardized reading tests: "well-educated people can and do exhibit a general proficiency in reading comprehension, and we can indeed reliably measure that proficiency on reading tests, just as the test-makers claim" (kd:ch2_C73, 2006).

His rhetoric asserts (1); his evidence supports (2); his defence of reading tests requires the denial of (3). Until this is fixed the dispute is partly verbal — see crux C1.

2.3 Other definitions to fix before appraising anything

2.4 Scope boundaries

In scope: the cognitive mechanism (schema, situation model, working memory); the empirical knowledge–comprehension relation; the measurement claim about reading tests; the negative claim about general skill; the immediate instructional inference about the reading block.

Out of scope, deliberately deferred:


3. Hirsch's steelman (298 words)

Written language is not speech written down. A writer addressing strangers must decide what to leave unsaid; efficiency in a literate culture depends on how much can be taken for granted (cl:ch1_C23, 1987). A text is therefore systematically incomplete by design: "the explicit meanings of a piece of writing are the tip of an iceberg of meaning; the larger part lies below the surface of the text and is composed of the reader's own relevant knowledge" (cl:ch2_C5).

Comprehension is consequently constructive. The reader assembles a mental model of the situation the text describes, supplying the unstated connections from memory (cl:ch2_C33; how-to-educate-a-citizen:ch5_C64). Whether this succeeds is not a matter of effort or technique, because it is bounded by a hard architectural limit: short-term memory holds four to seven chunks (cl:ch2_C6), so the needed schemata must already be in long-term memory and quickly available. "When the appropriate schemata are not quickly available… the limits of short-term memory are quickly reached, and the process has to be painfully restarted and restarted" (cl:ch2_C87). You cannot look it up; the lookup costs more working memory than it saves.

What is required is not general ability but the particular content the writer assumed. Hence the experimental result: poor readers who know baseball out-comprehend good readers who do not (kd:ch2_C77, 2006). Hence too the measurement consequence — a reading test is a knowledge test in disguise (ae:chpreface_C6, 2022) — and the finding that vocabulary size, a proxy for accumulated world knowledge, is "the single most reliable correlate to reading ability" (wkm:ch3_C29).

Comprehension-strategy instruction produces a real but small, quickly exhausted benefit: it teaches a child what reading is for, then stops paying (wkm:ch1_C73). Every further hour spent on it is an hour of knowledge not built. The efficient lever is a cumulative, coherent, shared body of knowledge, taught early (kd:ch4_C44).


4. Minimal argument backbone

Seven propositions. Each is a premise for the ones below it: D1-P1–P3 are the mechanism, D1-P4–P5 the empirical predictions, D1-P6 the negative conclusion, D1-P7 the policy bridge to Dossier 2.

Each row gives the occurrence id (linked to the raw atlas), the book and year, the canonical proposition it maps to in data/theses/mappings.json, and the quoted source_passage. All 38 quotes below were verified as literal substrings of the corpus record (verify_quotes.pyverified 38/38 quoted occurrences against source_passage (0 failed)).

4.1 How the dossier backbone relates to the canonical propositions

The canonical layer (25 propositions, data/theses/canonical_propositions.json) was drafted independently of this dossier. Reconciling the two is itself informative:

DossierCanonicalFit
D1-P1 text is incomplete≈ P024 same, 1 supports. Exact fit.
D1-P2 the knowledge is topic-specific and shared≈ P023 same, 2 supports. P02 does not separate P1 from P2 — the canonical layer merges "text is incomplete" with "the missing content is community-shared". Those are different claims with different evidence (Bransford for the first, Steffensen and Krauss for the second), and the merge hides that. Recommend splitting P02.
D1-P3 working memory binds→ P02 (supports) ×4, P03 (same) ×1, and one unmappedpoc:ch4_C116 (Miller's limit, 1977) maps to no canonical proposition at all. There is no canonical proposition for the working-memory constraint, although it is what makes knowledge non-substitutable by lookup and is therefore load-bearing for D1-P6. Recommend adding one.
D1-P4 knowledge outweighs measured ability≈ P011 same, 3 supports, 1 → P02.
D1-P5 tests measure knowledge; vocabulary is the best correlatesplits: P14, P06, P01The measurement claim (P14) and the vocabulary-proxy claim (P06) are genuinely different propositions that Hirsch runs together.
D1-P6 no general teachable skill≈ P01 (+P03)4 same on P01, 1 same on P03. The flagship T3 statement sk:ch12_C21 maps to P01 at confidence 1.0.
D1-P7 curriculum is the leversplits: P03, P01, P08, P24A bridge proposition, as expected. kd:ch4_C44 → P24 narrower; sk:ch12_C24 → P08 narrower; re:ch6_C51 → P08 broader.

The load-bearing inference is D1-P4/P5 → D1-P6, and it is the weakest joint in the chain (§5).

4.2 D1-P1 — Written text is systematically incomplete; comprehension requires the reader to supply unstated content from memory

OccurrenceBook / yearCanonicalQuoted source_passage
poc:ch4_C74PoC 1977P02 (same)"It consists of linguistic 'rules' and semantic conventions. It embraces large domains of tacitly shared knowledge, and it includes tacit suppositions about the theme and tendency of the text as a whole."
cl:ch2_C5CL 1987P02 (same)"The explicit meanings of a piece of writing are the tip of an iceberg of meaning; the larger part lies below the surface of the text and is composed of the reader's own relevant knowledge."
cl:ch2_C33CL 1987P02 (supports)"To make sense of what we read, we must use relevant prior knowledge to form a model of how sentence meanings hang together."
kd:chappendix_C19KD 2006P02 (same)"Sharing the unsaid makes it possible for them to comprehend the said. It is the very thing that makes them a speech community."
how-to-educate-a-citizen:ch5_C64HtEC 2020P02 (same)"much of the information needed to understand a text is not provided by the information expressed in the text itself but must be drawn from the language user's [prior] knowledge"

Note the 1977 date on the first row. The mechanism predates the education argument by a decade; in The Philosophy of Composition it is a theory of prose readability, not of schooling.

4.3 D1-P2 — The knowledge that must be supplied is topic-specific and shared — the unstated store a writer assumes in a speech community

OccurrenceBook / yearCanonicalQuoted source_passage
poc:ch4_C75PoC 1977P02 (supports)"A shrewd decision about the knowledge that the writer can tacitly assume in his audience may be the most important decision the writer makes."
cl:ch1_C23CL 1987P02 (supports)"If they can take a lot for granted, their communications can be short and efficient, subtle and complex. But if strangers share very little knowledge, their communications must be long and relatively rudimentary."
kd:ch4_C5KD 2006P02 (same)"Unless writers and their readers internalize the shared knowledge of the wider speech community (such as shared knowledge about baseball), they cannot expect the blanks to be filled in; they cannot be successful writers or proficient readers."
the-making-of-americans:ch1_C51MoA 2010P02 (same)"This means that communication depends on both sides, writer and reader, sharing a basis of unspoken knowledge."
sk:ch12_C26SK 2023P02 (same)"shared background knowledge is a necessity for comprehension which must include not just very general, nationally shared background knowledge but also specifically shared topic knowledge."

4.4 D1-P3 — Working-memory limits make possession and speed of access, not availability of information, the binding constraint

OccurrenceBook / yearCanonicalQuoted source_passage
cl:ch2_C6CL 1987P02 (supports)"The mind cannot reliably hold in short-term memory more than about four to seven separate items…"
cl:ch2_C87CL 1987P02 (supports)"When the appropriate schemata are not quickly available, and the reader is forced to do a lot of pondering to construct them at the time of reading, the limits of short-term memory are quickly reached, and the process has to be painfully restarted and restarted."
cl:ch2_C112CL 1987P02 (supports)"schemata perform two essential functions that are relevant to literacy. The first is storing knowledge in retrievable form; the second is organizing knowledge in more and more efficient ways, so that it can be applied rapidly and efficiently."
poc:ch4_C116PoC 1977— (none)"The basic discovery is the precisely limited capacity of short-term memory…"
wkm:ch4_C140WKM 2016P02 (supports)"verbal comprehension is a form of problem solving. Topic familiarity makes the task of forming a situation model fast and accurate, leaving more space in working memory to figure out word relations…"
how-to-educate-a-citizen:ch5_C84HtEC 2020P03 (same)"It was ingrained, specific factual knowledge, stored in long-term memory, not some general mental skill, that explained the skilled performance."

The "you can't look it up" corollary is carried in the corpus as an objection-and-response rather than as a claim: cl:ch2_OBJ11 ("Students do not need to possess specific facts because they can simply look them up…" → "Reference tools are impractical and too slow…") and the parallel reply in kd:ch4.

4.5 D1-P4 — Therefore comprehension is topic-variable within a reader, and topic knowledge can outweigh measured general reading ability

OccurrenceBook / yearCanonicalQuoted source_passage
cl:ch2_C57CL 1987P02 (supports)"among seven-year-olds who score the same on reading and IQ tests, those who have greater knowledge relevant to the text at hand show superior reading skills."
kd:ch2_C77KD 2006P01 (supports)"the reading comprehension of the low-skills, baseball-knowing group proved superior to the reading comprehension of the high-skills, baseball-ignorant group for that particular text."
the-making-of-americans:ch6_C22MoA 2010P01 (same)"It has been shown decisively that subject-matter knowledge trumps formal skill in reading"
wkm:ch1_C100WKM 2016P01 (supports)"topic familiarity, often indicated by word familiarity, is a stronger predictor of comprehension than formal ability to deal with difficult syntax."
sk:ch12_C29SK 2023P01 (supports)"as soon as one looks elsewhere in language comprehension—especially in second-language learning—one discovers that 'topic familiarity' is a dominant theme. (That is also true for our native speakers, as Recht and Leslie have shown in the baseball study.)"

Watch the escalation in the year column. 1987: "superior reading skills." 2006: the result stated for one text, "for that particular text". 2010: "shown decisively." 2023: "a dominant theme." The qualifier survives in 2006 and is gone by 2010. This is the single clearest instance of drift in the dossier, and it is visible only because the occurrences are dated.

4.6 D1-P5 — Therefore reading tests measure knowledge, and vocabulary breadth is the best single correlate of comprehension

OccurrenceBook / yearCanonicalQuoted source_passage
cl:ch1_C26CL 1987P06 (related)"According to John B. Carroll… the verbal SAT is essentially a test of 'advanced vocabulary knowledge,' which makes it a fairly sensitive instrument for measuring levels of literacy."
kd:ch6_C21KD 2006P14 (same)"because the tests have been presented as tests of formal comprehension skills, they are unwittingly unfair… The tests favor children who happen to have domain knowledge relevant to the passages in the test."
the-making-of-americans:ch6_C17MoA 2010P01 (supports)"Many educators see this question as probing the general skill of 'finding the main idea.' It does not."
wkm:ch3_C29WKM 2016P06 (same)"Vocabulary size is the single most reliable correlate to reading ability."
ae:chpreface_C6AE 2022P14 (supports)"An American reading test is an American ethnicity test in disguise. It's 'in disguise' because the shared background knowledge – the shared ethnicity required for comprehension – is unstated."

Internal tension to hold on to. kd:ch2_C73C74 (2006) defends the validity of standardized reading tests precisely because they sample many domains, and asserts a real, measurable general proficiency: "well-educated people can and do exhibit a general proficiency in reading comprehension… just as the test-makers claim." kd:ch6_C21 (same book) says the tests are "unwittingly unfair" because they measure domain knowledge. Both can be true only if "general proficiency" means breadth of knowledge — which is Hirsch's actual position (kd:ch1_C55: "The only thing that transforms reading skill and critical thinking skill into general all-purpose abilities is a person's possession of general, all-purpose knowledge"). That reading dissolves the tension and also drains T3 of its bite: the general skill exists; it is just made of knowledge.

4.7 D1-P6 — Therefore there is no general, teachable reading-comprehension skill; strategy instruction gives a small, one-time, quickly-exhausted gain

OccurrenceBook / yearCanonicalQuoted source_passage
swn:ch5_C129SWN 1996P01 (same)"There is no accurate way to describe reading ability as a purely formal skill, or to remove from it the information-based knowledge disparaged as 'factoids.'"
kd:ch2_C145KD 2006P01 (same)"It is not mainly comprehension strategies that young children lack in comprehending texts but knowledge—knowledge of formal language conventions and knowledge of the world."
wkm:ch1_C73WKM 2016P03 (supports)"Lessons in reading strategies offer an initial score boost for test taking, but are quickly learned; they plateau fast, and they don't have to be practiced."
wkm:ch4_C125WKM 2016P01 (same)"But reading comprehension, unlike decoding, is not a general skill for standard written English."
how-to-educate-a-citizen:ch5_C89HtEC 2020P03 (same)"All this was summarized in Anders Ericsson's remark 'There is no such thing as developing a general skill.'"
sk:ch12_C21SK 2023P01 (same)"In sum: there is no such thing as a general reading skill independent of specific unstated topic knowledge shared between writer and reader."

Hirsch also states the qualified version, which is the defensible one:

His route from D1-P4/P5 to D1-P6 runs mostly through expertise research — de Groot's chess players, Chase & Simon, Ericsson, Simon's "dormative power" jibe — not through comprehension research. That is an analogy from expert performance in a bounded task domain to comprehension of arbitrary prose, and it is the one inference in the chain the corpus never tests directly (§5, I5).

4.8 D1-P7 — Therefore the school's efficient lever is a cumulative, coherent, shared knowledge curriculum, not more strategy practice (bridge to Dossier 2)

OccurrenceBook / yearCanonicalQuoted source_passage
swn:ch7_C41SWN 1996P03 (same)"The best foundation for general skill is not the inculcation of abstract strategies but the inculcation of broad knowledge."
kd:ch2_C87KD 2006P01 (same)"Since relevant, domain-specific knowledge is an absolute requirement for reading comprehension, there is no way around the need for children to gain broad general knowledge in order to gain broad general proficiency in reading."
kd:ch4_C44KD 2006P24 (narrower)"Since at least ninety minutes per day are currently allotted to reading in early grades, about an hour could be devoted to the language and world knowledge that is most important for competence…"
the-making-of-americans:ch6_C35MoA 2010P01 (supports)"it is far more fruitful to teach children the broad array of domain-specific knowledge they will need to become mature readers than to practice reading strategies such as 'finding the main idea,' 'clarifying,' and 'summarizing.'"
sk:ch12_C24SK 2023P08 (narrower)"shared topic knowledge is not an educationally useful characteristic—unless and until specific and predictable topic knowledge has been taught and learned by all the readers in the class."
re:ch6_C51RE 2024P08 (broader)"supplying shared background knowledge systematically is a key principle of an effective curriculum in every modern nation and every modern literate language."

sk:ch12_C24 is the most important claim in that table and the least noticed: it concedes that individual topic knowledge is not an instructional lever. Only class-wide predictable knowledge is. That converts the mechanism argument into a curriculum-coordination argument — which is what Dossier 2 must evaluate, and which a purely cognitive defence of D1-P1–P6 does not establish.


5. Inference audit

One row per arrow in the backbone. "Corpus warrant" quotes a regenerated warrant (provenance: regenerated_2026-08-18_gemini3-flash, ids of the form W_R*) attached to one of the quoted occurrences; these are the warrants rebuilt after the sliding-window ID-remap repair, so they are wired to an evidence item that actually exists. Thirteen of the 38 backbone occurrences carry one. Older W* warrants on these claims are mostly flagged stale (evidence_id_is_a_claim_id or evidence_relinked_to_different_claim) and are not quoted here.


I1. D1-P1 → D1-P2 · text is incompletethe missing content is community-shared

Type: deductive-with-a-smuggled-premise. Required assumption: that the writer's assumptions about the reader are drawn from a common public store rather than from a guess about a particular reader, a genre convention, or the text's own internal build-up. Only the first reading licenses the curricular conclusion; the other two are equally consistent with the incompleteness premise. Corpus warrant (poc:ch4_W_R23, implicit, evidence: Carroll & Freedle, Language Comprehension and the Acquisition of Knowledge): "Successful communication relies on a vast substrate of unstated information that functions as a necessary condition for interpreting explicit signs." Recorded vulnerability: "The 'illusion of transparency' where speakers overestimate the extent to which their internal knowledge is actually shared by their audience." Assessment: the inference is sound for the class of texts addressed to strangers, which is the class Hirsch cares about, and it is genuinely well-evidenced (Steffensen 1979 is the clean test — §7, underuse A). The slide is that Hirsch runs the shared reading as if it followed from incompleteness alone.


I2. D1-P3 → possession, not access ⟹ lookup cannot substitute for knowledge

Type: deductive from a capacity premise. Required assumptions: (a) that the capacity limit applies during reading with the magnitude Miller reported; (b) that consulting a reference costs more capacity than it frees; (c) that the relevant knowledge cannot be supplied by the text itself in time. Corpus warrant (cl:ch2_W_R1, implicit, evidence: "George Miller's research on the 4-7 item limit of short-term memory across various sensory domains"): "Standardized psychological test results across multiple sensory modalities provide a reliable basis for defining the structural constraints of human cognition." Recorded vulnerability: "The 'magic number' may vary significantly depending on the complexity of the items or the specific task environment." Assessment: premise (a) is dated but directionally intact. Miller himself called the coincidence of the two sevens "only a pernicious, Pythagorean coincidence" and explicitly said the span of absolute judgment and the span of immediate memory "are quite different kinds of limitations"; Cowan (2001) puts the real limit nearer four chunks, and whether the limit is slots or a flexible resource is still disputed. The qualitative point survives; the number does not, and Hirsch keeps quoting the number. (b) and (c) are asserted, not tested — but Hirsch does convert them into an explicit objection-and-response (cl:ch2_OBJ11), which is more than he does for most of his opponents' positions.


I3. D1-P1–P3 → D1-P4 · mechanism ⟹ topic knowledge can outweigh measured reading ability

Type: inductive, from a small number of matched-topic experiments. Required assumptions: (a) that the experimental knowledge measure indexes the knowledge the mechanism says is needed; (b) that the comprehension measure is not itself knowledge-loaded in a way that inflates the result; (c) that the ability split is a real ability split. Corpus warrant (kd:ch2_W_R17, implicit, evidence: "An 'elegant experiment' (The Baseball Study)… Recht and Leslie (1988), implied by endnote 18"): "If a group with inferior technical skills outperforms a group with superior technical skills on a specific task when subject knowledge is the only variable favoring the former, then that knowledge is the primary driver of performance." Recorded vulnerability: "The 'high-skills' group may have had such low motivation for the specific topic (baseball) that they failed to apply the strategies they possessed." Assessment: the corpus's own recorded vulnerability is the weakest of the available objections. The real problems are assumptions (a) and (c). Reynolds, Hattan & Markham (2025) document that the baseball literature's knowledge measures are "vocabulary and baseball trivia"; Shanahan notes the sample excluded readers below the 30th percentile, which truncates the ability contrast the design depends on. The result also rests on main effects plus a null interaction, not on a demonstrated crossover. See §7, card 1.


I4. D1-P4 → D1-P5 · topic-variable comprehension ⟹ reading tests measure knowledge

Type: deductive, and largely correct — but it damages the rest of the chain. Required assumption: that test passages sample topics unevenly with respect to readers' knowledge. Uncontroversial. Assessment: valid, and Hirsch defends it well (kd:ch6_OBJ4/OBJ5: supplying the background inside the item fails, because unfamiliar students must assimilate new content while answering). The cost is crux C4. If reading tests are knowledge tests, then evidence that knowledge predicts reading-test scores is partly definitional rather than confirmatory — and most of the supporting evidence in §7 uses reading-test-like outcomes. Hirsch cannot both use test scores as the outcome that proves knowledge matters and characterise those tests as knowledge tests, without conceding that the causal claim needs an outcome that is not a knowledge test. The corpus records no engagement with this.


I5. D1-P4/P5 → D1-P6 · knowledge matters a great dealno general skill existsthe weak joint

Type: invalid as stated. A graded causal claim (T1) plus a measurement claim cannot entail a universal negative existence claim (T3). What would be required: that no comprehension variance survives controlling for topic knowledge, or that no content-general process is trainable. Neither is shown anywhere in the corpus. What the corpus offers instead: two substitutes.

Assessment: this is where the argument breaks. Everything above I5 is defensible. D1-P6 as stated is not derived; it is asserted, and — per the thesis layer — 106 of the 121 same occurrences of P01 carry no direct evidence item at all. The flagship sk:ch12_C21 is a summary sentence with no attached evidence.


I6. D1-P6 → D1-P7 · no general skill ⟹ knowledge instruction is the efficient use of the reading block

Type: practical/economic inference. Requires a marginal-return comparison, not an existence claim. Required assumptions: (a) the time is genuinely substitutable; (b) knowledge instruction's comprehension yield per hour exceeds strategy instruction's; (c) the knowledge taught transfers to unseen texts. Assessment: (a) is plausible; (b) is crux C2 and no study in the ledger measures it — no equal-time head-to-head trial exists; (c) is crux C3, and the best evidence is mixed in a specific, informative way: Kim et al. 2023 found ES = .22 on near-transfer passages, .17 on mid-transfer, and .04, n.s., on far-transfer passages with no lexical overlap. Note that the practical conclusion (D1-P7) would survive even if D1-P6 were abandoned entirely — T2 is enough for it. Hirsch's policy argument does not need his most contested claim.


I7. D1-P7 individual ⟹ class-wide · the hidden step

Type: a concession that functions as a premise change. sk:ch12_C24 (2023): "shared topic knowledge is not an educationally useful characteristic — unless and until specific and predictable topic knowledge has been taught and learned by all the readers in the class." Corpus warrant (sk:ch12_W_R11, explicit, evidence: "The Kim study shows that topic knowledge is not educationally useful unless it is specifically taught to the whole class"): "Individual variations in prior knowledge among students render topic-based comprehension gains useless for standardized classroom instruction unless the curriculum creates a uniform knowledge base." Recorded vulnerability: "Differentiated instruction or personalized learning pathways could leverage students' diverse existing knowledge bases without requiring a rigid, identical curriculum for all." Assessment: this is the most consequential sentence in the dossier's slice. It concedes that the cognitive mechanism by itself yields no instructional lever, and relocates the argument onto curriculum coordination — an organisational claim requiring intervention evidence, not laboratory evidence. Everything after this point belongs to Dossier 2.


6. Evidence ledger

source_id is the key in data/evidence_registry/sources.json. "Books" is the number of the ten books in which the registry records at least one occurrence. Directness is a distinction the corpus's evidence_type field does not encode and this dossier adds: tests (directly tests the proposition) / illustrates (analogy or demonstration) / context (trend or historical background) / program (outcome evidence for Core Knowledge).

#source_idSourceKindIndependenceBears onDirectnessBooksStatus
1S0027Recht & Leslie 1988, baseball studystudyprimary_independentD1-P4 (rhetorically P6)tests6Contested — never directly replicated; measures criticised
2S0021, S0070Bransford & Johnson 1972; Bransford, Barclay & Franks 1972studyprimary_independentD1-P1tests6Classic, robust
3S0033Kintsch & van Dijk; Kintsch 1988 CI modelstudyprimary_independentD1-P1, P3illustrates (cited as authority)4Established — and partly against Hirsch
4S0012Miller 1956, "Magical Number Seven"studyprimary_independentD1-P3illustrates7Classic, superseded in detail
5S0013, S0080de Groot 1946; Chase & Simon 1973studyprimary_independentD1-P6 (the bridge)illustrates (analogy)6Established for chess; the strong null is superseded
6S0019Ing, Lunney & Olsen, AFQT/NLSY79 reanalysisstudyprimary_independentD1-P5tests (psychometric)5Real, but not peer-reviewed, and misdated in the corpus
7S0362Rosenshine & Meister 1994, reciprocal teachingstudyprimary_independentD1-P6tests (meta-analytic)2Established, over-read
8S0020Willingham & Lovette 2014; Willingham 2006study (in fact commentary)primary_independent but see cardD1-P5, P6illustrates5Sympathetic, over-recruited
9S0105, S0374Cunningham & Stanovich; Stanovich Matthew effectsstudyprimary_independentD1-P4, P5, P7tests (correlational)3Replicated as association; the Matthew pattern is contested
10S0031Hart & Risley 1995, Meaningful Differencesstudyprimary_independentD1-P7 / equity framingcontext4Contested — magnitude not replicated
11S0005Grissmer et al. 2023, Colorado CK lotterystudyckf_affiliatedD1-P7 (→ Dossier 2)program5Not peer-reviewed; high attrition; author on the CKF board
12(not in registry)Cabell et al. 2025, CKLA kindergarten RCTsstudyprimary_independent (declared firewall)D1-P7 equity directionprogram0 — postdates the corpusPeer-reviewed; null on standardized outcomes
AS1348Steffensen, Joag-Dev & Anderson 1979studyprimary_independentD1-P1, P2tests1Underused — the cleanest test of D1-P2
BS0060Anderson, Reynolds, Schallert & Goetz 1977studyprimary_independentD1-P1tests3Underused
CS0366Schneider, Körkel & Weinert 1989, soccerstudyprimary_independentD1-P4tests1The conceptual replication Hirsch has and barely uses
DS0322Hwang, McMaster & Kendeou 2023studyprimary_independentT1 directionalitytests1Best directionality evidence; enters in 2023
ES0161, S0162Kim et al., MORE / content-literacy trialsstudyprimary_independentD1-P7, crux C3program2Best transfer evidence; enters in 2023–24

6.1 What the ledger shows before any single card is read


7. Evidence appraisal cards

Every card ends with verified_by. Where a number could not be seen in a primary or authoritative source it says unverified — that is a real state, not a placeholder to be filled in with a plausible figure.


Card 1 · Recht & Leslie 1988 — "the baseball study" · S0027


Card 2 · Bransford & Johnson 1972; Bransford, Barclay & Franks 1972 · S0021, S0070


Card 3 · Kintsch & van Dijk; Kintsch's construction–integration model · S0033


Card 4 · Miller 1956, "The Magical Number Seven" · S0012


Card 5 · de Groot 1946; Chase & Simon 1973 — the chess analogy · S0013, S0080


Card 6 · Ing, Lunney & Olsen — the AFQT/NLSY79 reanalysis · S0019


Card 7 · Rosenshine & Meister 1994, reciprocal teaching · S0362


Card 8 · Willingham & Lovette 2014; Willingham 2006 · S0020


Card 9 · Cunningham & Stanovich; Matthew effects · S0105, S0374


Card 10 · Hart & Risley 1995, Meaningful Differences · S0031


Card 11 · Grissmer et al. 2023, Colorado Core Knowledge lottery · S0005


Card 12 · Cabell et al. 2025, the CKLA kindergarten RCTs · (not in the registry — postdates the corpus)


Two notable underuses

Underuse A · Steffensen, Joag-Dev & Anderson 1979 — cited once, in 1987 · S1348

Steffensen, M. S., Joag-Dev, C., & Anderson, R. C. (1979). A cross-cultural perspective on reading comprehension. Reading Research Quarterly, 15(1), 10–29. DOI 10.2307/747429. Paywalled at JSTOR, but the identical study is free as Center for the Study of Reading Technical Report No. 97 (July 1978) at ideals.illinois.edu/items/18083.

American and Indian adults read letters describing an American and an Indian wedding, matched for readability (136 and 127 idea units). N ≈ 38–39 (19 Indian, 20 American, one excluded); note that reported F tests use df = (1,35), which does not cleanly match that N — the exact analytic N is unverified. Verbatim results: reading-time interaction F(1,35) = 10.09, p < .01; gist recall interaction F(1,35) = 39.84, p < .01; culturally appropriate elaborations F(1,35) = 208.67; culturally based distortions F(1,35) = 128.24. Cell means for Americans/American vs Americans/Indian: 52.4 vs 27.3 gist units, 5.7 vs 0.2 elaborations, 0.1 vs 5.5 distortions.

Why it matters more than its one citation suggests. This is the cleanest test of D1-P2 in the entire ledger — not "topic knowledge helps" but "the knowledge the writer assumed is culturally specific, and lacking it produces systematic distortion, not merely reduced recall." It is exactly the evidence Hirsch's shared-knowledge thesis needs, and he cites it once, in 1987, in a chapter about literacy and culture (cl:ch1_E30). The corpus also carries the misspelling "C. Joag-Des".

One caution the dossier should record: the recall interaction is driven almost entirely by the American readers collapsing on the Indian passage (52.4 → 27.3); the Indian readers recalled about the same from both (37.9 vs 37.6). The elaboration/distortion split is the cleaner cultural-schema signal. Also, the "Indian" readers were US-resident and reading in a second language. 574 citations; no direct replication found.

Underuse B · Anderson, Reynolds, Schallert & Goetz 1977 — the prison-break passage · S0060

Anderson, R. C., Reynolds, R. E., Schallert, D. L., & Goetz, E. T. (1977). Frameworks for comprehending discourse. American Educational Research Journal, 14(4), 367–381. DOI 10.3102/00028312014004367. Free as Technical Report No. 12 (July 1976) at ideals.illinois.edu/items/17896.

N = 60: "30 students from a section of an educational psychology course (all female) designed specifically for persons planning a career in music education, and 30 students from two weight-lifting classes (all male)." Two deliberately ambiguous passages (prison break / wrestling match; card game / woodwind rehearsal). Verbatim: t(58) = 5.60 and t(58) = 6.53 on the disambiguating multiple-choice tests; passage × background interaction F(1,58) = 48.61; and "62% of the subjects reported that another interpretation never occurred to them."

Why it matters. The 62% figure is the strongest single demonstration in the ledger that schema selection is not experienced as interpretation — readers do not know they have chosen. That is a stronger and more interesting claim than "knowledge helps recall," and Hirsch never uses it.

The caution that must travel with it: background is perfectly confounded with sex (music = all female; wrestling = all male), and the card/music passage turns partly on lexical puns ("hand," "diamonds," "recorder") rather than on background knowledge as such. 722 citations; replication status unverified. All statistics above are from the 1976 technical report; the AERJ version's numbers were not read.


Three sources that enter late and carry more weight than their citation count

C · Schneider, Körkel & Weinert 1989 — the soccer study · S0366. Journal of Educational Psychology 81(3), 306–312, DOI 10.1037/0022-0663.81.3.306. Paywalled; the JEP article's own N is unverified (the authors' book-chapter restatement of the same programme gives N = 576 and N = 185 across grades 3, 5 and 7). Soccer experts and novices, crossed with general aptitude, read a story about a soccer game. Verbatim from the authors' chapter: "High- and low-aptitude soccer experts performed equally well on all measures of text recall and comprehension," with no main effect of general ability and no significant interactions; and "Third-grade experts recalled significantly more text units than both fifth-grade and seventh-grade novices." This is the answer to Shanahan's "it's a one off" objection — a conceptual replication in a different language, sport, country and decade. It sits in the corpus (wkm:ch4_E52, E53, and a thinker record for Wolfgang Schneider) and is cited in one book, against six for Recht & Leslie. Recommendation: promote it. Caveats: quasi-experimental, single domain, single text, and the key result is a failure to reject a null on aptitude, with no equivalence test visible.

D · Hwang, McMaster & Kendeou 2023 · S0322. Reading Research Quarterly 58(1), 59–77, DOI 10.1002/rrq.481, open access (CC BY-NC-ND; free author copy at ERIC ED623594). N = 10,706, ECLS-K:2011, kindergarten through fifth grade, random-intercept cross-lagged panel models with working memory, cognitive flexibility, language proficiency and basic literacy as covariates. Verbatim: the relation is "bidirectional and positive throughout the elementary years"; but "all coefficients of the cross-lagged paths from science to reading were significantly greater than those from reading to science (18.13 < χ²[1] < 49.99, p < .001)." This is the best directionality evidence for T1 in existence, and it half-concedes the opposite direction. Domain knowledge is operationalised only as science. Correlational.

E · Kim et al., the content-literacy trials · S0161, S0162. Kim, J. S., et al. (2023). A longitudinal randomized trial of a sustained content literacy intervention from first to second grade. Journal of Educational Psychology, 115(1), 73–98, DOI 10.1037/edu0000751; and Kim, J. S., et al. (2024). Time to transfer. Developmental Psychology, 60(7), 1279–1297. The 2023 trial: 30 schools, N = 2,952, blocked cluster RCT, rated by the What Works Clearinghouse as "Meets WWC standards without reservations." The numbers that matter for crux C3: science content reading comprehension ES = .18; near-transfer passages .22 (p < .001), mid-transfer .17 (p < .01), far-transfer .04 (n.s.). WWC judged the general-literacy outcome (MAP Reading Total, n = 2,275) indeterminate. The 2024 spiralled trial (N = 2,870, grades 1–3 with a grade-4 follow-up) does reach a domain-general comprehension outcome (ES = .11) and mathematics (.12), sustained at 14 months (.12 reading, .16 maths). Read together: school-taught knowledge transfers, and transfer is graded by how much of the taught vocabulary appears in the passage. Where the lexical footholds are absent, the effect is .04 and not significant. That single number is the most informative datum in the dossier for anyone testing Hirsch's transfer claim — in either direction. Caveat: all three Kim trials draw on overlapping schools in one district, so they are extensions, not independent replications.


8. Named objections — and Hirsch's response

Each entry gives the citation, DOI, open-access status, an objection stated as precisely as possible against a named target (a backbone proposition, an inference id from §5, or a specific evidence card), the limits of the objection itself, and what the corpus records as Hirsch's reply.

A finding that governs this whole section. A name search across all 10,312 claims, 2,885 evidence items and 521 thinker records finds zero occurrences of Catts, Kamhi, McNamara, Gough, Tunmer, Hoover, Counsell, Okkinga, Elleman, Filderman, Sperry, Cowan or Pfost. Shanahan appears — but only as an ally on levelled readers (wkm:ch4_E24); Snow appears only as Catherine Snow, co-author of Chall's Families and Literacy (1982), not Pamela Snow of the 2021 critical review. So for nine of the ten objections below the answer to "does Hirsch respond?" is no response in the ten books — and where he does respond, it is to a reconstructed position, not to a person.


8.1 Reynolds, Hattan & Markham 2025 — the baseball study's measures → attacks Card 1, and through it I3

Reynolds, D., Hattan, C., & Markham, M. (2025). Fair or foul? Interrogating the role of baseball knowledge in studies of knowledge and comprehension. Reading Research Quarterly, 60(1), e575. DOI 10.1002/rrq.575. Open access (CC BY-NC-ND). Companion: Reynolds, D., & Hattan, C. (2024). Baseball, presidents, and state test passages: considering gendered knowledge. The Reading Teacher, 77(6), 997–1000. DOI 10.1002/trtr.2330. Open access. (⚠ SCOPE lists Reynolds as sole author of both; the RRQ paper has three authors, the Reading Teacher piece two.)

The objection, precisely. Not that knowledge is unimportant — that the construct in the canonical demonstration is not the construct Hirsch's curricular argument needs. Nineteen "baseball studies" 1978–2018; 13 used the same two knowledge measures; those measures "focused heavily on vocabulary and baseball trivia"; the standard comprehension text was "deceptively complex." If "topic knowledge" in the showpiece experiment is sport-specific vocabulary plus trivia, the result looks closer to a lexical-access effect on a jargon-loaded passage than to evidence for the broad conceptual background knowledge a curriculum could supply. The companion adds that the proxy is gender-loaded, so group differences may be confounded with sample composition.

Limits of the objection. It reviews one proxy literature. It produces no counter-effect-size and does not claim knowledge is unimportant; the authors' own recommendation is to stop leaning on baseball studies and diversify the evidence base. It leaves non-baseball knowledge studies untouched. Caveat: the RRQ full text returned HTTP 402 on every route; all content above is from Crossref/OpenAlex/Semantic Scholar metadata plus two secondary accounts. The editor should read it (acquisition P0).

Hirsch's response. None — the work postdates the corpus. But note that kd:ch2_OBJ8 ("Reading comprehension consists merely of a general working vocabulary plus the application of formal comprehension strategies") is the closest recorded objection, and Hirsch's reply ("Experiments show that knowledge of the subject matter is far more important than technical reading strategies") presupposes exactly the knowledge/vocabulary distinction Reynolds argues the experiments failed to draw.


8.2 The strategy-instruction evidence base → attacks D1-P6 directly

Four items, and they do not point the same way.

Net. Taken together these establish that D1-P6 fails as stated — strategy effects are non-zero and reach standardized measures at least sometimes — while vindicating Hirsch's practical claim that at realistic classroom scale the returns are small and measurement-inflated. That is T2, not T3.

Hirsch's response. He addresses the category, never an author. kd:ch2_OBJ15: benefits are "real but small and initial; they only serve to show children that reading is communicative. Once that is understood, additional sessions (beyond six) are wasteful and distracting." kd:chappendix_OBJ2: "Most interventions produce some positive data, but these improvements are small and long-term results are unimpressive." The sources are recorded as "A 'predictable chorus' of researchers" and "Proponents of strategy instruction."


8.3 The simple view of reading → attacks D1-P6's architecture, while agreeing with much of it

Gough, P. B., & Tunmer, W. E. (1986), RASE 7(1), 6–10, DOI 10.1177/074193258600700104; Hoover, W. A., & Gough, P. B. (1990), Reading and Writing 2(2), 127–160, DOI 10.1007/BF00401799; Catts, H. W. (2018). The simple view of reading: advancements and false impressions. Remedial and Special Education 39(5), 317–323, DOI 10.1177/0741932518767563free full text at files.eric.ed.gov/fulltext/EJ1191985.pdf; Catts, Adlof & Ellis Weismer (2006), JSLHR 49(2), 278–293, DOI 10.1044/1092-4388(2006/023); Kamhi (2009), LSHSS 40(2), 174–177; Catts & Kamhi (2017), LSHSS 48(2), 73–76.

The objection, precisely. SVR locates comprehension in linguistic comprehension — a modality-general language ability. That yields a reader-side general factor D1-P6 has no room for: children with intact decoding and no obvious topic-knowledge deficit fail comprehension stably across topics because of oral-language weakness. Catts (2018), verbatim: Lonigan et al. "found that a considerable amount of the variance in RC (40%–70%) was shared by decoding and language comprehension and suggested that this shared variance may be the result of one or more general cognitive-linguistic factors." Catts, Adlof & Ellis Weismer identified 57 poor comprehenders, 27 poor decoders and 98 typical readers in eighth grade, with the poor comprehenders showing concurrent language deficits and normal phonological processing.

Where it agrees with Hirsch — and this should be said plainly. (a) There is no reading-specific comprehension skill distinct from language comprehension. (b) Generic strategy instruction has a small payoff — Catts cites Scammacca et al. (2015): "the average effect size of interventions on standardized measures of RC was .19." (c) Catts endorses Willingham directly: "Adequate content knowledge is critical for comprehension and should be central to any instruction directed at improving it (Willingham, 2006). Given the centrality of content knowledge, it is always surprising how little attention has been devoted to it in comprehension intervention." (d) Catts (2021/22, American Educator): "reading comprehension is not a skill someone learns and can then apply in different reading contexts."

Where it disagrees. Knowledge is one input among several, not what comprehension is made of; the poor-comprehender profile is a person-level, topic-invariant deficit a knowledge-only account under-predicts; and Catts and Kamhi replace Hirsch's single moderator (topic knowledge) with a four-way dependency on reader × text × task × purpose.

Limits of the objection. Catts's shared-variance factor is a residual construct he explicitly says is not yet characterised: "It is still possible that the cognitive-linguistic factors that underlie this common variance are malleable, but it is not clear at this point what they are and how they might be changed." Poor comprehenders are a minority subgroup, and their oral-language deficits are themselves partly vocabulary and world-knowledge deficits.

Hirsch's response. No response in the ten books — zero mentions of Catts, Kamhi, Gough, Tunmer or Hoover anywhere in the corpus. This is the most surprising silence in the dossier, because these are the friendliest critics available to him.


8.4 McNamara and colleagues — strategy × knowledge interactions → the sharpest counter to D1-P6

O'Reilly, T., & McNamara, D. S. (2007), AERJ 44(1), 161–196, DOI 10.3102/0002831206298171; McNamara, D. S., Kintsch, E., Songer, N. B., & Kintsch, W. (1996). Are good texts always better? Cognition and Instruction 14(1), 1–43, DOI 10.1207/s1532690xci1401_1; McNamara (2004), SERT, Discourse Processes 38(1), 1–30, DOI 10.1207/s15326950dp3801_1.

The objection, precisely. These isolate a trainable, domain-general processing competence that changes comprehension while topic knowledge is held constant. O'Reilly & McNamara (n = 1,651 high-schoolers), verbatim from the abstract: "Reading skill helped the learner compensate for deficits in science knowledge for most measures of achievement and had a larger effect on achievement scores for higher knowledge than lower knowledge students." ⚠ Correction to SCOPE: the published abstract attributes the compensation to reading skill, not to reading-strategy knowledge. Whether strategy knowledge specifically compensates is unverified — the full text could not be opened. The stronger version of the point is McNamara's SERT training study (n = 42), whose protocol analysis found training "helped these participants to use logic, or domain-general knowledge, rather than domain-specific knowledge to make sense of the text" — the general mechanism D1-P6 denies.

"Are Good Texts Always Better?" adds a boundary condition that appears nowhere in ten books: "readers who know little about the domain of the text benefit from a coherent text, whereas high-knowledge readers benefit from a minimally coherent text… the rewards to be gained from active processing are primarily at the level of the situation model." More knowledge is not monotonically better; what produces deep understanding is inferential work, for which knowledge is the raw material.

Limits of the objection. The compensation widens gaps rather than closing them — reading skill paid off most for students who already had the knowledge, a Matthew pattern consistent with Hirsch. SERT's benefit for low-knowledge readers appeared only on text-based questions, not on bridging-inference questions — shallow gains, not deep ones. The reverse-cohesion effect requires high background knowledge, so it presupposes rather than displaces Hirsch's premise, and O'Reilly & McNamara (2007, Discourse Processes) later narrowed it to less skilled, high-knowledge readers.

Hirsch's response. No response in the ten books — zero mentions of McNamara.


8.5 Kintsch as internal dissent, and McCarthy & McNamara's four dimensions → attacks D1-P6 from inside Hirsch's own authority

Kintsch (1988) is treated in Card 3. Add: McCarthy, K. S., & McNamara, D. S. (2021). The multidimensional knowledge in text comprehension framework. Educational Psychologist 56(3), 196–214, DOI 10.1080/00461520.2021.1872379free at files.eric.ed.gov/fulltext/ED616077.pdf.

The objection, precisely. Two moves. (a) Kintsch's own model is a general-process account; citing him for the situation model while denying general processes takes the vocabulary and discards the architecture. (b) "Background knowledge" decomposes into amount, accuracy, specificity and coherence (verbatim: "Amount refers to how many relevant concepts the reader knows. Accuracy refers to the extent to which the reader's knowledge is correct. Specificity refers the degree to which the knowledge is related to information in the target text. Coherence refers to the interconnectedness of prior knowledge"). Hirsch operationalises amount only — and the framework notes that "amount often serves as the default metric but is often confounded with other dimensions."

Accuracy makes knowledge a signed quantity, and this is the part with no corpus counterpart at all. Verbatim: "Inaccurate knowledge interferes with comprehension and the acquisition of new knowledge"; "students who held misconceptions generated significantly more incorrect inferences during reading and significantly fewer correct inferences than their peers (Kendeou & van den Broek, 2007)"; and inaccurate information "is easy to acquire and often resistant to change." The framework also reports a non-linearity: O'Reilly et al. (2019) found a knowledge threshold — below ~59% on a prior-knowledge test, prior knowledge had no significant relationship with comprehension performance.

Limits of the objection. The framework is emphatically pro-knowledge: it opens by citing prior knowledge as predicting 30–60% of the variance in comprehension, which is at least as strong as anything Hirsch claims, and it supports the knowledge-specificity view. It is a conceptual framework, not new data.

Hirsch's response. No response in the ten books. He cites Kintsch as an authority (S0033, 9 occurrences across 4 books) and never engages the model's architecture.


8.6 Smith, Snow, Serry & Hammond 2021 — the critical review → attacks the magnitude in D1-P4

Smith, R., Snow, P., Serry, T., & Hammond, L. (2021). The role of background knowledge in reading comprehension: a critical review. Reading Psychology 42(3), 214–240. DOI 10.1080/02702711.2021.1888348. Open access (green, via the ECU repository; bronze at T&F).

The objection, precisely. 23 studies, mid-to-late primary children. Verbatim from the abstract: "higher levels of background knowledge have a range of effects that are influenced by the nature of the text, the quality of the situation model required, and the presence of reader misconceptions about the text… Readers with lower background knowledge appear to benefit more from text with high cohesion, while weaker readers were able to compensate somewhat for their relatively weak reading skills in the context of a high degree of background knowledge."

That one adverb is the objection. Partial, not complete, compensation — which is precisely how Recht & Leslie is usually reported, and how Hirsch reports it in kd:ch2_C77 and (with the qualifier removed) in the-making-of-americans:ch6_C22. The review also independently reproduces the McNamara cohesion × knowledge interaction (§8.4) in a school-age population, and names misconceptions as a moderator.

Limits of the objection. 23 studies, one age band, no pooled effect sizes — a critical review, not a meta-analysis; and it is pro-knowledge in orientation. "Somewhat" is an abstract-level summary phrase and the full text could not be opened (both OA routes returned 403), so which studies drove it is unverified.

Hirsch's response. No response in the ten books (Pamela Snow does not appear; the corpus's Snow is Catherine Snow, on Families and Literacy 1982).


8.7 Willingham as sympathetic critic → attacks the use Hirsch makes of him

Treated in Card 8. The objection in one line: Willingham & Lovette answer "Can reading comprehension be taught?" with "not really, beyond a point" — and then recommend teaching the strategies, briefly. Hirsch quotes the headline and drops the recommendation, converting a bounded-returns finding into "useless." Corpus receipt: wkm:ch1_E20 is the plateau citation; wkm:ch1_C87 is the escalation.

Hirsch's response. He does not treat Willingham as a critic at all. Notably, the corpus records Hirsch crediting Willingham with restraining his "excessive absolutism" (ae:chappendix) — a rare self-aware moment worth surfacing on the page.


8.8 Cabell et al. 2025 — the equity direction → attacks the equity claim inside D1-P7

Treated in Card 12. The objection in one line: a peer-reviewed, developer-firewalled cluster-RCT of the curriculum Hirsch's foundation produces found essentially zero transfer to standardized vocabulary, listening comprehension and social-studies knowledge, and an interaction in which higher-vocabulary children benefited more. It does not touch D1-P1–P4. It bears directly on the claim that knowledge-building is especially equalising — a claim stated in this dossier's chapters (ae:ch1_C80, wkm:ch8_C72) and therefore not quarantinable in Dossier 2.

Hirsch's response. None — the work postdates the corpus.


8.9 Shanahan — the methodological sceptic and the pedagogical paradox → attacks Card 1 and I6

Shanahan, T. (2020, 14 March). Prior knowledge, or he isn't going to pick on the baseball study. https://www.shanahanonliteracy.com/blog/prior-knowledge-or-he-isnt-going-to-pick-on-the-baseball-study Free; blog, not peer-reviewed.

Three points, verbatim. (a) "It's a one off. There aren't other studies with this kind of finding… interesting that this has not been replicated in more than 30 years"; the study "was conducted with only a single text and that one designed to be used in a research study," and the sample excluded students reading below the 30th percentile, "potentially eliminating effects of reading ability." (b) "research not only shows that knowledge contributes to comprehension, but to miscomprehension as well (a research result, not an opinion)." (c) The pedagogical paradox: "If I'm always providing kids with the appropriate background knowledge to understand each text used for instruction, then how do students ever learn to take on a text on their own?"

Why (c) is the sharpest. It is a direct challenge to I6 that no amount of mechanism evidence answers: even granting D1-P1–P5 entirely, the instructional inference may be self-undermining. Shanahan elsewhere ("Knowledge or comprehension strategies — what should we teach?", 2023) argues "it should not be a choice between the two."

Limits. A blog. Shanahan does not dispute that knowledge matters; (b) is asserted without a citation beyond Bartlett (1932). A "5% to 30%" variance figure circulating under his name refers to syntax, not prior knowledge — unverified and should not be used.

Hirsch's response. He cites Shanahan approvingly on levelled readers (wkm:ch4_E24, wkm:ch4_C80; thinker record "Tim Shanahan", stance agrees) and never engages him on this. Two thinker records exist for Shanahan and neither concerns the baseball study.


8.10 Counsell — "right, but incomplete" → attacks the sufficiency of D1-P7, not its truth

Counsell, C. (2011). Disciplinary knowledge for all, the secondary history curriculum and history teachers' achievement. The Curriculum Journal 22(2), 201–225. DOI 10.1080/09585176.2011.574951. Paywalled. Companion, open access CC BY-NC: Chapman, A. (ed.) (2021). Knowing History in Schools: Powerful Knowledge and the Powers of Knowledge. UCL Press, DOI 10.14324/111.9781787357303, free at uclpress.co.uk/book/knowing-history-in-schools/.

The objection, precisely. Counsell's stated target is genericism, and on that she is Hirsch's ally — verbatim from the abstract: "such genericism is, first, redundant… and, second, inadequate: disciplinary knowledge and concepts are necessary in order to reach or challenge claims about the past." The move against Hirsch is the corollary: how claims are warranted within a discipline is itself knowledge, distinct from propositional content, and a curriculum specified as a list of things-to-know under-specifies it. Hirsch's account is almost entirely substantive.

Limits. UK-specific, secondary history, conceptual not empirical: it does not show that disciplinary teaching produces better comprehension. The article's own substantive/disciplinary definitions could not be verified — the only accessible full text is an image scan.

Hirsch's response. No response in the ten books. Feeds Dossier 4.


9. Hirsch's responses, as the corpus records them

The corpus holds 892 objections_raised records. These are the ones bearing on this dossier, with the response text as extracted. All twelve were checked against all_objections_raised in data/corpus_consolidated.json.

Objection idThe objectionHirsch's responseAttributed to
cl:ch2_OBJ8The knowledge required is too vast or "profound" to teachThe schemata needed are abstract and elementary (Union won; Grant was a Union general) — cultural pointers, not expertise(none recorded)
cl:ch2_OBJ11Students can just look things upReference tools are "impractical and too slow for the integration required"; schemata must be in memory and quickly available"Implicit/Educational Traditionalists"
cl:ch2_OBJ13Expert performance reflects superior general abilityDe Groot: on random board positions experts "perform no better than novices""Common assumption/pre-de Groot psychology"
cl:ch2_OBJ4Is background knowledge part of meaning, or extralinguistic?Inferences from prior knowledge are integrated into meaning from the outset"Unattributed (posed as a conceptual question)"
kd:ch2_OBJ6Comprehension follows decoding automaticallyNAEP grade-4-to-12 divergence falsifies it"General educational assumption"
kd:ch2_OBJ7Comprehension transfers from simple to complex texts as decoding doesProgress is non-automatic and non-linear because it depends on domain knowledge"Current reading programs / Common assumptions"
kd:ch2_OBJ8Comprehension = working vocabulary + formal strategiesExperiments show domain knowledge matters far more than strategies"Current reading programs"
kd:ch2_OBJ10OBJ11Start from each child's own prior knowledge"Some validity" in the moment of reading, but proficiency needs a shared base"Bank Street School" / "American educational doctrine"
kd:ch2_OBJ14Young readers can't apply the knowledge they have, so strategies are neededEndorses "questioning the author" as a way of triggering application"Reading researchers who advocate for comprehension strategies"
kd:ch2_OBJ15Research shows strategy exercises benefit young readersBenefits are "real but small and initial"; sessions beyond six are "wasteful and distracting""A 'predictable chorus' of researchers"
kd:ch6_OBJ4OBJ5Tests can be levelled by supplying background inside the itemFails: unfamiliar students must assimilate new content while answering — a speed and overload effect"Test-makers (unnamed)" / "State test-makers"
kd:chappendix_OBJ2Strategy-instruction data are positive"Most interventions produce some positive data, but these improvements are small and long-term results are unimpressive""Proponents of strategy instruction"
the-making-of-americans:ch1_OBJ10The SAT decline reflects changed test-taker compositionJencks: scores fell as steeply in Iowa, then 98% white and middle class"Unnamed 'some' (common sociological claim)"

9.1 The structural finding

Not one of these responses is addressed to a named published critic. Every source field is a reconstructed position — "Proponents of strategy instruction," "Current reading programs," "A 'predictable chorus' of researchers." Across ten books and forty-seven years, on the question where he is strongest, Hirsch argues against positions he has assembled rather than against people who hold them.

Two consequences for how this dossier should be read:

  1. It is not evidence of evasion. The reconstructions are mostly fair, and cl:ch2_OBJ11, kd:ch6_OBJ4OBJ5 and the-making-of-americans:ch1_OBJ10 are genuinely good answers to genuinely serious objections.
  2. But it means the objections in §8 have never been answered, and cannot be assumed to be answerable by anything in the corpus. That is why §8 had to be acquired externally, and why the editor's verdict in §10 cannot lean on Hirsch's own replies.

10. Editorial assessment

The assessment has five axes and is recorded separately for T1, T2 and T3 (§2.1). A causal policy conclusion can be textually faithful and empirically underdetermined at the same time; a normative conclusion can be well argued without being proven. This vocabulary exists to keep those apart.

Axes.Textual fidelity — does the proposition fairly represent Hirsch? ② Inferential validity — do the stated premises support the conclusion? ③ Empirical support — how directly and strongly is it tested? ④ Robustness and independence — replication, counterevidence, source independence, scope. ⑤ Normative dependence — how much rests on values rather than findings?

10.1 TEMPLATE — for the editor

Leave the machine read in §10.2 alone; write here. Record author, date, status, rationale, decisive sources, strongest dissent, and what would trigger revision.

Assessed by: ______________________   Date: __________   Status: draft / provisional / settled

T1 — Prior knowledge is a large causal determinant of reading comprehension.
  ① textual fidelity        [ ]  rationale:
  ② inferential validity    [ ]  rationale:
  ③ empirical support       [ ]  rationale:
  ④ robustness/independence [ ]  rationale:
  ⑤ normative dependence    [ ]  rationale:
  decisive sources:
  strongest dissent:
  what would trigger revision:

T2 — At the margin of school time, knowledge buys more comprehension than strategy instruction.
  ① [ ]  ② [ ]  ③ [ ]  ④ [ ]  ⑤ [ ]
  rationale:
  decisive sources:
  strongest dissent:
  what would trigger revision:

T3 — There is no general reading-comprehension skill independent of topic knowledge.
  ① [ ]  ② [ ]  ③ [ ]  ④ [ ]  ⑤ [ ]
  rationale:
  decisive sources:
  strongest dissent:
  what would trigger revision:

Scale suggestion (not binding): established / well supported / plausible / underdetermined /
contested / not established / normative.

10.2 MACHINE PROVISIONAL READ — not an editorial verdict; written to be argued with

Authored by an LLM on 2026-08-18 from the evidence in §6–§8. It has no standing. Where it disagrees with the editor, the editor is right by definition; the point of writing it down is to make disagreement cheap.

T1 — knowledge is a large causal determinant of comprehension

AxisReadReason
① Textual fidelityHighT1 is what Hirsch says, everywhere, from 1977. No reconstruction needed.
② Inferential validityHighI1–I3 are sound; the premises support the conclusion.
③ Empirical supportStrong, with the outcome-measure caveatBransford (causal manipulation), Anderson 1977, Steffensen 1979, Schneider 1989, Hwang 2023 (directionality, N = 10,706, knowledge→reading paths significantly larger). McCarthy & McNamara put prior knowledge at 30–60% of comprehension variance.
④ Robustness/independenceGood, with one soft spotAll primary_independent. The soft spot is that the showpiece — Recht & Leslie — is the weakest item supporting it, and the strong items (Bransford, Steffensen, Anderson) are underused by Hirsch.
⑤ Normative dependenceNone

Provisional: established. The correct qualification is partial compensation (Smith et al. 2021: "somewhat"), graded effects (Kim: near .22 / mid .17 / far .04), and signed knowledge (McCarthy & McNamara: inaccurate knowledge hurts). None of those overturn T1; all of them make it less useful as a slogan.

T2 — knowledge beats strategy at the margin of school time

AxisReadReason
① Textual fidelityHighswn:ch7_C41, kd:ch4_C44, moa:ch6_C35 state it cleanly.
② Inferential validityMediumI6 requires a marginal-return comparison; Hirsch supplies an existence claim instead and lets it do the work.
③ Empirical supportUnderdetermined — no study in the ledger measures itNo equal-time head-to-head trial exists. The nearest evidence is indirect and points both ways: Okkinga (whole-class strategy d = .186 on standardized) supports it; Elleman 2017 (inference instruction d = .58 on general comprehension) and NRP 2000 cut against.
④ Robustness/independenceMediumProgram evidence is ckf_affiliated with an author chain through Willingham. Kim's trials are independent and give a real but small domain-general effect (.11).
⑤ Normative dependenceLow, but non-zero"Efficient" presupposes comprehension is the objective being maximised.

Provisional: plausible, underdetermined. This is the claim most worth a decisive study and the one least likely to get one.

T3 — no general reading-comprehension skill exists

AxisReadReason
① Textual fidelityHigh for the strong version, and that is the problemsk:ch12_C21 says it flatly. But kd:ch2_C72, swn:ch7_C41 and ae:ch7_C88 state the weak version, and kd:ch2_C73 affirms measurable general proficiency. He holds both.
② Inferential validityLow — this is where the argument breaks (I5)A graded causal claim cannot entail a universal negative. The two substitutes are a dose/plateau argument that does not establish T2's head-to-head allocation comparison and an untested analogy from chess.
③ Empirical supportWeak, and contradicted106 of 121 same occurrences of P01 carry no evidence item. Elleman 2017 (d = .58, general comprehension) and NRP 2000 (7 strategies with a scientific basis; combinations improving standardized scores) are counterexamples.
④ Robustness/independenceWeakThe bridge evidence is a three-subject 1973 study whose strong null was superseded in 1996 by authors Hirsch elsewhere cites. Kintsch, his own cited authority, is a general-process theorist.
⑤ Normative dependenceLowIt is an empirical claim, stated as one.

Provisional: not established as stated; the defensible neighbouring claim is plausible but its dose curve remains underdetermined. The reading on which T3 is true — no content-free skill is trainable with continuing returns — is consistent with the measurement-sensitive strategy literature, but no reviewed study manipulates enough doses to establish the continuing-return curve. The reading Hirsch's rhetoric asserts — no content-independent process operates in comprehension — is contradicted by the simple view of reading, by Kintsch's construction–integration model, and by non-zero trainable effects. Fixing C1 dissolves most of the dispute, and costs Hirsch nothing he needs.

Cross-cutting: overstatement in transmission

Four instances, each checkable against a dated occurrence:

  1. "for that particular text" (2006, kd:ch2_C77) → "shown decisively" (2010, the-making-of-americans:ch6_C22) → "a dominant theme" (2023, sk:ch12_C29).
  2. Willingham & Lovette's "ten sessions yield the same benefit as fifty" plus "RCS instruction should be explicit and brief" → "useless after ten lessons" (wkm:ch1_C87).
  3. Miller's own "pernicious, Pythagorean coincidence" → a hard architectural constant quoted across seven books.
  4. Chase & Simon's three-subject null → "experts perform no better than novices" (cl:ch2_OBJ13), retained after Gobet & Simon 1996 and after Hirsch cites Gobet & Simon 2000.

None of these is fabrication. Each is a real source read at its strongest and then re-quoted one notch stronger than it reads. That pattern is the single most useful thing this dossier has to say to a reader deciding how much to trust the rest of the corpus.

10.3 What would change this assessment


11. Cruxes and open questions

C1 — The definitional crux. Decides whether T3 is even a claim. Does "no general reading skill" mean (a) no content-independent process operates in comprehension, (b) no content-free skill is trainable with continuing returns, or (c) no measured comprehension ability generalises across topics? Hirsch's rhetoric asserts (a); his evidence supports (b); his defence of standardized tests (kd:ch2_C73) requires denying (c). Much of the heat in the reading debates is opponents attacking (a) while proponents defend (b). Resolvable by stipulation, not by evidence — and it is the cheapest thing on this list. Doing it costs Hirsch nothing he needs, and costs his critics their easiest target.

C2 — The marginal-return crux. The decisive empirical question. Per hour of grade K–5 instructional time, what is the comprehension return of knowledge-building content instruction versus comprehension-strategy instruction versus additional fluency work? This is what D1-P7 asserts and no study in the ledger measures it. Grissmer compares whole schools, not time allocations; the strategy meta-analyses have no knowledge-building arm; Kim's trials have no strategy arm; Filderman codes both as moderators in one model without a head-to-head. A well-powered, equal-time, three-arm trial in grades 2–5, with comprehension measured on both matched and unfamiliar topics, would settle more of this dispute than any other single study.

C3 — The transfer crux. Does school-taught knowledge raise comprehension of unseen topics, or only matched ones? Every mechanism study is matched-topic; D1-P7 needs breadth transfer. This crux now has a number. Kim et al. 2023: near-transfer .22, mid-transfer .17, far-transfer .04 (n.s.); the 2024 spiralled trial reaches a domain-general comprehension outcome at .11, sustained at 14 months at .12, and spills into mathematics at .12–.16. Cabell 2025 is the counter-evidence: zero on standardized vocabulary and listening comprehension. Current best reading: transfer is real, small, and graded by lexical overlap between what was taught and what is read.

C4 — The circularity crux. D1-P5 says reading tests are knowledge tests. If so, evidence that knowledge predicts reading-test scores is partly definitional rather than confirmatory. Hirsch cannot both use test scores as the outcome that proves knowledge matters and characterise those tests as knowledge tests, without conceding that the causal claim needs an outcome that is not a knowledge test. As far as the corpus shows this is unaddressed, and it is not merely a debating point: it determines which of the studies in §6 count as confirmations at all.

C5 — The normative bridge. Named here, argued in Dossier 4. Even with C1–C4 resolved in Hirsch's favour, which shared knowledge, chosen by whom, is not determined by any of it. One sentence in this dossier; the whole of Dossier 4.

11.1 Open questions this dossier could not close

  1. Independent replication and generalisation of Recht & Leslie. The primary statistics are now verified: the knowledge main effect is strong, the measured ability effect and interaction are null, and several poor-reader/high-knowledge cell means exceed good-reader/low-knowledge means. What remains open is whether this selected single-school, single-domain pattern replicates under preregistered sampling, balanced cells, unfamiliar topics and broader reading profiles.
  2. Whether reading-strategy knowledge specifically compensates for low domain knowledge, or only reading skill does (O'Reilly & McNamara 2007 full text). SCOPE assumed the former; the abstract says the latter. (P1.)
  3. The Cabell K–2 results from IES R305A170635 — the kindergarten paper is one semester of a multi-year trial. (Watch.)
  4. Whether the Grissmer paper's 47% enrolment and 52% compliance figures can be reconciled. As printed they are inconsistent.
  5. What Hirsch would say to Catts. The most interesting unwritten page in the corpus.

12. Evolution — how these propositions and their evidence changed across ten books

Counts are same-relation occurrences from data/theses/mappings.json, by book, in year order. A blank means the proposition is not asserted in that book.

1977 PoC1987 CL1996 SWN2006 KD2010 MoA2016 WKM2020 HtEC2022 AE2023 SK2024 REtotal
P02 — text is incomplete (D1-P1, P2)32342611128161815136
P01 — no general reading skill (D1-P4, P6)517199221020148124
P03 — domain specificity / no transfer (D1-P3, P6, P7)52732141294278
P06 — vocabulary as knowledge proxy (D1-P5)41611
P14 — reading tests are invalid (D1-P5)41115

Five things this table shows that no essay can.

  1. The mechanism is a decade older than the negative claim. P02 is asserted in The Philosophy of Composition (1977), where the subject is prose readability, not schooling. P01 has no occurrence in 1977 at all. It enters in Cultural Literacy (1987) and never leaves. The negative existence claim is not what the cognitive science gave him; it is what he added.
  2. Its first form in 1987 is mild. cl:ch1_C9: "Literacy is far more than a skill; it requires large amounts of specific information." cl:ch1_C46: reading and writing "cannot be treated as empty skills that are independent of specific knowledge." Both are compatible with T1 and T2 and commit to neither (a) nor (c) in C1.
  3. The absolutism arrives in 1996 and hardens. swn:ch3_C124: "It is a fallacy to assume that it is possible to teach abstract, generalized problem-solving ability." swn:ch4_C273: "Modern faith in general 'critical-thinking' skills lacks scientific justification." P03 peaks in the same book at 27 occurrences — the domain-specificity/transfer literature is what he acquires in the 1990s, and it is what makes the strong form say-able.
  4. By 2016–2022 it is a flat assertion. wkm:chprologue_C73: "A single, overarching general skill does not exist." ae:ch1_C78: "The idea that reading comprehension is a general skill is false." Note the drop of the word reading in the 2016 version — the claim has widened from reading to skill as such.
  5. Then it softens, quietly, in the same books. ae:ch7_C88 (2022): "It is more accurate to speak of 'reading skills' than of 'reading skill.'" sk:ch12_C24 (2023) concedes that individual topic knowledge is not an instructional lever. The late Hirsch is more careful than the middle Hirsch, and the corpus is the only place you can see that.

Evidence added over time. Miller and Kintsch are there from 1977. Bransford enters in 1977 and is still cited in 2022–24. De Groot enters in 1987 and peaks in 2020–22. Recht & Leslie enters in 2006 — eighteen years after publication, and after the strong claim it is used to support was already in print. The AFQT reanalysis enters in 2016. Schneider (1989) enters once, in 2016. Hwang (2023) and Kim (2022–24) enter in the last two books. The pattern is not "evidence accumulates and the claim follows"; it is "the claim is stated, then evidence is recruited to it."

Reframings, not retractions. Nothing in the table is retracted. What changes is framing: a theory of readability (1977) → a theory of cultural membership (1987) → a critique of educational ideology (1996) → a measurement critique (2006, 2016) → a theory of American ethnicity (2022) → a comparative-national argument (2023–24). The propositions are Lakatosian hard core; the surrounding argument is the auxiliary belt.


13. Raw atlas

Everything above is checkable against the corpus. These are the entry points.

13.1 The 38 backbone occurrences

Every occurrence id in §4 links to ../claim.html#<id>, which opens that claim with its full source_passage, page range, chapter, evidence, warrants and counter-arguments. Machine-readable source of truth: backbone.json in this directory; re-verify with .venv-hirsch/bin/python3 verify_quotes.py.

13.2 The thesis layer

13.3 Feedback to the registry and the thesis layer

Concrete, checkable corrections this dossier produced. All are unapplied — this dossier does not write to shared data.

A. Missing canonical propositions.

B. Registry field corrections (all 1,515 entries currently have doi: null, resolution_status: null):

source_idCorrection
S0027DOI 10.1037/0022-0663.80.1.16. resolutions.json has 10.1037//0022-0663.80.1.16 — double slash.
S0012DOI 10.1037/h0043158. resolutions.json resolved miller_1956_magical_number to 10.1037/11347-011, which is a biographical entry titled "George A. Miller.", not the paper.
S0020Kind is commentary/review, not study. No DOI (Teachers College Record ID 17701). resolutions.json resolved willingham_why_students to 10.1037/021535 — wrong work.
S0362DOI 10.3102/00346543064004479. resolutions.json resolved rosenshine_direct_instruction to 10.1007/bf00364747, a paper on HLA-DQ polymorphism — clearly wrong. Also: the registry title says "Nineteen Experimental Studies" (the technical report); the published article reviews 16.
S0019Year is wrong. Not 2007 — internal evidence dates it 2013 or later. Kind should be technical_report, not study. Not peer-reviewed. URL: nlsinfo.org, NLSY79 Appendix 24.
S0013, S0080Chase & Simon DOI 10.1016/0010-0285(73)90004-2. N = 3. Add a cross-reference to Gobet & Simon 1996 (10.3758/BF03212414), which supersedes the random-position null Hirsch relies on.
S0021, S0070DOIs 10.1016/S0022-5371(72)80006-9 and 10.1016/0010-0285(72)90003-5. Author order is Bransford, Barclay & Franks.
S0033Kintsch 1988 DOI 10.1037/0033-295X.95.2.163. resolutions.json resolved kintsch_comprehension to 10.1353/lan.1986.0098, which is a book review of van Dijk & Kintsch, not the work.
S1348DOI 10.2307/747429. Free full text of the identical study: CSR Technical Report No. 97, ideals.illinois.edu/items/18083. Corpus misspells the second author as "C. Joag-Des" (correct: Joag-Dev; the 1978 report spells it Jogdeo).
S0060Anderson et al. 1977 DOI 10.3102/00028312014004367; free precursor Technical Report No. 12 at ideals.illinois.edu/items/17896.
S0005DOI 10.26300/nsbq-hb21. Not peer-reviewed — add a field for this; it currently reads as a study like any other.
S0366DOI 10.1037/0022-0663.81.3.306.
S0322DOI 10.1002/rrq.481; open access; correct year 2023 (registry has 2022).
S0161Kim et al. 2023 DOI 10.1037/edu0000751.
(new)Cabell et al. 2025, DOI 10.1037/edu0000916 — postdates the corpus; needs a registry entry to be citable from Dossier 2.

C. Independence. The independence flag is per-source and cannot see chains. Willingham is primary_independent as S0020, a co-author of ckf_affiliated S0005, and a CKF trustee. Suggest a related_independence_note field, or an explicit person-level affiliation table.

D. One warrant to regenerate. how-to-educate-a-citizen:ch5_W1 — the warrant under the chess analogy, the single most load-bearing inference in the backbone — is still flagged stale: evidence_id_is_a_claim_id. It was not repaired by the regeneration pass.


14. Time estimates

Reading time. ~19,000 words. Read straight through: 75–95 minutes. Read as designed — §1–§5 in full, §6 ledger, §7 cards for the two or three sources you actually care about, §10.2 and §11: 25–30 minutes. §7 is deliberately reference material, not a chapter; skipping ten of the twelve cards on a first pass is the intended use.

Editor time to a verdict you would preserve and cite.

TaskEstimate
Read this draft and mark disagreements1 h
Read the P0 acquisitions (Recht & Leslie; Reynolds 2025; Smith et al. 2021; Willingham & Lovette; Grissmer) and correct the cards4–5 h
Read the P1 acquisitions (Kintsch 1988; McNamara et al. 1996; O'Reilly & McNamara; Cabell 2025; NRP ch. 4-II; Rosenshine & Meister; the AFQT memo)5–6 h
Read cl:2, kd:2, sk:12, wkm:1, ae:7, htec:5 against the backbone2–3 h
Adjudicate C1–C5 and write §10.12 h
Minimum viable (P0 + corpus chapters + assessment)≈ 8–9 h
Publishable (P0 + P1)≈ 14–17 h

The P0 set alone is enough for a defensible verdict on T1, T2 and T3. Everything after that sharpens the crux statements rather than changing the verdict.

What is not included: applying any of §13.3 to the shared registry, which is a separate decision and a separate session.


15. What this dossier has to beat

essay-control.md in this directory is the markdown-essay control condition (plan amendment 21.1 item 5), and baseline-questions.md is the protocol. The structured dossier earns its complexity only if, on that protocol, it beats the essay on:

  1. recalling the decisive evidence and its limits without reopening the page;
  2. naming the crux;
  3. keeping T1, T2 and T3 apart;
  4. being materially easier to revisit and update when Reynolds 2025 or the Cabell K–2 results are actually read.

A tie is a fail. On (4) specifically, the test is concrete: when the editor reads Recht & Leslie's primary text, there is exactly one place the new information goes — Card 1's effect line and the T1 ④ read — and it should take under a minute to find. If it does not, the structure is decoration.


Draft generated 2026-08-18. Corpus: data/corpus_consolidated.json (10,312 claims, 2,885 evidence items, 4,327 warrants, 892 objections). Thesis layer: data/theses/ (25 canonical propositions, 10,312 mappings, uncalibrated). Registry: data/evidence_registry/ (1,515 sources). Quote verification: verify_quotes.py → 38/38.