Dossier 1 — Does background knowledge drive reading comprehension?
And does it follow that there is no general, teachable reading-comprehension skill?
DRAFT — MACHINE-WRITTEN, UNREVIEWED BY THE EDITOR. Drafted 2026-08-18 by an LLM (Claude Opus) from
SCOPE.md,data/corpus_consolidated.json,data/theses/*,data/evidence_registry/*, and web verification of the external literature. Nothing here has been signed off by a human. The editorial assessment in §10 is supplied in two forms: an empty template that the editor owns, and a clearly-labelled machine provisional read that exists only to be argued with. Where a fact could not be verified against a primary source, it saysunverified— that word is load-bearing and should not be smoothed away in editing.What is checked: every quoted occurrence in §4 is a literal substring of that claim's
source_passagein the corpus (38/38 verified byverify_quotes.py; re-run it after any edit). Every occurrence→proposition correspondence in §4 is read offdata/theses/mappings.json. Every external citation in §6–§8 carries averified_bynote naming the URL actually opened.
1. Research question and stakes
The question. Does relevant prior knowledge drive reading comprehension — and does it follow that there is no general, teachable reading-comprehension skill?
Why it decides other things. This is the floor of the whole Hirsch project. If background knowledge is the dominant determinant of comprehension then reading tests are knowledge tests, the achievement gap is a knowledge gap, the elementary reading block is misallocated, curriculum content is not interchangeable, and content-neutral "skills" standards are a category error. Dossiers 2 (do knowledge-rich curricula work?), 3 (France) and 4 (what is worth knowing?) all inherit from here.
If the floor is only partly sound — if knowledge matters a great deal and general processes also exist and are partly trainable — then the policy conclusion becomes an allocation argument about marginal returns rather than an existence argument, and it has to be won on much narrower ground than Hirsch fights on.
It also matters because this is where Hirsch is strongest. An atlas that cannot state precisely what he has established here has no claim on a reader's attention for anything else.
What the corpus brings that a reading list does not. Three things, all checkable:
- The negative claim ("no general reading skill") and the mechanism claim ("text is incomplete") have different first dates. The mechanism is 1977; the negative claim does not appear as an asserted proposition until 1987 (§12).
- The qualifier attached to the decisive experiment — "for that particular text" — is present in 2006 and gone by 2010 (§4, D1-P4).
- Across ten books and 892 recorded objections, not one of Hirsch's replies on this question is addressed to a named living critic (§9).
2. Definitions and scope
2.1 Three theses that must not be run together
The central discipline of this dossier. These have different evidential requirements and different truth statuses, and most of the public heat in the reading debates comes from people attacking one while their opponents defend another.
| Thesis | Logical form | What would establish it | |
|---|---|---|---|
| T1 | Relevant prior knowledge is a large causal determinant of reading comprehension. | Causal, graded | Experimental manipulation of knowledge → comprehension; longitudinal directionality tests |
| T2 | At the margin of school time, adding knowledge buys more comprehension than adding strategy instruction. | Comparative, causal, policy-relevant | Head-to-head equal-time trials; or a well-identified time-substitution study |
| T3 | There is no general reading-comprehension skill independent of topic knowledge. | Existence claim (negative, universal) | Showing that no comprehension variance survives controlling for topic knowledge; or that no content-general process is trainable |
Hirsch states T3 in its strongest form — "there is no such thing as a general reading skill independent of specific unstated topic knowledge shared between writer and reader" (sk:ch12_C21, 2023) — while the evidence he marshals establishes T1 and gestures at T2. The job of §4–§7 is to show exactly where the argument crosses from T1 into T3, and on what.
2.2 "General skill" has at least three readings, and Hirsch uses all of them
- A content-free process that operates in comprehension (inference-making, coherence monitoring, constraint satisfaction). Denying this is a strong claim about cognitive architecture.
- A content-free ability trainable with continuing returns. Denying this is a claim about instructional yield. Brief strategy efficacy is supported, but the continuing-return curve is underdetermined by the reviewed dose evidence.
- A measured trait that generalises across topics for a given reader. Hirsch affirms this one when he defends standardized reading tests: "well-educated people can and do exhibit a general proficiency in reading comprehension, and we can indeed reliably measure that proficiency on reading tests, just as the test-makers claim" (
kd:ch2_C73, 2006).
His rhetoric asserts (1); his evidence supports (2); his defence of reading tests requires the denial of (3). Until this is fixed the dispute is partly verbal — see crux C1.
2.3 Other definitions to fix before appraising anything
- Background knowledge. Hirsch means the unstated content a writer assumes — "the tip of an iceberg" (
cl:ch2_C5). The current research framework (McCarthy & McNamara 2021) decomposes it into four dimensions: amount, accuracy, specificity, coherence. Hirsch measures only amount and treats specificity as binary. Accuracy is absent from the corpus: no claim in ten books registers that incorrect prior knowledge can make comprehension worse, which is a documented effect (§7, card 12; §8.5). - Comprehension. Recht & Leslie measured recall, summarization and sentence-importance sorting of one researcher-written text, partly via re-enactment on a model baseball field. Standardized tests measure multi-passage multiple-choice performance across topics. These are different constructs and effect sizes do not transfer between them.
- Transfer. The argument needs breadth transfer — knowledge taught in school raising comprehension of unseen passages — not merely matched-topic advantage (knowing baseball helps with a baseball passage). Almost all the mechanism evidence is matched-topic. This is crux C3, and the best current measurement of it is Kim et al. 2023: near-transfer ES = .22, mid-transfer ES = .17, far-transfer ES = .04, not significant.
2.4 Scope boundaries
In scope: the cognitive mechanism (schema, situation model, working memory); the empirical knowledge–comprehension relation; the measurement claim about reading tests; the negative claim about general skill; the immediate instructional inference about the reading block.
Out of scope, deliberately deferred:
- Whether knowledge-rich curricula raise outcomes and close gaps → Dossier 2. Grissmer 2023 and Cabell 2025 appear here only because Hirsch's equity claim is stated in this dossier's chapters, and because their independence structure bears on how his other citations should be read.
- Decoding and phonics. Hirsch concedes decoding is a genuine general skill (
wkm:ch4_C125); that concession is a premise here, not a subject. - The equity, civic and national-identity justifications → Dossiers 2 and 4.
- Which knowledge, chosen by whom → Dossier 4. Flagged once in crux C5.
3. Hirsch's steelman (298 words)
Written language is not speech written down. A writer addressing strangers must decide what to leave unsaid; efficiency in a literate culture depends on how much can be taken for granted (cl:ch1_C23, 1987). A text is therefore systematically incomplete by design: "the explicit meanings of a piece of writing are the tip of an iceberg of meaning; the larger part lies below the surface of the text and is composed of the reader's own relevant knowledge" (cl:ch2_C5).
Comprehension is consequently constructive. The reader assembles a mental model of the situation the text describes, supplying the unstated connections from memory (cl:ch2_C33; how-to-educate-a-citizen:ch5_C64). Whether this succeeds is not a matter of effort or technique, because it is bounded by a hard architectural limit: short-term memory holds four to seven chunks (cl:ch2_C6), so the needed schemata must already be in long-term memory and quickly available. "When the appropriate schemata are not quickly available… the limits of short-term memory are quickly reached, and the process has to be painfully restarted and restarted" (cl:ch2_C87). You cannot look it up; the lookup costs more working memory than it saves.
What is required is not general ability but the particular content the writer assumed. Hence the experimental result: poor readers who know baseball out-comprehend good readers who do not (kd:ch2_C77, 2006). Hence too the measurement consequence — a reading test is a knowledge test in disguise (ae:chpreface_C6, 2022) — and the finding that vocabulary size, a proxy for accumulated world knowledge, is "the single most reliable correlate to reading ability" (wkm:ch3_C29).
Comprehension-strategy instruction produces a real but small, quickly exhausted benefit: it teaches a child what reading is for, then stops paying (wkm:ch1_C73). Every further hour spent on it is an hour of knowledge not built. The efficient lever is a cumulative, coherent, shared body of knowledge, taught early (kd:ch4_C44).
4. Minimal argument backbone
Seven propositions. Each is a premise for the ones below it: D1-P1–P3 are the mechanism, D1-P4–P5 the empirical predictions, D1-P6 the negative conclusion, D1-P7 the policy bridge to Dossier 2.
Each row gives the occurrence id (linked to the raw atlas), the book and year, the canonical proposition it maps to in data/theses/mappings.json, and the quoted source_passage. All 38 quotes below were verified as literal substrings of the corpus record (verify_quotes.py → verified 38/38 quoted occurrences against source_passage (0 failed)).
4.1 How the dossier backbone relates to the canonical propositions
The canonical layer (25 propositions, data/theses/canonical_propositions.json) was drafted independently of this dossier. Reconciling the two is itself informative:
| Dossier | Canonical | Fit |
|---|---|---|
| D1-P1 text is incomplete | ≈ P02 | 4 same, 1 supports. Exact fit. |
| D1-P2 the knowledge is topic-specific and shared | ≈ P02 | 3 same, 2 supports. P02 does not separate P1 from P2 — the canonical layer merges "text is incomplete" with "the missing content is community-shared". Those are different claims with different evidence (Bransford for the first, Steffensen and Krauss for the second), and the merge hides that. Recommend splitting P02. |
| D1-P3 working memory binds | → P02 (supports) ×4, P03 (same) ×1, and one unmapped | poc:ch4_C116 (Miller's limit, 1977) maps to no canonical proposition at all. There is no canonical proposition for the working-memory constraint, although it is what makes knowledge non-substitutable by lookup and is therefore load-bearing for D1-P6. Recommend adding one. |
| D1-P4 knowledge outweighs measured ability | ≈ P01 | 1 same, 3 supports, 1 → P02. |
| D1-P5 tests measure knowledge; vocabulary is the best correlate | splits: P14, P06, P01 | The measurement claim (P14) and the vocabulary-proxy claim (P06) are genuinely different propositions that Hirsch runs together. |
| D1-P6 no general teachable skill | ≈ P01 (+P03) | 4 same on P01, 1 same on P03. The flagship T3 statement sk:ch12_C21 maps to P01 at confidence 1.0. |
| D1-P7 curriculum is the lever | splits: P03, P01, P08, P24 | A bridge proposition, as expected. kd:ch4_C44 → P24 narrower; sk:ch12_C24 → P08 narrower; re:ch6_C51 → P08 broader. |
The load-bearing inference is D1-P4/P5 → D1-P6, and it is the weakest joint in the chain (§5).
4.2 D1-P1 — Written text is systematically incomplete; comprehension requires the reader to supply unstated content from memory
| Occurrence | Book / year | Canonical | Quoted source_passage |
|---|---|---|---|
poc:ch4_C74 | PoC 1977 | P02 (same) | "It consists of linguistic 'rules' and semantic conventions. It embraces large domains of tacitly shared knowledge, and it includes tacit suppositions about the theme and tendency of the text as a whole." |
cl:ch2_C5 | CL 1987 | P02 (same) | "The explicit meanings of a piece of writing are the tip of an iceberg of meaning; the larger part lies below the surface of the text and is composed of the reader's own relevant knowledge." |
cl:ch2_C33 | CL 1987 | P02 (supports) | "To make sense of what we read, we must use relevant prior knowledge to form a model of how sentence meanings hang together." |
kd:chappendix_C19 | KD 2006 | P02 (same) | "Sharing the unsaid makes it possible for them to comprehend the said. It is the very thing that makes them a speech community." |
how-to-educate-a-citizen:ch5_C64 | HtEC 2020 | P02 (same) | "much of the information needed to understand a text is not provided by the information expressed in the text itself but must be drawn from the language user's [prior] knowledge" |
Note the 1977 date on the first row. The mechanism predates the education argument by a decade; in The Philosophy of Composition it is a theory of prose readability, not of schooling.
4.3 D1-P2 — The knowledge that must be supplied is topic-specific and shared — the unstated store a writer assumes in a speech community
| Occurrence | Book / year | Canonical | Quoted source_passage |
|---|---|---|---|
poc:ch4_C75 | PoC 1977 | P02 (supports) | "A shrewd decision about the knowledge that the writer can tacitly assume in his audience may be the most important decision the writer makes." |
cl:ch1_C23 | CL 1987 | P02 (supports) | "If they can take a lot for granted, their communications can be short and efficient, subtle and complex. But if strangers share very little knowledge, their communications must be long and relatively rudimentary." |
kd:ch4_C5 | KD 2006 | P02 (same) | "Unless writers and their readers internalize the shared knowledge of the wider speech community (such as shared knowledge about baseball), they cannot expect the blanks to be filled in; they cannot be successful writers or proficient readers." |
the-making-of-americans:ch1_C51 | MoA 2010 | P02 (same) | "This means that communication depends on both sides, writer and reader, sharing a basis of unspoken knowledge." |
sk:ch12_C26 | SK 2023 | P02 (same) | "shared background knowledge is a necessity for comprehension which must include not just very general, nationally shared background knowledge but also specifically shared topic knowledge." |
4.4 D1-P3 — Working-memory limits make possession and speed of access, not availability of information, the binding constraint
| Occurrence | Book / year | Canonical | Quoted source_passage |
|---|---|---|---|
cl:ch2_C6 | CL 1987 | P02 (supports) | "The mind cannot reliably hold in short-term memory more than about four to seven separate items…" |
cl:ch2_C87 | CL 1987 | P02 (supports) | "When the appropriate schemata are not quickly available, and the reader is forced to do a lot of pondering to construct them at the time of reading, the limits of short-term memory are quickly reached, and the process has to be painfully restarted and restarted." |
cl:ch2_C112 | CL 1987 | P02 (supports) | "schemata perform two essential functions that are relevant to literacy. The first is storing knowledge in retrievable form; the second is organizing knowledge in more and more efficient ways, so that it can be applied rapidly and efficiently." |
poc:ch4_C116 | PoC 1977 | — (none) | "The basic discovery is the precisely limited capacity of short-term memory…" |
wkm:ch4_C140 | WKM 2016 | P02 (supports) | "verbal comprehension is a form of problem solving. Topic familiarity makes the task of forming a situation model fast and accurate, leaving more space in working memory to figure out word relations…" |
how-to-educate-a-citizen:ch5_C84 | HtEC 2020 | P03 (same) | "It was ingrained, specific factual knowledge, stored in long-term memory, not some general mental skill, that explained the skilled performance." |
The "you can't look it up" corollary is carried in the corpus as an objection-and-response rather than as a claim: cl:ch2_OBJ11 ("Students do not need to possess specific facts because they can simply look them up…" → "Reference tools are impractical and too slow…") and the parallel reply in kd:ch4.
4.5 D1-P4 — Therefore comprehension is topic-variable within a reader, and topic knowledge can outweigh measured general reading ability
| Occurrence | Book / year | Canonical | Quoted source_passage |
|---|---|---|---|
cl:ch2_C57 | CL 1987 | P02 (supports) | "among seven-year-olds who score the same on reading and IQ tests, those who have greater knowledge relevant to the text at hand show superior reading skills." |
kd:ch2_C77 | KD 2006 | P01 (supports) | "the reading comprehension of the low-skills, baseball-knowing group proved superior to the reading comprehension of the high-skills, baseball-ignorant group for that particular text." |
the-making-of-americans:ch6_C22 | MoA 2010 | P01 (same) | "It has been shown decisively that subject-matter knowledge trumps formal skill in reading" |
wkm:ch1_C100 | WKM 2016 | P01 (supports) | "topic familiarity, often indicated by word familiarity, is a stronger predictor of comprehension than formal ability to deal with difficult syntax." |
sk:ch12_C29 | SK 2023 | P01 (supports) | "as soon as one looks elsewhere in language comprehension—especially in second-language learning—one discovers that 'topic familiarity' is a dominant theme. (That is also true for our native speakers, as Recht and Leslie have shown in the baseball study.)" |
Watch the escalation in the year column. 1987: "superior reading skills." 2006: the result stated for one text, "for that particular text". 2010: "shown decisively." 2023: "a dominant theme." The qualifier survives in 2006 and is gone by 2010. This is the single clearest instance of drift in the dossier, and it is visible only because the occurrences are dated.
4.6 D1-P5 — Therefore reading tests measure knowledge, and vocabulary breadth is the best single correlate of comprehension
| Occurrence | Book / year | Canonical | Quoted source_passage |
|---|---|---|---|
cl:ch1_C26 | CL 1987 | P06 (related) | "According to John B. Carroll… the verbal SAT is essentially a test of 'advanced vocabulary knowledge,' which makes it a fairly sensitive instrument for measuring levels of literacy." |
kd:ch6_C21 | KD 2006 | P14 (same) | "because the tests have been presented as tests of formal comprehension skills, they are unwittingly unfair… The tests favor children who happen to have domain knowledge relevant to the passages in the test." |
the-making-of-americans:ch6_C17 | MoA 2010 | P01 (supports) | "Many educators see this question as probing the general skill of 'finding the main idea.' It does not." |
wkm:ch3_C29 | WKM 2016 | P06 (same) | "Vocabulary size is the single most reliable correlate to reading ability." |
ae:chpreface_C6 | AE 2022 | P14 (supports) | "An American reading test is an American ethnicity test in disguise. It's 'in disguise' because the shared background knowledge – the shared ethnicity required for comprehension – is unstated." |
Internal tension to hold on to. kd:ch2_C73–C74 (2006) defends the validity of standardized reading tests precisely because they sample many domains, and asserts a real, measurable general proficiency: "well-educated people can and do exhibit a general proficiency in reading comprehension… just as the test-makers claim." kd:ch6_C21 (same book) says the tests are "unwittingly unfair" because they measure domain knowledge. Both can be true only if "general proficiency" means breadth of knowledge — which is Hirsch's actual position (kd:ch1_C55: "The only thing that transforms reading skill and critical thinking skill into general all-purpose abilities is a person's possession of general, all-purpose knowledge"). That reading dissolves the tension and also drains T3 of its bite: the general skill exists; it is just made of knowledge.
4.7 D1-P6 — Therefore there is no general, teachable reading-comprehension skill; strategy instruction gives a small, one-time, quickly-exhausted gain
| Occurrence | Book / year | Canonical | Quoted source_passage |
|---|---|---|---|
swn:ch5_C129 | SWN 1996 | P01 (same) | "There is no accurate way to describe reading ability as a purely formal skill, or to remove from it the information-based knowledge disparaged as 'factoids.'" |
kd:ch2_C145 | KD 2006 | P01 (same) | "It is not mainly comprehension strategies that young children lack in comprehending texts but knowledge—knowledge of formal language conventions and knowledge of the world." |
wkm:ch1_C73 | WKM 2016 | P03 (supports) | "Lessons in reading strategies offer an initial score boost for test taking, but are quickly learned; they plateau fast, and they don't have to be practiced." |
wkm:ch4_C125 | WKM 2016 | P01 (same) | "But reading comprehension, unlike decoding, is not a general skill for standard written English." |
how-to-educate-a-citizen:ch5_C89 | HtEC 2020 | P03 (same) | "All this was summarized in Anders Ericsson's remark 'There is no such thing as developing a general skill.'" |
sk:ch12_C21 | SK 2023 | P01 (same) | "In sum: there is no such thing as a general reading skill independent of specific unstated topic knowledge shared between writer and reader." |
Hirsch also states the qualified version, which is the defensible one:
kd:ch2_C72(2006): "general reading skill applies only to texts that are consciously directed to a general audience within a definite speech community."swn:ch7_C41(1996): "The best foundation for general skill is not the inculcation of abstract strategies but the inculcation of broad knowledge." — an allocation claim (T2), not an existence claim (T3).ae:ch7_C88(2022): "It is more accurate to speak of 'reading skills' than of 'reading skill.'"
His route from D1-P4/P5 to D1-P6 runs mostly through expertise research — de Groot's chess players, Chase & Simon, Ericsson, Simon's "dormative power" jibe — not through comprehension research. That is an analogy from expert performance in a bounded task domain to comprehension of arbitrary prose, and it is the one inference in the chain the corpus never tests directly (§5, I5).
4.8 D1-P7 — Therefore the school's efficient lever is a cumulative, coherent, shared knowledge curriculum, not more strategy practice (bridge to Dossier 2)
| Occurrence | Book / year | Canonical | Quoted source_passage |
|---|---|---|---|
swn:ch7_C41 | SWN 1996 | P03 (same) | "The best foundation for general skill is not the inculcation of abstract strategies but the inculcation of broad knowledge." |
kd:ch2_C87 | KD 2006 | P01 (same) | "Since relevant, domain-specific knowledge is an absolute requirement for reading comprehension, there is no way around the need for children to gain broad general knowledge in order to gain broad general proficiency in reading." |
kd:ch4_C44 | KD 2006 | P24 (narrower) | "Since at least ninety minutes per day are currently allotted to reading in early grades, about an hour could be devoted to the language and world knowledge that is most important for competence…" |
the-making-of-americans:ch6_C35 | MoA 2010 | P01 (supports) | "it is far more fruitful to teach children the broad array of domain-specific knowledge they will need to become mature readers than to practice reading strategies such as 'finding the main idea,' 'clarifying,' and 'summarizing.'" |
sk:ch12_C24 | SK 2023 | P08 (narrower) | "shared topic knowledge is not an educationally useful characteristic—unless and until specific and predictable topic knowledge has been taught and learned by all the readers in the class." |
re:ch6_C51 | RE 2024 | P08 (broader) | "supplying shared background knowledge systematically is a key principle of an effective curriculum in every modern nation and every modern literate language." |
sk:ch12_C24 is the most important claim in that table and the least noticed: it concedes that individual topic knowledge is not an instructional lever. Only class-wide predictable knowledge is. That converts the mechanism argument into a curriculum-coordination argument — which is what Dossier 2 must evaluate, and which a purely cognitive defence of D1-P1–P6 does not establish.
5. Inference audit
One row per arrow in the backbone. "Corpus warrant" quotes a regenerated warrant (provenance: regenerated_2026-08-18_gemini3-flash, ids of the form W_R*) attached to one of the quoted occurrences; these are the warrants rebuilt after the sliding-window ID-remap repair, so they are wired to an evidence item that actually exists. Thirteen of the 38 backbone occurrences carry one. Older W* warrants on these claims are mostly flagged stale (evidence_id_is_a_claim_id or evidence_relinked_to_different_claim) and are not quoted here.
I1. D1-P1 → D1-P2 · text is incomplete ⟹ the missing content is community-shared
Type: deductive-with-a-smuggled-premise. Required assumption: that the writer's assumptions about the reader are drawn from a common public store rather than from a guess about a particular reader, a genre convention, or the text's own internal build-up. Only the first reading licenses the curricular conclusion; the other two are equally consistent with the incompleteness premise. Corpus warrant (poc:ch4_W_R23, implicit, evidence: Carroll & Freedle, Language Comprehension and the Acquisition of Knowledge): "Successful communication relies on a vast substrate of unstated information that functions as a necessary condition for interpreting explicit signs." Recorded vulnerability: "The 'illusion of transparency' where speakers overestimate the extent to which their internal knowledge is actually shared by their audience." Assessment: the inference is sound for the class of texts addressed to strangers, which is the class Hirsch cares about, and it is genuinely well-evidenced (Steffensen 1979 is the clean test — §7, underuse A). The slide is that Hirsch runs the shared reading as if it followed from incompleteness alone.
I2. D1-P3 → possession, not access ⟹ lookup cannot substitute for knowledge
Type: deductive from a capacity premise. Required assumptions: (a) that the capacity limit applies during reading with the magnitude Miller reported; (b) that consulting a reference costs more capacity than it frees; (c) that the relevant knowledge cannot be supplied by the text itself in time. Corpus warrant (cl:ch2_W_R1, implicit, evidence: "George Miller's research on the 4-7 item limit of short-term memory across various sensory domains"): "Standardized psychological test results across multiple sensory modalities provide a reliable basis for defining the structural constraints of human cognition." Recorded vulnerability: "The 'magic number' may vary significantly depending on the complexity of the items or the specific task environment." Assessment: premise (a) is dated but directionally intact. Miller himself called the coincidence of the two sevens "only a pernicious, Pythagorean coincidence" and explicitly said the span of absolute judgment and the span of immediate memory "are quite different kinds of limitations"; Cowan (2001) puts the real limit nearer four chunks, and whether the limit is slots or a flexible resource is still disputed. The qualitative point survives; the number does not, and Hirsch keeps quoting the number. (b) and (c) are asserted, not tested — but Hirsch does convert them into an explicit objection-and-response (cl:ch2_OBJ11), which is more than he does for most of his opponents' positions.
I3. D1-P1–P3 → D1-P4 · mechanism ⟹ topic knowledge can outweigh measured reading ability
Type: inductive, from a small number of matched-topic experiments. Required assumptions: (a) that the experimental knowledge measure indexes the knowledge the mechanism says is needed; (b) that the comprehension measure is not itself knowledge-loaded in a way that inflates the result; (c) that the ability split is a real ability split. Corpus warrant (kd:ch2_W_R17, implicit, evidence: "An 'elegant experiment' (The Baseball Study)… Recht and Leslie (1988), implied by endnote 18"): "If a group with inferior technical skills outperforms a group with superior technical skills on a specific task when subject knowledge is the only variable favoring the former, then that knowledge is the primary driver of performance." Recorded vulnerability: "The 'high-skills' group may have had such low motivation for the specific topic (baseball) that they failed to apply the strategies they possessed." Assessment: the corpus's own recorded vulnerability is the weakest of the available objections. The real problems are assumptions (a) and (c). Reynolds, Hattan & Markham (2025) document that the baseball literature's knowledge measures are "vocabulary and baseball trivia"; Shanahan notes the sample excluded readers below the 30th percentile, which truncates the ability contrast the design depends on. The result also rests on main effects plus a null interaction, not on a demonstrated crossover. See §7, card 1.
I4. D1-P4 → D1-P5 · topic-variable comprehension ⟹ reading tests measure knowledge
Type: deductive, and largely correct — but it damages the rest of the chain. Required assumption: that test passages sample topics unevenly with respect to readers' knowledge. Uncontroversial. Assessment: valid, and Hirsch defends it well (kd:ch6_OBJ4/OBJ5: supplying the background inside the item fails, because unfamiliar students must assimilate new content while answering). The cost is crux C4. If reading tests are knowledge tests, then evidence that knowledge predicts reading-test scores is partly definitional rather than confirmatory — and most of the supporting evidence in §7 uses reading-test-like outcomes. Hirsch cannot both use test scores as the outcome that proves knowledge matters and characterise those tests as knowledge tests, without conceding that the causal claim needs an outcome that is not a knowledge test. The corpus records no engagement with this.
I5. D1-P4/P5 → D1-P6 · knowledge matters a great deal ⟹ no general skill exists — the weak joint
Type: invalid as stated. A graded causal claim (T1) plus a measurement claim cannot entail a universal negative existence claim (T3). What would be required: that no comprehension variance survives controlling for topic knowledge, or that no content-general process is trainable. Neither is shown anywhere in the corpus. What the corpus offers instead: two substitutes.
- I5a — the plateau argument.
wkm:ch1_C87: strategy exercises are "knowledge-displacing, soul-deadening, and useless after ten lessons";kd:ch2_OBJ15: sessions beyond six are "wasteful and distracting." The corpus attributes the plateau towkm:ch1_E20= Willingham & Lovette 2014. That attribution is fair: Willingham & Lovette do write "Ten sessions yield the same benefit as fifty sessions." What Hirsch drops is their conclusion, in the same piece: "RCS instruction should be explicit and brief" — i.e. teach them. And the SCOPE-level attribution of the plateau to Rosenshine & Meister 1994 does not hold: they report no relationship between number of sessions and study-level statistical-significance categories across a 6-to-100 session range. This is heterogeneous vote-counting, not a manipulated dose comparison; it cannot establish a plateau (§7, card 7). Even granted the saturation claim, it would support only one premise in T2; a head-to-head knowledge-versus-strategy comparison would still be missing. It cannot establish T3. - I5b — the expertise analogy. De Groot's and Chase & Simon's chess players: masters lose their advantage on random boards, so expertise is stored knowledge, not general capacity. Type: analogical. Required assumption: that comprehending arbitrary prose is relevantly like reconstructing a chess position. Corpus warrant —
how-to-educate-a-citizen:ch5_W1, and note it is flaggedstale(evidence_id_is_a_claim_id), i.e. it was never repaired by the regeneration pass and points at a claim rather than an evidence item: "If a professional's performance advantage disappears when domain-specific patterns are removed, the advantage must reside in the stored patterns (knowledge) rather than innate cognitive processing speed (skill)." Recorded vulnerability: "Recalling a board layout is a task of pattern recognition, which may not be perfectly analogous to other 'skills' like logical reasoning or creative synthesis." That the corpus's single most load-bearing analogy carries only a stale warrant is itself a finding. The analogy is never tested anywhere in ten books. It is also built on an outdated version of the chess result: Chase & Simon 1973 had three subjects, and Gobet & Simon (1996) showed across 13 experiments that masters do retain a small reliable advantage on random positions (≈1 piece per 400 Elo). Hirsch cites Gobet & Simon (2000) elsewhere (wkm:ch4_E_new58) — so he has read these authors, and still reports the strong null (§7, card 5).
Assessment: this is where the argument breaks. Everything above I5 is defensible. D1-P6 as stated is not derived; it is asserted, and — per the thesis layer — 106 of the 121 same occurrences of P01 carry no direct evidence item at all. The flagship sk:ch12_C21 is a summary sentence with no attached evidence.
I6. D1-P6 → D1-P7 · no general skill ⟹ knowledge instruction is the efficient use of the reading block
Type: practical/economic inference. Requires a marginal-return comparison, not an existence claim. Required assumptions: (a) the time is genuinely substitutable; (b) knowledge instruction's comprehension yield per hour exceeds strategy instruction's; (c) the knowledge taught transfers to unseen texts. Assessment: (a) is plausible; (b) is crux C2 and no study in the ledger measures it — no equal-time head-to-head trial exists; (c) is crux C3, and the best evidence is mixed in a specific, informative way: Kim et al. 2023 found ES = .22 on near-transfer passages, .17 on mid-transfer, and .04, n.s., on far-transfer passages with no lexical overlap. Note that the practical conclusion (D1-P7) would survive even if D1-P6 were abandoned entirely — T2 is enough for it. Hirsch's policy argument does not need his most contested claim.
I7. D1-P7 individual ⟹ class-wide · the hidden step
Type: a concession that functions as a premise change. sk:ch12_C24 (2023): "shared topic knowledge is not an educationally useful characteristic — unless and until specific and predictable topic knowledge has been taught and learned by all the readers in the class." Corpus warrant (sk:ch12_W_R11, explicit, evidence: "The Kim study shows that topic knowledge is not educationally useful unless it is specifically taught to the whole class"): "Individual variations in prior knowledge among students render topic-based comprehension gains useless for standardized classroom instruction unless the curriculum creates a uniform knowledge base." Recorded vulnerability: "Differentiated instruction or personalized learning pathways could leverage students' diverse existing knowledge bases without requiring a rigid, identical curriculum for all." Assessment: this is the most consequential sentence in the dossier's slice. It concedes that the cognitive mechanism by itself yields no instructional lever, and relocates the argument onto curriculum coordination — an organisational claim requiring intervention evidence, not laboratory evidence. Everything after this point belongs to Dossier 2.
6. Evidence ledger
source_id is the key in data/evidence_registry/sources.json. "Books" is the number of the ten books in which the registry records at least one occurrence. Directness is a distinction the corpus's evidence_type field does not encode and this dossier adds: tests (directly tests the proposition) / illustrates (analogy or demonstration) / context (trend or historical background) / program (outcome evidence for Core Knowledge).
| # | source_id | Source | Kind | Independence | Bears on | Directness | Books | Status |
|---|---|---|---|---|---|---|---|---|
| 1 | S0027 | Recht & Leslie 1988, baseball study | study | primary_independent | D1-P4 (rhetorically P6) | tests | 6 | Contested — never directly replicated; measures criticised |
| 2 | S0021, S0070 | Bransford & Johnson 1972; Bransford, Barclay & Franks 1972 | study | primary_independent | D1-P1 | tests | 6 | Classic, robust |
| 3 | S0033 | Kintsch & van Dijk; Kintsch 1988 CI model | study | primary_independent | D1-P1, P3 | illustrates (cited as authority) | 4 | Established — and partly against Hirsch |
| 4 | S0012 | Miller 1956, "Magical Number Seven" | study | primary_independent | D1-P3 | illustrates | 7 | Classic, superseded in detail |
| 5 | S0013, S0080 | de Groot 1946; Chase & Simon 1973 | study | primary_independent | D1-P6 (the bridge) | illustrates (analogy) | 6 | Established for chess; the strong null is superseded |
| 6 | S0019 | Ing, Lunney & Olsen, AFQT/NLSY79 reanalysis | study | primary_independent | D1-P5 | tests (psychometric) | 5 | Real, but not peer-reviewed, and misdated in the corpus |
| 7 | S0362 | Rosenshine & Meister 1994, reciprocal teaching | study | primary_independent | D1-P6 | tests (meta-analytic) | 2 | Established, over-read |
| 8 | S0020 | Willingham & Lovette 2014; Willingham 2006 | study (in fact commentary) | primary_independent but see card | D1-P5, P6 | illustrates | 5 | Sympathetic, over-recruited |
| 9 | S0105, S0374 | Cunningham & Stanovich; Stanovich Matthew effects | study | primary_independent | D1-P4, P5, P7 | tests (correlational) | 3 | Replicated as association; the Matthew pattern is contested |
| 10 | S0031 | Hart & Risley 1995, Meaningful Differences | study | primary_independent | D1-P7 / equity framing | context | 4 | Contested — magnitude not replicated |
| 11 | S0005 | Grissmer et al. 2023, Colorado CK lottery | study | ckf_affiliated | D1-P7 (→ Dossier 2) | program | 5 | Not peer-reviewed; high attrition; author on the CKF board |
| 12 | (not in registry) | Cabell et al. 2025, CKLA kindergarten RCTs | study | primary_independent (declared firewall) | D1-P7 equity direction | program | 0 — postdates the corpus | Peer-reviewed; null on standardized outcomes |
| A | S1348 | Steffensen, Joag-Dev & Anderson 1979 | study | primary_independent | D1-P1, P2 | tests | 1 | Underused — the cleanest test of D1-P2 |
| B | S0060 | Anderson, Reynolds, Schallert & Goetz 1977 | study | primary_independent | D1-P1 | tests | 3 | Underused |
| C | S0366 | Schneider, Körkel & Weinert 1989, soccer | study | primary_independent | D1-P4 | tests | 1 | The conceptual replication Hirsch has and barely uses |
| D | S0322 | Hwang, McMaster & Kendeou 2023 | study | primary_independent | T1 directionality | tests | 1 | Best directionality evidence; enters in 2023 |
| E | S0161, S0162 | Kim et al., MORE / content-literacy trials | study | primary_independent | D1-P7, crux C3 | program | 2 | Best transfer evidence; enters in 2023–24 |
6.1 What the ledger shows before any single card is read
- The mechanism base is narrow and old. Excluding program evidence, D1-P1–P4 rest on a handful of studies from 1946–1988 plus one undated psychometric memo. The most recent independent mechanism evidence in the corpus (Hwang 2023, Kim 2022–24) enters only in the last two books.
- Directness varies enormously and the corpus does not encode it. Bransford, Recht & Leslie, Steffensen and Anderson test the mechanism. De Groot, Miller and Ericsson illustrate an analogy. NAEP and SAT trends are context — and they are the most-cited numbers in the whole corpus (SAT: 76 evidence items across all ten books).
- Assertion outruns evidence. Per the thesis layer, 17% of P01's asserting occurrences carry any direct evidence item, and 106 of its 121
sameoccurrences carry none. For P02 the figure is 20% (120/136sameoccurrences carry none). - Independence is mostly clean, with one chain. Items 1–5, 9 and A–E are
primary_independent. Item 11 isckf_affiliated. Daniel Willingham is item 8 and a co-author of item 11 and currently listed on the Core Knowledge Foundation Board of Trustees — an independence flag on a chain, not a disqualification of any single piece. - Registry corrections this dossier produced. All 1,515 registry entries currently carry
doi: null/resolution_status: null, andresolutions.json(64 attempted) contains several wrong hits. Verified corrections are insources.jsonin this directory; the worst are listed in §13.3.
7. Evidence appraisal cards
Every card ends with verified_by. Where a number could not be seen in a primary or authoritative source it says unverified — that is a real state, not a placeholder to be filled in with a plausible figure.
Card 1 · Recht & Leslie 1988 — "the baseball study" · S0027
- Citation. Recht, D. R., & Leslie, L. (1988). Effect of prior knowledge on good and poor readers' memory of text. Journal of Educational Psychology, 80(1), 16–20. DOI 10.1037/0022-0663.80.1.16. Not open access.
- Design. 2×2 between-subjects: reading ability (high/low) × baseball knowledge (high/low), 16 per cell. Silent reading of a narrative account of a half-inning, then re-enactment with wooden figures on a scale model field with verbal description, a distractor task, written summarization, and sorting 22 sentences by importance.
- Sample. N = 64 US junior-high students — 32 seventh-graders, 32 eighth-graders.
- Effect. Primary-verified. The multivariate knowledge effect was significant, F(8,53) = 20.8, p < .001; the multivariate reading-ability effect was not, F(8,53) = 1.94, p > .05, nor was the interaction, F(8,53) = .43, p > .05. Every univariate knowledge test was significant. The rhetorically important cell comparison is real: low-ability/high-knowledge readers exceeded high-ability/low-knowledge readers on verbal recall (113.8 vs 68.4 propositions) and summary quantity (30.4 vs 14.5). It remains a comparison of selected pre-existing groups, not a randomized knowledge intervention.
- Bears on. D1-P4, and rhetorically D1-P6.
- Can establish. That among these 64 selected students reading one matched-topic text, high pretested baseball knowledge was associated with substantially more and more expert-like recall/summarization across both measured reading-ability groups.
- Cannot establish. That acquired knowledge caused the difference; that knowledge generally outweighs reading ability; that no general skill exists; or anything about unfamiliar texts, other ages/domains, transfer or standardized comprehension.
- Status: contested, and structurally weaker than its reputation. Three problems, in order of severity:
- Knowledge was measured, not manipulated. Eligible students were sampled from the top/bottom knowledge and comprehension bands. The paper therefore cannot identify a causal effect of acquiring knowledge, despite the size and consistency of the group differences.
- The headline rests on main effects and a null interaction, not on a significant crossover. "Poor readers beat good readers" accurately describes several cell means but is not itself a reported interaction effect. Failure to detect an ability effect in 16-person cells is not evidence that reading ability never matters.
- The low-ability group was restricted in a narrower way than secondary accounts imply. It scored below the 30th percentile on comprehension but had to score above the 30th percentile in vocabulary to avoid word-recognition problems. The result applies to poor comprehenders with adequate vocabulary, not to every weak reader. The cells are also strongly gender-imbalanced: high-knowledge cells contain 22 boys/10 girls, low-knowledge cells 10 boys/22 girls.
- The knowledge measure is a trivia/vocabulary measure. Reynolds, Hattan & Markham (2025) reviewed 19 "baseball studies" 1978–2018, found 13 used the same two knowledge measures, that these "focused heavily on vocabulary and baseball trivia", and that the standard comprehension text was "deceptively complex." Never directly replicated in 35+ years (Shanahan: "It's a one off"). The nearest thing is a conceptual replication in another language and sport — Schneider, Körkel & Weinert 1989 (card E).
- How Hirsch uses it. Cited in 6 of 10 books. The 2006 statement retains the qualifier "for that particular text" (
kd:ch2_C77); by 2010 it has become "shown decisively" (the-making-of-americans:ch6_C22); by 2023 "a dominant theme" (sk:ch12_C29). - verified_by. Full primary paper at
https://www.yesataretelearningtrust.net/Portals/0/Effect-of-Prior-Knowledge-on-Good-and-Poor-Readers-Memory-of-Text.pdf(read completely: selection criteria, procedure, all means/SDs, MANOVA and univariate tests, discussion); Crossrefhttps://api.crossref.org/works/10.1037/0022-0663.80.1.16; Semantic Scholar record for Reynolds et al.https://api.semanticscholar.org/graph/v1/paper/DOI:10.1002/rrq.575(abstract verbatim);https://www.shanahanonliteracy.com/blog/prior-knowledge-or-he-isnt-going-to-pick-on-the-baseball-study(non-replication argument; its percentile description is corrected above from the primary paper).
Card 2 · Bransford & Johnson 1972; Bransford, Barclay & Franks 1972 · S0021, S0070
- Citations. Bransford, J. D., & Johnson, M. K. (1972). Contextual prerequisites for understanding: Some investigations of comprehension and recall. Journal of Verbal Learning and Verbal Behavior, 11(6), 717–726. DOI 10.1016/S0022-5371(72)80006-9. Bransford, J. D., Barclay, J. R., & Franks, J. J. (1972). Sentence memory: A constructive versus interpretive approach. Cognitive Psychology, 3(2), 193–209. DOI 10.1016/0010-0285(72)90003-5. ⚠ Author order correction: it is Bransford, Barclay & Franks — not "Barclay, Bransford & Franks" as SCOPE has it. The corpus (
ae:ch7_E11) has it right. - Design. Four experiments manipulating when context is supplied (no context / context before / context after / partial context) for the "Washing Clothes" and balloons-serenade passages; dependent variables a subjective comprehension rating and free recall in idea units. The 1972 Cognitive Psychology paper is a recognition-memory study showing subjects "recognise" sentences describing the situation they inferred rather than the sentences presented.
- Sample. Primary-verified total N = 154. Experiment I: 50 high-school volunteers, five groups of ten; Experiment II: 52 university students; Experiment III: 21 high-school volunteers; Experiment IV: 31 high-school volunteers.
- Effect. Primary-verified. In Experiment I, context-before comprehension was 6.10/7 and recall 8.00/14 versus 2.30 and 3.60 with no context; both exceeded every comparison condition, p < .005. Topic-before likewise exceeded topic-after/no-topic on recall in Experiments II–IV and on comprehension in II–III (all reported p < .005; Experiment IV recall p < .05).
- Bears on. D1-P1.
- Can establish. That comprehension and memory are constructive and depend on an activated schema, and — because the title is a manipulated variable — that the relation is causal.
- Cannot establish. Effect sizes relevant to classroom instruction; how much school-taught knowledge matters; that strategies cannot help. The passages are engineered to be incomprehensible without the topic — maximal internal validity, minimal ecological validity.
- Status: classic and conceptually robust. The most secure item in the ledger. Routinely reproduced as a demonstration rather than as formal replication; an OSF replication project exists but its result could not be retrieved (unverified).
- verified_by. Full primary paper at
https://labs.la.utexas.edu/gilden/files/2016/04/Bransford_Johnson_1972.pdf(all ten pages visually read/OCR-checked, including methods, tables and tests); ERIChttps://eric.ed.gov/?id=ej068571; Crossrefhttps://api.crossref.org/works/10.1016/0010-0285(72)90003-5(second paper's author order, journal and pages). The second Bransford-Barclay-Franks paper was not read in full.
Card 3 · Kintsch & van Dijk; Kintsch's construction–integration model · S0033
- Citations. Kintsch, W. (1988). The role of knowledge in discourse comprehension: A construction-integration model. Psychological Review, 95(2), 163–182. DOI 10.1037/0033-295X.95.2.163. van Dijk, T. A., & Kintsch, W. (1983). Strategies of Discourse Comprehension. Academic Press. Kintsch, W., & Keenan, J. (1973). Cognitive Psychology, 5(3), 257–274. DOI 10.1016/0010-0285(73)90036-4.
- Kind. Theoretical model plus simulation; not an experiment on Hirsch's question.
- Bears on. D1-P1 and D1-P3 — it supplies the vocabulary ("situation model") for the mechanism.
- What it can establish. That a knowledge-integrated situation model is required for deep comprehension; that proposition density predicts reading time.
- What it cannot establish — and this is the point. The CI model is precisely a theory of general, content-independent processes. Construction is bottom-up and context-blind, using "weak" rules that fire whether or not they are relevant, producing an incoherent network; integration is a constraint-satisfaction relaxation over that network. Knowledge is the input to the mechanism, not the mechanism. The architecture is the same across domains — Kintsch applied it to word identification, arithmetic word problems and decision biases alike. The correct Kintschian statement is: there is a general comprehension mechanism whose output quality is knowledge-dependent — which is not the same as "there is no general comprehension process."
- Status: established, and partly dissenting from the use Hirsch makes of it. 4,153 citations; extended rather than overturned. Hirsch cites Kintsch for the situation model while the model posits the general machinery D1-P6 denies (§8.4).
- Caveat. The primary paper establishes what Kintsch's theory says, not that the model is the uniquely correct architecture or that its general processes are teachable as a freestanding proficiency. Its empirical demonstrations and simulations are restricted.
- verified_by. Full primary paper at
https://condor.depaul.edu/dallbrit/extra/hon207/readings/kintsch-1988-construction-integration.pdf(read completely, including architecture, simulations, limitations and conclusion); PubMedhttps://pubmed.ncbi.nlm.nih.gov/3375398/; Crossrefhttps://api.crossref.org/works/10.1016/0010-0285(73)90036-4for Kintsch & Keenan 1973.
Card 4 · Miller 1956, "The Magical Number Seven" · S0012
- Citation. Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63(2), 81–97. DOI 10.1037/h0043158. Free full text: https://psychclassics.yorku.ca/Miller/
- Kind. Review/synthesis of others' data — not a single experiment.
- Bears on. D1-P3 — the architectural constraint that makes knowledge non-substitutable by lookup.
- Can establish. That immediate memory is capacity-limited in chunks and that chunking mitigates it.
- Cannot establish. That the limit is 4–7 for reading comprehension specifically.
- Status: classic, superseded in detail — and Hirsch's use inverts Miller's own caution. Verbatim from the full text, Miller separates the two limits — "the span of absolute judgment and the span of immediate memory are quite different kinds of limitations that are imposed on our ability to process information" — and dismisses their numerical agreement: "I suspect that it is only a pernicious, Pythagorean coincidence." Cowan (2001, BBS 24(1), 87–114, DOI 10.1017/S0140525X01003922) states in his abstract that Miller's number "was meant more as a rough estimate and a rhetorical device than as a real capacity limit," and locates a central capacity limit "averaging about four chunks." Cowan (2010) gives "3 to 5 meaningful items in young adults." Whether the limit is discrete slots or a flexible resource is still actively disputed (Morra, Patella & Muscella 2024, Journal of Cognition 7(1):60, DOI 10.5334/joc.387).
- Net. The qualitative premise D1-P3 needs — comprehension is capacity-bound, so knowledge must be possessed and fast — survives intact. The specific number Hirsch quotes across seven books does not, and Miller said so first.
- verified_by.
https://psychclassics.yorku.ca/Miller/(full text read: the two-limits sentence, the 2.5-bit pitch capacity, the "Pythagorean coincidence" conclusion);https://api.crossref.org/works/10.1037/h0043158;https://api.crossref.org/works/10.1017/S0140525X01003922(Cowan 2001 abstract verbatim);https://pmc.ncbi.nlm.nih.gov/articles/PMC2864034/(Cowan 2010 full text);https://pmc.ncbi.nlm.nih.gov/articles/PMC11259112/(Morra et al. 2024 full text).
Card 5 · de Groot 1946; Chase & Simon 1973 — the chess analogy · S0013, S0080
- Citations. de Groot, A. D. (1946/1965). Thought and Choice in Chess. Chase, W. G., & Simon, H. A. (1973). Perception in chess. Cognitive Psychology, 4(1), 55–81. DOI 10.1016/0010-0285(73)90004-2. A public scan of the original is readable at
andymatuschak.org/prompts/Chase1973.pdf. - Design and sample. N = 3. Verbatim from the paper: "Three chess players, a master (M), a Class A player (A), and a beginner (B), were used as subjects." Two tasks (perception, 5-second memory), 10 middle-game and 10 end-game positions plus 8 random positions.
- Effect (verbatim). Middle-game: "M was able to place an average of about 16 pieces correctly on the first trial, while A and B placed about eight and four, respectively." Random: "In the random, unstructured positions there was no relation at all between memory of the position and playing strength."
- Bears on. D1-P6 — this is the bridge from expertise to comprehension.
- Can establish. That expert recall in chess is domain-knowledge-based, not general-memory-based.
- Cannot establish. That comprehension of arbitrary prose is relevantly like chess recall. The analogy is nowhere tested in the corpus.
- Status: established for chess, but the strong null Hirsch relies on is superseded. Gobet & Simon (1996), Psychonomic Bulletin & Review 3(2), 159–163, DOI 10.3758/BF03212414, pooled 13 experiments: "the strongest skill group outperforms the weakest in 12 cases out of 13… the probability of the strongest group outperforming the weakest in twelve or more cases is .0017" (F(3,17) = 10.35, p < .001). The advantage is small — "roughly one piece per additional 400 ELO points," against about five pieces for game positions — but it is real, and the single exception in the 13 is Chase & Simon's own three-subject study. So the master's advantage does not vanish on random boards; it shrinks by ~80%.
- Sharpest point for the editor. Hirsch cites Gobet & Simon (2000) elsewhere in the corpus (
wkm:ch4_E_new58, "Five Seconds or Sixty? Presentation Time in Expert Memory," Cognitive Science 24, 651–682). He is reading these authors. He continues to report the 1973 strong null, and stakes his central analogy on it (cl:ch2_OBJ13: "When experts are faced with random patterns … they perform no better than novices"). - verified_by.
https://andymatuschak.org/prompts/Chase1973.pdf(pp. 55–60 read: N = 3, stimulus construction, all piece counts, the random-position sentence);https://api.crossref.org/works/10.1016/0010-0285(73)90004-2;https://bura.brunel.ac.uk/bitstream/2438/1346/1/FullText.pdf(Gobet & Simon 1996 read: 12-of-13 count, F-values, the "one piece per 400 ELO" line, the Chase & Simon exception). de Groot 1946: metadata only, no title page fetched — publisher/edition unverified.
Card 6 · Ing, Lunney & Olsen — the AFQT/NLSY79 reanalysis · S0019
- Citation. Ing, P., Lunney, C. A., & Olsen, R. J. (n.d., internal evidence puts it at 2013 or later). Reanalysis of the 1980 AFQT Data from the NLSY79. Center for Human Resource Research, Ohio State University; published as NLSY79 Codebook Supplement, Appendix 24. Free PDF at
nlsinfo.org. No DOI. Not peer-reviewed — it describes itself as "this memo." - ⚠ Date correction. Both
wkm:ch1_E2and the registry date this 2007. It cannot be 2007: the document refers to a 2010 BLS decision, to the public data set "prior to 2013," to "the 2012 release of the item characteristics," and cites flexMIRT (Cai, 2012). The corpus reproduces Hirsch's own citation error. - Design and sample. Factor analysis (EFA/CFA, polychoric correlations, DWLS) plus 3PL IRT on recovered 1980 ASVAB item-level answer sheets. The ASVAB was administered to 11,914 of the 12,686 NLSY79 respondents; factor-analysis subsamples Word Knowledge n = 3,710 (35 items), Paragraph Comprehension n = 3,625 (15 items).
- Effect (verbatim). "when we combine the questions in Word Knowledge and Paragraph Comprehension, those pooled items are unidimensional. The Paragraph Comprehension items do not have notably high discriminant power, so they are not inherently superior to the simpler Word Knowledge items… the most efficient strategy is to drop the Paragraph Comprehension test and use whatever respondent time one is willing to devote to measuring verbal acuity to administering Word Knowledge items."
- Bears on. D1-P5.
- Can establish. That within a large national sample, vocabulary and paragraph-comprehension items load on one verbal factor, and that the comprehension items discriminate no better.
- Cannot establish. Causal direction — "vocabulary ← reading" is equally consistent. That school-taught knowledge moves the same construct. And note the finding is about test-scoring efficiency on a 15-item 1980 subtest, not about the nature of reading comprehension. Hirsch's reading of it — that vocabulary size is the strongest predictor of comprehension and general competence — goes beyond what the memo claims.
- Status: real document, correctly attributed to its authors, but weaker and later than cited. Not peer-reviewed; no replication or critique found.
- verified_by.
https://www.nlsinfo.org/sites/default/files/attachments/140116/AFQT%20Analysis%20Results%20Final%20Revised.pdf(pp. 1–10 read: authorship, "memo" self-description, 11,914/12,686, subsample n's, the quoted conclusion, the internal dating evidence);https://www.nlsinfo.org/content/cohorts/nlsy79/other-documentation/codebook-supplement/nlsy79-appendix-24-reanalysis-1980.
Card 7 · Rosenshine & Meister 1994, reciprocal teaching · S0362
- Citation. Rosenshine, B., & Meister, C. (1994). Reciprocal teaching: A review of the research. Review of Educational Research, 64(4), 479–530. DOI 10.3102/00346543064004479. Not open access. (The registry's title, "A Review of Nineteen Experimental Studies," is the earlier Technical Report No. 574; the published article reviews 16 studies. Don't conflate them.)
- Design. Quantitative research review of 16 studies, published and unpublished.
- Effect. The 1994 primary abstract reports median effect sizes .32 on standardized tests and .88 on experimenter-developed tests. The full 1993 author report's 19-study analysis gives the same .32 standardized median and .87 for experimenter-developed measures. Its overall median is .57, falling to .40 when restricted to studies with acceptable control-group activity.
- Bears on. D1-P6, via the plateau claim.
- What it actually reports about dose — and this is the appraisal. In the full author report, sessions ranged from six to 100. Significant and non-significant studies both had a 20-session median among good-poor/all students; below-average-student medians were 28 and 29. The ET/RT table reports medians of 20 (significant), 17 (mixed) and 20 (non-significant); RTO reports 14, 25 and 24. This is a comparison of heterogeneous studies grouped by statistical significance, not a dose-response experiment. A null on dose is not a plateau. It is equally inconsistent with "more is better" and with "benefit stops after six to ten lessons." Hirsch's
wkm:ch1_C87("useless after ten lessons") andkd:ch2_OBJ15("beyond six… wasteful and distracting") are an interpretation laid on top of a null. - Contrary primary evidence in the same literature. Westera & Moore (1995): 95% of an extended reciprocal-teaching group (12–16 sessions) gained, against 47% for a short group (6–8 sessions) and 45% for controls — the short arm was indistinguishable from control. Takala (2006) found 10 sessions less effective than 15.
- What it does support for Hirsch. The .32/.88 split is a genuine and serious construct-validity problem for the strategy literature: effects are much larger on tests aligned to what was practised. That point is Hirsch's to make, and it is well made.
- Status: established, over-read on dose, correctly read on measurement.
- verified_by. Full 1993 author technical report in the University of Illinois repository,
https://hdl.handle.net/2142/17744(all substantive sections and effect/session tables read); primary journal abstracthttps://doi.org/10.3102/00346543064004479(16 studies, .32/.88); ERIChttps://eric.ed.gov/?id=EJ500529;https://files.eric.ed.gov/fulltext/ED541081.pdf(Blazer 2007 — read: the verbatim "no relationship between the number of sessions… six to 100" and the Westera & Moore figures);https://blog.soton.ac.uk/edpsych/files/2021/06/Reciprocal-teaching-2019-Kirsty-Russell.pdf(read: the d = .88 / d = .32 split, Takala 2006). The full 1994 journal article was not read; its 16-study pool must not be treated as identical to the 19-study report.
Card 8 · Willingham & Lovette 2014; Willingham 2006 · S0020
- Citation. Willingham, D. T., & Lovette, G. (2014, September 26). Can reading comprehension be taught? Teachers College Record, content ID 17701. No DOI. Free author PDF at danielwillingham.com. Willingham, D. T. (2006). How knowledge helps. American Educator, Spring 2006,
https://www.aft.org/ae/spring2006/willingham(free). - Kind. ⚠ Commentary, not primary research and not a formal meta-analysis — a narrative review of eight prior quantitative reviews. The registry classifies it
kind: study; that is wrong and should be corrected tosecondary_retellingor a review kind. - What it actually says (verbatim from the PDF).
- "none of these reviews show that more practice with a strategy provides an advantage."
- "Ten sessions yield the same benefit as fifty sessions."
- "RCS instruction should be explicit and brief."
- "Far from a let-down, this strikes us as wonderful news."
- "When it comes to improving reading comprehension, strategy instruction may have an upper limit, but building background knowledge does not."
- Bears on. D1-P5, D1-P6.
- Appraisal. Hirsch's plateau claim is a fair reading of the saturation sentence — the corpus attributes it here (
wkm:ch1_E20) and the attribution holds. What he drops is the recommendation attached to it: teach the strategies, briefly. "Useless after ten lessons" (wkm:ch1_C87) is not what Willingham & Lovette wrote; they wrote that ten sessions buy what fifty buy, and that those ten are worth having. - Independence.
primary_independentin the registry, and that is defensible for this item on its own. But Willingham is a co-author of the Colorado lottery study (card 11) and is currently listed on the Core Knowledge Foundation Board of Trustees. The registry's independence flag is per-source and cannot see the chain; the dossier should. - Also note. The 2014 piece's own dose claim rests on reviews that did not manipulate dosage — no dose-response RCT is cited. So Hirsch's strongest source for the plateau inherits the same weakness as card 7.
- verified_by.
https://r.jina.ai/http://www.danielwillingham.com/uploads/…/willingham%26lovette_2014_can_reading_comprehension_be_taught_.pdf(full text read — all verbatim quotes above);https://www.aft.org/ae/spring2006/willingham(full text read);https://www.coreknowledge.org/board-staff/(Willingham listed as trustee; whether he held the seat in 2023 is unverified). Teachers College Record volume/page numbers: unverified.
Card 9 · Cunningham & Stanovich; Matthew effects · S0105, S0374
- Citations. Cunningham, A. E., & Stanovich, K. E. (1997). Early reading acquisition and its relation to reading experience and ability 10 years later. Developmental Psychology, 33(6), 934–945. DOI 10.1037/0012-1649.33.6.934. Stanovich, K. E. (1986). Matthew effects in reading. Reading Research Quarterly, 21(4), 360–407.
- Design and sample. Longitudinal follow-up of first-graders re-tested as eleventh-graders. N = 27 — a figure taken from Sparks et al. 2014's description; the primary participants section was not read. Cunningham & Stanovich's own account says "about one half of these students were available ten years later."
- Effect (ERIC abstract, verbatim). "first-grade reading ability predicted all 11th-grade outcomes — even when cognitive ability was partialed out — and was linked to print exposure, even after 11th-grade reading comprehension ability was partialed out. Print exposure predicted reading comprehension growth." Specific coefficients: unverified.
- Bears on. D1-P4, D1-P5, D1-P7 (cumulative advantage).
- Can establish. Strong predictive relations and a cumulative pattern in print exposure.
- Cannot establish. Causal direction; that curriculum-supplied knowledge produces the same trajectory. N = 27 is very small for a hierarchical regression with several partialed covariates.
- Status. The core longitudinal result replicates: Sparks, Patton & Murdoch (2014), Reading and Writing 27(2), 189–211, DOI 10.1007/s11145-013-9439-2, N = 54, same design, same conclusion. But the Matthew-effect framing is contested: Pfost, Hattie, Dörfler & Artelt (2014), Review of Educational Research 84(2), 203–244, DOI 10.3102/0034654313509492, found "no strong support for the general validity of a pattern of widening achievement differences," and their meta-analysis of baseline–growth correlations gives a negative mean (r = −0.214) — i.e. convergence, not fanning out, in those studies. Widening appears selectively (decoding efficiency, vocabulary, composite scores, when measurement precision is adequate).
- verified_by. ERIC
https://eric.ed.gov/?id=EJ561721;https://api.crossref.org/works/10.1037/0012-1649.33.6.934;https://www.msj.edu/…/Sparks,-Patterson,-Murdoch-2014.pdf(read: replication abstract and the "27 students" description);https://www.aft.org/sites/default/files/cunningham.pdf(read: authors' own 1998 account); ERIChttps://eric.ed.gov/?id=EJ1025240(Pfost et al. abstract verbatim incl. r = −0.214).
Card 10 · Hart & Risley 1995, Meaningful Differences · S0031
- Citation. Hart, B., & Risley, T. R. (1995). Meaningful Differences in the Everyday Experience of Young American Children. Baltimore: Brookes. (ERIC ED387210.)
- Design and sample. Longitudinal home observation of 42 families retained to the end — 13 professional, 10 middle-class, 13 lower-SES, 6 welfare — observed one hour per month from age 7–9 months to 36 months.
- The 30-million figure, as constructed. Hourly rates 2,153 / 1,251 / 616 words per hour; then "a linear extrapolation from the averages in the observational data to a 100-hour week (given a 14-hour waking day)" → 11.2M / 6.5M / 3.2M words a year → ~45M / 26M / 13M over four years. The 30-million gap is an arithmetic projection from roughly one recorded hour per month, not a measured cumulative count.
- Bears on. D1-P7 and the equity framing. Directness: context, not a test of the mechanism.
- Status: contested. Sperry, Sperry & Miller (2019), Child Development 90(4), 1303–1318, DOI 10.1111/cdev.13072, pooled five ethnographic community studies (42 children, 18–48 months) and did not support Hart & Risley's claim, finding wide within-SES variation and arguing that excluding multiple caregivers and overheard speech underestimates low-income children's exposure. A published exchange followed: Golinkoff, Hoff, Rowe, Tamis-LeMonda & Hirsh-Pasek (2019), DOI 10.1111/cdev.13128, and Sperry et al.'s reply, DOI 10.1111/cdev.13125. What is conceded by the critics of Sperry (verbatim): "SSM found variability within socioeconomic strata, a finding replicated by every study in the literature," and the 30-million-word gap is "a catchy phrase that let the public in on the research." What is contested: whether Sperry's design (no highly educated comparison group) can speak to the between-group claim, and whether overheard speech supports early language learning.
- How Hirsch uses it. Cited across four books; the corpus records no mention of the replication dispute — unsurprising for the three books that precede it, but he is still citing the figure in work published after 2019 elsewhere. This is a "keep an eye on it" item, not a refutation of anything in D1-P1–P6.
- verified_by.
https://www.aft.org/ae/spring2003/hart_risley(read: the 42 families, the 13/10/13/6 breakdown, the words-per-hour figures and the full extrapolation chain);https://api.crossref.org/works/10.1111/cdev.13072;https://files.eric.ed.gov/fulltext/ED617633.pdf(Golinkoff et al. read in full);https://api.crossref.org/works/10.1111/cdev.13125.
Card 11 · Grissmer et al. 2023, Colorado Core Knowledge lottery · S0005
- Citation. Grissmer, D., White, T., Buddin, R., Berends, M., Willingham, D., DeCoster, J., Duran, C., Hulleman, C., Murrah, W., & Evans, T. (2023). A kindergarten lottery evaluation of Core Knowledge charter schools. EdWorkingPaper No. 23-755, Annenberg Institute at Brown. DOI 10.26300/nsbq-hb21. Open access. ⚠ SCOPE's author list is wrong: Thomas White and Tanya Evans are authors; J. Soland is not.
- Peer review. None. The paper states: "This draft has not been submitted for publication nor been peer reviewed." No journal version was found as of August 2026.
- Design and sample. 14 oversubscribed kindergarten lotteries at 9 Core Knowledge charter schools in Colorado, entry years 2009–10 and 2010–11, in "predominately middle/high income school districts." 2,853 applications from 2,310 students. Outcomes: Colorado PARCC reading/ELA and maths, grades 3–6, plus a grade-5 science test. Confirmatory sample defined by dropping the four lotteries with the highest differential attrition and all young students in the delayed-entry window.
- Effect (verbatim, Table 12, confirmatory sample). Reading/ELA grades 3–6 combined: ITT 0.241
***, TOT 0.473***. Grade 4 ITT0.196**; grade 5 ITT0.281**; grade 6 ITT0.208*. Mathematics insignificant (all-gender ITT 0.081). Grade-5 science ITT0.154*, TOT0.300*. (Asterisks are the paper's significance markers, not emphasis.) Standard errors: unverified — only point estimates with significance asterisks were visible. - Attrition and take-up (verbatim). "The overall attrition rate for all 10349 applications from 3rd to 6th grade is 35.5%." "Lottery winners (32.5%) have lower attrition than lottery losers (37.1%) and a statistically significant level of differential attrition (
−4.6***)." "47.0% of applications resulted in enrollment, leaving a 53.0% rate of non-compliance." (The paper elsewhere describes a "52% compliance rate"; these two statements are inconsistent as printed and could not be reconciled.) - Bears on. D1-P7 — and properly belongs to Dossier 2. It is here because it is the empirical payoff Hirsch points to from this dossier's chapters, and because of its independence structure.
- Can establish. That attending these schools raised tested reading achievement relative to lottery losers.
- Cannot establish. That the curriculum rather than school culture, teacher selection, peer composition or parental commitment caused it; that CKLA at scale replicates it.
- Independence.
ckf_affiliated. No conflict-of-interest statement appears in the paper. Daniel Willingham is an author and is currently listed on the CKF Board of Trustees. No author lists a CKF affiliation. Funding: IES R305E090003, NSF 1252463, Arnold Foundation, Smith-Richardson. - Status: promising and seriously caveated. 35.5% attrition with significant differential attrition would not clear WWC standards without adjustment; the confirmatory sample was constructed by dropping the worst-attrition lotteries; 53% non-compliance makes the TOT a strong-assumption estimate.
- verified_by.
https://edworkingpapers.com/ai23-755(landing page: authors, DOI, abstract);https://r.jina.ai/https://edworkingpapers.com/sites/default/files/ai23-755.pdf(full paper read: Table 12, attrition, enrolment, funding, the no-peer-review line); ERIChttps://eric.ed.gov/?id=ED672259;https://www.coreknowledge.org/board-staff/.
Card 12 · Cabell et al. 2025, the CKLA kindergarten RCTs · (not in the registry — postdates the corpus)
- Citation. Cabell, S. Q., Kim, J. S., White, T. G., Gale, C. J., Edwards, A. A., Hwang, H., Petscher, Y., & Raines, R. M. (2025). Impact of a content-rich literacy curriculum on kindergarteners' vocabulary, listening comprehension, and content knowledge. Journal of Educational Psychology, 117(2), 153–175. DOI 10.1037/edu0000916. Peer-reviewed.
- Design and sample. Two school-level cluster-randomised trials (Study 2 replicating Study 1) of the CKLA: Knowledge Strand against business-as-usual, roughly one semester. Combined N = 1,194 students, 134 classrooms, 47 schools.
- Effect (Hedges' g, verbatim).
- Curriculum-specific vocabulary 0.63 (SE .06, p < .001); curriculum-specific knowledge: plants 0.26, Native Americans 0.93.
- Standardized measures: PPVT 0.01 (p = .874); WJ Picture Vocabulary 0.09; CELF Sentence Structure −0.05 (p = .301); Test of Narrative Language 0.02 (p = .725); WJ Science 0.11 (p = .054); WJ Social Studies 0.00 (p = .999).
- The interaction, verbatim from the abstract. "Significant interactions were found for vocabulary and content knowledge, such that children who began the year with relatively higher receptive vocabulary scores derived a greater benefit of learning the words and content knowledge taught in the curriculum." On plants knowledge the effect was significant at the mean (0.25) and at +1 SD (0.38) but not at −1 SD (0.13, p = .139).
- Bears on. The equity direction in D1-P7 — and it points the other way. A knowledge-building curriculum widened the vocabulary gap on the measures where it worked.
- Independence. Declared: IES R305A170635 and OSEP H325D190037; "no known conflicts of interest"; and explicitly, "There was a firewall between the research team and: (a) the developer (Core Knowledge Foundation) and (b) the publisher (Amplify Education) of CKLA: Knowledge." This is a materially stronger independence declaration than card 11's, and the contrast is the sharpest single appraisal point available in this dossier.
- Limits. One semester; proximal measures built by the research team; waitlist control; the K–2 results from the parent grant (IES R305A170635, ~1,440 students across 48 schools, 160 days of instruction) are not yet published — unverified.
- verified_by.
https://r.jina.ai/https://psycnet.apa.org/fulltext/2025-46446-001.pdf(full text read: citation, Ns, every effect size and SE, the interaction estimates, the funding and firewall statements);https://improvingliteracy.org/resource/boosting-early-literacy-through-content-rich-instruction/;https://ies.ed.gov/funding/grantsearch/details.asp?ID=1791(parent grant).
Two notable underuses
Underuse A · Steffensen, Joag-Dev & Anderson 1979 — cited once, in 1987 · S1348
Steffensen, M. S., Joag-Dev, C., & Anderson, R. C. (1979). A cross-cultural perspective on reading comprehension. Reading Research Quarterly, 15(1), 10–29. DOI 10.2307/747429. Paywalled at JSTOR, but the identical study is free as Center for the Study of Reading Technical Report No. 97 (July 1978) at ideals.illinois.edu/items/18083.
American and Indian adults read letters describing an American and an Indian wedding, matched for readability (136 and 127 idea units). N ≈ 38–39 (19 Indian, 20 American, one excluded); note that reported F tests use df = (1,35), which does not cleanly match that N — the exact analytic N is unverified. Verbatim results: reading-time interaction F(1,35) = 10.09, p < .01; gist recall interaction F(1,35) = 39.84, p < .01; culturally appropriate elaborations F(1,35) = 208.67; culturally based distortions F(1,35) = 128.24. Cell means for Americans/American vs Americans/Indian: 52.4 vs 27.3 gist units, 5.7 vs 0.2 elaborations, 0.1 vs 5.5 distortions.
Why it matters more than its one citation suggests. This is the cleanest test of D1-P2 in the entire ledger — not "topic knowledge helps" but "the knowledge the writer assumed is culturally specific, and lacking it produces systematic distortion, not merely reduced recall." It is exactly the evidence Hirsch's shared-knowledge thesis needs, and he cites it once, in 1987, in a chapter about literacy and culture (cl:ch1_E30). The corpus also carries the misspelling "C. Joag-Des".
One caution the dossier should record: the recall interaction is driven almost entirely by the American readers collapsing on the Indian passage (52.4 → 27.3); the Indian readers recalled about the same from both (37.9 vs 37.6). The elaboration/distortion split is the cleaner cultural-schema signal. Also, the "Indian" readers were US-resident and reading in a second language. 574 citations; no direct replication found.
Underuse B · Anderson, Reynolds, Schallert & Goetz 1977 — the prison-break passage · S0060
Anderson, R. C., Reynolds, R. E., Schallert, D. L., & Goetz, E. T. (1977). Frameworks for comprehending discourse. American Educational Research Journal, 14(4), 367–381. DOI 10.3102/00028312014004367. Free as Technical Report No. 12 (July 1976) at ideals.illinois.edu/items/17896.
N = 60: "30 students from a section of an educational psychology course (all female) designed specifically for persons planning a career in music education, and 30 students from two weight-lifting classes (all male)." Two deliberately ambiguous passages (prison break / wrestling match; card game / woodwind rehearsal). Verbatim: t(58) = 5.60 and t(58) = 6.53 on the disambiguating multiple-choice tests; passage × background interaction F(1,58) = 48.61; and "62% of the subjects reported that another interpretation never occurred to them."
Why it matters. The 62% figure is the strongest single demonstration in the ledger that schema selection is not experienced as interpretation — readers do not know they have chosen. That is a stronger and more interesting claim than "knowledge helps recall," and Hirsch never uses it.
The caution that must travel with it: background is perfectly confounded with sex (music = all female; wrestling = all male), and the card/music passage turns partly on lexical puns ("hand," "diamonds," "recorder") rather than on background knowledge as such. 722 citations; replication status unverified. All statistics above are from the 1976 technical report; the AERJ version's numbers were not read.
Three sources that enter late and carry more weight than their citation count
C · Schneider, Körkel & Weinert 1989 — the soccer study · S0366. Journal of Educational Psychology 81(3), 306–312, DOI 10.1037/0022-0663.81.3.306. Paywalled; the JEP article's own N is unverified (the authors' book-chapter restatement of the same programme gives N = 576 and N = 185 across grades 3, 5 and 7). Soccer experts and novices, crossed with general aptitude, read a story about a soccer game. Verbatim from the authors' chapter: "High- and low-aptitude soccer experts performed equally well on all measures of text recall and comprehension," with no main effect of general ability and no significant interactions; and "Third-grade experts recalled significantly more text units than both fifth-grade and seventh-grade novices." This is the answer to Shanahan's "it's a one off" objection — a conceptual replication in a different language, sport, country and decade. It sits in the corpus (wkm:ch4_E52, E53, and a thinker record for Wolfgang Schneider) and is cited in one book, against six for Recht & Leslie. Recommendation: promote it. Caveats: quasi-experimental, single domain, single text, and the key result is a failure to reject a null on aptitude, with no equivalence test visible.
D · Hwang, McMaster & Kendeou 2023 · S0322. Reading Research Quarterly 58(1), 59–77, DOI 10.1002/rrq.481, open access (CC BY-NC-ND; free author copy at ERIC ED623594). N = 10,706, ECLS-K:2011, kindergarten through fifth grade, random-intercept cross-lagged panel models with working memory, cognitive flexibility, language proficiency and basic literacy as covariates. Verbatim: the relation is "bidirectional and positive throughout the elementary years"; but "all coefficients of the cross-lagged paths from science to reading were significantly greater than those from reading to science (18.13 < χ²[1] < 49.99, p < .001)." This is the best directionality evidence for T1 in existence, and it half-concedes the opposite direction. Domain knowledge is operationalised only as science. Correlational.
E · Kim et al., the content-literacy trials · S0161, S0162. Kim, J. S., et al. (2023). A longitudinal randomized trial of a sustained content literacy intervention from first to second grade. Journal of Educational Psychology, 115(1), 73–98, DOI 10.1037/edu0000751; and Kim, J. S., et al. (2024). Time to transfer. Developmental Psychology, 60(7), 1279–1297. The 2023 trial: 30 schools, N = 2,952, blocked cluster RCT, rated by the What Works Clearinghouse as "Meets WWC standards without reservations." The numbers that matter for crux C3: science content reading comprehension ES = .18; near-transfer passages .22 (p < .001), mid-transfer .17 (p < .01), far-transfer .04 (n.s.). WWC judged the general-literacy outcome (MAP Reading Total, n = 2,275) indeterminate. The 2024 spiralled trial (N = 2,870, grades 1–3 with a grade-4 follow-up) does reach a domain-general comprehension outcome (ES = .11) and mathematics (.12), sustained at 14 months (.12 reading, .16 maths). Read together: school-taught knowledge transfers, and transfer is graded by how much of the taught vocabulary appears in the passage. Where the lexical footholds are absent, the effect is .04 and not significant. That single number is the most informative datum in the dossier for anyone testing Hirsch's transfer claim — in either direction. Caveat: all three Kim trials draw on overlapping schools in one district, so they are extensions, not independent replications.
8. Named objections — and Hirsch's response
Each entry gives the citation, DOI, open-access status, an objection stated as precisely as possible against a named target (a backbone proposition, an inference id from §5, or a specific evidence card), the limits of the objection itself, and what the corpus records as Hirsch's reply.
A finding that governs this whole section. A name search across all 10,312 claims, 2,885 evidence items and 521 thinker records finds zero occurrences of Catts, Kamhi, McNamara, Gough, Tunmer, Hoover, Counsell, Okkinga, Elleman, Filderman, Sperry, Cowan or Pfost. Shanahan appears — but only as an ally on levelled readers (wkm:ch4_E24); Snow appears only as Catherine Snow, co-author of Chall's Families and Literacy (1982), not Pamela Snow of the 2021 critical review. So for nine of the ten objections below the answer to "does Hirsch respond?" is no response in the ten books — and where he does respond, it is to a reconstructed position, not to a person.
8.1 Reynolds, Hattan & Markham 2025 — the baseball study's measures → attacks Card 1, and through it I3
Reynolds, D., Hattan, C., & Markham, M. (2025). Fair or foul? Interrogating the role of baseball knowledge in studies of knowledge and comprehension. Reading Research Quarterly, 60(1), e575. DOI 10.1002/rrq.575. Open access (CC BY-NC-ND). Companion: Reynolds, D., & Hattan, C. (2024). Baseball, presidents, and state test passages: considering gendered knowledge. The Reading Teacher, 77(6), 997–1000. DOI 10.1002/trtr.2330. Open access. (⚠ SCOPE lists Reynolds as sole author of both; the RRQ paper has three authors, the Reading Teacher piece two.)
The objection, precisely. Not that knowledge is unimportant — that the construct in the canonical demonstration is not the construct Hirsch's curricular argument needs. Nineteen "baseball studies" 1978–2018; 13 used the same two knowledge measures; those measures "focused heavily on vocabulary and baseball trivia"; the standard comprehension text was "deceptively complex." If "topic knowledge" in the showpiece experiment is sport-specific vocabulary plus trivia, the result looks closer to a lexical-access effect on a jargon-loaded passage than to evidence for the broad conceptual background knowledge a curriculum could supply. The companion adds that the proxy is gender-loaded, so group differences may be confounded with sample composition.
Limits of the objection. It reviews one proxy literature. It produces no counter-effect-size and does not claim knowledge is unimportant; the authors' own recommendation is to stop leaning on baseball studies and diversify the evidence base. It leaves non-baseball knowledge studies untouched. Caveat: the RRQ full text returned HTTP 402 on every route; all content above is from Crossref/OpenAlex/Semantic Scholar metadata plus two secondary accounts. The editor should read it (acquisition P0).
Hirsch's response. None — the work postdates the corpus. But note that kd:ch2_OBJ8 ("Reading comprehension consists merely of a general working vocabulary plus the application of formal comprehension strategies") is the closest recorded objection, and Hirsch's reply ("Experiments show that knowledge of the subject matter is far more important than technical reading strategies") presupposes exactly the knowledge/vocabulary distinction Reynolds argues the experiments failed to draw.
8.2 The strategy-instruction evidence base → attacks D1-P6 directly
Four items, and they do not point the same way.
- National Reading Panel (2000), Teaching Children to Read, Ch. 4 Part II. Free at
nichd.nih.gov/sites/default/files/publications/pubs/nrp/Documents/ch4-II.pdf. 481 studies reviewed in the comprehension area; 16 categories of strategy instruction identified, 7 judged to have (verbatim from NICHD's own summary) "a solid scientific basis for concluding that these types of instruction improve comprehension in non-impaired readers"; and "When used in combination, these techniques can improve results in standardized comprehension tests." That last clause is the direct denial of D1-P6. Caveat:ch4-II.pdfreturned unparsed binary; the quotations are from two NICHD-authored summary pages, not the chapter. The report also drew a formal Minority View and is 26 years old. - Filderman et al. (2022), Exceptional Children 88(2), 163–184, DOI 10.1177/00144029211050860 (restricted). Meta-analysis, 64 studies, struggling readers grades 3–12: overall g = .59, 95% CI [0.47, 0.74], with "significantly higher effects associated with researcher-developed measures, background knowledge instruction, and strategy instruction." It cuts both ways: strategy instruction is not inert, and the measurement-inflation problem Hirsch presses is confirmed inside a pro-intervention meta-analysis.
- Elleman (2017), Journal of Educational Psychology 109(6), 761–781, DOI 10.1037/edu0000180. 25 inference-instruction studies, K–12: "inference instruction was effective for increasing students' general comprehension, d = 0.58, inferential comprehension, d = 0.68, and literal comprehension, d = 0.28," with less skilled readers benefiting more on literal outcomes (d = 0.97 vs 0.06). A content-general intervention moving general comprehension by d = .58 is a direct empirical counterexample to D1-P6. Whether that holds specifically on standardized measures is unverified. See also Elleman & Oslund (2019), Policy Insights 6(1), 3–11, DOI 10.1177/2372732218816339 — a both/and position from authors who fully concede the knowledge half.
- Okkinga et al. (2018), Educational Psychology Review 30(4), 1215–1239, DOI 10.1007/s10648-018-9445-7, free to read (bronze OA). ⚠ Journal correction: this is Educational Psychology Review, not the Journal of Research in Reading as SCOPE has it. 52 studies, K = 125 effect sizes, whole-class delivery: "a very small effect on reading comprehension (Cohen's d = .186) for standardized tests and a small effect (d = .431) on researcher-developed reading comprehension tests. A medium overall effect was found for strategic ability (d = .786)," with larger effects "when the trainer was the researcher as opposed to teachers." This one cuts Hirsch's way, and hard: taught at classroom scale, strategy instruction reliably makes children better at performing the strategy and barely moves standardized comprehension.
Net. Taken together these establish that D1-P6 fails as stated — strategy effects are non-zero and reach standardized measures at least sometimes — while vindicating Hirsch's practical claim that at realistic classroom scale the returns are small and measurement-inflated. That is T2, not T3.
Hirsch's response. He addresses the category, never an author. kd:ch2_OBJ15: benefits are "real but small and initial; they only serve to show children that reading is communicative. Once that is understood, additional sessions (beyond six) are wasteful and distracting." kd:chappendix_OBJ2: "Most interventions produce some positive data, but these improvements are small and long-term results are unimpressive." The sources are recorded as "A 'predictable chorus' of researchers" and "Proponents of strategy instruction."
8.3 The simple view of reading → attacks D1-P6's architecture, while agreeing with much of it
Gough, P. B., & Tunmer, W. E. (1986), RASE 7(1), 6–10, DOI 10.1177/074193258600700104; Hoover, W. A., & Gough, P. B. (1990), Reading and Writing 2(2), 127–160, DOI 10.1007/BF00401799; Catts, H. W. (2018). The simple view of reading: advancements and false impressions. Remedial and Special Education 39(5), 317–323, DOI 10.1177/0741932518767563 — free full text at files.eric.ed.gov/fulltext/EJ1191985.pdf; Catts, Adlof & Ellis Weismer (2006), JSLHR 49(2), 278–293, DOI 10.1044/1092-4388(2006/023); Kamhi (2009), LSHSS 40(2), 174–177; Catts & Kamhi (2017), LSHSS 48(2), 73–76.
The objection, precisely. SVR locates comprehension in linguistic comprehension — a modality-general language ability. That yields a reader-side general factor D1-P6 has no room for: children with intact decoding and no obvious topic-knowledge deficit fail comprehension stably across topics because of oral-language weakness. Catts (2018), verbatim: Lonigan et al. "found that a considerable amount of the variance in RC (40%–70%) was shared by decoding and language comprehension and suggested that this shared variance may be the result of one or more general cognitive-linguistic factors." Catts, Adlof & Ellis Weismer identified 57 poor comprehenders, 27 poor decoders and 98 typical readers in eighth grade, with the poor comprehenders showing concurrent language deficits and normal phonological processing.
Where it agrees with Hirsch — and this should be said plainly. (a) There is no reading-specific comprehension skill distinct from language comprehension. (b) Generic strategy instruction has a small payoff — Catts cites Scammacca et al. (2015): "the average effect size of interventions on standardized measures of RC was .19." (c) Catts endorses Willingham directly: "Adequate content knowledge is critical for comprehension and should be central to any instruction directed at improving it (Willingham, 2006). Given the centrality of content knowledge, it is always surprising how little attention has been devoted to it in comprehension intervention." (d) Catts (2021/22, American Educator): "reading comprehension is not a skill someone learns and can then apply in different reading contexts."
Where it disagrees. Knowledge is one input among several, not what comprehension is made of; the poor-comprehender profile is a person-level, topic-invariant deficit a knowledge-only account under-predicts; and Catts and Kamhi replace Hirsch's single moderator (topic knowledge) with a four-way dependency on reader × text × task × purpose.
Limits of the objection. Catts's shared-variance factor is a residual construct he explicitly says is not yet characterised: "It is still possible that the cognitive-linguistic factors that underlie this common variance are malleable, but it is not clear at this point what they are and how they might be changed." Poor comprehenders are a minority subgroup, and their oral-language deficits are themselves partly vocabulary and world-knowledge deficits.
Hirsch's response. No response in the ten books — zero mentions of Catts, Kamhi, Gough, Tunmer or Hoover anywhere in the corpus. This is the most surprising silence in the dossier, because these are the friendliest critics available to him.
8.4 McNamara and colleagues — strategy × knowledge interactions → the sharpest counter to D1-P6
O'Reilly, T., & McNamara, D. S. (2007), AERJ 44(1), 161–196, DOI 10.3102/0002831206298171; McNamara, D. S., Kintsch, E., Songer, N. B., & Kintsch, W. (1996). Are good texts always better? Cognition and Instruction 14(1), 1–43, DOI 10.1207/s1532690xci1401_1; McNamara (2004), SERT, Discourse Processes 38(1), 1–30, DOI 10.1207/s15326950dp3801_1.
The objection, precisely. These isolate a trainable, domain-general processing competence that changes comprehension while topic knowledge is held constant. O'Reilly & McNamara (n = 1,651 high-schoolers), verbatim from the abstract: "Reading skill helped the learner compensate for deficits in science knowledge for most measures of achievement and had a larger effect on achievement scores for higher knowledge than lower knowledge students." ⚠ Correction to SCOPE: the published abstract attributes the compensation to reading skill, not to reading-strategy knowledge. Whether strategy knowledge specifically compensates is unverified — the full text could not be opened. The stronger version of the point is McNamara's SERT training study (n = 42), whose protocol analysis found training "helped these participants to use logic, or domain-general knowledge, rather than domain-specific knowledge to make sense of the text" — the general mechanism D1-P6 denies.
"Are Good Texts Always Better?" adds a boundary condition that appears nowhere in ten books: "readers who know little about the domain of the text benefit from a coherent text, whereas high-knowledge readers benefit from a minimally coherent text… the rewards to be gained from active processing are primarily at the level of the situation model." More knowledge is not monotonically better; what produces deep understanding is inferential work, for which knowledge is the raw material.
Limits of the objection. The compensation widens gaps rather than closing them — reading skill paid off most for students who already had the knowledge, a Matthew pattern consistent with Hirsch. SERT's benefit for low-knowledge readers appeared only on text-based questions, not on bridging-inference questions — shallow gains, not deep ones. The reverse-cohesion effect requires high background knowledge, so it presupposes rather than displaces Hirsch's premise, and O'Reilly & McNamara (2007, Discourse Processes) later narrowed it to less skilled, high-knowledge readers.
Hirsch's response. No response in the ten books — zero mentions of McNamara.
8.5 Kintsch as internal dissent, and McCarthy & McNamara's four dimensions → attacks D1-P6 from inside Hirsch's own authority
Kintsch (1988) is treated in Card 3. Add: McCarthy, K. S., & McNamara, D. S. (2021). The multidimensional knowledge in text comprehension framework. Educational Psychologist 56(3), 196–214, DOI 10.1080/00461520.2021.1872379 — free at files.eric.ed.gov/fulltext/ED616077.pdf.
The objection, precisely. Two moves. (a) Kintsch's own model is a general-process account; citing him for the situation model while denying general processes takes the vocabulary and discards the architecture. (b) "Background knowledge" decomposes into amount, accuracy, specificity and coherence (verbatim: "Amount refers to how many relevant concepts the reader knows. Accuracy refers to the extent to which the reader's knowledge is correct. Specificity refers the degree to which the knowledge is related to information in the target text. Coherence refers to the interconnectedness of prior knowledge"). Hirsch operationalises amount only — and the framework notes that "amount often serves as the default metric but is often confounded with other dimensions."
Accuracy makes knowledge a signed quantity, and this is the part with no corpus counterpart at all. Verbatim: "Inaccurate knowledge interferes with comprehension and the acquisition of new knowledge"; "students who held misconceptions generated significantly more incorrect inferences during reading and significantly fewer correct inferences than their peers (Kendeou & van den Broek, 2007)"; and inaccurate information "is easy to acquire and often resistant to change." The framework also reports a non-linearity: O'Reilly et al. (2019) found a knowledge threshold — below ~59% on a prior-knowledge test, prior knowledge had no significant relationship with comprehension performance.
Limits of the objection. The framework is emphatically pro-knowledge: it opens by citing prior knowledge as predicting 30–60% of the variance in comprehension, which is at least as strong as anything Hirsch claims, and it supports the knowledge-specificity view. It is a conceptual framework, not new data.
Hirsch's response. No response in the ten books. He cites Kintsch as an authority (S0033, 9 occurrences across 4 books) and never engages the model's architecture.
8.6 Smith, Snow, Serry & Hammond 2021 — the critical review → attacks the magnitude in D1-P4
Smith, R., Snow, P., Serry, T., & Hammond, L. (2021). The role of background knowledge in reading comprehension: a critical review. Reading Psychology 42(3), 214–240. DOI 10.1080/02702711.2021.1888348. Open access (green, via the ECU repository; bronze at T&F).
The objection, precisely. 23 studies, mid-to-late primary children. Verbatim from the abstract: "higher levels of background knowledge have a range of effects that are influenced by the nature of the text, the quality of the situation model required, and the presence of reader misconceptions about the text… Readers with lower background knowledge appear to benefit more from text with high cohesion, while weaker readers were able to compensate somewhat for their relatively weak reading skills in the context of a high degree of background knowledge."
That one adverb is the objection. Partial, not complete, compensation — which is precisely how Recht & Leslie is usually reported, and how Hirsch reports it in kd:ch2_C77 and (with the qualifier removed) in the-making-of-americans:ch6_C22. The review also independently reproduces the McNamara cohesion × knowledge interaction (§8.4) in a school-age population, and names misconceptions as a moderator.
Limits of the objection. 23 studies, one age band, no pooled effect sizes — a critical review, not a meta-analysis; and it is pro-knowledge in orientation. "Somewhat" is an abstract-level summary phrase and the full text could not be opened (both OA routes returned 403), so which studies drove it is unverified.
Hirsch's response. No response in the ten books (Pamela Snow does not appear; the corpus's Snow is Catherine Snow, on Families and Literacy 1982).
8.7 Willingham as sympathetic critic → attacks the use Hirsch makes of him
Treated in Card 8. The objection in one line: Willingham & Lovette answer "Can reading comprehension be taught?" with "not really, beyond a point" — and then recommend teaching the strategies, briefly. Hirsch quotes the headline and drops the recommendation, converting a bounded-returns finding into "useless." Corpus receipt: wkm:ch1_E20 is the plateau citation; wkm:ch1_C87 is the escalation.
Hirsch's response. He does not treat Willingham as a critic at all. Notably, the corpus records Hirsch crediting Willingham with restraining his "excessive absolutism" (ae:chappendix) — a rare self-aware moment worth surfacing on the page.
8.8 Cabell et al. 2025 — the equity direction → attacks the equity claim inside D1-P7
Treated in Card 12. The objection in one line: a peer-reviewed, developer-firewalled cluster-RCT of the curriculum Hirsch's foundation produces found essentially zero transfer to standardized vocabulary, listening comprehension and social-studies knowledge, and an interaction in which higher-vocabulary children benefited more. It does not touch D1-P1–P4. It bears directly on the claim that knowledge-building is especially equalising — a claim stated in this dossier's chapters (ae:ch1_C80, wkm:ch8_C72) and therefore not quarantinable in Dossier 2.
Hirsch's response. None — the work postdates the corpus.
8.9 Shanahan — the methodological sceptic and the pedagogical paradox → attacks Card 1 and I6
Shanahan, T. (2020, 14 March). Prior knowledge, or he isn't going to pick on the baseball study. https://www.shanahanonliteracy.com/blog/prior-knowledge-or-he-isnt-going-to-pick-on-the-baseball-study Free; blog, not peer-reviewed.
Three points, verbatim. (a) "It's a one off. There aren't other studies with this kind of finding… interesting that this has not been replicated in more than 30 years"; the study "was conducted with only a single text and that one designed to be used in a research study," and the sample excluded students reading below the 30th percentile, "potentially eliminating effects of reading ability." (b) "research not only shows that knowledge contributes to comprehension, but to miscomprehension as well (a research result, not an opinion)." (c) The pedagogical paradox: "If I'm always providing kids with the appropriate background knowledge to understand each text used for instruction, then how do students ever learn to take on a text on their own?"
Why (c) is the sharpest. It is a direct challenge to I6 that no amount of mechanism evidence answers: even granting D1-P1–P5 entirely, the instructional inference may be self-undermining. Shanahan elsewhere ("Knowledge or comprehension strategies — what should we teach?", 2023) argues "it should not be a choice between the two."
Limits. A blog. Shanahan does not dispute that knowledge matters; (b) is asserted without a citation beyond Bartlett (1932). A "5% to 30%" variance figure circulating under his name refers to syntax, not prior knowledge — unverified and should not be used.
Hirsch's response. He cites Shanahan approvingly on levelled readers (wkm:ch4_E24, wkm:ch4_C80; thinker record "Tim Shanahan", stance agrees) and never engages him on this. Two thinker records exist for Shanahan and neither concerns the baseball study.
8.10 Counsell — "right, but incomplete" → attacks the sufficiency of D1-P7, not its truth
Counsell, C. (2011). Disciplinary knowledge for all, the secondary history curriculum and history teachers' achievement. The Curriculum Journal 22(2), 201–225. DOI 10.1080/09585176.2011.574951. Paywalled. Companion, open access CC BY-NC: Chapman, A. (ed.) (2021). Knowing History in Schools: Powerful Knowledge and the Powers of Knowledge. UCL Press, DOI 10.14324/111.9781787357303, free at uclpress.co.uk/book/knowing-history-in-schools/.
The objection, precisely. Counsell's stated target is genericism, and on that she is Hirsch's ally — verbatim from the abstract: "such genericism is, first, redundant… and, second, inadequate: disciplinary knowledge and concepts are necessary in order to reach or challenge claims about the past." The move against Hirsch is the corollary: how claims are warranted within a discipline is itself knowledge, distinct from propositional content, and a curriculum specified as a list of things-to-know under-specifies it. Hirsch's account is almost entirely substantive.
Limits. UK-specific, secondary history, conceptual not empirical: it does not show that disciplinary teaching produces better comprehension. The article's own substantive/disciplinary definitions could not be verified — the only accessible full text is an image scan.
Hirsch's response. No response in the ten books. Feeds Dossier 4.
9. Hirsch's responses, as the corpus records them
The corpus holds 892 objections_raised records. These are the ones bearing on this dossier, with the response text as extracted. All twelve were checked against all_objections_raised in data/corpus_consolidated.json.
| Objection id | The objection | Hirsch's response | Attributed to |
|---|---|---|---|
cl:ch2_OBJ8 | The knowledge required is too vast or "profound" to teach | The schemata needed are abstract and elementary (Union won; Grant was a Union general) — cultural pointers, not expertise | (none recorded) |
cl:ch2_OBJ11 | Students can just look things up | Reference tools are "impractical and too slow for the integration required"; schemata must be in memory and quickly available | "Implicit/Educational Traditionalists" |
cl:ch2_OBJ13 | Expert performance reflects superior general ability | De Groot: on random board positions experts "perform no better than novices" | "Common assumption/pre-de Groot psychology" |
cl:ch2_OBJ4 | Is background knowledge part of meaning, or extralinguistic? | Inferences from prior knowledge are integrated into meaning from the outset | "Unattributed (posed as a conceptual question)" |
kd:ch2_OBJ6 | Comprehension follows decoding automatically | NAEP grade-4-to-12 divergence falsifies it | "General educational assumption" |
kd:ch2_OBJ7 | Comprehension transfers from simple to complex texts as decoding does | Progress is non-automatic and non-linear because it depends on domain knowledge | "Current reading programs / Common assumptions" |
kd:ch2_OBJ8 | Comprehension = working vocabulary + formal strategies | Experiments show domain knowledge matters far more than strategies | "Current reading programs" |
kd:ch2_OBJ10–OBJ11 | Start from each child's own prior knowledge | "Some validity" in the moment of reading, but proficiency needs a shared base | "Bank Street School" / "American educational doctrine" |
kd:ch2_OBJ14 | Young readers can't apply the knowledge they have, so strategies are needed | Endorses "questioning the author" as a way of triggering application | "Reading researchers who advocate for comprehension strategies" |
kd:ch2_OBJ15 | Research shows strategy exercises benefit young readers | Benefits are "real but small and initial"; sessions beyond six are "wasteful and distracting" | "A 'predictable chorus' of researchers" |
kd:ch6_OBJ4–OBJ5 | Tests can be levelled by supplying background inside the item | Fails: unfamiliar students must assimilate new content while answering — a speed and overload effect | "Test-makers (unnamed)" / "State test-makers" |
kd:chappendix_OBJ2 | Strategy-instruction data are positive | "Most interventions produce some positive data, but these improvements are small and long-term results are unimpressive" | "Proponents of strategy instruction" |
the-making-of-americans:ch1_OBJ10 | The SAT decline reflects changed test-taker composition | Jencks: scores fell as steeply in Iowa, then 98% white and middle class | "Unnamed 'some' (common sociological claim)" |
9.1 The structural finding
Not one of these responses is addressed to a named published critic. Every source field is a reconstructed position — "Proponents of strategy instruction," "Current reading programs," "A 'predictable chorus' of researchers." Across ten books and forty-seven years, on the question where he is strongest, Hirsch argues against positions he has assembled rather than against people who hold them.
Two consequences for how this dossier should be read:
- It is not evidence of evasion. The reconstructions are mostly fair, and
cl:ch2_OBJ11,kd:ch6_OBJ4–OBJ5andthe-making-of-americans:ch1_OBJ10are genuinely good answers to genuinely serious objections. - But it means the objections in §8 have never been answered, and cannot be assumed to be answerable by anything in the corpus. That is why §8 had to be acquired externally, and why the editor's verdict in §10 cannot lean on Hirsch's own replies.
10. Editorial assessment
The assessment has five axes and is recorded separately for T1, T2 and T3 (§2.1). A causal policy conclusion can be textually faithful and empirically underdetermined at the same time; a normative conclusion can be well argued without being proven. This vocabulary exists to keep those apart.
Axes. ① Textual fidelity — does the proposition fairly represent Hirsch? ② Inferential validity — do the stated premises support the conclusion? ③ Empirical support — how directly and strongly is it tested? ④ Robustness and independence — replication, counterevidence, source independence, scope. ⑤ Normative dependence — how much rests on values rather than findings?
10.1 TEMPLATE — for the editor
Leave the machine read in §10.2 alone; write here. Record author, date, status, rationale, decisive sources, strongest dissent, and what would trigger revision.
Assessed by: ______________________ Date: __________ Status: draft / provisional / settled
T1 — Prior knowledge is a large causal determinant of reading comprehension.
① textual fidelity [ ] rationale:
② inferential validity [ ] rationale:
③ empirical support [ ] rationale:
④ robustness/independence [ ] rationale:
⑤ normative dependence [ ] rationale:
decisive sources:
strongest dissent:
what would trigger revision:
T2 — At the margin of school time, knowledge buys more comprehension than strategy instruction.
① [ ] ② [ ] ③ [ ] ④ [ ] ⑤ [ ]
rationale:
decisive sources:
strongest dissent:
what would trigger revision:
T3 — There is no general reading-comprehension skill independent of topic knowledge.
① [ ] ② [ ] ③ [ ] ④ [ ] ⑤ [ ]
rationale:
decisive sources:
strongest dissent:
what would trigger revision:
Scale suggestion (not binding): established / well supported / plausible / underdetermined /
contested / not established / normative.
10.2 MACHINE PROVISIONAL READ — not an editorial verdict; written to be argued with
Authored by an LLM on 2026-08-18 from the evidence in §6–§8. It has no standing. Where it disagrees with the editor, the editor is right by definition; the point of writing it down is to make disagreement cheap.
T1 — knowledge is a large causal determinant of comprehension
| Axis | Read | Reason |
|---|---|---|
| ① Textual fidelity | High | T1 is what Hirsch says, everywhere, from 1977. No reconstruction needed. |
| ② Inferential validity | High | I1–I3 are sound; the premises support the conclusion. |
| ③ Empirical support | Strong, with the outcome-measure caveat | Bransford (causal manipulation), Anderson 1977, Steffensen 1979, Schneider 1989, Hwang 2023 (directionality, N = 10,706, knowledge→reading paths significantly larger). McCarthy & McNamara put prior knowledge at 30–60% of comprehension variance. |
| ④ Robustness/independence | Good, with one soft spot | All primary_independent. The soft spot is that the showpiece — Recht & Leslie — is the weakest item supporting it, and the strong items (Bransford, Steffensen, Anderson) are underused by Hirsch. |
| ⑤ Normative dependence | None |
Provisional: established. The correct qualification is partial compensation (Smith et al. 2021: "somewhat"), graded effects (Kim: near .22 / mid .17 / far .04), and signed knowledge (McCarthy & McNamara: inaccurate knowledge hurts). None of those overturn T1; all of them make it less useful as a slogan.
T2 — knowledge beats strategy at the margin of school time
| Axis | Read | Reason |
|---|---|---|
| ① Textual fidelity | High | swn:ch7_C41, kd:ch4_C44, moa:ch6_C35 state it cleanly. |
| ② Inferential validity | Medium | I6 requires a marginal-return comparison; Hirsch supplies an existence claim instead and lets it do the work. |
| ③ Empirical support | Underdetermined — no study in the ledger measures it | No equal-time head-to-head trial exists. The nearest evidence is indirect and points both ways: Okkinga (whole-class strategy d = .186 on standardized) supports it; Elleman 2017 (inference instruction d = .58 on general comprehension) and NRP 2000 cut against. |
| ④ Robustness/independence | Medium | Program evidence is ckf_affiliated with an author chain through Willingham. Kim's trials are independent and give a real but small domain-general effect (.11). |
| ⑤ Normative dependence | Low, but non-zero | "Efficient" presupposes comprehension is the objective being maximised. |
Provisional: plausible, underdetermined. This is the claim most worth a decisive study and the one least likely to get one.
T3 — no general reading-comprehension skill exists
| Axis | Read | Reason |
|---|---|---|
| ① Textual fidelity | High for the strong version, and that is the problem | sk:ch12_C21 says it flatly. But kd:ch2_C72, swn:ch7_C41 and ae:ch7_C88 state the weak version, and kd:ch2_C73 affirms measurable general proficiency. He holds both. |
| ② Inferential validity | Low — this is where the argument breaks (I5) | A graded causal claim cannot entail a universal negative. The two substitutes are a dose/plateau argument that does not establish T2's head-to-head allocation comparison and an untested analogy from chess. |
| ③ Empirical support | Weak, and contradicted | 106 of 121 same occurrences of P01 carry no evidence item. Elleman 2017 (d = .58, general comprehension) and NRP 2000 (7 strategies with a scientific basis; combinations improving standardized scores) are counterexamples. |
| ④ Robustness/independence | Weak | The bridge evidence is a three-subject 1973 study whose strong null was superseded in 1996 by authors Hirsch elsewhere cites. Kintsch, his own cited authority, is a general-process theorist. |
| ⑤ Normative dependence | Low | It is an empirical claim, stated as one. |
Provisional: not established as stated; the defensible neighbouring claim is plausible but its dose curve remains underdetermined. The reading on which T3 is true — no content-free skill is trainable with continuing returns — is consistent with the measurement-sensitive strategy literature, but no reviewed study manipulates enough doses to establish the continuing-return curve. The reading Hirsch's rhetoric asserts — no content-independent process operates in comprehension — is contradicted by the simple view of reading, by Kintsch's construction–integration model, and by non-zero trainable effects. Fixing C1 dissolves most of the dispute, and costs Hirsch nothing he needs.
Cross-cutting: overstatement in transmission
Four instances, each checkable against a dated occurrence:
- "for that particular text" (2006,
kd:ch2_C77) → "shown decisively" (2010,the-making-of-americans:ch6_C22) → "a dominant theme" (2023,sk:ch12_C29). - Willingham & Lovette's "ten sessions yield the same benefit as fifty" plus "RCS instruction should be explicit and brief" → "useless after ten lessons" (
wkm:ch1_C87). - Miller's own "pernicious, Pythagorean coincidence" → a hard architectural constant quoted across seven books.
- Chase & Simon's three-subject null → "experts perform no better than novices" (
cl:ch2_OBJ13), retained after Gobet & Simon 1996 and after Hirsch cites Gobet & Simon 2000.
None of these is fabrication. Each is a real source read at its strongest and then re-quoted one notch stronger than it reads. That pattern is the single most useful thing this dossier has to say to a reader deciding how much to trust the rest of the corpus.
10.3 What would change this assessment
- Toward Hirsch: a replication of the Recht–Leslie ordering with a validated, non-trivia, non-gender-confounded knowledge measure; a demonstration of breadth transfer to genuinely unseen topics (Kim's far-transfer cell moving off .04); CKLA-scale trials replicating the Colorado effect with the equity direction intact; a variance decomposition showing nothing content-general survives controlling for topic knowledge.
- Away from him: confirmation that strategy effects are dose-dependent well past ten lessons (Westera & Moore points that way already); robust replication of the McNamara compensation finding; further nulls like Cabell 2025 on standardized outcomes; publication of the Cabell K–2 results with the same pattern.
- On the circularity (C4): any well-designed test of the knowledge effect using an outcome that is not itself a knowledge-loaded reading test.
- On the corpus itself: if reading Recht & Leslie's primary text shows a reported interaction after all, Card 1 and the T1 ④ read both improve.
11. Cruxes and open questions
C1 — The definitional crux. Decides whether T3 is even a claim. Does "no general reading skill" mean (a) no content-independent process operates in comprehension, (b) no content-free skill is trainable with continuing returns, or (c) no measured comprehension ability generalises across topics? Hirsch's rhetoric asserts (a); his evidence supports (b); his defence of standardized tests (kd:ch2_C73) requires denying (c). Much of the heat in the reading debates is opponents attacking (a) while proponents defend (b). Resolvable by stipulation, not by evidence — and it is the cheapest thing on this list. Doing it costs Hirsch nothing he needs, and costs his critics their easiest target.
C2 — The marginal-return crux. The decisive empirical question. Per hour of grade K–5 instructional time, what is the comprehension return of knowledge-building content instruction versus comprehension-strategy instruction versus additional fluency work? This is what D1-P7 asserts and no study in the ledger measures it. Grissmer compares whole schools, not time allocations; the strategy meta-analyses have no knowledge-building arm; Kim's trials have no strategy arm; Filderman codes both as moderators in one model without a head-to-head. A well-powered, equal-time, three-arm trial in grades 2–5, with comprehension measured on both matched and unfamiliar topics, would settle more of this dispute than any other single study.
C3 — The transfer crux. Does school-taught knowledge raise comprehension of unseen topics, or only matched ones? Every mechanism study is matched-topic; D1-P7 needs breadth transfer. This crux now has a number. Kim et al. 2023: near-transfer .22, mid-transfer .17, far-transfer .04 (n.s.); the 2024 spiralled trial reaches a domain-general comprehension outcome at .11, sustained at 14 months at .12, and spills into mathematics at .12–.16. Cabell 2025 is the counter-evidence: zero on standardized vocabulary and listening comprehension. Current best reading: transfer is real, small, and graded by lexical overlap between what was taught and what is read.
C4 — The circularity crux. D1-P5 says reading tests are knowledge tests. If so, evidence that knowledge predicts reading-test scores is partly definitional rather than confirmatory. Hirsch cannot both use test scores as the outcome that proves knowledge matters and characterise those tests as knowledge tests, without conceding that the causal claim needs an outcome that is not a knowledge test. As far as the corpus shows this is unaddressed, and it is not merely a debating point: it determines which of the studies in §6 count as confirmations at all.
C5 — The normative bridge. Named here, argued in Dossier 4. Even with C1–C4 resolved in Hirsch's favour, which shared knowledge, chosen by whom, is not determined by any of it. One sentence in this dossier; the whole of Dossier 4.
11.1 Open questions this dossier could not close
- Independent replication and generalisation of Recht & Leslie. The primary statistics are now verified: the knowledge main effect is strong, the measured ability effect and interaction are null, and several poor-reader/high-knowledge cell means exceed good-reader/low-knowledge means. What remains open is whether this selected single-school, single-domain pattern replicates under preregistered sampling, balanced cells, unfamiliar topics and broader reading profiles.
- Whether reading-strategy knowledge specifically compensates for low domain knowledge, or only reading skill does (O'Reilly & McNamara 2007 full text). SCOPE assumed the former; the abstract says the latter. (P1.)
- The Cabell K–2 results from IES R305A170635 — the kindergarten paper is one semester of a multi-year trial. (Watch.)
- Whether the Grissmer paper's 47% enrolment and 52% compliance figures can be reconciled. As printed they are inconsistent.
- What Hirsch would say to Catts. The most interesting unwritten page in the corpus.
12. Evolution — how these propositions and their evidence changed across ten books
Counts are same-relation occurrences from data/theses/mappings.json, by book, in year order. A blank means the proposition is not asserted in that book.
| 1977 PoC | 1987 CL | 1996 SWN | 2006 KD | 2010 MoA | 2016 WKM | 2020 HtEC | 2022 AE | 2023 SK | 2024 RE | total | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| P02 — text is incomplete (D1-P1, P2) | 3 | 23 | 4 | 26 | 11 | 12 | 8 | 16 | 18 | 15 | 136 |
| P01 — no general reading skill (D1-P4, P6) | — | 5 | 17 | 19 | 9 | 22 | 10 | 20 | 14 | 8 | 124 |
| P03 — domain specificity / no transfer (D1-P3, P6, P7) | — | 5 | 27 | 3 | 2 | 14 | 12 | 9 | 4 | 2 | 78 |
| P06 — vocabulary as knowledge proxy (D1-P5) | — | — | — | 4 | 1 | 6 | — | — | — | — | 11 |
| P14 — reading tests are invalid (D1-P5) | — | — | — | 4 | — | 11 | — | — | — | — | 15 |
Five things this table shows that no essay can.
- The mechanism is a decade older than the negative claim. P02 is asserted in The Philosophy of Composition (1977), where the subject is prose readability, not schooling. P01 has no occurrence in 1977 at all. It enters in Cultural Literacy (1987) and never leaves. The negative existence claim is not what the cognitive science gave him; it is what he added.
- Its first form in 1987 is mild.
cl:ch1_C9: "Literacy is far more than a skill; it requires large amounts of specific information."cl:ch1_C46: reading and writing "cannot be treated as empty skills that are independent of specific knowledge." Both are compatible with T1 and T2 and commit to neither (a) nor (c) in C1. - The absolutism arrives in 1996 and hardens.
swn:ch3_C124: "It is a fallacy to assume that it is possible to teach abstract, generalized problem-solving ability."swn:ch4_C273: "Modern faith in general 'critical-thinking' skills lacks scientific justification." P03 peaks in the same book at 27 occurrences — the domain-specificity/transfer literature is what he acquires in the 1990s, and it is what makes the strong form say-able. - By 2016–2022 it is a flat assertion.
wkm:chprologue_C73: "A single, overarching general skill does not exist."ae:ch1_C78: "The idea that reading comprehension is a general skill is false." Note the drop of the word reading in the 2016 version — the claim has widened from reading to skill as such. - Then it softens, quietly, in the same books.
ae:ch7_C88(2022): "It is more accurate to speak of 'reading skills' than of 'reading skill.'"sk:ch12_C24(2023) concedes that individual topic knowledge is not an instructional lever. The late Hirsch is more careful than the middle Hirsch, and the corpus is the only place you can see that.
Evidence added over time. Miller and Kintsch are there from 1977. Bransford enters in 1977 and is still cited in 2022–24. De Groot enters in 1987 and peaks in 2020–22. Recht & Leslie enters in 2006 — eighteen years after publication, and after the strong claim it is used to support was already in print. The AFQT reanalysis enters in 2016. Schneider (1989) enters once, in 2016. Hwang (2023) and Kim (2022–24) enter in the last two books. The pattern is not "evidence accumulates and the claim follows"; it is "the claim is stated, then evidence is recruited to it."
Reframings, not retractions. Nothing in the table is retracted. What changes is framing: a theory of readability (1977) → a theory of cultural membership (1987) → a critique of educational ideology (1996) → a measurement critique (2006, 2016) → a theory of American ethnicity (2022) → a comparative-national argument (2023–24). The propositions are Lakatosian hard core; the surrounding argument is the auxiliary belt.
13. Raw atlas
Everything above is checkable against the corpus. These are the entry points.
13.1 The 38 backbone occurrences
Every occurrence id in §4 links to ../claim.html#<id>, which opens that claim with its full source_passage, page range, chapter, evidence, warrants and counter-arguments. Machine-readable source of truth: backbone.json in this directory; re-verify with .venv-hirsch/bin/python3 verify_quotes.py.
13.2 The thesis layer
- Spine and all 25 propositions:
../theses/index.html - The propositions this dossier draws on: P01 · P02 · P03 · P04 · P05 · P06 · P14 · P15
- Research-programme timeline:
../theses/timeline.html
13.3 Feedback to the registry and the thesis layer
Concrete, checkable corrections this dossier produced. All are unapplied — this dossier does not write to shared data.
A. Missing canonical propositions.
- No canonical proposition covers the working-memory constraint.
poc:ch4_C116(Miller's limit, 1977) maps tonone. It is load-bearing for D1-P3 and therefore for D1-P6. Add one. - P02 conflates two claims — "text is systematically incomplete" and "the missing content is community-shared" — which have different evidence (Bransford vs Steffensen/Krauss). Split it.
B. Registry field corrections (all 1,515 entries currently have doi: null, resolution_status: null):
| source_id | Correction |
|---|---|
S0027 | DOI 10.1037/0022-0663.80.1.16. resolutions.json has 10.1037//0022-0663.80.1.16 — double slash. |
S0012 | DOI 10.1037/h0043158. resolutions.json resolved miller_1956_magical_number to 10.1037/11347-011, which is a biographical entry titled "George A. Miller.", not the paper. |
S0020 | Kind is commentary/review, not study. No DOI (Teachers College Record ID 17701). resolutions.json resolved willingham_why_students to 10.1037/021535 — wrong work. |
S0362 | DOI 10.3102/00346543064004479. resolutions.json resolved rosenshine_direct_instruction to 10.1007/bf00364747, a paper on HLA-DQ polymorphism — clearly wrong. Also: the registry title says "Nineteen Experimental Studies" (the technical report); the published article reviews 16. |
S0019 | Year is wrong. Not 2007 — internal evidence dates it 2013 or later. Kind should be technical_report, not study. Not peer-reviewed. URL: nlsinfo.org, NLSY79 Appendix 24. |
S0013, S0080 | Chase & Simon DOI 10.1016/0010-0285(73)90004-2. N = 3. Add a cross-reference to Gobet & Simon 1996 (10.3758/BF03212414), which supersedes the random-position null Hirsch relies on. |
S0021, S0070 | DOIs 10.1016/S0022-5371(72)80006-9 and 10.1016/0010-0285(72)90003-5. Author order is Bransford, Barclay & Franks. |
S0033 | Kintsch 1988 DOI 10.1037/0033-295X.95.2.163. resolutions.json resolved kintsch_comprehension to 10.1353/lan.1986.0098, which is a book review of van Dijk & Kintsch, not the work. |
S1348 | DOI 10.2307/747429. Free full text of the identical study: CSR Technical Report No. 97, ideals.illinois.edu/items/18083. Corpus misspells the second author as "C. Joag-Des" (correct: Joag-Dev; the 1978 report spells it Jogdeo). |
S0060 | Anderson et al. 1977 DOI 10.3102/00028312014004367; free precursor Technical Report No. 12 at ideals.illinois.edu/items/17896. |
S0005 | DOI 10.26300/nsbq-hb21. Not peer-reviewed — add a field for this; it currently reads as a study like any other. |
S0366 | DOI 10.1037/0022-0663.81.3.306. |
S0322 | DOI 10.1002/rrq.481; open access; correct year 2023 (registry has 2022). |
S0161 | Kim et al. 2023 DOI 10.1037/edu0000751. |
| (new) | Cabell et al. 2025, DOI 10.1037/edu0000916 — postdates the corpus; needs a registry entry to be citable from Dossier 2. |
C. Independence. The independence flag is per-source and cannot see chains. Willingham is primary_independent as S0020, a co-author of ckf_affiliated S0005, and a CKF trustee. Suggest a related_independence_note field, or an explicit person-level affiliation table.
D. One warrant to regenerate. how-to-educate-a-citizen:ch5_W1 — the warrant under the chess analogy, the single most load-bearing inference in the backbone — is still flagged stale: evidence_id_is_a_claim_id. It was not repaired by the regeneration pass.
14. Time estimates
Reading time. ~19,000 words. Read straight through: 75–95 minutes. Read as designed — §1–§5 in full, §6 ledger, §7 cards for the two or three sources you actually care about, §10.2 and §11: 25–30 minutes. §7 is deliberately reference material, not a chapter; skipping ten of the twelve cards on a first pass is the intended use.
Editor time to a verdict you would preserve and cite.
| Task | Estimate |
|---|---|
| Read this draft and mark disagreements | 1 h |
| Read the P0 acquisitions (Recht & Leslie; Reynolds 2025; Smith et al. 2021; Willingham & Lovette; Grissmer) and correct the cards | 4–5 h |
| Read the P1 acquisitions (Kintsch 1988; McNamara et al. 1996; O'Reilly & McNamara; Cabell 2025; NRP ch. 4-II; Rosenshine & Meister; the AFQT memo) | 5–6 h |
Read cl:2, kd:2, sk:12, wkm:1, ae:7, htec:5 against the backbone | 2–3 h |
| Adjudicate C1–C5 and write §10.1 | 2 h |
| Minimum viable (P0 + corpus chapters + assessment) | ≈ 8–9 h |
| Publishable (P0 + P1) | ≈ 14–17 h |
The P0 set alone is enough for a defensible verdict on T1, T2 and T3. Everything after that sharpens the crux statements rather than changing the verdict.
What is not included: applying any of §13.3 to the shared registry, which is a separate decision and a separate session.
15. What this dossier has to beat
essay-control.md in this directory is the markdown-essay control condition (plan amendment 21.1 item 5), and baseline-questions.md is the protocol. The structured dossier earns its complexity only if, on that protocol, it beats the essay on:
- recalling the decisive evidence and its limits without reopening the page;
- naming the crux;
- keeping T1, T2 and T3 apart;
- being materially easier to revisit and update when Reynolds 2025 or the Cabell K–2 results are actually read.
A tie is a fail. On (4) specifically, the test is concrete: when the editor reads Recht & Leslie's primary text, there is exactly one place the new information goes — Card 1's effect line and the T1 ④ read — and it should take under a minute to find. If it does not, the structure is decoration.
Draft generated 2026-08-18. Corpus: data/corpus_consolidated.json (10,312 claims, 2,885 evidence items, 4,327 warrants, 892 objections). Thesis layer: data/theses/ (25 canonical propositions, 10,312 mappings, uncalibrated). Registry: data/evidence_registry/ (1,515 sources). Quote verification: verify_quotes.py → 38/38.