Three layers, three states of review. The proposition list is machine-drafted (2026-08-18) and under human review. The mappings and evidence joins beneath each proposition are unreviewed. The accounts are agent-written from bounded packets; each proposition page shows its own audit status.
P14 · methodological · empirical line

Standardised reading tests that purport to measure a general comprehension skill are invalid instruments for guiding instruction or judging schools; valid and fair assessment must be curriculum-based, testing what was actually taught.

auxiliaryunder T5: The measuring instruments of skills-based schooling — general reading tests, readability levels — are invalid, which is why the failure they are supposed to detect stays invisible.
Used in Dossier 1

This proposition is made concrete in the argument map at A5. The dossier supplies the source appraisals, inference audit and objections.

methodological positionThe account · what the corpus lets you say

The proposition is narrower than it sounds: Hirsch repeatedly concedes that the better standardised reading tests are valid measures of general reading ability and defends them against performance assessment; his real claim is that they are unfair and unproductive as guides to instruction and early-grade progress because they are knowledge tests in disguise. The knowledge-loading is adjudicated in the dossier as plausible with a circularity risk and no stable estimate of its size; the causal 'consequential invalidity' story rests on journalism and anecdote; and the prescription that tests must be curriculum-based is never evaluated anywhere in the packet.

Why it matters: If reading tests cannot separate school-taught from out-of-school knowledge, then attributing flat or contrary test results to a curriculum or a teacher requires care, and the case for a specified curriculum acquires a measurement rationale (thesis T5).

The statement, in its strongest form

Standardised reading tests built on the premise that they measure a general, content-independent comprehension skill are, whatever their reliability as measures of large-group reading ability, unsuitable for guiding instruction, for measuring yearly progress in the early grades, or for attributing value to individual teachers, because what they mainly register on Hirsch's account is passage-relevant knowledge, and they cannot separate the knowledge a school taught from knowledge acquired elsewhere. The fair and productive test, at least in the early grades, is one whose passages draw on knowledge the curriculum actually assigned.

Scope: A methodological claim about the validity of a test for a purpose (instruction, accountability, teacher evaluation, early-grade progress), not about whether reading tests measure anything at all. Hirsch explicitly defends objective standardised testing against performance assessment and concedes that the better reading tests are reliable and valid measures of average reading ability; the attack is on skills-framed reading tests used as guides to schooling. The early-grade scope is present from 2006 (yearly progress in early schooling), made concrete as grades one to four in 2010, and restated in 2024.

What the corpus establishes

Hirsch's target is skills-framed reading tests, not standardised or objective testing as such: he defends multiple-choice standardised tests as more reliable and fairer than performance-based assessment for large-scale use.
Stated as a main conclusion in The Schools We Need and maintained in the objections he answers in 1996 and 2006. The only independent study in the packet bearing on this half of the argument is Koretz's analysis of the Vermont portfolio programme, which found scoring too unreliable for most intended uses.
Hirsch concedes, in every decade from 1996 to 2016, that the better standardised reading tests (Gates-MacGinitie, ITBS, NAEP) are reliable and valid measures of the average reading ability of large groups and correlate with real-world reading. 'Invalid' in the proposition therefore means invalid for guiding instruction, measuring early-grade yearly progress, and attributing teacher value-added, not invalid as tests.
Eight concessions in the packet, all in Hirsch's own voice, plus his explicit denial in 2006 that the critique is an attack on tests like the ITBS. The narrowing to 'yearly progress in the early grades' is a main conclusion in The Knowledge Deficit.
The reviewed dossier regards passage-relevant knowledge as contributing materially to what reading-comprehension tests measure, with a circularity risk and no stable estimate of how dominant that contribution is (D1:A5, plausible_with_circularity_risk).
In this packet the node is carried by Willingham's commentary that a reading test is a knowledge test, Hirsch's reading of PARCC grade-5 practice items, a Norwegian parliamentary statement to the same effect, and Hirsch's own Civil War hypothetical, which illustrates school-taught knowledge improving comprehension of related passages even when the taught words do not appear on the test: commentary, item reading, testimony and illustration, not a variance study.
Standardised reading tests differ from one another in what they demand: individual-level tests intercorrelate at only about .70, and they vary substantially in how much they depend on decoding versus oral comprehension.
Two independent studies cited in Why Knowledge Matters (Morsy 2010; Keenan 2008), known here only through registry summaries. Incomplete intertest agreement and differing test demands can coexist with a shared general component, which Hirsch concedes, and the packet supplies no construct-validation or dimensionality analysis.
S0345 S1067 wkm:ch1_C36
In one urban district students scored significantly lower on an unfamiliar standardised test than on their usual one, and journalistic and trade-press accounts describe history, science and the arts being displaced by strategy drill under test pressure.
Koretz 1991 as summarised in the registry (the summary calls it an experiment but gives no design details); the narrowing picture comes from the Perlstein registry records, Walker's 2014 article and reported teacher testimony. These describe narrowing alongside testing, not test design rather than school practice causing it, and the two Perlstein records carry different dates and titles for what may be one book.
Un-normed, ad hoc content tests can reach a reliability coefficient of roughly .8, so a curriculum-based test is not ruled out on reliability grounds.
Sireci's validity-theory work as cited in Why Knowledge Matters. This is the only item in the packet that bears on the feasibility of the prescriptive half of the proposition, and it speaks to reliability only, not to fairness or to what such a test predicts.
S0370

Asserted without evidence reaching it

Valid and fair assessment must be curriculum-based; there is 'no other way' of making tests fair and productive.
A prescription stated as a necessity. The packet contains no evaluation of a curriculum-based reading test in use (no curriculum-based-measurement literature, no trial in grades one to four), only Sireci's reliability figure. Hirsch also holds that a diversity of passages on quite different subjects is 'absolutely critical' to a reading test's validity and reliability; whether a test confined to assigned curriculum content can keep that diversity while being fair and generalising is a question the packet raises but does not settle.
Skills-framed reading tests are consequentially invalid: they cause schools to engage in self-defeating practices.
The causal claim is supported by description and anecdote (Perlstein, Walker, teacher testimony) and by Hirsch's own characterisation of NCLB's failure, with no study isolating test design from school implementation. Hirsch himself is ambivalent about where the fault lies: in 1996 and 2006 he answers the teaching-to-the-test objection by blaming mindless application and poor instruction rather than the tests, which is the opposite attribution from the 2016 consequential-validity argument.
The tests are 'unwittingly unfair' to disadvantaged children because they reward knowledge acquired outside school.
In this packet the fairness claim is carried by Hirsch's Appalachian Trail contrast between an advantaged and a disadvantaged fourth-grader and by his assertions in 2022 and 2024, not by data. The evidence the drafter says would count (score variance by passage topic) is not in the packet. The step from 'knowledge-loaded' to 'unfair' needs the premise that the loaded knowledge is disproportionately acquired outside school, which Hirsch argues in reply to value-added proponents rather than shows.
State 'criterion-referenced' reading tests have empty criteria ('find the main idea') that could be applied to any state's test.
Rests on Hirsch's own comparison of standards and tests across Texas, New York, Florida and Michigan, presented as analysis rather than as a documented study; no independent comparison of state test content is in the packet.
Current reading tests cannot reliably or validly gauge the value a teacher has added in a year.
The pipeline flags the warrant as missing (that accountability instruments must isolate what is within the teacher's control), and the packet holds no value-added-modelling study or analysis of within-year score variance; the claim is inferred from the knowledge-loading argument rather than evidenced.
Standardised reading tests are culturally biased.
Hirsch's position is inconsistent within the packet: in 2006 he calls the cultural-bias charge 'certainly true'; in 2010 he answers the same charge as unfounded because the tests are accurate indexes of real-world reading with standardised scoring. The account cannot resolve which he holds; both are asserted without evidence here.
Evidence profile

Only 17% of asserting occurrences carry direct evidence, and 14 of the 15 'same' occurrences carry none. The registry lists 13 entries as primary-independent, and their summaries are mostly adjacent to the proposition: Koretz on portfolio unreliability and test familiarity, Morsy and Keenan on test heterogeneity, Sireci on content-test reliability, Perlstein and Walker on curriculum narrowing, NAEP and SAT trend data; underlying identities and methods were not verified from the packet, and the two Perlstein records (S0254, S1246) carry different dates and titles for what may be one book. The core 'knowledge test in disguise' claim is carried by Willingham's commentary, Hirsch's own PARCC item reading and a Norwegian parliamentary statement; no Core Knowledge Foundation report appears. The independence flags look mostly right, with the caveat that PARCC items (S0350) and Willingham (S0020) are independent sources on which the inferential work is Hirsch's.

Pivotal sourceIndependenceWhat it showsLimits
S0333 Koretz 1994, Vermont Portfolio Assessment ProgramindependentScoring of large-scale portfolio assessment was too unreliable for most intended uses of the scores.Bears on Hirsch's defence of objective tests over performance assessment, not on whether skills-framed standardised tests are invalid for guiding instruction.
S1081 Koretz 1991, effects of high-stakes testingindependentAs summarised in the registry: students in one urban district, given an unfamiliar standardised test alongside their usual one, scored significantly lower on the unfamiliar test.The summary gives no design details; the familiarity reading is Hirsch's use of it. It does not distinguish skills-framed from curriculum-based tests, and whether a curriculum-based test that schools drilled for would show the same pattern is a question the packet does not answer.
S0345 Morsy 2010, Measure for Measure
S0345
independentIndividual-level reading comprehension tests intercorrelate at only about .70.Establishes that tests differ, not what the difference is; .70 is also compatible with a substantial shared general component, which Hirsch elsewhere concedes. Its anchor occurrence is outside the citable set.
S1067 Keenan 2008, comprehension tests and decoding dependence
S1067
independentReading comprehension tests vary significantly in how much they depend on decoding versus oral comprehension and knowledge.The heterogeneity is about decoding load; it does not directly measure passage-knowledge load. Its anchor occurrence is outside the citable set, so the registry entry is the only anchor.
S0350 PARCC 2015, grade-5 practice test itemsindependentItems framed as strategy questions (find the clue for a word's meaning) which Hirsch reads as answerable only from passage-relevant knowledge.The items are independent; the reading of them is Hirsch's, and the packet holds no demonstration that answers depend only on passage knowledge. wkm:ch6_C23 is a timeline anchor with no passage in the packet.
S0020 Willingham 2014, reading-comprehension commentaryindependentA cognitive psychologist's principle, quoted in three books, that a reading test is a knowledge test in disguise.Commentary, not a study; it lends authority to the knowledge-loading claim rather than measuring it.
S0370 Sireci 2007, validity theory and test validationindependentValidity-theory research and an empirical analysis showing un-normed, ad hoc content tests reach a reliability of roughly .8.The only evidence in the packet on the feasibility of curriculum-based testing, and it addresses reliability only. The packet does not identify where Hirsch's 'consequential validity' vocabulary comes from.
S0254 Perlstein, Tested (registry records S0254 dated 2003 and S1246 dated 2007)independentRegistry summaries of journalistic documentation of classrooms where comprehension-strategy drill on thin texts displaced history, science and the arts under test pressure.Descriptive case reporting from a few schools, held as two registry records that were not verified as distinct studies; cannot attribute the narrowing to test design rather than to district or school choices.

Absent from the corpus: Absent from the packet are the three kinds of evidence the drafter named: a study partitioning reading-test score variance by passage topic familiarity, the curriculum-based-measurement literature on what such tests predict, and NCLB-era studies that identify curriculum narrowing causally. Also absent is any evaluation of a curriculum-based reading test actually administered in grades one to four.

Strongest challenge

From okkinga2018-p1

Claim Okkinga et al.'s meta-analysis of whole-class reading-strategy instruction (52 studies, 125 effect sizes) finds a very small effect on standardised reading-comprehension tests (d = .186) and a small effect on researcher-developed tests (d = .431).

Hirsch’s response not recorded

Assessment The small but positive standardised-test effect bounds Hirsch's categorical claim that these tests do not test comprehension strategies, though it does not show that they measure a general strategy skill. The gap between the two effect sizes is unexplained in the packet: alignment of researcher-developed tests to what was taught, score inflation of the kind Koretz 1991 reports, and limited transfer are candidate explanations, not findings. If alignment is the explanation, the same question falls on the curriculum-based tests Hirsch prescribes, which the packet never evaluates. Within Hirsch's own text the sharpest pressure is his repeated concession that the better tests are valid measures of average reading ability, which forces 'invalid' down to 'invalid for instructional guidance and early-grade progress'.

Evolution

First stated: 2006 · KD · kd:ch6_C21

The drafter dates the first statement to The Schools We Need (1996) but the first mapped 'same' occurrence is The Knowledge Deficit (2006); the packet supports the later date, with 1987 (educational formalism in writing assessment) and 1996 (defending standardised tests against performance assessment while conceding their validity) supplying precursors rather than the claim itself. The 2006 statement ('unwittingly unfair', empty criteria, shortcomings for early-grade yearly progress) is the claim proper; 2010 adds the curriculum-narrowing framing and the concrete proposal of curriculum-based tests for grades one to four. 2016 recasts it in validity-theory vocabulary (consequential validity), extends it to teacher value-added, and adds the PARCC item analysis; 2024 restates the early-grade prescription, and the selected passages show no new restriction or retraction.

1987CL
Precursor: objection to testing writing ability independent of subject knowledge, and to reading texts screened by readability and skill sequence rather than content.
1996SWN
Hirsch positions himself as a defender of standardised tests against performance assessment, citing Koretz on Vermont, and concedes that even the worst standardised reading tests correlate with real-world reading.
2006KD
First explicit statement: state reading tests are unwittingly unfair because they measure domain knowledge; 'criterion-referenced' criteria are empty; the critique is limited to yearly progress in early schooling, not to tests like the ITBS.
2010THE-
The complaint is reframed as curriculum narrowing rather than the tests themselves, and the positive proposal appears: curriculum-based reading tests in grades one to four.
2016WKM
Validity-theory vocabulary ('consequentially invalid'), extension to teacher value-added evaluation, PARCC item analysis (wkm:ch6_C23 is a timeline anchor without a passage in the packet), and the flat formulation that only curriculum-based tests can be fair and productive.
2024RE
Restatement for the early grades: the only fair and productive tests there are those based on background knowledge assigned in the curriculum; Willingham's dictum is paired with Tomasello and the 'Norway Principle'. The scope matches 2006 and 2010 rather than narrowing them.
Dependencies
Requires P15
Feeds none
Dossier nodes D1:A5
What would change it
  • A study partitioning standardised reading-test score variance by passage-topic familiarity versus general skill (the drafter's own criterion) would either establish the knowledge-loading claim with a magnitude or bound it; D1:A5 already flags that no stable dominance estimate exists.
  • An evaluation of a curriculum-based reading test administered in grades one to four, reporting reliability, predictive validity for later general reading, passage diversity, and effects on the advantaged-disadvantaged gap, would move the prescriptive half from stated necessity to tested proposal.
  • A study that separates test-design effects from school-implementation effects on curriculum narrowing would settle Hirsch's own unresolved attribution (1996 and 2006 blame application; 2016 blames the tests).
  • Evidence on whether gains registered by curriculum-aligned tests transfer to unfamiliar passages would show whether the Okkinga gap (d = .186 versus d = .431) reflects alignment, inflation or limited transfer, and whether a curriculum-based test can be fair and valid at once.
45 of 142 mapped occurrences were in the packet1 critic positions in the packet
Gaps: Passages are excerpts and 97 of 142 mapped occurrences are not shown; the registry's evidence items anchor mostly to occurrences outside the citable set (Morsy, Keenan, Sireci, Koretz 1991, Perlstein, Willingham), so source-to-claim links were not verified, and the two Perlstein records (S0254, S1246) look like one book under two dates and should go to the registry pass. The single critic position (Okkinga) targets P14 only obliquely; no critic of curriculum-based testing and no defender of the general-skill construct is recorded. The passage-topic variance evidence lives in other propositions' packets, so knowledge-loading is credited only via D1:A5, and the knowledge-to-comprehension proposition P14 presupposes is not in the packet's proposition list, so 'requires' is incomplete. The drafter and mapper disagree on first statement (1996 vs 2006), and the self-critique flag on wkm:ch1_C99 refers to an occurrence not in the citable set.
Account written by claude-fable-5.1 on 2026-09-13 (v0.2); audit: pass with fixes by codex gpt-6 on 2026-09-13
142occurrences
15stated as such
9books
22direct evidence
19sources
13independent

77879606101620222324

First appears 1987; first asserted as the proposition itself 2006; restated as such in 2006, 2016.

Drafted first_book was swn; the mapping's earliest asserting book is kd. Unresolved.

Scope: Note that Hirsch defends objective testing per se (against the progressive critique of standardised tests) while attacking skills-based reading tests specifically — do not collapse the two. Sharpened after NCLB into a consequential-validity argument (the test narrows the curriculum it measures).

What would count as evidence: Score variance by passage topic; curriculum-based measurement literature; NCLB curriculum-narrowing studies.

First stated: 2006 · The Knowledge Deficit (ch 6) — same

State reading tests are unwittingly unfair because they primarily measure domain knowledge rather than the formal comprehension skills they claim to test.
“But because the tests have been presented as tests of formal comprehension skills, they are unwittingly unfair, because these skills are not what they are really testing. The tests favor children who happen to have domain knowledge relevant to the passages in the test.”

Machine-classified from the book-by-book counts, quotes, sources and objections. One Gemini 3 Flash pass, unreviewed; the model was told not to read development into repetition, and whether it obeyed is exactly what a reviewer should check.

2006KD
first statedHirsch first explicitly argues that state reading tests are 'unwittingly unfair' because they measure domain knowledge while claiming to measure formal comprehension skills.
kd:ch6_C21 · confidence 0.95
2006KD
answers objectionHe rebuts the claim that state tests are 'criterion-referenced,' arguing that abstract skills like 'finding the main idea' are empty criteria that do not reflect a knowledge-based curriculum.
kd:ch6_OBJ6 · kd:ch6_OBJ7 · confidence 0.90
2010MoA
reframing · scopeThe argument shifts toward curriculum narrowing, asserting that the problem is not the tests themselves but their failure to be based on specific grade-level knowledge.
2016WKM
restatedHirsch provides a sharp, concise formulation of his position: only curriculum-based reading tests can be fair and productive.
wkm:ch6_C25 · confidence 0.95
2016WKM
new evidenceHirsch introduces formal validity theory and analysis of specific test items (PARCC) to demonstrate that reading tests are 'knowledge tests in disguise.'
source: Sireci — Validity Theory and Test Validation (S0370) · wkm:ch6_C23 · wkm:ch1_OBJ6 · confidence 0.90
2024RE
narrowingThe proposition is specified for the early grades, where tests must be based exclusively on the assigned curriculum to be fair.
re:ch6_C63 · confidence 0.80
technical unreliability -> lack of curriculum fit -> consequential validity/equity
Hirsch's justification moves from a critique of 'educational formalism' and performance-based assessments toward a psychometric argument regarding 'consequential validity.' While early books emphasize the technical unreliability of portfolios, his later work argues that skills-based tests are inherently unfair to low-SES students because they reward outside-of-school knowledge rather than what was actually taught.
1987-1996 — Educational formalism and performance-based assessments are theoretically flawed and technically unreliable. cl:ch5_C52 swn:ch6_C40
2006-2010 — Tests are unfair because their 'empty' criteria do not match the domain-specific nature of reading comprehension. kd:ch6_C21 kd:ch6_OBJ7 the-making-of-americans:ch6_C45
2016-2024 — Testing 'strategies' rather than taught content creates an equity gap and makes teacher evaluation (VAMs) impossible. wkm:ch2_C27 ae:ch1_C53 re:ch6_C39
confidence 0.90

See this proposition in the research-programme view →

Up to three occurrences per book that the mapping marked as restating the proposition. Unreviewed.

2006 · The Knowledge Deficit
State reading tests are unwittingly unfair because they primarily measure domain knowledge rather than the formal comprehension skills they claim to test.
“But because the tests have been presented as tests of formal comprehension skills, they are unwittingly unfair, because these skills are not what they are really testing. The tests favor children who happen to have domain knowledge relevant to the passages in the test.”
ch 6 · kd:ch6_C21
State reading tests are inadequate for guiding schooling because they are conceived as tests of empty processes rather than knowledge-based comprehension.
“Even if these tests were valid and reliable (an issue somewhat in doubt), they would still be inadequate when conceived as criterion-referenced tests that could productively guide schooling.”
ch 6 · kd:ch6_C32
Standard reading tests fail to positively influence instruction because they are unrelated to any specific content curriculum.
“Their two most damaging flaws are, first, that they do not positively influence instruction, since they are unrelated to any content curriculum”
ch 6 · kd:ch6_C51
2016 · Why Knowledge Matters
Only curriculum-based reading tests can be fair and educationally productive.
“If the test is not curriculum based, it will not be productive.”
ch 6, pp. 115-118 · wkm:ch6_C25
The only way to make standardized tests fair and productive is to base them on well-defined, knowledge-based curriculums.
“There’s just one way to do that—to base them on well-defined, knowledge-based curriculums. There is no other way of making tests fair and productive.”
ch 1, pp. 25-28 · wkm:ch1_C1
Current standardized reading tests are consequentially invalid because they lead schools to engage in self-defeating educational practices.
“to engage in self-defeating practices. They are consequentially invalid.”
ch 1, pp. 32-35 · wkm:ch1_C37

Chain: this proposition → its asserting occurrences → their direct evidence items → the normalised source registry. The independence flag is the registry's, not a judgement of study quality.

SourceYearKindIndependenceItemsBooks
Koretz — The Vermont Portfolio Assessment Program1994studyindependent2swn
PARCC — PARCC Practice Test Question Grade 52015otherindependent2wkm
Sireci — Validity Theory and Test Validation2007studyindependent2wkm
College Board — SAT Verbal Score Trends?datasetindependent1re
NAEP — NAEP Long-Term Trend assessment?datasetindependent1ae
Willingham and collaborators — Willingham's Reading-Comprehension Commentary and Reviews2014commentaryindependent1ae
Hirsch — Civil War passage background knowledge analysis?otherHirsch's own1kd
Hirsch — Test reliability and passage diversity analysis?otherHirsch's own1the-making-of-americans
Hirsch — Appalachian Trail Reading Comparison?anecdoteHirsch's own1the-making-of-americans
Perlstein — Tested: One American School in the Age of Accountability2003studyindependent1kd
Morsy — Measure for Measure: Reading Comprehension Assessments2010studyindependent1wkm
Hirsch — Superintendent multiple-choice test abuse?anecdoteHirsch's own1swn
Hirsch — Analysis of State Education Standards Interchangeability2006otherHirsch's own1kd
Keenan — Reading comprehension tests and decoding dependence2008studyindependent1wkm
Koretz — Effects of High-Stakes Testing on Achievement1991studyindependent1swn

… and 4 further sources.

Example, with its regenerated warrant

Evidence: Daniel Koretz's analysis of Vermont's portfolio assessment program showing fatal unreliability in scoring.

For the claim: The unreliability of scoring in large-scale performance-based systems (like Vermont's) renders them useless for most intended educational purposes.

Warrant: A large-scale assessment program cannot fulfill its educational mission if its scoring system fails to meet standard statistical thresholds for inter-rater reliability.

Vulnerability A program might still provide valuable qualitative feedback to teachers and students for formative purposes even if the scores are too unreliable for 'high-stakes' accountability.

14 of 15 occurrences that restate the proposition carry no direct evidence of their own.

Occurrences the mapping marked as contradicting the proposition, or as positions Hirsch concedes. Some are views he reports in order to reject them — check the passage.

Even the lowest-quality standardized reading tests correlate well with high-quality ones and real-world reading abilities.
1996 · ch 6 · contradicts / concedes · swn:ch6_C102
Standardized tests show a consistent positive correlation with real academic competencies.
1996 · ch 1 · contradicts / concedes · swn:ch1_C18
The better a person reads, the higher they tend to score on standardized reading tests.
1996 · ch 1 · contradicts / concedes · swn:ch1_C19
Standardized reading tests are accurate measures of general reading proficiency provided they include a diversity of subjects.
2006 · ch 2 · contradicts / concedes · kd:ch2_C73
Standardized reading tests must include a diversity of subject matters to maintain validity and reliability.
2006 · ch 2 · related / concedes · kd:ch2_C74
The author concedes that the argument that standardized reading tests are culturally biased is certainly true, even though test-makers attempt to achieve knowledge-neutrality.
2006 · ch 6 · supports / concedes · kd:ch6_C73
Standardized reading tests are fast and accurate indexes of real-world reading ability.
2010 · ch 6 · contradicts / concedes · the-making-of-americans:ch6_C10
Well-established reading tests like the Gates-MacGinitie and NAEP are technically reliable and valid measures of average reading ability.
2016 · ch 1 · related / concedes · wkm:ch1_C36
1996 Specialists/Proponents of 'Authentic Assessment'

Performance-based assessments are superior because they are more 'authentic' and motivating for students.

His reply: They do not authentically duplicate real-world performance and are much less fair and reliable for high-stakes testing than objective tests.

1996 Progressive educators (implied)

Performance-based tests (portfolios/essays) are necessary to overcome the unfairness and inaccuracy of multiple-choice tests.

His reply: Empirical data shows that performance-based tests are significantly less reliable and less fair unless they are conducted at an impossibly high financial cost; multiple-choice tests are actually more accurate indicators of average ability.

1996 unattributed

Tests debase educational achievement by fragmenting instruction and encouraging 'teaching to the test.'

His reply: This is a significant abuse, but the fault lies in the mindless application of tests and the failure to set precise goals, not in testing itself.

2006 Implicit critics of Hirsch's position

The author's critique of standardized reading tests is an attack on the tests themselves or their necessity in education.

His reply: It is not an attack on tests like the ITBS, which are reliable for measuring general ability; rather, the critique is that they are inappropriate for measuring yearly progress in early schooling.

2006 Unattributed (Many of the complaints)

Standardized reading tests exert a harmful influence through intensive test preparation.

His reply: The prep as currently conducted is indeed harmful, but this is due to the school's failure to use effective reading instruction methods, not the tests themselves.

2006 General technical jargon/educational establishment

Criterion-referenced tests are inherently fair because schools can 'teach to the criterion' and students can study for it.

His reply: In reading, the criteria are 'empty' (e.g., 'find the main idea'), making the term misleading because they don't reflect a specific knowledge-based curriculum.

2006 State education departments (implied)

State reading tests are 'criterion-referenced' and therefore meaningful assessments of the local curriculum.

His reply: This is misleading because the 'criteria' are abstract, empty processes that do not reflect the knowledge-based nature of reading and could be applied to any state's test.

2010 Unattributed 'attacks'

Reading tests are culturally biased.

His reply: These complaints are unfounded because tests are accurate indexes of real-world reading ability and they use standardized scoring to ensure fairness.

… and 2 further objections.

Flags the extraction-critique pass raised against the very occurrences that restate this proposition. They bear on extraction quality, not on whether Hirsch is right.

level_correction · low wkm:ch1_C1

This is a prescription for how education *should* be structured to achieve fairness, which is a value-based goal.

granularity · medium wkm:ch1_C37

Both claims define current tests as 'consequentially invalid' for the same reasons.

granularity · medium wkm:ch1_C99

Both claims define current tests as 'consequentially invalid' for the same reasons.

missing_warrant · low wkm:ch2_C30

Accountability measures are only ethically and technically valid if they isolate variables within the agent's (teacher's) direct control; since reading tests reflect out-of-school knowledge, they fail this requirement.

Arguments the critique pass found in the text but that the extraction never captured, whose suggested claim maps here. Leads for a future pass, not occurrences.

Schools are in a Kafkaesque state where they follow all official pedagogical instructions (strategy instruction) yet remain unable to meet test requirements because the instructions are based on a false premise of what the tests measure.
“Operating under the how-to conception of reading, the schools find themselves in a Kafkaesque situation in relation to reading tests. Like Josef K in The Trial, teachers and students are dutifully trying to obey all the instructions by practicing the time-consuming strategy exercises demanded by the authorities, but they have still not managed to fulfill the mysterious requirements of the test.”
2010 · ch 6 · significant
The framing of standardized test questions (e.g., 'What is the main idea?') constitutes a deceptive communication that misleads educators into prioritizing strategy instruction over the knowledge required to actually answer the questions.
“No doubt unintentionally... the test makers are implying a lie. By the form of their questions they suggest that they are probing formal skills. But... the student with the smaller relevant vocabulary and knowledge is the one who will fare worse on the test.”
2016 · ch 1 · critical
The maintenance of reading tests as 'skill-based' is a necessary institutional pretense used to hide the unfairness of testing unknown content.
“Test makers must not acknowledge the disguise. They have to pretend that reading tests are about skills. Lifting the disguise will require devising knowledge-based curricula and reading tests based on that known school curriculum—the only kind of reading tests that can be fair and productive.”
2016 · ch 6 · significant