Evidence-source registry, machine-normalized 2026-08-24. Identity and independence flag are unreviewed; occurrence→claim links were machine-relinked. ← Back to the registry
S0056 · godshalk_1966_writing_measurement · registry identity

The Measurement of Writing Ability

Godshalk · 1966
study independent8 occurrences across 2 books · load 16

Resolution status: resolved

Roundtable Review: The Measurement of Writing Ability, by F. I. Godshalk, Frances Swineford, and W. E. Coffman
Becker, Steinmann, Godshalk · 1967 · Research in the Teaching of English · journal-article

Every occurrence, grouped by book and chapter, with the claim it was linked to support and (where the extraction regenerated one) the underlying warrant.

1977 · The Philosophy of Composition
Report indicating that student performance varies significantly based on the topic assigned.
direct Students show more variation in the quality of their ideas and aims across different topics than they do in their quality of presentation.
ch 6 · statistical_data · poc:ch6_E20
Warrant (implicit): Observed fluctuations in student test scores across different subject matter prompts reflect underlying differences in cognitive mastery of content rather than fluctuations in the student's mastery of linguistic mechanics.
could fail ifInterest-based motivation effects, where a student's engagement with a specific topic improves their attention to presentation and syntax beyond their baseline ability.
Reports of performance variation in students across different topics.
direct Students show more variation in the quality of their ideas and aims across different topics than they do in their quality of presentation.
ch 6 · empirical_study · poc:ch6_E21
Warrant (implicit): Large-scale statistical patterns of performance variance across different writing prompts allow for the analytical separation of 'content' skills from 'formal' writing skills.
could fail ifThe interdependence of thought and language, where complex ideas may inherently require more complex (and thus more error-prone) syntactic structures.
1996 · The Schools We Need
College Board research finding that essay grades depended more on the year or reader than the student's writing, prompting the shift to multiple-choice.
direct Performance-based test scores often depend more on the specific grader or the year of the exam than on the quality of the student's work.
ch 6 · historical_example · swn:ch6_E14
Warrant (implicit): Any assessment system where the identity of the judge or the timing of the test accounts for more variance than the student's output fails the basic requirement of procedural justice.
could fail ifHistorical findings from the College Board's early essay tests may not apply to modern performance assessments that use more sophisticated rubrics and digital moderation.
A study of 1,300 students across 15 states comparing five essays scored by five different graders with multiple-choice tests.
direct A multiple-choice test is a more valid and reliable measure of writing ability than a performance-based test of equal length.
ch 6 · statistical_data · swn:ch6_E18
Warrant (implicit): Higher correlations between a specific testing format and a comprehensive multi-sample criterion indicate that the format is a more accurate proxy for the underlying skill being measured.
could fail ifThe multiple-choice test might only be measuring the specific sub-skills (like grammar or vocabulary) that overlap with the criterion, while failing to measure the generative aspects of writing.
Correlation data showing that multiple-choice sections correlated at .775 with student writing ability, compared to .640 for essays read three times.
direct Multiple-choice sections achieve higher accuracy and fairness at a lower cost than multiple essay readings.
ch 6 · statistical_data · swn:ch6_E19
Warrant (implicit): Statistical correlation coefficients against a gold-standard criterion are the primary measure of a test's fairness and accuracy.
could fail ifStatistical reliability does not equal construct validity; a test can consistently measure the wrong thing more reliably than it measures the right thing.
A validity study showing that a combination of multiple-choice and writing samples achieved a .784 coefficient against the criterion, compared to .775 for the objective test alone.
direct A test consisting of multiple-choice segments combined with a writing sample yields higher validity than an objective test alone.
ch 6 · empirical_study · swn:ch6_E20
Warrant (implicit): Combining different assessment methods that address various facets of a skill will marginally increase the predictive validity of the total score compared to using a single method.
could fail ifThe marginal increase in validity (.009) might be statistically insignificant or not cost-justified given the added complexity of grading writing samples.
Examples of three types of multiple-choice items (Usage, Sentence Correction, Construction) that were proven to be highly informative in sampling writing ability.
direct Objective English questions do not focus exclusively on superficial aspects of writing ability.
ch 6 · empirical_study · swn:ch6_E22
Warrant (implicit): Specific question types that effectively discriminate between high and low performers in a complex skill area must be tapping into non-superficial components of that skill.
could fail ifThe discriminative power of the questions could stem from correlations with socioeconomic status or general test-taking strategy rather than deep linguistic competence.
Godshalk, Swineford, and Coffman's data showing objective English questions effectively discriminate ability levels.
direct Objective English questions do not focus exclusively on superficial aspects of writing ability.
ch 6 · empirical_study · swn:ch6_E23
Warrant (implicit): The failure to find empirical evidence supporting the hypothesis that objective questions are superficial justifies the rejection of that hypothesis.
could fail ifThe research instruments used in 1966 might not have been sensitive enough to detect the specific nuances of 'depth' that modern critics are concerned with.

Raw book:chapter:endnote references collected during endnote backfill.

poc:6:10swn:6:11swn:6:15swn:6:16swn:6:17swn:6:19

As they appeared before normalization/consolidation into this source identity.

  • F. I. Godshalk et al., The Measurement of Writing Ability
  • Godshalk, F. I. et al., The Measurement of Writing Ability
  • Godshalk, Swineford, and Coffman
  • Godshalk, Swineford, and Coffman, 'The Measurement of Writing Ability'
  • Godshalk, Swineford, and Coffman, The Measurement of Writing Ability
  • Godshalk, Swineford, and Coffman, The Measurement of Writing Ability (1966)