S0049 · diederich_1961_essay_grading · registry identity
Factors in Judgments of Writing Ability
External resolution
Resolution status: ambiguous
How Hirsch uses it
Every occurrence, grouped by book and chapter, with the claim it was linked to support and (where the extraction regenerated one) the underlying warrant.
1977 · The Philosophy of Composition
Paul Diederich’s experiment where 16 readers prioritized ideas while 9 prioritized wording and phrasing.
direct Most readers are 'Platonists' in their actual judgments, prioritizing the quality of ideas (extrinsic) over wording and phrasing (intrinsic).
Warrant (implicit): The distribution of priorities among a sample of expert readers reflects the general psychological tendencies of the reading population at large.
could fail ifThe preference for 'ideas' over 'wording' might be a selection effect of the specific professional backgrounds of the readers in Diederich's study.
Experiment using 300 college student essays graded by 60 distinguished readers across six professional fields (English teachers, social scientists, natural scientists, writers/editors, lawyers, executives).
direct There is currently no holistic agreement on what constitutes 'good writing' among English teachers or the general public.
Warrant (implicit): If a diverse group of experts across several intellectual fields cannot reach a consensus on the quality of specific writing samples, then no objective social standard for 'good writing' currently exists.
could fail ifThe experts might share a standard but apply it differently due to the lack of training in a specific, shared assessment framework for the experiment.
In the Diederich study, 101 out of 300 essays received every possible grade (1 to 9) from the pool of readers.
direct Readers disagree widely in their holistic judgments of writing quality.
Warrant (implicit): A high frequency of instances where a single item receives the full range of possible scores from different evaluators is a definitive indicator of wide disagreement in judgment.
could fail ifThe 'full range' of grades might result from a small number of outlier raters who misunderstood the scale, rather than a representative distribution of disagreement.
94 percent of essays in the Diederich experiment received seven, eight, or nine different grades from the fifty-three readers.
direct Readers disagree widely in their holistic judgments of writing quality.
Warrant (implicit): When nearly every sample in a large-scale experiment triggers a high number of distinct score levels from a pool of readers, the variability is systemic rather than incidental.
could fail ifReaders may be using the same criteria but interpreting the numerical scale differently (e.g., one reader's '7' is another reader's '9').
The highest correlation found within professional grading groups was .40.
direct English teachers exhibit the same low level of correlation (.40) in holistic writing judgment as non-academic professionals like lawyers and business executives.
Warrant (implicit): Specialized training in a field like English does not provide a more unified standard of judgment if those professionals show no higher correlation than non-specialists.
could fail ifThe specific professional groups chosen (like lawyers or executives) may have their own rigorous but different internal standards for clarity that are as consistent as those of English teachers.
An analysis of 3,357 papers and 11,018 comments showing that readers naturally fall into five clusters based on different primary criteria.
direct Readers of student writing naturally cluster into groups based on the predominant weight they give to specific criteria.
Warrant (implicit): Patterned correlations between specific textual comments and assigned grades allow for the categorization of readers into distinct evaluative schools of thought.
could fail ifThe identified 'clusters' might be artifacts of the 55 coding headings provided by the researchers rather than the natural cognitive categories of the readers.
Diederich's analytical weighting system for composition.
related To achieve reliability in an analytical approach, a uniform system of weighting must be imposed on the scoring categories.
1996 · The Schools We Need
Paul Diederich's study of 300 student papers graded by 53 different readers, showing massive grading variance.
direct Experienced teachers grading essays rarely achieve a correlation greater than .40 in their grading.
Warrant (implicit): Low statistical correlation between the judgments of professional evaluators indicates that the subjective nature of the assessment makes it an unreliable measure of student ability.
could fail ifLow correlation may result from a lack of specific, shared rubrics among the 53 graders rather than an inherent impossibility of grading essays objectively.
Paul Diederich study showing 300 papers graded by 53 graders resulted in 1/3 of papers receiving every possible grade from A to D.
direct Even experienced teachers demonstrate extremely low correlation in the grades they assign to the same student essays.
Warrant (implicit): When a large sample of the same work receives the full range of possible marks from different experts, the assessment tool lacks the consistency required for meaningful measurement.
could fail ifThe extreme variance could be an artifact of the specific Diederich study's design, such as intentionally selecting ambiguous papers or using graders from widely different professional backgrounds.
ETS internal findings on the instability of grader agreement and the necessity of constant recalibration by 'table leaders.'
direct The agreement achieved between graders in professional performance assessment sessions is arbitrary and unstable.
Warrant (explicit): If a consensus on grading standards requires constant external intervention and social pressure to maintain, that consensus is a manufactured artifact rather than a reflection of the work's quality.
could fail ifConstant recalibration by 'table leaders' could be viewed as a successful quality-control mechanism that ensures fairness, similar to how scientific instruments require regular calibration.
Endnote numbers seen
Raw book:chapter:endnote references collected during endnote backfill.
poc:6:2poc:6:3swn:6:10
Raw source_reference strings seen
As they appeared before normalization/consolidation into this source identity.
Diederich (implied context from previous sections)Diederich et al. (1961)Diederich in A. Jewett and C. Bush, eds., Improving English Composition (1965)Diederich, French, and Carlton, 'Factors in Judgments of Writing Ability' (1961)Diederich, P., Measuring Growth in English, NCTE (1974)P. Diederich, J. W. French, S. Carlton, 'Factors in Judgments of Writing Ability,' Educational Testing Service Research Bulletin (1961)