Introduction. Content analysis is the technique of making replicable, valid inferences from texts to their contexts; Krippendorff’s textbook gives its design logic (unitising, sampling, coding, reducing, inferring) and its treatment of reliability. Reliability among coders is the degree to which independent coders, working from the same instructions, assign the same categories to the same units. Percent agreement overstates it because coders agree by chance. Krippendorff’s alpha corrects for chance agreement, handles any number of coders, any metric (nominal, ordinal, interval, ratio), missing data and small samples, and is therefore the general-purpose coefficient; Cohen’s kappa is the two-coder nominal special case and intraclass correlation the interval case. Alpha is a point estimate and its sampling distribution is not simple, so confidence intervals are obtained by bootstrapping, and the width of the interval depends on the number of units coded. A reliability study is therefore designed backwards from the precision required: state the smallest alpha that would be acceptable, state the interval width needed to distinguish it from alternatives, and derive the sample. Reliability is necessary and not sufficient for validity: coders can agree reliably on a category that tracks nothing real.
Important authors. Klaus Krippendorff (1932–2022) was the Gregory Bateson Professor of Communication at the Annenberg School, University of Pennsylvania, a past president of the International Communication Association and of the American Society for Cybernetics; trained at Ulm and at Illinois in von Foerster’s circle, he was a second-order cybernetician as well as a methodologist. Andrew Hayes (Ohio State, later Calgary) co-authored the standard argument for alpha as the default coefficient and supplied its software. Jacob Cohen introduced kappa in 1960.
Importance for cybernetics and the VSM. VSM diagnosis is a coding task: an analyst assigns units of an organisation to Systems One to Five and to recursion levels. No reliability study of that assignment exists. Krippendorff’s authorship is a small irony: the coefficient the VSM tradition needs was developed by a cybernetician, and the tradition never used it. The distinction between reliability and validity is also the distinction between the instrumental and constitutive readings in §2.2: two trained analysts agreeing does not show the levels are there.
Importance for the article. §4.4 requires that a reliability receipt “names the reliability statistic (Krippendorff’s alpha with reported confidence intervals†), a target precision, and the sample size that precision requires.” §8.3 applies it to IIIa: alpha estimated separately for named units, functional relations and diagnostic predictions, “at a preregistered target precision that determines the sample; six to ten cases and three to five analysts will not reach it, and the receipt says so.” The import is from C1 §5 (“you also don’t name a reliability statistic”). IIIa rival (3), analysts disagreeing on boundaries but converging on predicted vulnerabilities, requires reliability to be estimated on predictions as well as partitions; that is the informative endpoint. Reviewer pressure: what alpha threshold counts as “stated reliability” and why; whether units, relations and predictions are themselves codable; and the reliability–validity gap, which is why IIIb and IIIc exist. Krippendorff (2004) is Tier C and † in v02; the 2018 fourth edition is current.
Sources in the reading list.
- the reliability chapters: alpha’s definition, its confidence intervals, and the argument that reliability does not establish validity
Other important sources and authors.
- Hayes, A. F., & Krippendorff, K. (2007). Answering the call for a standard reliability measure for coding data. Communication Methods and Measures, 1(1), 77–89. — the case for alpha over kappa and the bootstrapped confidence interval procedure §8.3 presupposes
- Krippendorff, K. (2004). Reliability in content analysis: Some common misconceptions and recommendations. Human Communication Research, 30(3), 411–433. — the reply to critics of alpha; useful when a reviewer proposes a different coefficient
- Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. — kappa, the coefficient most reviewers will expect; know why alpha is preferred
- Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86(2), 420–428. — the interval-scale case and the forms of ICC, relevant if diagnostic predictions are rated on scales
- Gwet, K. L. (2014). Handbook of Inter-Rater Reliability (4th ed.). Advanced Analytics, Gaithersburg. — the comprehensive reference on coefficients, their paradoxes and sample-size planning
- Flake, J. K., & Fried, E. I. (2020). Measurement schmeasurement: Questionable measurement practices and how to avoid them. Advances in Methods and Practices in Psychological Science, 3(4), 456–465. — the metascience of measurement; the frame for §5.2’s criticism of the viability proxy
