When Label Disagreement Is Part of the Evidence
A majority label can conceal both a correctable mistake and a defensible difference of interpretation. Before scoring a model, decide which judgments the test is meant to predict—and what the disagreement says about the test itself.
The same label, different evidence
Six readers find a joke sarcastic; four read it literally. In another panel, all ten choose sarcasm. These are hypothetical annotation results. A dataset retaining only the majority label stores the same answer for both. An evaluator using that answer cannot distinguish a model that anticipates divided interpretation from one that confidently expects agreement.
That lost distinction matters if the task is to anticipate readers' reactions. It may matter less if the task is to apply a deliberately narrow editorial rule. The problem begins when those purposes share an answer key without sharing a definition of success.
This essay develops a practical interpretation of research published from 2022 through 2024, checked on October 2, 2026. Our proposal is to diagnose disagreement before compressing it, then choose an evaluation target that matches the intended use. Keeping every label and forcing consensus are both inadequate as universal instructions.
Diagnose the split
Barbara Plank's position paper on human label variation distinguishes plausible variation from errors such as lapses of attention. It treats ambiguity, uncertainty, and multiple reasonable answers as problems for data collection, modeling, and evaluation together. It is a research agenda, not evidence that every disagreement improves a model.
A more specific investigation by Nan-Jiang Jiang and Marie-Catherine de Marneffe examines natural language inference: deciding what follows from a sentence. Their taxonomy separates uncertainty in sentence meaning, underspecified guidelines, and annotator behavior. Its importance here is diagnostic: the same spread of labels can arise through different mechanisms. The study concerns inference datasets, so its categories should not be assumed to exhaust disagreement in other tasks.
For a working annotation project, we propose asking what intervention could resolve the split. In a hypothetical sarcasm task, an accidental click calls for correction. A missing preceding message calls for restoring context. A disputed definition of sarcasm calls for clarifying the question. None is repaired merely by collecting more votes under unchanged conditions.
A residual difference may remain after those repairs. One careful reader may interpret understated praise as sincere; another may hear mockery. If both can explain their reading within the supplied context and task, the project has reason to preserve both judgments. That is a test of defensibility, not a declaration that all answers are equivalent.
Reviewers should also be allowed to leave the cause unresolved. A label count cannot diagnose itself. Brief explanations and targeted reannotation can help investigate a disputed item, but demanding an explanation does not guarantee that the explanation is accurate. Diagnosis remains work, with uncertainty of its own.
What adjudication settles
Adjudication helps when the promised output is a consistent application of a specified rule. Consider a hypothetical editorial dataset whose instructions explicitly exclude quoted sarcasm and ask only about the author's own wording. An adjudicator can resolve a disagreement by checking that boundary. A single reference label is useful because the question has been narrowed enough to support it.
But narrowing the question changes what the result means. The adjudicated answer describes compliance with that editorial definition. It does not establish how an ordinary reader will experience the quotation. A model can follow the rule correctly while predicting audience reactions poorly.
Our recommendation is to preserve the initial judgments alongside the adjudicated result when feasible. That allows later evaluation of both rule application and reader variation. Record whether the adjudicator corrected a mistake, supplied context, or selected a convention. Those interventions should not all disappear under the word “cleaning.”
A clarified rule can legitimately raise agreement. The useful question is whether it makes the task more faithful to its purpose. Agreement obtained by deleting every difficult example would instead leave a test that says little about the ambiguous material the system is expected to handle.
What the distribution estimates
The 2023 Learning with Disagreements shared task made this distinction operational. It evaluated subjective language tasks using cross-entropy against annotation distributions as its primary measure, while reporting F1 as additional information. The released format included individual annotations and their counts, as well as hard and soft labels. This demonstrates an evaluation design; it does not establish that every application should use its metric.
A soft label in our hypothetical split records six sarcasm judgments out of ten. It does not mean the sentence contains a measurable quantity of sarcasm equal to sixty percent. Nor does it automatically estimate the views of everyone who might read it. It describes a panel responding to particular wording, context, and instructions.
The practical consequence is that the panel belongs in the definition of the target. If the intended readers change, the old distribution may answer the wrong question. If each item is rated by a different mix of readers, apparent differences between items may partly reflect that mix.
Even identical proportions can carry different amounts of evidence. Six of ten and sixty of a hundred both yield sixty percent. Under comparable independent sampling, the larger panel estimates the response share more precisely. Additional judgments cannot, however, repair a systematically unrepresentative panel. Keeping counts prevents a tidy decimal from concealing these distinctions.
Give the score a meaning
For the hypothetical six-to-four item, models assigning sarcasm probabilities of 0.60 and 0.99 choose the same winning label. Majority-label accuracy treats them alike. Cross-entropy against the observed split favors 0.60 because it leaves appropriate probability for the literal readings. This is an illustrative consequence of the scoring rule, not an experimental result.
Metric choice still needs scrutiny. Giulia Rizzi and colleagues assessed properties of distributional metrics and recalculated LeWiDi rankings. Rankings changed with the measure. Their preferred distance measures satisfied their proposed properties for binary tasks, while their multiclass analysis exposed further limitations. One benchmark exercise cannot establish a universally best metric.
Our reading is that a score needs an interpretation before it needs a leaderboard. Cross-entropy includes the uncertainty already present in the target distribution: a perfect match to divided judgments still has positive loss. That does not invalidate comparison between models on the same fixed test. It does complicate comparing raw scores across tests with different levels of disagreement. Report performance on clear and divided items separately so readers can see where improvement occurs.
An evaluation that keeps the distinction
The evaluation we propose starts by writing down whether it predicts a rule-based decision, the panel's response distribution, or particular readers' judgments. Retain the corresponding references. Score rule compliance against adjudicated labels and distribution matching against preserved responses. If both matter, report both rather than letting success on one silently stand in for the other.
Before treating a gain as useful, inspect whether it comes from resolving genuine confusion, reproducing a panel's habits, or merely becoming less confident everywhere. The last strategy can help on divided items while weakening predictions on clear ones. This is why the comparison needs examples with strong agreement as well as disputed examples.
A distribution also leaves a decision unfinished. An editor still has to decide whether a joke fits the intended audience. The model's estimate can inform that judgment without dictating it. The stronger evaluation claim is precise: the system reproduces these judgments, under these conditions, with these remaining errors. That gives the person using it evidence they can interpret.
Sources
- Barbara Plank. The “Problem” of Human Label Variation: On Ground Truth in Data, Modeling and Evaluation. EMNLP, December 2022.
- Nan-Jiang Jiang and Marie-Catherine de Marneffe. Investigating Reasons for Disagreement in Natural Language Inference. TACL, December 2022.
- Elisa Leonardelli and colleagues. SemEval-2023 Task 11: Learning With Disagreements (LeWiDi). July 2023.
- Giulia Rizzi and colleagues. Soft metrics for evaluation with disagreements: an assessment. May 2024.
Related reading
- The Annotator Disagreement Becomes the Ensemble Receipt
- Who Sets the Test for an Underserved Language?
- The Annotation Tool Becomes the Labor Meter
Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.