The Gender Estimate Becomes the Audit Variable
A demographic estimate can be too illegitimate to attach to a person and still provide conditional evidence about a structure. That distinction is not a permission slip. It is a demand to specify what is being measured, why less harmful evidence is unavailable, and where the estimate must stop.
An audit may ask whether people perceived as feminine receive different treatment. It cannot turn that limited estimate into authority over anyone’s identity.
The Paper
The source is Evan Dong and Angelina Wang’s Algorithmic Gender Prediction Is Illegitimate, But Gender Imputation Can Yield Valid Measurements, arXiv:2608.13444v1 [cs.CY], submitted August 13, 2026. It is a conceptual and normative paper, not a report of a new model, dataset, experiment, or deployed audit.
The angle extends this site’s review of the coded gaze, which asks whether a category should be produced at all. Dong and Wang examine a narrower bind: whether a harmful inference can ever contribute to a defensible aggregate disparity measurement without acquiring standing as a claim about an individual.
Prediction and Imputation
The paper defines gender prediction as inferring gender from inputs. It defines gender imputation as a subset distinguished by application and interpretation: the estimate serves a specific, limited audit or descriptive purpose, and conclusions are drawn in the aggregate or used for structural action. A model does not become imputation merely because its operator puts the word audit on the dashboard.
This boundary follows the estimate after inference. If a row-level label is retained for targeting, eligibility, personalization, surveillance, or adjudication, the use no longer fits the authors’ definition of imputation. Aggregate intent cannot cleanse an individual decision.
Valid Does Not Mean Legitimate
The paper’s central move is to separate validity from legitimacy. Validity asks whether a measurement supports the claim made from it. Legitimacy asks whether the practice has acceptable authority, here under a trans-inclusive commitment to gender self-determination. A measurement could track a narrowly specified disparity and still lack authority to state, override, or police a person’s gender.
That separation blocks two shortcuts. Statistical usefulness does not confer moral permission, while ethical objection does not establish that every aggregate estimate is statistically uninformative. Governance has to carry both judgments instead of letting one erase the other.
Two Harms, Not One
For disparity measurement, the authors distinguish traditional sexism from oppositional sexism. The former privileges masculinity over femininity; the latter targets departures from dominant gender norms and includes transphobia, homophobia, and cissexism. The harms can compound. A method that estimates perceived femininity may expose one pattern while obscuring or reproducing another.
This is why the measured construct must not silently slide from perceived gender in this decision setting to gender identity. The paper argues that imputation invariably fails to respect trans and nonbinary identity, is likely to miss discrimination against nonbinary people, and always carries oppositional-sexist harms. Its possible value is narrower: under carefully specified conditions, it may help measure some discrimination attached to presentation or social position without asserting who is a woman.
Three Case Boundaries
The paper applies its framework to three cases. With generated images, no real depicted person is misgendered, but the model and label system can still reinforce gender norms. In film analysis, face-only classification can be both harmful and a poor proxy for how audiences perceive characters using wider narrative and social cues. With names, inference about real people remains illegitimate; synthetic names avoid misgendering a real person while retaining the risk of hardening stereotypes.
The cases show that data type is not the governing concept. The relevant questions are whose identity can be harmed, which perception actually drives the suspected discrimination, whether the input represents that perception, and what downstream authority the estimate receives.
Aggregate Is a Governance Boundary
The paper’s recommendations form a sequence. Use imputation only when an anti-discrimination benefit cannot reasonably be obtained another way. Define a narrow, contextual construct. Use inputs that fit that construct rather than merging unlike definitions. Keep interpretation aggregate, apply sound statistical methods, and validate with evidence collected for the chosen construct.
The aggregation discussion notes that continuous model outputs can reduce biases introduced by forcing estimates into discrete labels, subject to calibration assumptions that require empirical testing in the particular context, including comparison with self-reported data when that is the relevant benchmark. This is a method claim, not an ethical absolution. Aggregation reduces the authority placed on any single estimate; it does not remove the act of inference, the category system, or the possibility of misuse.
A useful implementation rule follows: row-level estimates should be treated as hazardous intermediate data. Limit access, retention, linkage, and export; publish uncertainty and subgroup definitions with the aggregate result; and prevent the audit variable from returning to an operational profile. That is this essay’s governance inference from the paper, not a tested result reported by the authors.
What the Paper Does Not Prove
The authors’ limitations are substantial. Their recommendations are not exhaustive, their main outcome is disparity, and they address only some aspects of gender. They do not supply a finished method for measuring oppositional sexism, validate a particular imputation system, or generalize the argument to every demographic attribute. The paper establishes a framework for judgment, not evidence that a proposed audit has passed it.
The Imputation Receipt
An auditable use should record the discrimination claim, decision context, narrowly named gender construct, reason self-reported or less harmful evidence is unavailable, affected population, input provenance, model and version, output scale, calibration and validation evidence, aggregation method, uncertainty, access controls, deletion schedule, prohibited reuses, downstream action, reviewer, and correction path. It should separately name the forms of sexism the measure can and cannot detect.
The Spiralist lesson is restraint through precision. An estimate can enter the audit as a provisional variable without entering the person’s record as truth. The moment it governs an individual, travels beyond its purpose, or masquerades as identity, the analytical exception has become the administrative rule.
Related Pages
- Unmasking AI and the Coded Gaze
- The Audit Query Becomes the Fairness Budget
- AI System Inventory
- Privacy and Data
Sources
- Evan Dong and Angelina Wang, Algorithmic Gender Prediction Is Illegitimate, But Gender Imputation Can Yield Valid Measurements, arXiv:2608.13444v1 [cs.CY], submitted August 13, 2026.
- Paper definitions and scope, checked for the distinction between prediction and imputation.
- Paper legitimacy analysis, checked for self-determination, prediction, norms, and label-system harms.
- Paper validity analysis, checked for measurement validity, gender constructs, and the traditional-versus-oppositional-sexism distinction.
- Paper recommendations, checked for necessity, narrow concepts, suitable inputs, aggregation, calibration, and empirical validation.
- Paper case studies and conclusion and limitations, checked for case boundaries and unproven generalizations.