The Regulatory Obligation Becomes the Responsiveness Ledger
Jianing Fan and Yue Yao trace public comments through individual obligations in Environmental Protection Agency rules, asking what changed between proposal and final text.
A responsiveness ledger should preserve that chain while exposing each classifier, proxy, threshold, human correction, and unresolved uncertainty used to construct it.
The Paper
The paper is Jianing Fan and Yue Yao's Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking, arXiv:2608.10329v1 [cs.CY], cross-listed in cs.CL and submitted August 11, 2026. The 14-page version 1 PDF lists both authors at Columbia University; the arXiv record says it was accepted as an EAAMO 2026 full paper for oral presentation.
The study assembles 786,197 comments from 6,145 EPA dockets spanning 2010–2022, then selects 36 anchor rulemakings through stratified random and extreme-case sampling; the resulting set is biased toward high-attention dockets. Its analytic anchor corpus contains 70,075 comments. Because 56.8 percent of the full corpus was attachment-only, the authors sampled attachments and recovered 1,634 of 1,769 attempted files. Those choices turn a seemingly simple question—did government listen?—into a problem of document recovery, legal-unit extraction, semantic matching, and outcome classification.
The Obligation Is the Unit
Section 553(c) of Title 5 requires an agency, after considering relevant matter presented in rulemaking, to incorporate a concise general statement of a final rule's basis and purpose. Counting comments cannot reveal how that consideration travels into operative text. A comment may address one deadline, monitoring duty, exemption, or threshold inside a much larger rule.
Fan and Yao therefore represent each obligation as an actor, modal, action, and object: who must or may do what. They extract 12,730 obligations from 29 of the 36 anchors; seven yielded no verified obligations under the pipeline. The proposed and final versions can then be compared at the level where compliance changes, while comments can be matched to the particular duty they support, oppose, or otherwise address.
The Measurement Pipeline
The source records came from the Regulations.gov and Federal Register APIs. A structural parser found candidate deontic language, and GPT-5 verified candidate obligations. Another model stage classified whether retrieved comments substantively addressed an obligation and, if so, their stance. A sentence-embedding classifier then assigned outcome states such as unchanged, edited, modified, dropped, or newly added. A 23-indicator rhetorical schema remained descriptive because the paper does not claim it was fully validated.
This decomposition is more informative than one end-to-end score. It also creates several distinct places where a measurement can fail: missing attachments, omitted obligations, retrieval misses, stance errors, outcome errors, and submitter-type misclassification. An audit of the final chart cannot substitute for audits of those components.
What the Results Say
Among obligations with at least five addressing commenters, 69.0 percent were revised, compared with 49.5 percent among less-addressed obligations. The within-docket odds ratio was 1.93, with a 95 percent confidence interval of 1.04–3.57 and p=.037. The authors call the association modest: many revisions had no addressing commenter, and many addressed obligations were not revised.
Direction did not supply a simple influence story. Opposition-addressed obligations were revised 62.3 percent of the time and support-addressed obligations 68.3 percent, a difference the analysis did not distinguish statistically (p=.155). The same obligation could appear in both groups. These are observational associations, not proof that a comment caused a textual change or that an agency accepted its argument.
The Audit That Failed Productively
The study's most useful result may be a broken instrument. The embedding classifier was poor at separating editorial refinement from a substantively changed compliance burden: its agreement with a blind human audit was only κ=.137. Surface similarity could not reliably answer whether regulated behavior had changed. The authors corrected the audited labels instead of hiding the failure.
That episode is a governance lesson. A plausible model, a familiar similarity metric, and neat outcome categories can still miss the load-bearing distinction. “Human validated” is not enough detail; a public account should show which boundary was audited, how cases were sampled, what agreement was observed, and which published results changed after correction.
Participation Has a Capacity Layer
The paper also investigates who participates. Regulations.gov bulk data universally redacted the Organization Name field in this corpus, so the authors inferred organizational submissions from the Title field. A permissive rule labeled 28.6 percent organizational, while a stricter alternative labeled 14.4 percent. After the outcome audit, organizational-majority composition was associated across dockets with editorial-refinement outcomes (odds ratio 4.90, p=.044, n=115), but the interval was wide and the proxy was not independently audited.
This is not evidence that agencies favored organizations over individuals within the same docket, and it is not a causal estimate. The sharper question is upstream: which participants have the time and expertise to locate, interpret, and contest a particular obligation? Equal access to a submission form does not equal equal capacity to produce obligation-specific engagement.
The Responsiveness Ledger
A responsiveness ledger should record the docket and document identifiers; each proposed and final obligation; the comment identifiers or privacy-preserving aggregates matched to it; retrieval rules and thresholds; addressing and stance labels; model, prompt, and version; proposed-final outcome; human-audit sample and correction; submitter proxy and confidence; mass-comment clustering policy; reviewer; uncertainty; and release version. It should retain the distinction between an edit and a changed duty.
Such a ledger would not automate the legal judgment of whether an agency adequately responded. It would make the empirical path contestable. A commenter, agency reviewer, researcher, or court could locate where a record disappeared, where a proxy entered, where a model disagreed with a person, and whether a later correction altered the public conclusion.
What the Study Does Not Establish
The 36 anchors are biased toward high-attention rulemakings. Mass-comment campaigns were not deduplicated, retrieval recall was not measured, all three completed human audits used only the two authors as raters, the submitter proxy remains uncertain, and a methodological choice about matching Code of Federal Regulations sections materially changed the measured modification rate. Small clusters and overlapping support and opposition subsets further limit strong inference.
The version 1 paper repeatedly refers to a supplementary repository, but no repository URL appears in the arXiv abstract, HTML, PDF, or submitted source archive. The available paper is detailed enough to audit claims and reported checks, but not to reproduce the full pipeline from linked artifacts. The study supports a bounded finding about measured engagement and revision, plus a stronger case for component-level auditability; it does not establish causation, capture, or a universal pattern across agencies.
Source Discipline
The factual record was checked against the arXiv abstract, complete version 1 PDF and submitted TeX source; the official Regulations.gov and Federal Register API documentation; and the Office of the Law Revision Counsel's text of 5 U.S.C. § 553. This essay reports the paper's measures with their qualifications, labels the ledger as a proposal, and uses no interview or source-text quotations.
Related Pages
- The Public-Comment Bot Becomes the Rulemaking Participant
- The Evaluation Schema Becomes the Public Ledger
- The LLM Judge Becomes the Annotation Budget
- The Manuscript Becomes the Build Artifact
- AI Audits and Assurance
- Transparency and Public Registers
Sources
- Jianing Fan and Yue Yao, Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking, arXiv:2608.10329v1 [cs.CY], cross-listed in cs.CL, submitted August 11, 2026, DOI 10.48550/arXiv.2608.10329.
- Fan and Yao, version 1 PDF, reviewed in full for the 14-page study's corpus, pipeline, validation, findings, ethics, limitations, and author-declared model use.
- Fan and Yao, version 1 submitted source archive, checked for the paper's TeX, references, figures, and the absence of a supplementary-repository URL.
- U.S. General Services Administration, Regulations.gov API documentation, including the document, comment, and docket endpoints and documented attachment and field constraints.
- Office of the Federal Register, Federal Register API documentation.
- Office of the Law Revision Counsel, U.S. House of Representatives, 5 U.S.C. § 553.