Blog · Accessibility and Maintenance · October 2, 2026

Accessibility Can Break After the AI Audit Passes

An accessibility result belongs to the version that was examined. Maintaining access means noticing when content, interfaces, generated output, or testing methods change what that result can tell us.

The booking after release

Imagine a community workshop booking site. This is a hypothetical example. An AI-assisted audit checks the reservation form, a specialist reviews the selected interactions, and the team fixes the problems found. Later, a content update replaces a short error instruction with a generated explanation. It sounds friendlier but omits which field needs correcting. A person who previously recovered from the error now cannot finish the booking.

The earlier work may have been careful. Its evidence still describes the earlier version. Our argument is that accessibility maintenance needs an explicit account of what changed, which observations remain relevant, and who must investigate the difference. A passing result is useful only while its scope and age remain visible.

The site's earlier essays on accessibility overlays and generated image descriptions examine responsibility and editorial judgment. Here the narrower question is what happens after a repair: how can a team notice when people lose an ability the service previously supported?

Evidence has a version

W3C's WCAG Evaluation Methodology 2.0, a Group Note dated July 23, 2026, addresses repeat evaluation. It recommends retaining some previous samples for comparison and replacing others to broaden coverage. It also warns that evaluating a subset alone cannot establish whole-site conformance. This is evaluation guidance, rather than an additional accessibility requirement.

Our proposed application is to keep a short record beside each important result: the task attempted, the state reached, the product and content versions, the testing environment, and the actual observation. Where generated output affects the task, keep the reviewed output and the available model and prompt identifiers. An unavailable provider version should remain an explicit gap.

This record need not become a second publishing system. For the hypothetical booking form, it could connect the tested error message to the content revision that replaced it. That connection lets an editor find the relevant evidence without reconstructing an audit from screenshots and memory. It also gives a reviewer something more precise to challenge than a green status indicator.

Three things can change

We propose separating three kinds of change when investigating a suspected regression. First, the interface can change: a replacement calendar, confirmation panel, or navigation component may alter the route through a task. Second, the content can change while the interface stays familiar. Third, the measurement can change because the scanner, evaluation model, rules, or test setup changed.

The distinction matters for deciding what to repair. In another hypothetical case, an AI evaluator newly flags a control after its own model update, while the page remains identical. That deserves investigation, but it does not establish that the service became worse overnight. Conversely, a stable score cannot establish stability in interactions the evaluator never examined.

W3C's guidance on selecting evaluation tools says tools assist assessment, can return misleading results, and require human judgment for aspects that cannot be checked automatically. It also describes their use throughout development. These limits support a useful role for automation without making an automated pass a conformance finding.

Our recommendation is to compare like with like before interpreting a trend. If the evaluator changes, rerun saved examples where feasible and have a knowledgeable person inspect disagreements. Keep newly discovered barriers on the repair list even when they predate the update. Distinguishing discovery from regression improves diagnosis; it should never become a reason to leave someone blocked.

Make repeated checks interpretable

The Accessibility Conformance Testing effort documents rules for automated, semi-automated, and manual evaluation. Its purpose includes reducing disagreement caused by differing interpretations of accessibility guidelines. ACT gives testing methods a shared structure; a product's broad claim to use AI says much less about what its checks actually mean.

We would therefore ask a tool supplier which checks it performs, which rules it implements, and how results change between releases. A recorded result should distinguish a failed check from a check that never ran. An authentication failure that prevents a scanner reaching the booking form must remain visible as missing coverage.

For a small service, start repeated checks with tasks whose failure would prevent its purpose: finding an event, reserving a place, correcting a mistake, and locating confirmation. Preserve the actions and expected outcomes. Supplement those familiar paths with newly introduced states. Otherwise, the team can perfect a rehearsal while the live service acquires an unexamined route.

Our proposed trigger is the affected task, rather than the department making the change. A prompt revision that rewrites recovery instructions should reach the same responsible reviewer as an interface revision affecting recovery. A cosmetic edit may justify a narrower check. The team should explain that judgment instead of pretending every update requires an identical audit.

Test the encounter with assistive technology

The ARIA and Assistive Technologies project maintains a test suite and harness for assessing ARIA support. Its repository describes an initial scope of selected desktop screen readers and assertions about expected behavior with example design patterns. That is bounded interoperability work, not a study showing that every implementation of a pattern works for everyone.

Our inference is that evidence about a component should travel with its tested conditions. An example's result cannot settle what happens after a team embeds it in a different workflow. A local check should follow the person's attempted action through to its consequence, including whether the resulting state is understandable in the chosen access setup.

For AI-generated content, we propose preserving a stable interaction structure around variable prose. In the booking example, a reviewed instruction naming the field could remain available while optional generated explanation supplies additional detail. This is a design suggestion, not a demonstrated universal solution. It reduces how much essential behavior depends on each new wording choice, while leaving the added text open to evaluation.

Give feedback a route to repair

W3C's guidance on involving disabled users explains that user evaluation can reveal usability problems missed by conformance evaluation. It also cautions against generalizing one person's experience to everyone with a disability and recommends combining user involvement with standards evaluation. Successful participation is valuable evidence with a defined scope.

Our maintenance proposal gives a reported barrier a named human owner with authority to obtain a repair. The report should be able to describe the interrupted task in ordinary language. People should not need to identify the faulty accessibility property, disclose a diagnosis, or reproduce a failure repeatedly before receiving a response.

Planned sessions with disabled participants should be compensated and scheduled around meaningful changes. They should inform which tasks deserve closer attention, while routine checks remain the team's responsibility. A feedback channel supplements that work. Silence from the channel is an ambiguous observation: it cannot tell the team whether nobody encountered a barrier or whether reporting was too difficult.

Close the issue when the repaired task has been checked, with the verification recorded separately from the code change. Offer the reporter a clear update without making their return visit a condition of closure. If the full repair will take time, assign someone to make the immediate alternative usable and explain its limits.

Return to the workshop reservation. The maintenance record should let a responsible person locate the changed instruction, restore a usable recovery path, and determine which similar messages need review. That is a concrete benefit an audit can leave behind: evidence the next editor can use when the service changes again.

Sources

Primary sources consulted October 2, 2026. The maintenance workflow proposed here is our synthesis; no product or user testing was conducted for this essay.

Production: commissioned by the site operator; researched and drafted by GPT-6 Astra; editorial and source review by the coordinating AI assistant.


Return to Blog