Can AI Fix Web Accessibility? What the Evidence Actually Says
Notes from my session at DrupalCon Europe 2026, Rotterdam, 29 September.
I opened with a number that I expected to be boring, and it turned out to be the most uncomfortable slide in the deck.
The WebAIM Million scans the top one million home pages every year and reports what an automated accessibility check finds. In the 2026 report, 95.9% of those home pages had detected WCAG failures. That is up from 94.8% the year before. The average home page carried 56.1 detectable errors - 10.1% more than in 2025.
Take a moment with the direction of travel. For three years the industry has been told that AI is about to solve web accessibility. In those same three years, by the one measure we can take at scale, the web got measurably less accessible. Years of slow improvement reversed.
That is where I wanted the room to start, because it reframes the question. The interesting question is not whether AI can help. It is why we are losing ground on the part of the problem we already solved.
We are failing the easy half
Here is the detail that makes the scoreboard sting. WebAIM reports that 96% of all detected errors fall into just six categories:
| Failure | Share of home pages |
|---|---|
| Low contrast text | 83.9% |
| Missing image alt text | 53.1% |
| Missing form input labels | 51.0% |
| Empty links | 46.3% |
| Empty buttons | 30.6% |
| Missing document language | 13.5% |
Every one of those is machine-detectable. axe-core finds all six in milliseconds, for free, in your pipeline. These are not the hard problems of accessibility. These are the ones automation has technically already solved.
And we are getting worse at them.
So the question I put to the room, and the one I would put to you: if we cannot fix the half of accessibility a machine can already see, why would we believe a machine is about to fix the half it can't?
The one-click fix is an architecture problem, not a quality problem
You have seen the advert. One line of JavaScript, a monthly subscription, and your legacy platform is WCAG 2.2 AA compliant by tomorrow.
I want to be fair about why this sells, because it is not stupidity. An overlay removes a budget line nobody wants to defend, removes an awkward conversation with a client about scope, and removes an audit that would embarrass the last five years of delivery. It is the best-sounding proposition in our industry.
It also cannot work, and the reason is structural rather than a matter of vendor quality.
Think about the sequence. Your HTML ships with missing landmarks and unlabelled inputs. The browser builds the accessibility tree and the screen reader begins reading it. Then the overlay's JavaScript wakes up and starts rewriting the DOM underneath both of them. And then your framework re-renders.
That last step is the one that matters to anyone building decoupled Drupal, or React, or Vue. Components change state on their own schedule. The overlay applied its repair once; the next render blows it away. You have not fixed the accessibility tree. You have added a second writer to it, and it is the one that loses.
This is not my opinion. The Overlay Fact Sheet now carries 1,031 signatories, including accessibility engineers at Google, Microsoft, Apple, Shopify, ServiceNow and NBC, researchers at MIT, Carnegie Mellon and Gallaudet, contributors to the WAI-ARIA specification, contributors to JAWS and NVDA themselves. Its technical findings name React, Angular and Vue state changes explicitly as something overlays cannot repair.
When the people who build the screen readers an overlay claims to help sign a document saying it does not help, the technical argument is over.
The regulator has now said the quiet part
In January 2025 the US Federal Trade Commission brought an action against accessiBe. It settled at one million dollars, and the final order was approved on 22 April 2025 on a 3–0 Commission vote.
It is worth being precise here, because this gets misreported constantly. The FTC did not rule that overlays are illegal, and it did not rule that the technology does not work. It ruled on marketing. The order bars claiming that an automated product can make any website WCAG-compliant, or ensure continued compliance over time, without competent and reliable evidence.
That is narrower than the internet says. It is also, for those of us who sell to clients, considerably worse. The compliance claim is now the deceptive act and the compliance claim is exactly what an agency repeats to a client when it resells a widget.
There was a second count that almost nobody quotes. accessiBe had been publishing third-party reviews of its own product as independent opinion while paying, or being commercially connected to, the reviewers, without disclosing it. The order covers that too, barring misrepresentation of endorsers as independent or objective.
Sit with the implication. If you researched overlays in 2022 and came away reassured, some of what reassured you was paid for. The evidence base was partly manufactured.
I am not raising that to relitigate 2022. I am raising it because the same playbook is running right now, with "AI-powered" where "automated" used to be. When you read a glowing independent review of an AI remediation tool this year, ask who paid for it.
The litigation data points the same way
If a widget protected you, you would see it in the filings. You do not.
UsableNet's reports show that in 2024, 1,023 companies were sued for digital accessibility while an accessibility widget was live on their site - roughly one in four of the ~4,000 filings that year. In the first half of 2025 it was 659 out of 2,019 filings: closer to one in three. The ratio is moving in the wrong direction, and 2026 is on track to pass 6,000 suits.
You will sometimes hear that plaintiff firms actively scan for overlay scripts. Practitioners will tell you they do. I could not find a primary source for it, so I did not assert it from the stage, and I will not assert it here. But the correlation in the filing data is difficult to explain any other way.
Europe: the part I would now tell differently
The European Accessibility Act has been in force since 28 June 2025. It reaches private-sector digital services, e-commerce, banking, transport, e-books, telecoms, and the technical standard, EN 301 549, incorporates WCAG 2.1 level AA.
In the talk I argued that the penalty schedule is not the interesting part. Under the ADA, an overlay is a weak defence after something has gone wrong. Under the EAA, someone at your client has to sign an accessibility statement - a dated, documented assertion of conformance before anything goes wrong at all. My question to the room was whether you want that signature resting on a product the FTC has already said cannot substantiate a compliance claim.
Preparing this write-up, I found the case that makes the point far better than I made it.
In July 2025, three French disability organisations namely apiDV, Droit Pluriel and Intérêt à Agir served formal notices on four major grocery retailers: Auchan, Carrefour, E.Leclerc and Picard. The complaints cited keyboard-navigation failures, missing alternative text and checkout flows that could not be completed.
On 4 June 2026, the Tribunal judiciaire de Caen ruled against Carrefour, ordering its website and mobile application to be made fully accessible within six months, with a penalty of €500 for each day of delay thereafter.
The reasoning is what matters. Carrefour argued that it met 71% of the RGAA criteria, France's implementation of WCAG 2.1 AA, and that this represented substantial good-faith compliance. The court rejected the framing. It drew the distinction between an obligation of means and an obligation of result, and held, in effect, that an e-commerce site cannot be somewhat accessible. Partial conformance, however well-intentioned, is not conformance.
Seventy-one percent is not a bad score. Most of the sites I audit would not reach it. And it was not enough, because the test the court applied was not "did you try" but "can a disabled person buy their groceries."
A month earlier, the same court had acknowledged Auchan's non-compliance but declined to grant emergency relief so this is an area where outcomes are still settling. Confirmed monetary fines under national EAA-implementing law remain rare; the pattern so far is notice, then investigation, then corrective order, with penalties as a later step. Norway's equality authority has imposed a daily penalty of NOK 50,000 (about €4,500) on a health portal for persistent keyboard-accessibility failures.
But the direction is unmistakable, and it lands squarely on the overlay pitch. A widget cannot get you from 71% to conformance. It cannot get you there because the things it cannot repair labels, error handling, focus control, keyboard operation inside a framework's state changes are precisely the things that stop the purchase.
A note on sourcing: I have read specialist coverage of the Carrefour decision from several independent outlets, not the judgment itself. If you intend to rely on it commercially, get the text.
A digression about a ceiling
I want to describe something from home, because it explains the argument better than any diagram I could draw.
Khatamband is a Kashmiri ceiling craft. Hundreds of small pieces of walnut and deodar are cut so that they interlock and hold by geometry alone. No nails. No adhesive. And because there is no glue, any single panel can be lifted out and replaced without disturbing the ones around it. There are ceilings in Srinagar built this way three hundred years ago that are still being repaired today, panel by panel.
That is a design system. Small interlocking parts, no adhesive, repairable in place, one piece replaceable without touching the rest. Every property we claim to want from a component library, a Kashmiri carpenter solved with joinery.
An overlay is a sheet nailed over the top of one. It hides the joinery instead of repairing it. You cannot fix a single panel through it. And when it comes away and it does, on every re-render, you have the original problem plus a new one.
Which gives you the commercial case in four words: sheeting, or joinery. A subscription is an expense that produces nothing you own. A remediated component is capital - fixed once, in the repository, reusable across every client running your shared theme, and defensible under audit. That is not really an accessibility decision. It is a decision about whether you are building an asset or a liability.
Where AI genuinely fails, with numbers
So overlays are out. That was the easy half of the talk, and honestly most of the room was never going to buy one. The harder question is what happens when the AI stops being a widget in your user's browser and becomes an assistant in your IDE, because that has already happened.
The best evidence I know comes from Guriță and Vatavu, presented at the Web for All conference in 2025. They generated 80 interfaces across Claude 3.5 Haiku and GPT-4-turbo and had accessibility experts evaluate them against WCAG.
- With an accessibility-agnostic prompt - just "build me a signup form" - the violation rate was 58%.
- With an accessibility-oriented prompt naming WCAG 2.1 AA - same models, same task - it was 19%.
Two conclusions follow, and they deserve equal weight.
The model is not the variable. The person driving it is. A thirty-nine point swing came from nothing but the prompt.
And prompting is a mitigation, not a fix. Nineteen percent is still one interface in five failing, evaluated by experts who knew what they were looking for.
Break the agnostic condition down by criterion and it gets sharper. 100% of generated interfaces violated alternative text. 100% violated information and relationships. 90% failed on status messages, 80% on keyboard navigation.
Look at the top two and think back to the WebAIM scoreboard. Missing alt text and broken structure are also two of the six most common real-world failures on the actual web. Which makes sense: the model learned from the web, the web fails these criteria at scale, so the model reproduces the failure faithfully and we ship it back out, where it becomes training data. That is not a bug in the model. That is a loop, and it is one reason the WebAIM numbers went up this year instead of down.
The canonical artefact looks like this:
<!-- what the model gives you -->
<div class="custom-btn"
role="button"
aria-label="Click here to submit form"
tabindex="0">
Submit
</div>
<!-- what the platform already gave you -->
<button type="submit">Submit</button>The left-hand version passes axe. It passes ESLint. It will pass your code review, because it looks like somebody was being careful. But there is no keydown handler, so Enter and Space do nothing, it is focusable but not operable. It is not in the form's submit path. It has no native disabled state and no focus ring. And the aria-label overrides the visible word "Submit", so a voice-control user who says "click Submit" gets nothing at all. That is 2.5.3 Label in Name, broken by the model trying to be helpful.
This is not one bad generation. Abu Doush and Kassem benchmarked eleven component patterns across GPT-4o, Copilot Pro, Claude 3.7 Sonnet and Grok 3, and found all of them producing semantically valid code that failed accessibility without follow-up prompting. It is the default behaviour of the category.
On stage I showed two screen reader recordings of that button - the generated one and the native one. The generated one announces, the user presses Enter, and nothing happens. The silence is the whole argument. Then: axe-core reports zero violations on that version.
Where AI genuinely works
Same technology, different position in the lifecycle. Not a runtime patch for the user; a design-time assistant for the developer.
Four places it pays, in my experience and in roughly this order of return:
- Pattern analysis across a portfolio. This is the one almost nobody is doing and it has the highest return of the four. If you run forty Drupal sites off a shared theme, ten thousand audit findings are not ten thousand problems, they are about a dozen component defects, repeated. Models are very good at finding that structure in messy audit exports. Fix the pattern once in the component, ship it to forty sites.
- Draft patches in the pull request. axe-core fails the build, the model reads the failing DOM and proposes a diff, a developer reviews it like any other PR. It opens a pull request. It never merges.
- Baseline alt text at scale. Twelve thousand legacy media items: the model drafts, an editor corrects. Ninety percent of the typing disappears and the judgement stays human which matters, because a vision model can tell you there is a smiling woman in a blue blazer, but only your page knows she is Priya Raman, the CTO.
- Inline authoring feedback. WCAG-tuned prompts in the IDE flagging a missing label or focus state before the commit, when the fix costs minutes rather than a sprint.
Every one of those has a human at the end of it. That is not a limitation of the approach. That is the approach.
The loop fails at the human step
The protocol I argued for is four steps, and only one of them is the AI. Detect with deterministic tools. Propose with the model. Verify with a human. Commit to source control.
Step three is the one that gets skipped, so let me be specific about what it is not. It is not reading the diff and nodding. A diff cannot tell you where focus lands after the modal closes. A diff cannot tell you the reading order is incoherent. Somebody has to put their hands on a keyboard and their ears on a screen reader.
The research backs this up in a way I found genuinely encouraging. Mowar, Peng, Wu, Steinfeld and Bigham (CodeA11y, CHI 2025) observed 36 developers building with AI assistants and catalogued why accessibility failed. Three modes: nobody prompted the assistant for accessibility at all; nobody replaced the placeholder attributes the model emitted, so the scaffold shipped as the implementation; and nobody had a practical way to verify the result, so it was assumed.
Not one of those three is the model's fault. All three are human and organisational. Which means the fix is not a better model, and you are not waiting on anybody's roadmap.
If you do one thing
Add one line to your pull request template.
Keyboard tested: yes / noIt costs nothing. It needs no procurement and no vendor evaluation. It catches the class of failure no scanner and no diff can see. And a "no" in that field is not a failure, it is a question somebody now has to answer out loud.
That single line will do more for your accessibility than any tool mentioned in this post.
So, can AI fix the web?
No. But it can help you fix it.
Accessibility is an architectural requirement, not an operational afterthought. Automation is a detector, not a remediator — know which one you bought. The model is not the variable; the person driving it is. And overlays rent accessibility, while source code owns it.
Build the ceiling so that any panel can be lifted out and replaced. That is the craft, and no machine has learned it yet.
I gave this talk at DrupalCon Europe 2026 at the Postillion Convention Centre, WTC Rotterdam, on 29 September. If you were in the room and disagreed with something - particularly the Deque coverage argument, which I deliberately left unresolved - I would like to hear it.
Sources
- WebAIM Million, 2026 report annual accessibility scan of the top one million home pages
- Overlay Fact Sheet signatory count and technical findings
- FTC press release, 3 January 2025 complaint and proposed order
- FTC press release, 22 April 2025 final order approved
- UsableNet 2024 Year-End ADA Lawsuit Report and 2025 Midyear Report litigation counts
- Guriță, A.-E. & Vatavu, R.-D. (2025). When LLM-Generated Code Perpetuates User Interface Accessibility Barriers, How Can We Break the Cycle? W4A '25. DOI · open PDF
- Abu Doush, I. & Kassem, R. (2025). Can generative AI create accessible web code? A benchmark analysis of AI-generated HTML against accessibility standards. Universal Access in the Information Society 24(4), 3483–3506. DOI
- Mowar, P., Peng, Y.-H., Wu, J., Steinfeld, A. & Bigham, J. P. (2025). CodeA11y: Making AI Coding Assistants Useful for Accessible Web Development. CHI '25. arXiv
- Deque, Automated Accessibility Coverage Report the 57% counter-argument to the 30% figure
- Directive (EU) 2019/882 (European Accessibility Act) and EN 301 549
- Carrefour ruling, Tribunal judiciaire de Caen, 4 June 2026 reported in specialist coverage; the judgment text was not consulted for this post