A digital services director for a hundred-site enterprise, a state university system, a hospital network, a multi-agency government portfolio, rarely loses sleep over a single broken page. The harder problem is knowing which of several hundred subdomains, each with its own content management system instance, its own vendor-built form, and its own long-departed developer, is quietly failing the people who depend on it, and knowing that before a regulator, a plaintiff's attorney, or an angry constituent finds out first.
Why detection accuracy stops being the deciding factor
At the scale of a handful of websites, the differentiator between accessibility tools is how well each one catches errors: missing alternative text, poor color contrast, unlabeled form fields. At the scale of a hundred sites, that differentiator narrows. In my view, most vendors' automated engines now catch a similar set of repeatable defects, which makes detection accuracy a minor differentiator rather than a decisive one. The problem shifts from detection to coordination: who owns the finding, whether it gets fixed, and whether the fix holds the next time a template changes.
This matters because the Department of Justice's 2024 final rule on web and mobile accessibility for state and local governments, issued under Title II of the Americans with Disabilities Act (ADA), ties compliance to the Web Content Accessibility Guidelines (WCAG) 2.1 at Level AA, the mid-range of the three conformance tiers WCAG defines (A, AA, and AAA, running from minimum to most stringent), on a timeline phased by jurisdiction size. A tool that produces an accurate scan of one property and a useless report at portfolio scale will not help an enterprise meet that deadline, because the obligation is organizational, not technical.
What an enterprise tool actually needs to do
Choosing well starts with separating three categories that vendors often blur together. The first is a scanning engine: the software that crawls pages and flags likely violations against WCAG success criteria. The second is a monitoring platform: the layer that runs scans across many domains on a schedule, tracks findings over time, and assigns them to owners. The third is manual testing and audit services: people, often including users of assistive technology, working through representative tasks that automated scans cannot judge.
A hundred-site enterprise needs all three, but the second category is where most procurement decisions actually get made, because it is the layer that determines whether the organization can answer basic governance questions. Before evaluating any product, a team should be able to state what it needs the tool to do:
- Portfolio-wide inventory: a current, automatically refreshed list of every site, subdomain, and content type the organization is responsible for, since obligations attach to properties a team has often forgotten it owns.
- Ownership assignment: the ability to route a finding to the specific team, vendor, or content owner responsible for the affected template, rather than one undifferentiated backlog.
- Retest tracking: a record of when a finding was fixed and whether the same path was retested, not just that a ticket changed status.
- Regression detection: the ability to notice when a previously fixed template breaks again after a redesign, a plugin update, or a vendor change.
- Reporting mapped to success criteria: output that names the WCAG criterion at issue, not a generic severity score, so legal, engineering, and program staff read the same finding correctly.
Vendors in this space include automated-engine providers such as Deque, whose axe-core accessibility testing engine is among the most widely adopted open-source engines of its kind, alongside separate portfolio-monitoring platforms built specifically for organizations managing many properties at once. None of these categories substitutes for the others, and an enterprise evaluating a single scanning tool in isolation is very likely solving the wrong part of the problem.
What automated tools catch, and what a person still has to check
Automated scanning is efficient at catching structural and pattern-based defects: missing labels, insufficient contrast, absent alternative text, malformed heading order, and other issues that can be inferred from code without understanding intent. The World Wide Web Consortium's own guidance on evaluating web accessibility reflects what most practitioners hold to be true: automated testing alone cannot confirm whether instructions make sense or whether a form's error message actually helps someone recover. Those questions require a person working through the task, ideally someone using assistive technology as part of daily practice, rather than a crawler inferring intent from markup.
I think the instinct to treat a clean automated scan as proof of compliance is one of the more expensive mistakes an enterprise team can make, because it substitutes a defensible-looking report for the harder work of confirming the task actually works. A hundred-site rollout should budget for periodic manual review of representative journeys, not only continuous automated coverage of every page.
Matching the tool to the regulation actually in force
Regulatory context should shape which standard a tool is configured against, and enterprise teams frequently get this wrong. The U.S. Access Board's revised Section 508 standards incorporate WCAG 2.0 at Level AA for federal information and communication technology. The 2024 Title II rule under the ADA applies WCAG 2.1 Level AA to state and local government web and mobile content, phased by jurisdiction size. Neither framework currently mandates WCAG 2.2, the most recent published version, though adopting it voluntarily tends to be less disruptive than waiting for the next regulatory update to force the change. A tool configured against the wrong version produces findings that do not match the obligation the organization is actually under, which is a governance failure dressed up as a technical one.
None of this is legal advice, and no vendor's report substitutes for counsel's judgment about a specific institution's exposure. What a well-configured tool can do is keep the technical record aligned with the standard that governs the organization, so that when legal or compliance staff need an answer, the evidence already speaks the right regulatory language.
A rollout that survives past the pilot
Enterprise accessibility programs tend to fail in a specific, avoidable way: a pilot succeeds on a handful of flagship sites, and the tool never scales past them because ownership was never assigned before the rollout expanded. A sequence that holds up better starts by inventorying every property the organization actually controls, including subdomains run by departments that procured their own content management system years ago. It continues by naming an owner for each template family and vendor-hosted component before turning on portfolio-wide scanning, since a finding with nowhere to go simply accumulates. It proceeds by setting a retest cadence tied to actual task completion, not ticket closure, and it closes by building the monitoring vendor's contract around that cadence, so renewal terms reinforce the retest habit rather than only scan volume.
An enterprise that gets this sequence right ends up with something more durable than a clean scan: a living record of which services work, who is responsible when they stop working, and proof that the fix was tested against the task that originally failed. That record protects the institution and the people who depend on its services, and it is also the record regulators and courts have increasingly come to expect.
Key takeaways
- At enterprise scale, coordination across many properties differentiates accessibility tools more than detection accuracy, since most automated engines now catch a similar set of defects.
- A hundred-site rollout needs three distinct capabilities: a scanning engine, a portfolio-wide monitoring platform, and periodic manual testing by people who use assistive technology.
- The monitoring platform is where procurement decisions matter most, because it determines whether an organization can assign ownership, track retests, and catch regressions across every subdomain.
- Federal technology falls under Section 508's reference to WCAG 2.0 Level AA, while the Department of Justice's 2024 rule applies WCAG 2.1 Level AA to state and local government sites.
- A clean automated scan is not proof of compliance; pair it with manual review of representative tasks before treating a property as accessible.
Questions readers ask
Which WCAG version applies to a public-sector website portfolio?
It depends on the governing regulation: Section 508 references WCAG 2.0 Level AA for federal technology, while the Department of Justice's 2024 Title II rule applies WCAG 2.1 Level AA to state and local government sites on a timeline phased by jurisdiction size.
Does a clean automated scan mean a website is compliant?
No. Automated tools reliably catch structural defects like missing labels and poor contrast, but they cannot confirm whether a task is actually usable, so a rollout should pair continuous scanning with periodic manual testing.
What is the difference between a scanning engine and a monitoring platform?
A scanning engine crawls pages and flags likely violations against WCAG success criteria, while a monitoring platform runs those scans across many domains on a schedule and tracks whether fixes hold over time.
How should a large organization sequence a tool rollout?
Start with a full inventory of every property the organization controls, assign an owner to each template family before scanning begins, and tie the monitoring vendor's contract to a retest cadence rather than scan volume.