August 27, 2026 · 11 min read
Fix ADA Captioning Fast: 30 Day Compliance Checklist for U.S. Orgs
For U.S. organizations: an ADA captioning guide that maps DOJ Title II and WCAG 2.1 to caption quality standards, enforcement risks, and a 30 Day...

Under the ADA’s effective-communication standard, organizations providing synchronized audio-visual content to the public generally must supply accurate, synchronized captions. The Department of Justice ties that obligation directly to WCAG 2.1 Level AA, specifically the success criteria covering prerecorded video (1.2.2) and live content (1.2.4). Captions that meet the letter of the law also have to be accurate, properly timed, and include speaker identification and non-speech sound cues, not just a rough transcript scrolling underneath the picture, as outlined in caption quality practices.
TL;DR:
- Automated captioning generally requires human review to meet accuracy standards, especially for names, jargon, and accents.
- Synchronized captions are mandatory for all prerecorded videos and live content, with proper speaker labels and non-speech sound cues.
- Content like public meetings, webinars, and virtual events should normally be captioned unless the organization faces a genuine undue burden.
- Implementing a repeatable workflow, including automated transcription, human correction, and thorough testing, is crucial for compliance.
- Live captioning tools like Live Caption AI offer affordable, multi-language, real-time captions as a practical access solution, but do not guarantee ADA compliance alone.
Table of Contents
- What the ADA actually says about ada captioning requirements
- What “good” captions actually look like
- Which videos and events actually need captions
- How to build a captioning workflow that actually holds up
- Enforcement patterns and what triggers a complaint
- Your compliance checklist for the next 30 days
- The tradeoff nobody talks about enough
- Adding live captions without the stenographer bill
- Sources
What the ADA actually says about ada captioning requirements
The ADA doesn’t hand you a caption style guide. It hands you a principle: effective communication. Titles II and III require covered entities to provide auxiliary aids and services, including closed captioning and real-time captioning, whenever they’re necessary to communicate as effectively with people who are deaf or hard of hearing as with everyone else, according to Ada. That standard applies broadly, but the technical bar for meeting it has gotten a lot more specific in the last two years.
In April 2024, the DOJ finalized a rule requiring state and local governments to make their websites and mobile apps meet WCAG 2.1 Level AA, which folds captioning directly into Title II compliance. The rule phases in on a staggered timeline based on population size, but the technical expectations it sets are already shaping how courts and regulators read Title III obligations for public accommodations, even where no equivalent rule has been finalized yet.
Three WCAG success criteria matter most here:
- 1.2.2 (Captions, prerecorded): synchronized captions required for any prerecorded video with audio.
- 1.2.4 (Captions, live): synchronized captions required for live audio content delivered through synchronized media.
- 1.2.5 (Audio description, prerecorded): a related criterion covering description of visual content, relevant when video conveys information no caption track alone can carry.
Courts applying Title III’s effective-communication standard have repeatedly treated websites and streamed content as covered “services,” which means the WCAG framework built for Title II entities is becoming the de facto technical benchmark everywhere, not a government-only rule.
What “good” captions actually look like
Meeting the letter of ADA compliance standards means more than turning on auto-generated subtitles and calling it done. Section508 states plainly that captions need to include dialogue and relevant non-speech audio, stay synchronized with the audio track, and that automatically generated captions typically require review before they’re accessible. That last point trips up more organizations than any other.
Accuracy means the caption text matches what’s actually said, including names, technical terms, and industry jargon your speech recognition engine has never encountered. A caption track riddled with misheard words doesn’t satisfy effective communication just because it exists.
Synchronization matters because a two-second lag between speech and text breaks comprehension for viewers relying on captions, especially during fast dialogue or Q&A segments.
Non-speech audio and speaker identification round out a compliant caption track:
- Bracketed notations like [applause], [phone ringing], or [dog barking] when they carry meaning.
- Speaker labels (SPEAKER 1:, DR. CHEN:) when multiple people talk, especially in panels or interviews.
- Sound effects that signal a scene change or emotional cue in narrative content.
WCAG’s own guidance is blunt on one point worth calling out as a statistic-adjacent fact: automatically generated captions usually need human confirmation to be considered accessible. That’s not a suggestion; it’s the standard’s own caveat about relying on raw automatic speech recognition output.
On format: readable captions typically run no more than two lines on screen at once, with each line kept short enough to read at a natural pace, and they should sit clear of any on-screen graphics or lower-third text. SRT and WebVTT remain the most widely supported caption file formats for web players. Burned-in (open) captions work for social clips where viewers can’t toggle a track, but they strip away the option to turn captions off and often can’t be styled to individual reading preferences.
Which videos and events actually need captions
Not every piece of content carries the same obligation, but the default assumption should lean toward “caption it” rather than “wait and see.” WCAG 1.2.2 requires captions for prerecorded synchronized media, and WCAG 1.2.4 extends that to live content. In practice, that covers:
- Public meetings and city council sessions streamed or archived online.
- Recorded webinars and training videos posted for staff or public access.
- Marketing and promotional video on websites or social channels.
- Virtual events, town halls, and panel discussions with live audio.
Two limits show up often in real disputes. Undue burden and fundamental alteration are legal defenses, not blanket exemptions. An organization has to show the specific cost or difficulty is genuinely disproportionate, not just inconvenient, and even then it typically owes an alternative form of effective communication rather than nothing at all. Archived content that’s rarely accessed doesn’t get a free pass either; DOJ guidance treats posted video the same whether it was recorded yesterday or three years ago, since the accessibility barrier persists as long as the content stays live.
How to build a captioning workflow that actually holds up
Meeting legal captioning requirements comes down to a repeatable process, not a one-time fix. Here’s the workflow that tends to hold up under scrutiny:
- Generate a base transcript. Use automated speech-to-text as your starting draft, never your final product.
- Produce a timed caption file. Convert the transcript into a synchronized format like SRT with proper timing codes.
- Run human review. Correct misheard terms, add speaker labels, and insert non-speech audio notations the automated pass missed.
- Publish and verify. Check playback across devices, confirm the caption toggle works, and test synchronization at normal playback speed.
The hybrid approach, draft with automation, finish with human eyes, is what WAI’s own recommendations point toward. Automated tools have gotten good at general speech, but they still stumble on medical terminology, legal jargon, regional accents, and multi-speaker crosstalk. A five-minute human pass on a ten-minute video usually catches the errors that matter most.
Pro Tip: Keep a simple log of which videos have been captioned, who reviewed them, and when. If a complaint ever comes in, that log is your best evidence you took effective communication seriously rather than as an afterthought.

Developers building or buying a video player should confirm a few things beyond the caption file itself: keyboard operability for the caption toggle, ARIA labels on player controls so screen reader users can find the caption button, and persistence of a viewer’s caption preference across sessions. The U.S. Web Design System’s accessibility guidance covers these player-level expectations well for teams building custom video interfaces.
On file handling, most organizations export captions in SRT for video platforms, plain text for transcripts, or PDF for archival records. Whatever tool you use, check what happens to the source audio after processing. Some platforms retain uploaded audio indefinitely; others discard it once the caption file is generated. That distinction matters more for medical, legal, and HR content than for a marketing reel.
For high-stakes events, live legal proceedings, medical consultations, or complex multi-speaker panels, professional CART (Communication Access Realtime Translation) services or trained human captioners still outperform automated tools on accuracy. They cost more per hour, but for content where a misheard word carries real consequences, that tradeoff often makes sense.
Enforcement patterns and what triggers a complaint
Most captioning complaints don’t start with a lawsuit. They start with a single person who couldn’t access a public meeting recording or a training video and filed a request that went unanswered. DOJ’s web accessibility guidance shows a consistent pattern in enforcement: missing or unreviewed captions on public-facing video frequently lead to negotiated remediation agreements rather than immediate penalties, provided the organization responds and fixes the gap.
FCC rules add a separate layer for broadcast television and certain internet-distributed video that originated on TV. The FCC’s captioning guidance emphasizes accuracy, synchronization, completeness, and placement, and industry conversations often reference a 99% accuracy benchmark drawn from that framework. That FCC standard doesn’t automatically apply to ADA-covered web content, but regulators and plaintiffs’ attorneys increasingly point to it as evidence of what “accurate” should mean in practice.
If a complaint lands on your desk:
- Document the request and your response timeline immediately.
- Identify the specific content at issue and assign a remediation deadline.
- Keep records of the auxiliary aid you offered and why you chose it.
Your compliance checklist for the next 30 days
Run through this before your next audit or board meeting, not after a complaint arrives.
- Inventory your video assets and flag which ones are public-facing, archived, or restricted to internal audiences.
- Check each caption track against WCAG’s core requirements: present, accurate, synchronized, non-speech audio included, speakers labeled.
- Document your QA process for any automated captions, including who reviewed them and when.
- Test your player for keyboard access, ARIA labeling, and caption toggle persistence.
- Update internal policy so staff know when captions are required and how to request accommodations.
| Checklist item | Why it matters | Quick fix |
|---|---|---|
| Caption accuracy review | Automated tools miss names, jargon, accents | Add a human review pass before publishing |
| Synchronization timing | Lag breaks comprehension for caption users | Test playback at normal speed, not just preview |
| Non-speech audio notation | Meaning gets lost without sound cues | Add bracketed tags for key sound events |
| Player accessibility | Screen reader users need working controls | Confirm ARIA labels and keyboard toggles |
| Request documentation | Protects you if a complaint is filed | Keep a dated log of aids offered and chosen |
The tradeoff nobody talks about enough
Everyone treats accessibility like a binary: compliant or not. The more useful question is what you’re optimizing for on any given piece of content. A city council recording that ten people watch a year doesn’t need the same review rigor as a training video every new hire has to complete. Budget and staff time are real constraints, and pretending otherwise doesn’t help anyone.
What actually moves the needle is prioritization: caption your highest-traffic and highest-risk content first, document why you made that call, and treat individual accommodation requests as the trigger to move something up the queue immediately. The organizations that get into trouble aren’t usually the ones with imperfect captions. They’re the ones with no process and no paper trail showing they tried.
— Ryan
Adding live captions without the stenographer bill
If your organization runs live events, church services, medical consultations, or legal proceedings, Live Caption AI gives you a way to put captions in front of an audience in real time without renting hardware or booking a stenographer weeks in advance. Attendees scan a QR code, and their phone becomes a caption receiver, with support for 29 languages translated from a single source language per session.

The domain-specific models are tuned to catch terminology that generic speech-to-text tools miss, whether that’s medical shorthand or legal phrasing, and every session exports to TXT, SRT, JSON, or PDF for your records. Audio is never stored, and transcripts stay visible only to the organization that ran the session. Live Caption AI’s Professional plan runs $19.99 a month, well under what a single stenographer session often costs. It’s worth repeating clearly: using Live Caption AI doesn’t make your organization ADA-compliant on its own. Whether captions satisfy your specific effective-communication obligation is a legal determination for your organization to make, not a claim any captioning tool can make for you. What it does offer is a practical way to widen access and keep more of your audience engaged in real time. Start with a free trial and see how it fits your next live event.