Skip to research content
Research direction

Evaluating AI-assisted accessibility workflows.

AccessAudit is developed as a usable SaaS product and as a platform for studying hybrid accessibility evaluation: automated detection, AI-assisted explanations, visual evidence and guided human verification.

Benchmark evidence

Preliminary reports, not final proof.

Current benchmark reports show early detection differences against axe-core. They are useful for proof of concept and research planning, but final results require instance-level expert validation across more pages.

Known accessibility training page+35.8 pp

Duke Accessible-U

AccessAudit F1
96.3%
axe-core F1
72.7%
Coverage delta
+85%
Open PDF report
Accessibility demo page+23.0 pp

Deque Mars Demo

AccessAudit F1
66.7%
axe-core F1
50.0%
Coverage delta
+69%
Open PDF report
Before/after WAI demo website+11.1 pp

W3C BAD Demo

AccessAudit F1
51.9%
axe-core F1
41.7%
Coverage delta
+88%
Open PDF report
Research questions

A small, realistic research scope.

The project is intentionally framed around measurable support for accessibility work, not broad claims that automation can replace expert review.

RQ1

Detection coverage

How much verified WCAG coverage can a hybrid scanner provide compared with an axe-core baseline?

RQ2

AI remediation support

Do AI explanations and draft fix suggestions help users understand and resolve accessibility issues more effectively?

RQ3

Guided verification

Can structured manual checks reduce uncertainty for findings that require interaction, context or human judgement?

Methodology

Evidence should be separated from product claims.

The safest research direction is to separate what the tool detects automatically, what AI explains, and what still needs human review.

Criterion-level benchmark

Preliminary results compare detected WCAG criteria with known or manually verified issues. This is useful for early validation but is not the final research dataset.

Instance-level validation

The next step is expert review of individual findings, including false positives, false negatives and severity differences.

User evaluation

A small user study can compare issue understanding, time-to-fix, confidence and perceived usefulness with and without AI guidance.

AI transparency

AI is used as assistance, not as the main scanner.

This distinction matters for research, product honesty and user trust. Missing alt attributes can be detected by rules; poor or misleading alt text requires interpretation and review.

Primary detection

Most findings come from axe-core and AccessAudit custom rules. These checks are deterministic and easier to evaluate.

AI explanations

Claude Haiku is used for plain-language explanations, WCAG context, user impact and draft remediation guidance.

Vision assistance

Claude Sonnet vision can support selected semantic checks, such as reviewing whether existing alt text is meaningful.

Validation collaboration

Interested in reviewing the prototype?

Feedback from developers, researchers and accessibility practitioners can help improve the benchmark design and the guided verification workflow.