Duke Accessible-U
- AccessAudit F1
- 96.3%
- axe-core F1
- 72.7%
- Coverage delta
- +85%
AccessAudit is developed as a usable SaaS product and as a platform for studying hybrid accessibility evaluation: automated detection, AI-assisted explanations, visual evidence and guided human verification.
Current benchmark reports show early detection differences against axe-core. They are useful for proof of concept and research planning, but final results require instance-level expert validation across more pages.
The project is intentionally framed around measurable support for accessibility work, not broad claims that automation can replace expert review.
How much verified WCAG coverage can a hybrid scanner provide compared with an axe-core baseline?
Do AI explanations and draft fix suggestions help users understand and resolve accessibility issues more effectively?
Can structured manual checks reduce uncertainty for findings that require interaction, context or human judgement?
The safest research direction is to separate what the tool detects automatically, what AI explains, and what still needs human review.
Preliminary results compare detected WCAG criteria with known or manually verified issues. This is useful for early validation but is not the final research dataset.
The next step is expert review of individual findings, including false positives, false negatives and severity differences.
A small user study can compare issue understanding, time-to-fix, confidence and perceived usefulness with and without AI guidance.
This distinction matters for research, product honesty and user trust. Missing alt attributes can be detected by rules; poor or misleading alt text requires interpretation and review.
Most findings come from axe-core and AccessAudit custom rules. These checks are deterministic and easier to evaluate.
Claude Haiku is used for plain-language explanations, WCAG context, user impact and draft remediation guidance.
Claude Sonnet vision can support selected semantic checks, such as reviewing whether existing alt text is meaningful.
Feedback from developers, researchers and accessibility practitioners can help improve the benchmark design and the guided verification workflow.