AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries?

Sep 20, 2026·
Jiaxi Yang
,
Chaewan Chun
Jason S. Lucas, Ph.D., MPH, M.Sc.
Jason S. Lucas, Ph.D., MPH, M.Sc.
,
Yuchen Yang
,
Dongwon Lee
· 1 min read
Abstract
Safety alignment teaches models to refuse harmful requests, but an over-aligned model also refuses benign ones that merely resemble harmful queries. AOR-Bench examines this over-refusal behaviour in large audio language models, where the speech channel adds failure modes that text-only evaluation cannot surface: accent, prosody and acoustic ambiguity can all push a pseudo-harmful query across a refusal boundary. The benchmark measures how often audio models decline queries that are in fact safe, and what acoustic and linguistic factors drive those refusals.
Type
Publication
In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP ‘26)
publication

Accepted to EMNLP 2026, Budapest, Hungary, 24–29 October 2026.

AOR-Bench asks a question that safety evaluation usually skips: not whether a model refuses harmful requests, but whether it refuses safe ones that happen to look harmful.

Over-refusal is a real cost of alignment. A model that declines benign queries is less useful, and the burden does not fall evenly — speakers whose accent, dialect or phrasing sits further from the training distribution are more likely to be refused. In the audio setting that risk grows, because prosody and acoustic ambiguity give the model more ways to misread intent than text alone provides.

Authors
Jason S. Lucas, Ph.D., MPH, M.Sc.
Authors
Tenure-Track Assistant Professor & Director, Secure and Ethical AI Lab (SEAL) — CU Boulder

I completed my Ph.D. in Informatics at Penn State University (defended May 2026; formal conferral August 2026), where I conducted research at the PIKE Research Lab under Dr. Dongwon Lee and the College of IST. Starting August 2026, I will join the Department of Information Science at the College of Communication, Media, Design, and Information (CMDI), University of Colorado Boulder, as a Tenure-Track Assistant Professor and founding Director of the Secure and Ethical AI Lab (SEAL). My research advances trustworthy, safe, and equitable AI for the world’s languages and communities — spanning multilingual NLP, low-resource and dialectal language technology, AI safety, and information integrity, with work extending across 70+ languages. I have authored 14+ peer-reviewed papers with 315+ citations in premier venues including ACL, EMNLP, NAACL, ICML, KDD, and IEEE.

My doctoral research focuses on bridging the digital language divide through transfer learning, classification (NLU), generation (NLG), adversarial attacks, and developing end-to-end AI pipelines using RAG and Agentic AI workflows for combating multilingual threats. Drawing from my Grenadian background and knowledge of local Creole languages, I bring a global perspective to AI challenges, working to democratize state-of-the-art AI capabilities for underserved linguistic communities worldwide. My mission is to develop robust multilingual multimodal systems and mitigate evolving security vulnerabilities while enhancing access to human language technology through cutting-edge solutions.

As an NSF LinDiv Fellow, I conduct transdisciplinary research advancing human-AI language interaction for social good. I actively mentor 5+ research interns and teach Applied Generative AI courses. Through industry experience at Lawrence Livermore National Lab, Interaction LLC, and Coalfire, I bridge academic research with practical applications in combating evolving security threats and enhancing global AI accessibility. I see multilingual advances and interdisciplinary collaboration as a competitive advantage, not a communication challenge. Beyond research, I stay active through dance, fitness, martial arts, and community service.