Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts
Trilingual spoken hallucination detection, comparing what is detectable from raw audio against what survives transcription.
Trilingual spoken hallucination detection, comparing what is detectable from raw audio against what survives transcription.
A benchmark for over-refusal in large audio language models — how often they decline pseudo-harmful but benign spoken queries, and what drives it.
A position paper arguing that the dual curse of multilingual AI — 35% harmful generation and near-random reward-model accuracy in low-resource languages — cannot be fixed by …
DIA-HARM evaluates 16 harmful content detection models across 50 English dialects using 195K+ samples, revealing 1.4–3.6% F1 drops for fine-tuned models and up to 27% for zero-shot …
BLUFF is the largest multilingual fake news detection benchmark, spanning 79 languages with 202K+ samples. It introduces AXL-CoI for adversarial generation and mPURIFY for quality …
Chain-of-Interactions (CoI) introduces a novel multi-step framework that leverages LLMs' in-context learning capabilities for abstractive task-oriented dialogue summarization. …
GAMIC introduces a novel self-supervised learning approach for molecular in-context learning that combines graph neural networks with Morgan fingerprints to better capture …
Beemo introduces a novel benchmark featuring 6.5k expert-edited machine-generated texts across diverse domains from creative writing to summarization. Through comprehensive …
This work introduces semantic captioning for SQL queries, addressing the reverse operation of semantic parsing by translating SQL code into natural language explanations. Using …
This research from Penn State and KiNiT, benchmarks the effectiveness of 10 authorship obfuscation (AO) techniques against 37 machine-generated text (MGT) detection methods across …