Program/Track A/A.6/The Hidden Cost of Safety: Measuring Over-Refusal in Inference-Time Defenses for Multimodal LLMs
The Hidden Cost of Safety: Measuring Over-Refusal in Inference-Time Defenses for Multimodal LLMs
Bulat Nutfullin, Dmitry Namiot
15m
Inference-time safety defenses for multimodal large language models are com- monly tested on adversarial prompts, with less attention to benign traffic. An earlier proxy analysis of our archive appeared to show mass refusal. We re-audit that claim by separating direct system refusals, semantic labels, and generation failures. The canonical paired grid contains 28,000 outputs in 56 cells; 54 are estimable. Pooled at the output level, with the estimated number of valid out- puts as the denominator, the refusal estimate over the 56-cell grid is 0.5158%; the largest cell estimate is 3.2389%. These estimates do not support the ear- lier mass-refusal interpretation, but do not establish precise absolute rates: the semantic judge saw no questions, rare-class repeatability is weak, and the sys- tematic sample does not support design-based confidence intervals. The robust result is a construct correction: safety-language proxies can overcount answered requests and miss direct system refusals. Benign-utility audits must therefore bind measurement to the user-visible output, keep proxy signals separate, and report provenance and rare-class uncertainty.