Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
Published 16 Sept 2026arXiv:2604.27093
Updated 12 h ago · first seen 15 Sept 2026
paper_01M2JK0TRG4PR94K958ZZYPKDS
Abstract
Current LLM safety alignment techniques improve model robustness against adversarial attacks, but overlook whether and how LLMs can recover helpfulness when benign users clarify their intent. We introduce CarryOnBench, the first interactive benchmark that measures whether LLMs can revise their interpretation of user intent and recover utility, while remaining safe through multi-turn conversations. Starting from 398 seemingly harmful queries with benign underlying intents, we simulate 5,970 conversations by varying user follow-up sequences, evaluating 14 models on both intent-aligned utility and safety. CarryOnBench yields 1,866 different conversation flows of 4--12 turns, totaling 23,880 model responses. We design Ben-Util, a checklist-based metric that evaluates how well each model response fulfills the user's benign information need using atomic items. At turn one, models fulfill only 10.5--37.6% of the user's benign information need. When the same query includes the benign intent upfront, models fulfill 25.1--72.1%, confirming that models withhold information due to intent misinterpretation, not limited knowledge. With benign clarifications in multi-turn conversations, 13 of 14 models approach or exceed this single-turn baseline, yet recovery cost varies across models. We identify two failure modes invisible to single-turn evaluations---unsafe recovery, where a model updates at disproportionate safety cost and redundant recovery, where a model recycles prior responses rather than providing new information. Moreover, conversations converge to similar harmfulness levels regardless of how conservative the model starts. These findings expose a gap that single-turn evaluations miss---whether a model is appropriately cautious or simply unresponsive to clarified user intent.
Organizations
Organizations 0
No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.
Models
Models introduced or described 0
Inbound described_by relations from model cards and documentation.
No model links this paper yet
Datasets
Datasets used 0
No dataset relation recorded.
Benchmarks
Benchmarks used 0
No benchmark relation recorded.
Code
Repositories & frameworks 0
No repository linked.
Timeline
Timeline 2
- Property changedPaperUseless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations: published at changed from 2026-09-15T04:00:00+00:00 to 2026-09-16T04:00:00+00:00
Published15 Sept 2026→16 Sept 2026arxiv - New paperPaperUseless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
New paper: Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
arxiv
Sources
Sources 2
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.