Penn State and Seoul National University's CTAI ran 599 U.S. adults through five conditions of an AI chatbot that varied how often it refused to answer. Users were most satisfied with correct answers, then with confident wrong ones they could catch, and least satisfied with the refusals.
The safety consensus billed you and did not tell you
Every enterprise LLM deployment guide I read in 2026 gives the same instruction. If the model isn't sure, make it refuse. Refusal is treated as the free safety move. Nobody in those guides measures what it costs.
This paper measured it. All within-condition pairwise gaps were significant. Users knew the hallucinations were less accurate. They still found them more useful than "I cannot answer that."
Then the paper killed the usual fallback. Adding an explanation to the refusal rescued satisfaction when refusals were rare (Δ=.68, p<.001). When refusals were frequent, the explanation did nothing measurable (Δ=.18, n.s.).
A CFO in Zurich, in June
I was coaching a CFO at a mid-sized asset manager this June. Their internal knowledge assistant had been live for a month with the standard safety-first system prompt. The one that tells the model to refuse when unsure.
He pulled up a query in front of me. He asked whether a specific carve-out applied to a deal his team was structuring. The assistant refused and offered to help him rephrase.
He sighed. He opened a chatbot on his personal laptop and typed the same question. Confident answer. Wrong on one clause, right on the rest. He caught the wrong clause because he already knew that part. He used the right part as a starting point.
"That one gave me something," he said, pointing at the personal window. "This one keeps sending me away."
I did not have a clean answer for his frustration at the time. I thought the internal assistant was doing the right thing.
The Penn State study calls his profile high need for cognitive closure. High-NFCC users rated no-refusal systems Δ=.41 higher overall (p=.003). Their satisfaction gap between genuine and hallucinated answers also narrowed the most.
That is what a senior professional looks like on that scale. They came for a decision. A refusal is a subtraction from the work they came to do.
"These findings reveal a tension between hallucination avoidance and user satisfaction and highlight the importance of designing balanced refusal strategies."
Every refusal your assistant makes is a satisfaction cost. Your safety prompt is charging you one you did not budget for.