USC and Vanderbilt put 164 people through a scripted sales dispute with one of three mediators: none, an LLM, or an untrained crowd worker. The LLM cooled negative emotion turn by turn. The human in the same seat did not.
What that finding actually is
The paper is careful about the comparison. Its human group is 50 Prolific workers with no specialised mediation training. The abstract calls them novice human mediators and means it.
That happens to be the honest match for what most enterprises are actually staffing into dispute contexts. The L1 support agent. The HR generalist. The store manager taking the escalation call. Nobody in that queue has mediator credentials. So when Hale and Gratch report F(2,154)=4.57, p=.01 on the condition-by-time interaction for negative affect, and a within-AI paired t of 3.20, p<.01, they are saying something plain. The untrained human in the seat did not do the cooling we assume is the whole reason we put a human there.
The louder number is the trade-off suggestions. AI mediators sent M = 0.38 per message versus 0.07 for the human, t = 8.00, p < .001. Roughly five and a half times as many. The joint-gains interaction was only directional at β = 0.37, p = .056, and impasse rates did not differ across conditions. So the question the paper made me care about is who is doing the empathy-adjacent work when the room gets hot.
The Tuesday I stopped believing the routing rule
A VP of Customer Support at a B2B SaaS firm in Munich showed me his escalation routing in June. AI was allowed on tier-two triage and blocked from tier-one. The reason was written into the policy: customer needs to feel heard.
I asked him to open his last ten tier-one transcripts. We read them together in silence.
Not one trade-off proposed. Not one line that named the customer's frustration and offered a smaller thing back. Polite acknowledgements. Next-ticket handoffs. His L1 people were exhausted and quota-bound and doing the thing exhausted people do, which is close the ticket.
He looked up and said the quiet part out loud. The chatbot would have written the trade-off. His team, on paper the emotional layer of the product, was not producing the emotion work the policy was protecting.
The consensus story is empathy versus scale. The paper puts numbers on the harder story: the L1 human in the dispute chair is not doing the empathy work at all.
The policy called it empathy. The transcript called it ticket-closing.