Microsoft Research and coauthors just tested an AI-use policy inside ICML 2026, a machine-learning conference with over 24,000 submissions and 17,886 reviewers. Assignment to a policy prohibiting LLM use had near-zero effect on paper decisions, scores or reviewer confidence versus a policy permitting limited assistance.
Why the null result is the finding
I have written an AI-use policy for a client. I have also watched what people type into a chat window at 11pm on a Thursday when Friday is the deadline. The policy in the deck and the behaviour in the chat window are two different objects.
The paper puts numbers on the gap. On an anonymous post-survey of 1,486 reviewers, 22.5% of the prohibited group admitted using an LLM anyway. Among the permitted group, 36.5% admitted at least one use their own policy explicitly banned. Both are self-report on an anonymous survey, so read them as floors.
The detection-floor number is starker. The paper's own estimate is that the probability of catching a single false positive across roughly 53,000 conservative-policy reviews, under family-wise error control, is approximately 1 in 10,000. No single-review classifier can promise more at that N.
What the COO in Rotterdam already knew
The scene sitting in my head is a COO at a mid-cap manufacturer, in April. She had rolled out an AI-use policy in January. Two pages, three tiers, a training module. I asked how she would know if it was working.
She listed three tools her team had bought to monitor traffic. Then she stopped and said the quiet part: none of them tell me who did the thinking.
She had understood the mechanism a few months before the paper arrived. Her policy was moving the compliers. Among the ICML reviewers who did not use an LLM, 61.7% named policy compliance as their reason. So the policy does move behaviour. It moves the third of the workforce whose default is to follow the rule on the wall. It stays invisible to the people who use the tool silently, and irrelevant to the people who never wanted it.
The paper's own reviewers proposed the reframe unprompted.
"many proposed that enforcement should focus on holding reviewers accountable for the reviews they sign their name to, rather than policing LLM use."
The review is what the person signed. The email is what the person sent. The plan is what the person approved. That is what accountability can attach to at scale.