>human element
← All writing

The Audit Mirror You Already Own

By

Run an AI baseline in the shadow of your senior people and you have built an instrument that measures their drift.

Every executive I coach already suspects that their Friday afternoon calls look different from their Monday morning ones. They just cannot prove it. Nobody has ever put a number on the drift in their own judgment, because they had nothing clean to measure against.

That changed last week. And the number is bigger than I would have guessed.

The head of credit who could not explain her Fridays

A head of credit at a mid-market specialty lender in Amsterdam, on a coaching call two weeks ago. She had spent a fortnight trying to explain why her team's late-week approvals were defaulting at a higher rate than their Tuesday morning approvals. Same underwriters. Same policy. Same file templates.

She said the sentence I hear most from senior people about their own judgment. "We are the same people all week. It should not move."

I told her that nobody had ever put a number on it because nobody had a clean baseline. Same job, same call, done by something that does not have a Friday.

Then I sent her the paper.


What Yonsei measured

Kichang Lee and JeongGil Ko at Yonsei University used the Korean Baseball Organization's 2024 switch from human umpires to a fully automated ball-strike system. Same rulebook. Same strike zone. Same players. 1,216,246 pitches across 2021 to 2026, with the two human-only seasons as the baseline and the machine seasons as the benchmark.

They zeroed in on where discretion lives: 209,976 taken pitches within a quarter-foot of the rule-zone edge. Then they asked whether the count on the batter shifted the human call.

0–2 was associated with a -17.17 percentage-point effect and 3–0 with a +6.61 percentage-point effect.

Under the machine the same effects were 0.33 and 0.37 percentage points, and neither survived false-discovery-rate correction. The catcher-framing signal that has driven a decade of MLB analytics: 1.13 percentage points under humans, 0.00 under the machine.

The umpires were not cheating. They were under pressure and they moved. They did not see themselves moving.

That is the part senior people find hardest to sit with. The pressure moves the call, in numbers this big, and none of them see it happening.

My client in Amsterdam is now sitting with a different question. Whether to run a model baseline beside her human underwriters, and read the delta every Friday. Because the moment she stands up that baseline, she owns an instrument that measures pressure-shaped drift in her own team's judgment. The stomach for that reading is the actual leadership skill.

Source · Auditing Contextual Bias in Human Ball-Strike Calls Using KBO's Automated Umpiring Transition · Lee, Kichang · Ko, JeongGil · Yonsei University · 2026
Fatjon Tony Kalemaj is an AI Strategist and Consultant who helps organisations become AI-enabled. He is also the founder of Human Element, a space for practitioners and thinkers navigating the AI era. He has been using AI in production work since 2023 and believes the most valuable thing in the AI era is knowing what to ask of it.
> the writing

Ideas, observations,
and honest takes.

No hype. No tools-of-the-week. Just the work, explained clearly, from someone who uses this every day and cares whether it is accurate.

No spam. Unsubscribe any time.

>human element
© 2026 Human Elementhello@humanelement.tech