Two questions about the AI you’re already using.
Out of the last twenty things it handed you (a draft, a summary, a recommendation, a report), how many did you push back on, rewrite, or throw out entirely?
And when you pushed back, were you right?
Most people can’t answer either one. I couldn’t have answered them either, and AI does a large share of the production work in my own practice. That inability is the finding. Those two numbers are close to the best available measure of whether your team is working with the tool or just passing its output along.
Both failure modes feel like good judgment
Teams get this wrong in two directions, and neither direction feels wrong from the inside.
Some accept whatever the system produces. Decision researchers call that automation bias. From the inside it feels like efficiency. The draft was fine. Why spend another hour on it?
Others reject machine output close to reflexively. That one’s called algorithm aversion, and from the inside it feels like rigor. I’m the professional here.
Both feel like judgment. Both degrade the work.
That’s why a single number won’t tell you anything. This spring a group of management researchers had a paper accepted at Academy of Management Perspectives proposing a pair of indicators instead: the override rate, meaning what share of the AI’s output your team reverses, and override accuracy, meaning how often the reversal turns out to have been correct. Fair warning on scope. That’s a proposed routine synthesizing existing evidence, not a field experiment with results attached. Treat it as a lens, not a verdict.
The lens works because neither number means much alone. A team that never overrides sounds efficient and might be asleep at the wheel. A team that overrides constantly sounds rigorous and might be defending turf. High override rate with low accuracy says you’re fighting the tool for no reason. An override rate near zero says nobody is actually looking.
Skills that stop getting exercised don’t hold
A group of Polish endoscopists, specialists with at least two thousand procedures behind each of them, were finding precancerous growths in roughly 28 percent of colonoscopies before an AI detection tool arrived. After three months working alongside it, on the days the tool was switched off, that number was closer to 22 percent. The study ran in The Lancet‘s gastroenterology journal.
Hold it loosely. One study, one country, one specialty, three months. One of the study’s own authors says it needs replication before anyone treats it as settled. But it’s the clearest picture available of what happens to a capability that stops getting used, and it comes from a field where somebody bothered to measure.
Marketing has nobody measuring. Nobody is tracking whether your team can still write a subject line without the assist, or whether anyone left can read a report and spot the number that doesn’t add up. That’s the quiet cost underneath the question of whether AI replaces marketing work: long before anything gets replaced, the judgment that used to check the work goes soft.
There’s a related finding worth carrying. Cognitive scientists at UC Irvine ran about two hundred people through fifty rounds of judging whether an AI was right, with its confidence signal rigged differently for different groups. People learned. They built an internal correction for whatever bias they were seeing. Except one group: the people whose AI sounded most certain exactly when it was most likely to be wrong. A substantial share of them never got there. One caveat, because it matters here. That was a lab task with feedback after every single round. Your Monday doesn’t work that way.
Which is the uncomfortable part, right? Today’s tools lean confident. Fluent, certain, wrong at a rate the tone never admits. The lab says you can learn to discount that. The lab also says fluency is the signal people have the hardest time discounting.
Build the log
Start the dumbest possible version of this. Three columns. What the AI produced. What you did instead. And the column everyone skips: a check, a week later, on which of you turned out to be right.
Pick one recurring task rather than everything at once. The weekly report. The draft emails. The ad copy that goes out on repeat. Repetition is what makes the numbers mean anything.
Two patterns to watch for. If the second column stays empty, you’re not collaborating, you’re forwarding. And if the second column is always full but the third keeps saying the machine had it right, that’s aversion wearing rigor as a costume. Both are fixable, but only once the numbers exist.
I’ll name my own gap here. The content system I run at Auspicious operates under a standing rule: verify state before claiming it. No agent gets to say a page is live or a ranking moved because an earlier report said so. Fresh pull, every time. That rule is calibration infrastructure. But everything the system produces crosses a human review gate, and nobody counts how often the machine gets reversed at that gate, or checks a week later whether the reversal was right. I’ve been calibrating by feel. The exact thing worth not trusting.
One habit on top of the log. When the model underneath your tools gets swapped, and it will, probably without an announcement, your map expires. The discount rate you learned on last quarter’s version is a guess now. Run the check again. It’s the same reason speed from AI isn’t the same thing as clarity: the tool moves, and only a deliberate check tells you where you stand relative to it.
That’s what continuous actually means. It was never a setting you lock in once.
If the review gate in your business is one person’s instinct and nobody has ever written down what it catches, that’s worth a conversation. Here’s how I work with businesses on it.
