Most companies still talk about AI as if the main question is technical: Is the model accurate enough to trust? A more important question is human: Does the worker know when not to trust themselves?
That is the sharper lesson from a recent Management Science peer-reviewed study, which found that AI delivers the biggest gains not simply when people are less skilled, but when they are better calibrated about their own ability. The workers who benefit most are the ones who can judge, with some honesty, where their own judgment is likely to fail.
The study tested 732 people on a simple but revealing task: deciding whether faces in photos were over age 21, sometimes with an AI confidence score and sometimes without it. Average performance improved with AI help, and lower-ability participants gained more than stronger ones. That result lines up with a growing body of evidence showing that AI can act as a leveling technology.
In a widely discussed Science paper on generative AI and professional writing, it was found that ChatGPT meaningfully increased productivity and quality, with the largest gains going to weaker initial performers. In a large field experiment published in the Quarterly Journal of Economics, it was discovered that AI assistance raised customer-support productivity by 15% on average and helped novice and low-skilled workers far more than top performers.
What makes the Management Science paper especially useful is that it separates ability from self-knowledge. Two workers can have the same baseline skill and still get very different value from the same AI system, because one is better at recognizing when the system is likely to outperform them.
That sounds obvious until you look at how most organizations deploy AI. They benchmark models, buy licenses, redesign workflows, and train people on prompts. Then they assume workers will naturally learn when to defer to the tool and when to push back. The evidence suggests that assumption is shaky.
This matters because miscalibration cuts in both directions. Overconfident workers ignore useful AI advice. Underconfident workers defer when they should trust their own judgment. In both cases, the value of augmentation leaks away. The result is that companies can install a capable system and still see mediocre gains because the real bottleneck is not the model but the user’s sense of their own competence.
That is a very different diagnosis from the usual story that AI disappoints only when the technology is immature or the workflow is clumsy. It also changes the debate about inequality.
One of the most promising ideas in the economics of AI is that these systems could spread expert performance more widely instead of simply rewarding the already advantaged. One NBER paper argues that AI could help restore middle-skill work by embedding expertise into tools that broaden access to capability.
The Management Science findings support that view, but they add an important condition. AI can narrow performance gaps, yet it does not do so automatically. The study shows that when people use AI with their actual, imperfect self-beliefs, inequality falls. With perfect calibration, it would fall much more. The gains are real, but the full equalizing effect remains trapped behind a human judgment problem.
That insight should push leaders to rethink training. For years, companies have treated worker development mainly as a matter of adding skill: more instruction, more certification, more exposure to best practices. AI makes a different target newly valuable. Workers need help estimating uncertainty, reading signals, and recognizing when they are in an edge case. That is less glamorous than frontier-model talk, but it may be far more practical.
A call center, insurer, hospital, or legal team may not be able to turn average employees into experts overnight. It may, however, be able to make them much better at knowing when the machine is likely right and when it deserves skepticism.
There is reason to think calibration can improve. A recent study in Futures & Foresight Science found that an interactive training app reduced overconfidence and improved calibration in under 30 minutes. The underlying idea is not new. Decades ago, Sarah Lichtenstein and Baruch Fischhoff showed in classic research on calibration training that feedback can make people better judges of their own probabilities.
More recent work on automated calibration training for forecasters points in the same direction. Calibration is not magic, and it is not a cure-all, but it does not look like a fixed trait either.
That creates a more useful agenda for management than the stale choice between “train people” and “deploy AI.” The better answer is to train people to work with AI. Teach them how to assess confidence, how to compare their judgment against model output, how to recognize repeated blind spots, and how to change behavior when feedback shows they are off. The goal is not blind trust in the machine — it is disciplined partnership.
AI is often described as a force multiplier, but that phrase hides the real mechanism. In practice, these systems reward workers who can tell the difference between confidence and competence. As companies race to embed AI into everyday work, the most underrated advantage may be brutally simple: knowing when you are probably wrong.












