Between 2 and 5 August, five authorities arrived at the same sentence. Brian Solis, ServiceNow’s innovation head, in his own essay [1]. David Lancefield in Harvard Business Review [2]. Tomas Chamorro-Premuzic in Forbes [3]. Keith Ferrazzi and Wendy Smith in Fortune [4]. The World Economic Forum’s governance team [5]. Each named judgment the human function AI cannot replace. You have likely scrolled past at least one of them.
Solis wrote the sharpest version — the sixth of seven questions he says every CEO now faces — judgement — recapping a Drucker Forum session from last November [1]: “How will we protect and expand human judgment?” [1]. His answer only addresses “judgment” — he defined what AI can recommend, what it can decide, what it can execute [1]. For expand, he offered nothing. He answered half the question. He left the other half blank.
Four of the five ship frameworks. Lancefield maps four modes of AI collaboration [2]. Chamorro-Premuzic gives a selection model — hire for judgment [3]. Ferrazzi and Smith give five mental models for working with agents [4]. The Forum hands boards five priorities [5]. Every one tells you what to do with the judgment you already have: guard it, hire for it, calibrate trust around it. None explained how to build it. The research below prices that missing half.
📬 Every week I translate research into actionable plays you can run in your organization to discover the ROI hidden in your AI spend. Receive the AI Playbook in your inbox and use it as part of your strategic plan for the week.
The Erosion Has Numbers Now
Two research teams measured judgment loss under AI directly — and turned the anecdote into numbers:
The Wharton experiments (2026). Shaw & Nave ran 1,372 participants through 9,593 preregistered reasoning trials — a preprint, meaning other scientists have not yet formally reviewed it [6]. When the AI answered correctly, participants’ accuracy rose 25 points. When the AI answered incorrectly, their accuracy fell 15 points [6]. Access to the AI also lifted their confidence from 65.3% to 77.0% — an 11.7-point jump — and their confidence held even as the AI’s errors piled up [6].
The Radiology reader study (2023). Dratsch and colleagues, in the journal Radiology, asked 27 radiologists to evaluate 50 mammograms each with a purpose-built, simulated AI [8]. On the 12 sabotaged cases, inexperienced readers fell from 79.7% correct to 19.8% — a 59.9-point drop [8]. The very experienced fell too — 82.3% to 45.5%, a 36.8-point drop [8]. Experience softened the collapse. It did not stop it.
When the machine was right, people looked sharper. When it was wrong, they had stopped checking.
Human-factors researchers documented this long before today’s tools. Parasuraman & Manzey’s canonical 2010 review concluded complacency is “found in both naive and expert participants and cannot be overcome with simple practice” [10].
Hospital clinicians now report the same experience on the job. A CHI 2026 study — CHI, the leading conference on how humans interact with computers — followed a real AI deployment for 12 months across a multi-site hospital system in five North American locations: 42 participants, 15 of them radiation oncologists [9]. The authors watched early efficiency gains hide the erosion they came to call intuition rust — the gradual dulling of expert judgment [9]. A senior radiation oncologist supplied the first of two remarks the paper presents side by side, without attributing both to one speaker: “My intuition is rusting,” and “I’ve no idea why the AI suggested it, but I accept it anyway.” [9].
Executives report seeing it too. BCG surveyed 70 C-suite leaders and senior executives worldwide, and the leaders ranked judgment and decision making as carrying the highest de-skilling risk score of all skills [7]. Half report already observing de-skilling. Almost 90% cite overreliance on AI outputs without stress testing or challenge, and 53% cite slower development of junior talent [7]. Only one in ten companies has an organisation-wide strategy to address any of it [7]. Every figure is what leaders report seeing — perception, not measured incidence [7].
Twenty-Three Years of Measuring Spend Never Measured Judgment
Executives tell the same story with their budgets. 69% of firms actively use AI [11]. Yet when the Bank of England, the Atlanta Fed, Stanford and Germany’s Bundesbank jointly asked roughly 6,000 senior executives — more than 90% of them C-suite, across the US, UK, Germany and Australia — more than 90% reported no impact of AI on their firm’s employment over the past 3 years, and 89% reported no impact on labour productivity [11]. The same executives still predict AI will lift their productivity 1.4% over the next 3 years [11].
The Federal Reserve Bank of Atlanta measured the gap between belief and reality. CFOs told the Fed that AI raised their productivity 1.8% in 2025. The Fed then computed what those same CFOs’ own revenue and headcount figures implied: 0.6% [12]. Executive belief runs three times ahead of what the executives’ own numbers show. That is the gap this Call names: everyone feels the return; almost nobody can find it in the arithmetic.
The US Census asked firms using AI what they actually changed after buying the technology [13]:
64% changed nothing at all — no new training, no new workflows, no institutional adjustments [13].
About 15% trained staff, and about the same share developed new workflows [13].
7–8% made the deeper shifts, including reorganising their data so AI can use it, and the supporting systems and equipment AI needs to do its work [13].
Firms bought the tool. Most never retooled.
If you cannot say whether your people’s judgment improved, you cannot say what your AI spend returned.
Economists have known this for 23 years. In 2003, Brynjolfsson & Hitt measured it across 527 large US firms: the organisational investment that makes computing pay “may be up to 10 times as large as the direct investments in computers” [14]. The payoff also compounds slowly — up to 5 times greater when the researchers measured over 5-7-year horizons rather than 1 [14].
In 2012, the American Economic Review published the explanation. Bloom, Sadun and Van Reenen studied more than 11,000 UK workplaces and asked why US-owned firms got more out of computers [15]. Management explained it. When a US multinational doubled its computing, productivity rose 6.3%; when a non-US peer did the same, 4.6% — a 1.7-point gap [15]. Tougher people-management practices explained it; once the researchers accounted for management quality, the American ownership advantage disappeared entirely [15].
The 2026 evidence repeats the lesson for AI adoption. The OECD — the Organisation for Economic Co-operation and Development, the research body for the world’s advanced economies — compared AI adopters against non-adopters across 15 countries [16]. The raw advantage looks impressive: 7.7% higher productivity in France, up to 31% in Belgium [16]. Then the researchers controlled for the quality of each firm’s people and technology — and in 8 of 10 countries the AI advantage vanished. Only 2 of 10 kept a significant edge [16].
The same European firm data puts a price on the fix. When a firm moved 1 extra percentage point of spending into training, AI’s productivity effect rose about 5.9%. The same point spent on software and data bought 2.4% — training beat it by more than double. Spent on R&D or machinery: nothing measurable [17]. One honest clause: the researchers measured this as an interaction across firms, not a controlled split, and the effect is statistically zero for firms under 50 employees [17].
Here is my analysis of that record: The ROI literature has been pointing at the human variable for 23 years. It measures training spend, specialist headcount, management indices, workflow redesign. Not one study measures whether the judgment improved — or what it returns when it does. [14][15][16][17]
The Best Causal Evidence: Gains Concentrate Where Judgment Matters Least
The strongest objection deserves the floor before you hear the answer. Brynjolfsson, Li and Raymond published the best causal study in the field in the Quarterly Journal of Economics, 2025. They followed 5,172 customer-support agents — human agents, people answering customer chats, not AI agents — at one Fortune 500 software firm as it rolled out an AI assistant team by team [18].
Because the firm staggered the rollout, the researchers could compare teams working with AI against near-identical teams still waiting for it — a natural experiment rather than a randomised trial [18]. Their finding: “Less experienced and lower-skilled workers improve both the speed and quality of their output, while the most experienced and highest-skilled workers see small gains in speed and small declines in quality.” [18].
The average human agent resolved 15% more issues per hour. The least experienced resolved 30% more — and a newcomer with 2 months of tenure performed like a veteran of more than 6 months [18]. The Bank of Korea heard the same from its survey of 5,512 workers: the less-experienced saved the most time [19]. And AI pays — the European researchers tie adoption to roughly 4% higher labour productivity [17]. Nothing in this Call argues otherwise.
AI made consultants 25.1% faster on the work it handles well. On the one task beyond its reach, their accuracy fell 19 points.
Those gains live in routine, well-mapped tasks. The sharpest test came from 758 consultants at BCG, the Boston Consulting Group, in a preregistered randomised experiment [20]. The researchers built two kinds of tasks and named the boundary between them the frontier. Inside the frontier means work AI handles well — idea generation, drafting, analysis. Outside the frontier means work that looks similar but sits beyond what the tool does reliably — in the experiment, a business-problem-solving task [20].
On inside-the-frontier work, consultants using AI completed 12.2% more tasks, worked 25.1% faster, and scored more than 30% higher on quality [20]. On the outside-the-frontier task, consultants without AI answered correctly 84.5% of the time; consultants with AI, 60% and 70.6% — a 19-point drop, because they trusted the tool past the line where it stops working [20].
The authors’ conclusion should be required reading for anyone designing and testing AI transformation strategies: “The effectiveness of AI in knowledge work will critically depend on human judgment—particularly to discern which tasks within the workflow are suited to leveraging AI augmentation and where human expertise should be prioritized.” [20].
In other words: the tool cannot tell you where its own competence ends. A person decides which work goes to the AI and which work a human must own — and that person’s judgment decides whether your firm collects the 25.1% or eats the 19-point drop [20]. Novices gain speed and quality on the routine work; your senior people’s judgment gates the non-routine, high-stakes decisions — which is why judgment is the lead lever of AI return, never the sole one.
The Missing Study: Restore Judgment, Then Measure the AI Return
Every strand of this literature stops at the same wall. I searched the corpus in ten languages, as of publication date, and four strands come back empty:
Restore judgment, then measure the dollar return: no study, in any language searched, as of publication date [20].
The erosion priced in currency: no researcher has put a dollar figure on dulled judgment — the closest proxies stop at accuracy collapse [8] and leader-perceived threat [7].
Reversibility: every study measures one moment in time; nobody has watched judgment come back and value follow [21].
Profit tied to decision quality: the field measures capability as money spent, heads counted, indices scored [14][15][16][17] — never as decisions improved.
One team came closest to testing the fix itself. In 2021, Buçinca, Malaya and Gajos ran a 199-person controlled experiment on a simple question: if you force people to pause and think before accepting an AI’s answer, do they stop over-trusting it? They do. The forced-pause checks “significantly reduced over-reliance” [21]. Participants disliked the extra effort, and the participants who most enjoy hard thinking gained the most [21]. Here is what that means for you: the training fix works — and science has never connected it to money. The experiment measured better decisions, on nutrition questions, with laypeople. No one has run the next study: the same intervention, inside a company, measured in how better judgment moves your P&L [21].
No study has watched judgment come back and value follow. The absence of the number is not the absence of the cost.
Researchers can show that firms investing in people get more from AI — the training multiplier [17], the vanishing premium once you account for human capital [16]. No researcher has shown the last step: rebuild the judgment, then watch the return arrive. The five authorities walked up to exactly that missing step and stopped [1][2][3][4][5]. I have priced this layer before, in The Judgement Premium, and counted who occupies it, in The Sophistication Gap.
Is the Judgment Behind Your AI Spend Improving — or Quietly Eroding?
That is the Monday question, and your current reporting most likely cannot answer it. Your AI dashboard names licences, adoption rates, training hours — every proxy the literature has measured for 23 years [14][15][16][17]. Nowhere on it: whether the judgment your spend depends on got sharper this quarter, or duller. Solis’s unanswered half of the question is sitting in your own deck right now [1].
Building judgment is a different discipline from guarding it, and the evidence already names its composite in these three moves:
Give every non-routine, high-stakes decision a named owner and a written standard. The consultant experiment showed where accuracy collapses when nobody owns the call: on outside-the-frontier work, the decisions beyond what the tool does reliably [20].
Make humans deliberate and challenge AI outputs before accepting them — at the very least as part of training. It is the one intervention measured to cut overreliance and a cost-effective way to keep reviewers sharp thereafter [21].
Put a judgment line in the AI budget and instrument it. Training investment multiplied AI’s productivity effect more than any other spending the European researchers tested [17].
You can answer the half Solis left blank — how will we expand human judgment? — by doing exactly this: name the owners and score their calls against the written standard, quarter over quarter [20]. Train the challenge habit and count how often your reviewers catch the machine’s misses [21]. Fund the judgment line and measure it like any other investment [17]. How to run that measurement inside your own P&L — the study your firm can run on itself — is territory for a coming Call.
Untrained judgment forfeits the AI return. Sharpened judgment captures it.
The AI Leadership Playbook
Strategic Questions (copy-paste ready for an email to your CFO and CHRO):
Are we in the 64% that changed nothing — what institutional adjustments have we actually made since deploying AI, and who owns the complete list? [13]
Which decisions in our AI-touched workflows are non-routine and high-stakes — outside the frontier — and have we named the person accountable for each one? [20]
What share of our AI budget develops judgment at all, and what would tell us this quarter whether the judgment improved? [17]
Your Next Plays (copy-paste ready for a direct report):
Run the 64% check. Inventory what changed since AI arrived against the Census categories — training, new workflows, data management [13]. Every empty row gets an owner.
Draw the frontier line. For one revenue-critical workflow, list every decision AI touches and mark each one inside the frontier (work AI handles well) or outside it (work beyond what the tool does reliably) [20]. Every outside-the-frontier decision gets a named human owner and a written standard.
Put a judgment line in the AI budget. Reclassify a slice of AI spend to judgment development and instrument it: challenge-before-accept checks measurably cut overreliance [21], and training carries the largest multiplier the European evidence found [17]. This is The People Bet, carried to the judgment layer.
📅 Book a complimentary 1:1 Strategy Session—45 minutes to start the conversation about mapping the revenue hidden in your AI spend.
📬 Every week I translate research into actionable plays you can run in your organization to discover the ROI hidden in your AI spend. Receive the AI Playbook in your inbox and use it as part of your strategic plan for the week.
Sources
Brian Solis, ServiceNow — The CEO Guide to Using AI: Don’t Automate Your Way Out of the Future, 2 August 2026 — https://briansolis.com/2026/08/the-ceo-guide-to-using-ai-dont-automate-your-way-out-of-the-future/
David Lancefield, Harvard Business Review — Don’t Let AI Flatten Your Leadership Style, 3 August 2026 — https://hbr.org/2026/08/dont-let-ai-flatten-your-leadership-style
Tomas Chamorro-Premuzic, Forbes — Developing Leaders For The Human-AI Age, 4 August 2026 — https://www.forbes.com/sites/tomaspremuzic/2026/08/04/developing-leaders-for-the-human-ai-age-why-potential-remains-key/
Keith Ferrazzi and Wendy Smith, Fortune — Your AI agent can be a teammate. But it still needs a boss, 4 August 2026 — https://fortune.com/2026/08/04/your-ai-agent-needs-a-boss-is-a-teammate/
Helle Bank Jørgensen and Marlen Heide, World Economic Forum — The hybrid boardroom, 5 August 2026 — https://www.weforum.org/stories/artificial-intelligence/hybrid-boardroom-ai-changing-role-directors/
Shaw & Nave, Wharton School Research Paper — Thinking—Fast, Slow, and Artificial, posted 2 February 2026 — https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6097646
Goel, Martin and Kaffe, BCG Institute — When Everyone Uses AI, Companies Risk Losing Critical Skills, 17 June 2026 — https://web-assets.bcg.com/pdf-src/prod-live/when-everyone-uses-ai-companies-risk-critical-skills.pdf
Dratsch et al., Radiology — Automation Bias in Mammography, 2023 — https://pubs.rsna.org/doi/10.1148/radiol.222176
Ehsan et al., CHI 2026 — From Future of Work to Future of Workers, 2026 — https://dl.acm.org/doi/10.1145/3772318.3791081
Parasuraman & Manzey, Human Factors — Complacency and Bias in Human Use of Automation, 2010 — https://journals.sagepub.com/doi/10.1177/0018720810376055
Yotzov et al., NBER Working Paper 34836 — Firm Data on AI, February 2026 — https://www.nber.org/system/files/working_papers/w34836/w34836.pdf
Baslandze et al., Federal Reserve Bank of Atlanta Working Paper 2026-4, 25 March 2026 — https://www.atlantafed.org/research-and-data/publications/working-papers/2026/03/25/04-artificial-intelligence-productivity-and-the-workforce-evidence-from-corporate-executives
Bonney et al., US Census Bureau CES Working Paper CES-26-25 — The Microstructure of AI Diffusion, April 2026 — https://www2.census.gov/library/working-papers/2026/adrm/ces/CES-WP-26-25.pdf
Brynjolfsson & Hitt, The Review of Economics and Statistics — Computing Productivity: Firm-Level Evidence, November 2003 — https://www.iecon.net/wp-content/uploads/2015/01/cpg.pdf
Bloom, Sadun and Van Reenen, American Economic Review — Americans Do IT Better, February 2012 — https://www.aeaweb.org/articles?id=10.1257/aer.102.1.167
Calvino, Costa and Haerle, OECD STI Working Papers — Digital technology diffusion in the age of AI, 21 January 2026 — https://www.oecd.org/content/dam/oecd/en/publications/reports/2026/01/digital-technology-diffusion-in-the-age-of-ai_7f11be5d/ebc2debe-en.pdf
Aldasoro et al., BIS Working Papers No 1325 — AI adoption, productivity and employment, January 2026 — https://www.bis.org/publ/work1325.pdf
Brynjolfsson, Li and Raymond, The Quarterly Journal of Economics — Generative AI at Work, 2025 — https://doi.org/10.1093/qje/qjae044
Suh, Oh and Kim, Bank of Korea Issue Note 2025-22 — Rapid Adoption of Artificial Intelligence and Its Productivity Effects, 2025 — https://www.bok.or.kr/eng/bbs/B0000354/view.do?nttId=10094689
Dell’Acqua et al., Organization Science — Navigating the Jagged Technological Frontier, March 2026 — https://doi.org/10.1287/orsc.2025.21838
Buçinca, Malaya and Gajos, CSCW — To Trust or to Think, April 2021 — https://www.eecs.harvard.edu/~kgajos/papers/2021/bucinca21trust.pdf


