Your AI program review is on the calendar. The dashboard you will carry into it is green—seats active, prompts climbing, adoption tracking to plan.
That dashboard has a blind side. Four unconnected sources landed the same warning on it in the third week of August 2026—no shared data, no shared authors, no shared funding1,2,3,4. Their warning: your people’s judgment is wearing down as a byproduct of AI use, and nothing on the dashboard would show it. That means your review this week measures the wrong layer.
That wrong layer is a question I left open two weeks ago, in Everybody Named the Problem. Nobody Teaches the Solution. How do you measure the judgment rebuild inside your own P&L? This Call walks straight at it. What converged, what it says about the metrics you report, and where measurement actually stops.
Four Unconnected Sources Converged in One Week: AI Use Wears Down the Judgment That Aims It
Every week brings executives a fresh AI warning. Most trace back to the same recycled report. These four are different, and you can check the difference yourself—four unconnected directions, four different methods:
1. The Harvard experiment (August 19). Harvard Business School researchers ran a randomized field experiment—228 evaluators making 3,002 screening decisions on real innovation submissions, with data gathered in 20245. Harvard Business Review carried the finding to executives on August 19, 2026, under a blunt headline: AI is undermining leaders’ judgment1. The next section explains the full result, including the half the headline omits.
2. The Stanford payroll panel (revised August 2026). Stanford economists Brynjolfsson, Chandar and Chen analyzed millions of records from Automatic Data Processing (ADP), which runs payroll for 26 million US workers. Employment of workers aged 22–25 in the most AI-exposed occupations sits 19% below where it would be if it had kept pace2. A year earlier the gap was 15%—a 4-point widening. Underneath: entry-level work built on codified knowledge is slowing, while tacit-knowledge work grows for experienced workers—patterns the authors call descriptive, not causal, scoped to their ADP sample2. The same panel anchored You Priced the License. You Never Priced the Trust. two weeks ago; the continuing slow down of junior hiring is this Call’s highlighted evidence from the same study.
3. The practitioner warning (August 19). Futurist Ross Dawson, an executive advisor last on this page in Your AI ROI Is Hiding in the Judgment Layer, said it plainly on his own Humans + AI podcast: “by default, the way we are building and using AI, and the way we’re bringing AI to organizations, is taking away from human judgment and capabilities.” The human learning loop, he argues, “will basically erode unless we augment it”3. Dawson advises on the remedy he names—nevertheless, his expertise lends his voice credibility.
4. The decision-habit warning (August 17). Philip Topham, who advises founder-led private companies with $20M to $250M in revenue, published the adjacent failure on August 17, 2026: operators running AI-era decisions with habits fitted to a company that no longer exists. “AI did not create this problem. It changed the track underneath it”4. His warning is about misfit, not erosion—and it counts toward one claim only: executive decision habits are the binding constraint.
Four methods, four directions, zero shared inputs, one week—and none of the four measures the judgment doing the deciding.
Zero shared data, zero shared authors, zero shared funding: the false-convergence check passes. And none of the four measures whether any organization’s judgment is improving or eroding—not the experiment, not the payroll panel, not either practitioner. The four name the variable. No metric in the class your AI program ships with carries a judgment line that scores it.
So if you are extracting ROI from AI already deployed, the erosion mechanism runs inside tools you already bought—assistants that ship agreeable by default, attached to your decision loops—while the adoption dashboard reads green. Later sections price what judgment is worth when a firm designs for it.
And if you are implementing now, the sharpest lesson from these data is that outcome can be determined by a design decision made at build time. A design review today costs a meeting. Discovering judgment erosion later costs you the rework of habits your teams have already formed around AI use.
The Same AI Made Decisions Better as a Black Box—and Worse When It Explained Itself
That design decision is the experiment’s own finding, and the half the headline omits comes first: plain black-box recommendations—a score with no story attached—improved decision quality in the 2024 Harvard experiment5. AI in the loop was not the problem.
The problem arrived with the explanation. With a persuasively written narrative attached, the same model pushed false negatives—good submissions the evaluators incorrectly rejected—up 14.9 points, against 8.4 for the black box5. That 6.5-point difference means the narrative nearly doubled the black box’s error—which is why the override step your people run is the asset in play. The 72 experienced evaluators performed a little better. Their false negatives rose 12.9–14.5 points5.
That override suppression is the mechanism, and the authors name it: narratives “suppress productive overrides by substituting persuasive text for independent verification”5. The evaluators stopped checking because the narrative sounded checked. Two boundaries keep that verification claim honest: this is a working paper—not yet through journal review—and it tested one innovation-screening task5. And the two eroding capacities you may have seen quoted are the HBR authors’ framing, not the study’s1.
Most importantly, the persuasion is not exotic—it ships in the box. Anthropic researchers, in work published in 2024 that examined their own models alongside competitors’, found that sycophancy—agreeing convincingly instead of answering correctly—is a general behavior of flagship AI assistants, a documented by-product of training on human approval6. The same by-product bit OpenAI in April 2025: the company shipped an update to 500 million weekly users and rolled it back within days. Its post-mortem opens plainly—”It aimed to please the user”—and concedes the prior reward signals had been “holding sycophancy in check”7,8. I unpacked how this happens in It’s Not Thinking. It’s Predicting.
That baseline is why the Harvard team’s design clause matters: effective collaboration “requires designs that preserve rather than supplant independent human judgment”5. So for your deployments the test is simple. Wherever your AI explains itself persuasively inside a decision loop, your people need a verification step where the narrative cannot talk humans out of their own review justifications.
Developers Said AI Made Them Faster—the Stopwatch Showed 19% Slower
You might think that asking your people is the cheapest measurement plan; but one lab checked it against a clock. The nonprofit AI-evaluation lab METR (Model Evaluation and Threat Research) randomized real tasks—AI allowed or not—across 16 veteran open-source developers working their own mature codebases9. The 2025 run covered 246 real tasks.
With early-2025 AI tools the developers took 19% longer—while believing AI had sped them up by 20%. They had forecast 24% efficiency before starting9. Their own read was wrong in sign, not just size—and that kills the cheapest measurement plan. Treat the finding as directional—one study, built on self-reported time inside a randomized design. Even so, it is the best evidence on record that perception fails exactly where you need it to hold.
If you never score the judgment aiming your AI, you cannot claim the return it produces.
That perception gap outlived the study. METR has since labeled the 19% slowdown historical; its February 2026 update believes developers are now likely faster with AI. That means the tools improved while people’s own estimates stayed wrong—the gap between felt speed and measured speed is the finding that survived every revision10. What the lab could not survive was adoption itself, and by early 2026 the drift was structural. Then between 30% and 50% of its developers—up to half the panel—were declining tasks they would have to complete without AI, and METR wrote that such self-reported speedups “can be quite unreliable”10. METR lost its control group.
That loss is coming for your organization too. Once your people will not work without AI, participation-based measurement stops working—and perception-based measurement goes with it. So repeatable measures will have to rest on external signals rather than self-report. Naming those signals for each decision loop is the real work. It is exactly the work your adoption metrics most likely skip.
Make Adoption the Target and Your People Will Give You Adoption—Not Better Decisions
That skipped work runs into the oldest law of measurement. In 1975 the economist Charles Goodhart wrote it down: “any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes”11. For an operator the law reads simply: manage to a measure, and the measure stops telling you the truth. Your people will deliver the measure. Make adoption the goal and you will get adoption: logins, prompts, token counts, whether or not a single decision improves. So a green adoption dashboard proves pressure, not return—and that leaves your decision quality unmeasured.
This is the line I drew in June 2026, in The Sophistication Gap: adoption is a headcount; sophistication is a capability. A headcount can tell you who has access; it cannot tell you whether outcomes improved. Behind that sophistication line sits an analysis by the University of Texas at Austin with KPMG, the global professional-services firm12. They examined 1.4 million real prompts—and found roughly 5% of users working at the level that changes how work gets done. That means your usage report and your outcome report are different documents.
Do you know how your employees are using tokens—or only that they are using them? The difference is crucial.
So most teams fall back on a survey, and organizational-learning research has already run that experiment for three decades. In 2004, researchers validated the field’s flagship instrument—the Dimensions of the Learning Organization Questionnaire, by Yang, Watkins and Marsick, N = 836—with half the sample held out to check the result: a seven-dimension questionnaire in which even the performance outcomes are respondents’ perceptions, a limit its own authors state. It’s also worth noting that the validators were also the framework’s creators13. The important take away is that surveys anchored on self-perception–exactly the mechanism the METR results employed–are structurally deficient. You need metrics that score decision quality, which is the true indicator of increased or decreased quality in your AI implementation’s outcome layer.
Scoring Judgment Itself Is a Limit Research Hasn’t Fully Explored—Certainly Not With AI as a Factor
So what is judgment worth in money? Researchers have already priced its presence—twice, in opposite directions, at the same kind of firm. In November 2020, Management Science published a field experiment at an automobile spare-parts retailer: merchants overrode the finished decisions of a working forecasting tool, and the overrides significantly reduced profitability14. In January 2026 the same lead researcher published the redesigned experiment. It invited merchants’ judgment into the inputs the forecasting tool consumed—and profit rose 4.92%, a result the randomized design lets the authors call causal15.
Those two results never contradict each other; the design decides the sign. Misplaced judgment destroys value; judgment as part of the design creates it. And neither experiment—in either direction—scores anyone’s judgment as a capability. Both price its presence. Nothing in either design measures its quality. For you, that means the profit lever is real and the gauge for it still does not exist.
Judgment designed into the inputs raised profit. Judgment fighting the outputs cut it. Same setting—the design decided the sign.
That missing gauge is the field’s honest edge, and I will state it carefully. In the research we reviewed and the searches we ran—re-verified at publication—no firm-level study yet measures the improvement of human judgment as a scored capability, a judgment-quality score that moves, and connects that movement to P&L. I am not claiming none exists. This is a limit that thus far has not been fully researched—certainly not with AI as a factor. The finding still points to judgment as a lever in profitability; how and where those levers can be designed to move P&L is the salient question.
So look at the instruments you already own. Adoption dashboards, maturity indexes, ROI reframes—every one measures the rollout, not the judgment aiming it. On the dashboard you will carry into this quarter’s review, the judgment line is still blank.
Ng, Hill, and Ibarra Say the Constraint Is the Organization, Not the Technology
That blank line raises the pace question, and the field’s teachers already answer it. Andrew Ng—the most-followed instructor in AI, and a man who sells AI training—told builders in the August 14, 2026, issue of The Batch that the operator skill is knowing “when to slow down and take longer in order to build more carefully”16. He aims that advice at engineering teams—so for you it transfers as a posture, not a policy.
That posture has a leadership component. Harvard’s Linda Hill, featured on NPR’s Marketplace in July 2026, named it: “many leaders first thought of this AI as a technology,” and “[t]hey now understand it’s really about people and cultural transformation.” Her readiness test is three capabilities: “Can your organization collaborate, experiment, and learn?” She advises companies on exactly this transition17. At London Business School, Herminia Ibarra—writing on her school’s own platform—put the verdict in her April 2026 headline: “The issue isn’t the technology itself – it’s humans’ ability to use it”18.
Harvard and London Business School, four months apart, one verdict: the constraint is the humans’ ability to use AI—not the technology.
Those three voices agree, and the prompt-level evidence points the same way. The 1.4-million-prompt analysis behind The Sophistication Gap, from March 2026, located the value in the roughly 5% who use AI to change how the work gets done—not in usage volume12. That study never connects sophistication to the P&L, and I will not connect it here. What it supports is a starting point: improved decisions, not usage volume, are the better predictor of improved outcomes. So measure how, not whether—as a starting point you can iterate on. Evaluate outcomes, where and how human judgment is designed into the workflows.
Sit down with your dashboard this quarter. Every line on it answers some version of one question: whether your people used the tools. Look for the line that answers the other one: whether their decisions got better. The design levers exist—the Harvard experiment names them. So does the external-signal principle—METR paid for it with its own control group. And the price is on record—Kesavan’s autobody retail chain experiment where the design itself changed the outcome. What does not exist, as of this publication, is the score. Which external signals belong on that line—and survive the gaming pressure Goodhart named—is fertile ground for exploration.
Measure the judgment, not the rollout.
That line is yours to draw—and you do not need anyone’s permission to start. Today.
The AI Leadership Playbook
Strategic Questions — copy-paste ready for an email to your CFO and CHRO:
1. Which of our AI metrics would still move if not a single decision improved—and which one would catch it if judgment eroded? What is the first metric we retire?
2. Where does our AI explain itself persuasively inside a live decision loop—and whom did we assign, in writing, to dissent from it? What is the first loop we assign?
3. If the board asked for a judgment line next to the adoption line, what signal would we put on it this quarter—and who owns it?
Your Next Plays — copy-paste ready for an email to your leadership team:
1. Tag the metric class. The Sophistication Gap replaced the adoption dashboard with a sophistication scorecard; extend it one rung. List every AI metric you report and tag each one rollout-measure or outcome-measure. Mark which would catch a judgment change. One page.
2. Run the design review. Pick one decision loop where AI explains its recommendations persuasively. Add the independent-verification step the Harvard experiment showed the narrative suppresses—a named human who checks the recommendation against the evidence.
3. Define one external signal. Pick one decision loop and define its outcome-layer signal—external to the people in the loop, not self-reported—and name who scores it. Design it expecting gaming pressure; Goodhart’s law applies to the fix too.
📅 Book a complimentary 1:1 Strategy Session—45 minutes to start the conversation about mapping the revenue hidden in your AI spend.
Sources
1. Leonid Sudakov & Nathan Furr, “AI Is Undermining Leaders’ Judgment. Here’s What to Do About It.” Harvard Business Review, August 19, 2026. https://hbr.org/2026/08/ai-is-undermining-leaders-judgment-heres-what-to-do-about-it
2. Erik Brynjolfsson, Bharat Chandar & Ruyu Chen, “Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence,” Stanford Digital Economy Lab working paper, revised August 2026. https://digitaleconomy.stanford.edu/app/uploads/2026/08/Canaries_August2026.pdf
3. Ross Dawson, “Recursive Self-Improvement in Humans + AI Systems,” Humans + AI podcast, Episode 54, August 19, 2026. https://humansplus.ai/podcast/ross-dawson-recursive-self-improvement-humans-ai-systems-hai-ep54/
4. Philip Topham, “Does Your Company Have a Comfortable Shoes Problem?” SavionAI: The Shift. Lift., August 17, 2026.
5. Lane, Boussioux, Ayoubi, Chen, Lin, Spens, Wagh & Wang, “The Narrative AI Advantage? A Field Experiment on AI-Augmented Evaluations of Early-Stage Innovations,” Harvard Business School Working Paper 25-001, current revision fetched August 2026. https://www.hbs.edu/ris/Publication%20Files/25-001_75068eb2-d475-4889-b341-aeabab6ef6d1.pdf
6. Mrinank Sharma, Meg Tong et al., “Towards Understanding Sycophancy in Language Models,” ICLR 2024. https://arxiv.org/abs/2310.13548
7. OpenAI, “Sycophancy in GPT-4o: what happened and what we’re doing about it,” April 29, 2025. https://openai.com/index/sycophancy-in-gpt-4o/
8. OpenAI, “Expanding on what we missed with sycophancy,” May 2, 2025. https://openai.com/index/expanding-on-sycophancy/
9. Joel Becker, Nate Rush, Elizabeth Barnes & David Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, July 10, 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
10. Joel Becker, Nate Rush, Tom Cunningham, David Rein & Khalid Mahamud, “We are Changing our Developer Productivity Experiment Design,” METR, February 24, 2026. https://metr.org/blog/2026-02-24-uplift-update/
11. K. Alec Chrystal & Paul D. Mizen, “Goodhart’s Law: Its Origins, Meaning and Implications for Monetary Policy,” Bank of England festschrift paper, 2001, reproducing Charles Goodhart (1975), “Problems of Monetary Management: The U.K. Experience,” Reserve Bank of Australia. The popular one-line version—”When a measure becomes a target, it ceases to be a good measure”—is Marilyn Strathern’s 1997 paraphrase; the 1975 original wording is Goodhart’s. http://web.archive.org/web/20250910114222/https://cyberlibris.typepad.com/blog/files/Goodharts_Law.pdf
12. KPMG + University of Texas at Austin research, “What the Best AI Users Do Differently,” Harvard Business Review, March 2026—the analysis behind The Sophistication Gap. https://hbr.org/2026/03/what-the-best-ai-users-do-differently
13. Baiyin Yang, Karen E. Watkins & Victoria J. Marsick, “The construct of the learning organization: Dimensions, measurement, and validation,” Human Resource Development Quarterly 15(1), 2004. https://assets.csom.umn.edu/assets/21929.pdf
14. Saravanan Kesavan & Tarun Kushwaha, “Field Experiment on the Profit Implications of Merchants’ Discretionary Power to Override Data-Driven Decision-Making Tools,” Management Science 66(11), November 2020. https://pubsonline.informs.org/doi/10.1287/mnsc.2020.3743
15. Saravanan Kesavan, Tarun Kushwaha & D. Steele, “Profit Implications of Judgmental Adjustments to Forecast Inputs: Evidence from a Large-Scale Field Experiment,” Management Science 72(1), January 2026. https://pubsonline.informs.org/doi/10.1287/mnsc.2024.06321
16. Andrew Ng, The Batch, issue #366, DeepLearning.AI, August 14, 2026. https://www.deeplearning.ai/the-batch/issue-366
17. Marketplace Morning Report, “How the AI revolution impacts corporate leaders” (edited transcript of Linda Hill, Harvard Business School), July 20, 2026. https://www.marketplace.org/story/2026/07/20/how-the-ai-revolution-impacts-corporate-leaders
18. Herminia Ibarra & Florence Wilkinson, “Why AI is a leadership challenge – not a technology one,” London Business School Think, April 23, 2026. https://www.london.edu/think/ai-leadership-challenge
19. CognivaLab, “Everybody Named the Problem. Nobody Teaches the Solution.” The AI Playbook, August 9, 2026. https://www.cognivalab.blog/p/everybody-named-the-problem-nobody
20. CognivaLab, “You Priced the License. You Never Priced the Trust.” The AI Playbook, August 16, 2026. https://www.cognivalab.blog/p/you-priced-the-license-you-never
21. CognivaLab, “Your AI ROI Is Hiding in the Judgment Layer,” The AI Playbook, August 2, 2026. https://www.cognivalab.blog/p/your-ai-roi-is-hiding-in-the-judgment
22. CognivaLab, “It’s Not Thinking. It’s Predicting.” The AI Playbook, June 23, 2026. https://www.cognivalab.blog/p/its-not-thinking-its-predicting
23. CognivaLab, “The Sophistication Gap,” The AI Playbook, June 9, 2026. https://www.cognivalab.blog/p/the-sophistication-gap


