In an experiment published in 2023, Lucía Vicente and Helena Matute asked participants to make decisions with recommendations from an artificial intelligence system. The system sometimes gave flawed advice. Among the participants, 80.7% detected errors and still followed the recommendations. They knew that a problem existed, but detection did not give them enough confidence or authority to overrule the machine.
Nolan Lovett uses that finding in The tragedy of the cognitive commons to explain the Validation Tether. People can supervise artificial intelligence only when they possess enough independent knowledge to judge its work. Yet organizations often introduce artificial intelligence by removing the tasks through which people acquire that knowledge.
The Tether describes a dependency. It is not a skill, attitude, or procedure. An organization cannot add it to a competency model and teach it in a workshop. It survives when people continue to develop knowledge that does not depend on the system they must evaluate.
Expertise develops through work
Lovett distinguishes between Internalized Mastery and Distributed Mastery.
- Internalized Mastery is knowledge held by the practitioner. It includes the mental models that allow a clinician to connect an unusual symptom to a diagnosis, an engineer to anticipate how a component will fail, or a programme manager to see why a technically sound recommendation will not work in a particular district. Practitioners develop this knowledge through repeated work on increasingly difficult problems, feedback from more experienced colleagues, and opportunities to correct mistakes.
- Distributed Mastery is the ability to produce results by coordinating knowledge held across people and tools. It includes selecting an artificial intelligence system, writing useful instructions, comparing outputs, arranging a workflow, and combining machine output with material from colleagues or databases. These abilities are real forms of professional competence. They can produce strong work without giving the operator an independent model of the subject.
The distinction becomes important during review. Someone with Distributed Mastery may produce a polished clinical protocol in an afternoon. Determining whether the protocol fits the available staff, referral routes, disease patterns, and medicine supply requires Internalized Mastery. Production can draw on the machine. Validation cannot depend entirely on the system under review.
This dependency creates the tether. Distributed Mastery remains useful as long as Internalized Mastery is available somewhere in the workflow. If organizations stop creating people with Internalized Mastery, assisted performance may continue to improve while the capacity to check it declines.
Two reviews can produce the same approval
Lovett separates surface validation from substantive validation.
- Surface validation checks features visible in the output. A reviewer can find a fabricated citation, an arithmetic error, a contradiction, missing sections, or formatting that does not follow the template. Artificial intelligence can perform much of this work, and people can learn to do it without extensive domain knowledge.
- Substantive validation asks whether the output is correct for the decision at hand. A recommendation can cite genuine research and still rest on an unsuitable assumption. A risk model can be mathematically sound and omit the variable that determines the outcome. A clinical suggestion can match the recorded symptoms and ignore something specific about the patient.
The missing element causes the danger. Readers who lack a working model of the domain do not know what the output has left out. Fluency makes the omission harder to notice because the document contains no visible defect.
An approval label rarely records which review took place. A dashboard may show that every document passed quality assurance even when reviewers checked only citations, completeness, and style. The organization then has evidence that review occurred, but no evidence that anyone tested the recommendation against domain knowledge and local conditions.
Assisted performance can conceal lost capability
Several studies discussed by Lovett separate performance with artificial intelligence from performance without it.
- Kacper Budzyń and colleagues compared endoscopists before and after exposure to artificial intelligence assisted detection. The endoscopists showed lower independent accuracy after adopting the system.
- Fabrizio Dell’Acqua and colleagues found that consultants performed better on tasks within a model’s capabilities and worse on tasks beyond that boundary.
- Emma Wiles and colleagues found that participants performed better while using generative artificial intelligence during training, but the advantage disappeared after the researchers withdrew assistance.
These studies do not establish that every use of artificial intelligence causes deskilling. Rather, they show that successful assisted performance does not prove that the person using the system has learned the underlying task. Organizations that measure output while assistance is available cannot infer independent capability from those results.
Other evidence identifies conditions that may preserve learning.
- A randomized trial by Sarah Everett and colleagues found that clinicians retained or improved diagnostic performance when the workflow required active engagement with the system’s reasoning.
- Research by Anders Humlum and Emilie Vestergaard found no effect on earnings or recorded hours during the first two years of adoption among Danish workers. The available evidence supports a conditional mechanism rather than an inevitable decline.
Lovett describes two possible forms of decline.
- Stock depletion occurs when fewer people enter a profession and develop deep expertise.
- Functionality degradation occurs when practitioners continue to advance through senior roles but develop shallower mental models than earlier cohorts.
Employment data may eventually reveal the first. Error patterns, escalation failures, and poor performance without assistance may reveal the second sooner.
Knowledge does not guarantee an override
The Vicente and Matute experiment adds another requirement. Participants often detected the flawed advice and followed it anyway. Substantive validation therefore depends on knowledge and on whether a person treats machine output as a claim open to challenge.
Professional standing affects that decision. A junior analyst may identify a false assumption and still defer to a system selected by senior management. A district health officer may see that a national protocol does not fit local conditions and lack a route for challenging it. The problem is no longer error detection alone. It includes who may stop the workflow and whose knowledge counts when sources disagree.
Reda Sadki’s analysis of the tragedy of the cognitive commons connects this problem to global health. Evidence hierarchies often classify frontline observation as anecdote while assigning authority to standardized guidance produced far from the settings where practitioners will use it. Artificial intelligence enters a system that already distributes credibility unevenly. A health worker may recognize that a recommendation conflicts with local experience and still assume that the institution or machine knows better.
The Validation Tether will fail if a knowledgeable person has no standing to act. Review arrangements therefore need an escalation path, not only access to an expert.
Organizations can test the Validation Tether
Counts of prompts, documents, or machine corrections do not measure the Validation Tether. Evaluation needs to examine independent judgment.
| Practice | What it tests or protects |
|---|---|
| Ask staff to complete a draft, estimate, differential diagnosis, or analysis before using artificial intelligence. | The comparison reveals whether the practitioner can form an independent account of the problem. It also preserves practice on the underlying task. |
| Record the type of review performed. | The review log should state whether the reviewer checked presentation and internal consistency or tested the recommendation against evidence, domain knowledge, and local conditions. |
| Insert plausible errors or omissions into test outputs. | Seeded cases can measure who detects the problem, who explains it, and who stops or escalates the workflow. Compare results by role and experience. |
| Require both machines and people to abstain. | A workflow needs explicit conditions for stopping, requesting more information, or handing the decision to someone with greater expertise. |
| Reserve developmental work for junior staff. | Some tasks must remain available for practice under supervision, even when a machine could complete them faster. |
Seeded tests should resemble the errors the organization is likely to face. A spelling mistake measures attention. A plausible recommendation based on the wrong population tests substantive validation.
The result should also distinguish between detection, explanation, and action. Vicente and Matute showed that these are separate outcomes.
Abstention policies deserve the same precision. At the 2026 Agentic AI Summit in Berkeley, Emily Zhu asked whether an agent knows what it does not know. A useful policy names the missing information, risk threshold, or disagreement that triggers human review. It also identifies who receives the case and who has authority to decide.
The Validation Tether is a workforce question
Artificial intelligence changes the economics of apprenticeship. Junior tasks once served two purposes. They produced work for the organization and gave less experienced staff repeated contact with the problems of the profession. When a machine takes over those tasks, the organization receives the immediate saving. The loss of future expertise appears years later and may affect employers that had no part in the original decision.
Lovett calls this wider process a tragedy of the cognitive commons. Each organization benefits from reducing its training burden while depending on a professional labour market that other organizations continue to replenish. If many employers make the same decision, the shared supply of expert judgment declines.
Exposure will differ by occupation and institution.
- Strong professional regulation can protect supervised practice and define who may approve high risk decisions.
- Safety rules can require independent checks.
- Professional bodies can preserve developmental pathways.
- Employers can automate work more easily when they can split it into discrete deliverables, which also detaches that work from apprenticeship.
Those protections often weaken below national institutions in global health. District and facility staff face immediate consequences from inappropriate guidance while receiving less protected time, supervision, and professional recognition. A system that relies on local validation must invest in the people expected to perform it.
What the Validation Tether concept changes
The Validation Tether changes the question asked about human oversight. Naming a human reviewer does not establish that oversight exists. The reviewer needs independent domain knowledge, time to examine the work, access to relevant evidence, and authority to stop the decision.
The concept also changes how organizations evaluate productivity.
- Faster output may represent genuine efficiency, or it may represent borrowed capability that the organization is no longer rebuilding.
- Performance measures need to track what staff can do without assistance, whether reviewers catch substantive errors, and whether junior staff continue to encounter progressively harder work.
A machine can generate a plausible recommendation before anyone in the room has developed the knowledge required to challenge it. The Validation Tether names the dependence between that recommendation and the professional judgment needed to use it safely. Maintaining it requires choices about staffing, supervision, review, and authority long before a visible failure provides proof that it has been lost.
References
Benzing, R., et al. (2025). Employee confidence and verification behaviour in generative artificial intelligence use. Cited in Lovett (2026). https://doi.org/10.1177/15344843261470602
Budzyń, K., Romańczyk, M., Kitala, D., et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. The Lancet Gastroenterology & Hepatology, 10(10), 896-903. https://doi.org/10.1016/S2468-1253(25)00133-5
Dell’Acqua, F., McFowland, E. III, Mollick, E., et al. (2026). Navigating the jagged technological frontier: field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organization Science, 37(2), 403-423. https://doi.org/10.1287/orsc.2025.21838
Everett, S. S., Bunning, B. J., Jain, P., et al. (2025). From tool to teammate: a randomized controlled trial of clinician and artificial intelligence collaborative workflows for diagnosis. medRxiv. https://doi.org/10.1101/2025.06.07.25329176
Hofer, B. K., and Pintrich, P. R. (1997). The development of epistemological theories: beliefs about knowledge and knowing and their relation to learning. Review of Educational Research, 67(1), 88-140. https://doi.org/10.3102/00346543067001088
Humlum, A., and Vestergaard, E. (2025). The unequal adoption of ChatGPT exacerbates existing inequalities among workers. Cited in Lovett (2026). https://doi.org/10.1177/15344843261470602
Lovett, N. (2026). The tragedy of the cognitive commons: how artificial intelligence could disrupt the regeneration of professional expertise. Human Resource Development Review. Advance online publication. https://doi.org/10.1177/15344843261470602
Niederhoffer, K., et al. (2025). Flawed artificial intelligence generated content at work and its cost to recipients. Cited in Lovett (2026). https://doi.org/10.1177/15344843261470602
Sadki, R. (2026). Health workforce development in the Age of Intelligence: a tragedy of the cognitive commons? Reda Sadki: Learning to make a difference. https://doi.org/10.59350/t0630-9wt12
Sadki, R. (2026). You cannot send an AI agent to jail: the scariest quotes from the Agentic AI Summit in Berkeley. Reda Sadki: Learning to make a difference. https://redasadki.me/2026/08/14/you-cannot-send-an-ai-agent-to-jail-the-scariest-quotes-from-the-agentic-ai-summit-in-berkeley/
Vicente, L., and Matute, H. (2023). Humans inherit artificial intelligence biases. Scientific Reports, 13(1), 15737. https://doi.org/10.1038/s41598-023-42384-8
Wiles, E., Krayer, L., Abbadi, M., et al. (2024). Generative artificial intelligence as an exoskeleton: experimental evidence on knowledge workers using generative artificial intelligence on new skills. SSRN. https://doi.org/10.2139/ssrn.4944588
