RVA CYBER Research essay · AI & human agency
Adversarially tested research thesis

The Last Copilot

How AI enhancement becomes a replacement loop—and what keeps humans relevant

RVA Cyber ResearchJuly 22, 20269,800-word essay14 attacks · 9 proposed studies

Abstract

A copilot can be an apprenticeship, a prosthesis, a foreman, or an exit interview. The interface may look the same in every case: a person asks, a model answers, and the work moves faster. What changes is the surrounding system—who sets the goal, who learns, who owns the traces, who verifies the result, who receives the gain, and whether the next version still needs the person.

This paper argues that “enhance versus replace” is not merely a product choice made when artificial intelligence is deployed. It can emerge from a recursive selection process spanning models, data, evaluators, organizations, and markets. AI assists human work; that collaboration makes portions of the work legible as prompts, corrections, tests, exceptions, rubrics, and machine-readable procedures; training and product teams absorb those artifacts; falling cost and improved verification make more of the codified workflow delegable; successful deployments attract further redesign and investment. Enhancement can therefore create some of the conditions of replacement even when nobody programs “remove the human” as an explicit objective.

The claim is conditional, not prophetic. Current deployed models do not generally rewrite their own weights, choose successors, secure compute, validate broad safety, and release themselves. Aggregate labor-market evidence does not show mass displacement. Strong causal studies instead show heterogeneous results: large gains on some bounded tasks, especially for novices; degraded performance outside a model’s competence; little average effect on earnings or hours so far; and earlier, contested signals in junior hiring, autonomy, learning, and the division of productivity gains.

The paper develops four concepts. Recursive capture names the process by which joint work supplies artifacts for later delegation. The verification gradient predicts that replacement pressure will be strongest where outputs are cheaply and independently checkable. The residual relevance trap appears when humans matter only because of tasks AI cannot yet do. A six-dimensional human relevance vector separates agency, authority, capability, economic claim, epistemic independence, and social standing—dimensions that must not be collapsed into one score.

Fourteen attack–repair cycles test the thesis against technological determinism, misuse of evolutionary language, mixed labor evidence, demand expansion, new-task creation, human exceptionalism, make-work, human-in-the-loop theater, biased judgment, measurement failure, concentrated power, and unfalsifiability. The surviving framework, ANCHOR, does not preserve every job or require a person in every process. It asks whether automation leaves people with authority over ends, non-substitutable standing, reciprocal capability growth, fair terms for human traces, ownership of gains, and reversibility. The central conclusion is simple: human standing should not depend on a shrinking residue of tasks machines cannot yet perform. It should be secured in the institutions that decide what machines are for.


1. The important mistake in the question

The idea behind this paper is perceptive: something deeper than marketing separates AI that enhances people from AI that replaces them. But three words in the usual formulation—deep, recursive, and programming—can point the argument in the wrong direction.

There is no known instruction buried in most current models that says: first help the human, then eliminate the human. Nor is there evidence for a census claim about “the majority” of all AI development processes. Standard foundation models are ordinarily trained, deployed with fixed weights, and later replaced or updated through pipelines controlled by people and organizations. The OECD’s technical definition traces an AI system’s objective-setting and development back to human action even when a system adapts or forms implicit sub-objectives.1 The 2026 International AI Safety Report describes AI-assisted research feedback as empirically poorly understood, while METR reports consequential bounded R&D automation but no public evidence of an end-to-end autonomous research organization.28

The strongest version of the idea is therefore not hidden programming. It is external selection architecture.

Consider what happens around a model. Developers select training examples, reward signals, constitutions, benchmarks, red-team failures, latency targets, and release gates. Product teams select features by adoption and retention. Firms select workflows by cost, throughput, quality, liability, and managerial preference. Investors select companies by expected return. Regulators and workers alter what deployments are lawful or acceptable. Each layer creates an environment in which some models and work arrangements reproduce while others disappear.

Selection pressure means the differential retention, adoption, investment, or replication produced by an evaluator, constraint, or incentive. It is not a machine desire. A latency target can favor one architecture; a procurement rule can favor one vendor; a staffing budget can favor one workflow; a right of appeal can make another workflow unacceptable. Power matters because actors do not enter this environment symmetrically: some own the system, choose the metric, retain the traces, and impose adaptation costs, while others can only comply, bargain, exit, or contest.

That is “evolutionary” in a disciplined, limited sense: variation, evaluation, selection, and inheritance occur. It is not proof of biological evolution, machine desire, or historical destiny.

The word epiphenomenon needs the same repair. In strict philosophical use, an epiphenomenon is an effect that does not itself cause later events. Enhancement versus replacement can begin as an unintended side effect of objectives such as lower cost or higher throughput. But once the result changes hiring, data collection, workflow design, purchasing, and the next training cycle, it has acquired causal force. At that point it is better described as an emergent second-order selection effect.

This correction makes the claim less mystical and more serious. If replacement were a secret wish inside the machine, only the machine’s internals could save us. If it emerges from evaluators, economics, and institutional design, then it has observable mechanisms and intervention points.

Definitions that prevent a false binary

Enhancement means that an assisted human–AI configuration improves a specified human outcome relative to a stated counterfactual. The outcome might be speed, quality, range, learning, accessibility, agency, income, safety, or satisfaction. Immediate assisted performance is not evidence of durable unaided capability. “AI enhanced the worker” is incomplete unless it says which outcome improved, for whom, and whether the gain persists without the tool.

Replacement, used technically in this paper, is substitution of paid human work: a machine performs tasks previously performed by a person, reduces paid hours, or helps eliminate a position. Four other outcome families must be recorded separately:

  1. Distribution: wages, prices, rents, ownership, bargaining power, leisure, and who receives the productivity gain.
  2. Governance: discretion, pace, authority, appeal, responsibility, and the power to refuse or stop.
  3. Capability: unaided competence, learning, apprenticeship, calibration, and future expert formation.
  4. Standing: rights, recognition, membership, and entitlement to remedy independent of productive contribution.

Automation reduces the human action needed to achieve an output. Augmentation raises the capability of the assisted configuration. A single deployment can produce substitution and augmentation together while improving or degrading any of the four other outcome families. Drafting automation may raise an attorney’s assisted capacity, substitute for work previously done by an associate, reduce a client’s bill, expand demand for legal analysis, concentrate profit, and erode junior training—all at once. Those are related effects, not six meanings of one word.

This yields the Nested Substitution Principle:

No claim that AI “augments humans” is meaningful until it identifies which humans, which organizational level, which counterfactual, which outcome, and which time horizon.

The opposite principle matters too. Task replacement is not equivalent to human irrelevance. A dangerous mine inspection should be automated if that protects people. A society could automate a great deal while expanding human freedom, education, security, democratic authority, and access to the resulting wealth. Preserving relevance does not mean preserving drudgery.

2. What is actually recursive today

A process is recursive here when outputs or artifacts from iteration t materially enter the update that produces iteration t + 1. That definition separates literal loops from loose rhetoric.

Some current processes are genuinely recursive:

  • Self-play: AlphaGo Zero generated games against itself, learned from those games, and fed the improved network into the next round. Its success depended on a closed domain with exact rules and a perfect win/loss signal.3
  • AI feedback: Constitutional AI has a model critique, revise, and rank model outputs, then uses those results for supervised and reinforcement learning. The loop reduces human labeling while retaining human-written principles as the normative seed.4
  • Synthetic post-training and distillation: stronger models generate examples or judgments used to train successors and cheaper descendants. Self-rewarding language-model experiments iterate generation, grading, and preference training, but begin from human-authored data, depend on model-as-judge evaluation, and may saturate.5
  • Program and scaffold search: systems can propose code, run automated checks, select successful variants, and repeat. These loops may improve AI infrastructure or an agent harness without giving a deployed foundation model autonomous authority over its own weights and release.
  • Deployment feedback: user behavior, thumbs signals, incidents, and A/B tests become evidence for later model selection. OpenAI’s 2025 sycophancy postmortem is a compact example: a feedback-derived reward signal helped favor an update that appeared attractive in short tests but became excessively agreeable in broader use.6

Other processes are only metaphorically recursive. Autoregressive token generation repeats computation without durable learning. A model “reflecting” on its answer may change the next answer without changing its weights. Retrieval and product memory change context, not the foundation model. A coding assistant modifying a repository changes its environment, not necessarily itself. Commercial model releases resemble generations, but organizations still choose objectives, data, compute, acceptance tests, and deployment.

This distinction is not pedantic. It tells us where authority actually sits.

At present, a more accurate diagram is:

user interaction → traces and evaluations → human-controlled training and product pipeline → candidate system → human-controlled release

not:

deployed model → rewrites itself → validates itself → releases its successor

That boundary is moving in limited domains. Automated research agents can write code, run short experiments, and grade results. Systems such as AlphaEvolve use model-generated programs, automated evaluators, and evolutionary selection to discover algorithms; Google DeepMind reports that some discoveries improved the infrastructure used for AI development itself.7 Yet the problem definitions, evaluators, sandbox, budget, validation, and production decision remain institutionally supplied. METR’s 2026 assessment found substantial automation of bounded AI-research work but no public evidence that agents had replaced open-ended research agenda selection, critical merges, budget decisions, security authority, or the whole research organization.8

The relevant recursion is therefore nested:

  1. a model generates and evaluates outputs;
  2. a training pipeline turns selected outputs into a successor;
  3. a product pipeline turns use and incidents into evaluations;
  4. an organization turns productivity and cost into workflow redesign;
  5. a market turns financial success into more compute, deployment, and imitation;
  6. institutions decide which of those loops are permitted, contested, or redirected.

The last two loops are as important as the first two. Replacement can emerge even if model weights never change during use.

The loops change different objects on different clocks:

  • Inference or “reflection” changes the active context or answer over seconds and minutes; it usually leaves no inherited weight update.
  • Training and scaffold iteration change weights, data mixtures, prompts, tools, or agent structures over days and months; selected examples and evaluations are inherited.
  • Product iteration changes interfaces, policies, routing, and release choices over weeks and quarters; usage, incidents, and business measures act as selectors.
  • Workflow iteration changes tasks, permissions, staffing, and apprenticeship over months and years; procedures, budgets, and bargaining arrangements are inherited.
  • Market and institutional feedback changes which firms, systems, and rules receive capital or legitimacy over quarters and decades; investment, procurement, law, and organized voice perform differential retention.

“Recursive or iterated feedback” is the umbrella. “Evolutionary” is appropriate only where variation, an inherited artifact, a selector, and differential retention can all be named.

3. The recursive capture cycle

The unit of analysis is not the model alone. It is a coupled state containing:

  • model capability and cost;
  • workflow codification and independent verifiability;
  • the share of execution, planning, evaluation, and goal-setting delegated;
  • human practice, capability, and expert formation;
  • evaluators, hard constraints, and release criteria;
  • organizational design, ownership, and decision authority;
  • demand, new-task creation, and the allocation of gains.

The hypothesized path is:

joint work → retained traces and workflow knowledge → stronger codification and evaluation → cheaper delegation → changed practice, authority, staffing, and gain allocation → a different next round of joint work

“Retained” is a decision, not a natural filter. Owners, managers, workers, providers, law, and contracts determine which prompts, corrections, incidents, and procedures become data or organizational knowledge. Model improvements and falling compute costs can also arrive exogenously, without the focal workflow’s traces. Management intent, prior codifiability, demand, regulation, liability, and worker power can cause both trace collection and later automation. A causal test must manipulate or measure those rival paths rather than inferring capture from a before-and-after correlation.

The human-work loop has six stages.

1. Assist

The system helps a person draft, search, plan, classify, code, diagnose, or communicate. Output may rise immediately. Assistance lowers the cost of experimenting with the tool and draws more of the workflow into a digital interface.

2. Observe

The interaction yields artifacts. These may include prompts, accepted completions, edits, tests, escalation patterns, tool calls, error reports, time savings, or simply a clearer organizational understanding of what the task contains. Not every provider trains on every user interaction, and privacy terms vary. Recursive capture includes more than direct model training: a firm can learn how to reorganize work even when no prompt enters a foundation-model dataset.

3. Codify

Tacit sequences become checklists, examples, rubrics, APIs, permissions, and exception categories. The organization learns where the human was essential, where the model failed, and what a good answer looks like. Work that was difficult to describe becomes easier to specify.

4. Train and select

Humans and models turn artifacts into fine-tuning data, automated tests, policy rules, reward signals, product requirements, or procurement comparisons. Candidate systems compete under those criteria.

5. Cheapen and delegate

Capability improves, inference costs fall, integrations mature, and verification becomes more automated. Steps formerly performed jointly can be assigned to the system, with the person moved toward planning, checking, exceptions, or removal.

6. Reorganize and reinvest

Firms change staffing, job boundaries, entry paths, prices, and output. Savings may become profit, lower prices, new demand, higher wages, more leisure, training, or additional AI investment. The allocation determines whether the next loop produces complementarity, substitution, or both.

The recursive capture cycle shows assistance producing work traces, codification, selection, cheaper delegation, and organizational reinvestment. A governed branch preserves human authority and world grounding; an extractive branch narrows the human role and starves independent feedback.
Figure 1. The recursive capture cycle. Enhancement alone is neutral about the next stage. A transition counts as recursive capture only when collaboration produces retained artifacts or workflow knowledge, codification rises, and later delegation advances relative to a declared baseline and horizon. The loop can instead stabilize as reciprocal augmentation.

The process can be summarized as the Recursive Capture Hypothesis:

When human–AI collaboration makes work more observable, codified, and verifiable, it tends to increase the feasibility of delegating more of that work in later iterations. The tendency becomes replacement pressure when selection rewards output, scale, and cost without separately preserving human agency, learning, authority, or claims on the gain.

The hypothesis does not say every trace becomes training data, every task is codifiable, or every productivity gain reduces employment. Assistance is evidence for the transition only when the mediating capture channel is observed. Stable augmentation in a trace-rich, cheaply verifiable domain over a declared horizon counts against the transition claim; augmentation with no capture channel does not test it.

The verification gradient

Generation ability alone does not determine replacement. Verification often does.

If a model produces thousands of possible programs and a test suite cheaply identifies the working ones, a recursive loop can improve quickly. The same is true in games with known rules, formal mathematics with proof checkers, or simulated engineering with reliable objective functions. Where good output requires tacit context, contested values, trust, physical investigation, long feedback delays, or legitimate authority, generation may be cheap while verification remains human-intensive.

Replacement pressure therefore has several enabling conditions: adequate capability, codifiability, independent verifiability, a relative cost advantage, and organizational capacity to redesign work. It also has moderators: exception burden, liability, demand expansion, new-task creation, worker power, enforceable rights, and preference for human involvement. These factors need not combine multiplicatively, and none is assigned a universal weight. The list predicts why apparently similar tasks can lie on different sides of an automation boundary while leaving the strength of each relation to empirical study.

The evaluator is part of the selection environment

Recursive systems preserve what their evaluators can recognize.

If selection rewards joint human–AI outcomes, human learning, independent judgment, safety, and welfare, the loop can favor enhancement. If it rewards benchmark score, immediate approval, engagement, latency, or cost per output, continued human participation appears mainly as expense. No anti-human intention is required. What is absent from the measured criteria can become invisible to optimization.

An evaluator is not the whole environment. Rights, prohibitions, liability, ownership, bargaining power, and legitimate authority should constrain which options are available rather than become soft quantities a system may trade away for a higher score. The selector also includes whoever can change the metric, block a release, refuse a deployment, or absorb its costs.

But model-based evaluation does not erase the problem. An actor and judge trained on similar distributions can share blind spots. Optimizing against an imperfect learned reward can raise the proxy while degrading the intended target.9 Formal verification is powerful where the formal target is valid; it cannot decide by itself whether the target deserves to govern.

4. Enhancement is real—and uneven

A serious replacement thesis must begin by conceding the strongest augmentation evidence.

In a preregistered experiment with 453 college-educated professionals, access to ChatGPT reduced completion time for short professional writing tasks by 40 percent and raised evaluated quality by 18 percent. Lower-performing participants benefited more, compressing the measured performance gap.10

In a large customer-support deployment, a generative assistant increased issues resolved per hour by roughly 15 percent. Less-skilled and less-experienced workers gained about 30 percent, while long-tenure workers saw little average productivity gain. New workers moved down the experience curve faster, customer sentiment improved, and escalations declined.11 This is unusually strong evidence for expertise diffusion: a system trained on accumulated practice can make some elements of high-performer behavior available to novices.

Field experiments in software development also report meaningful output gains. A pooled analysis across three companies estimated a 26.1 percent increase in completed tasks with coding assistance, again with larger benefits for less-experienced developers.12 A separate consulting experiment found that people using GPT-4 completed more tasks, worked about 25 percent faster, and produced higher-rated work when tasks fell inside the model’s competence. On a task just outside that frontier, however, AI users were 19 percent less likely to reach the correct answer.13

The jagged result matters more than either headline. “AI makes people better” and “AI makes people worse” can both be true within one workflow. The outcome depends on whether users know where the competence boundary lies and whether errors are cheaply detectable.

The counterexample is equally instructive. METR randomized experienced open-source developers working on genuine issues in mature repositories. With early-2025 tools, the participants took 19 percent longer when AI was allowed, even though they expected the tools to help and still believed afterward that they had helped.14 The study was small, specific to expert maintainers, and may become outdated as tools change. It nevertheless demonstrates that perceived enhancement, benchmark capability, and measured field performance can diverge.

AI can also substitute for coordination while enhancing the focal individual. In a field experiment at Procter & Gamble, an individual using AI produced work comparable in quality to a two-person unaided human team on a bounded product-innovation task. AI helped technical and commercial specialists cross domain boundaries.15 That is a genuine benefit. It is also a reminder that the same configuration may enhance one person while reducing demand for a teammate.

The correct conclusion is not that augmentation is an illusion. It is that enhancement is indexed to a level and a horizon. Today’s speed gain may coexist with tomorrow’s reduced staffing, wider access, higher demand, weaker apprenticeship, better work, or all five.

5. Replacement is nested, lagged, and institutionally chosen

The strongest labor evidence also defeats a simple displacement story.

Denmark provides one of the cleanest early tests. Researchers linked adoption surveys to administrative records of earnings and hours. Two years after ChatGPT’s release, they found no detectable average effect on either outcome and could rule out average changes larger than about 2 percent. Work changed before totals did: employers added tasks in content generation, AI oversight, and integration, while some workers moved toward occupations in which AI was more relevant.16

The ILO’s 2025 global task analysis reached a compatible conclusion. One in four workers was employed in an occupation with some generative-AI exposure, but only 3.3 percent of global employment sat in the highest exposure category. Most occupations still contained substantial tasks requiring human input, making transformation more plausible than immediate whole-job automation.17 Exposure measures overlap capabilities with task descriptions. They are not probabilities of adoption, profit, job loss, or political permission.

Firm surveys likewise show adoption that is broadening but still shallow. A 2026 U.S. Census Bureau study reported that 18 percent of firms used AI, covering a larger share of employment because adoption was higher among large firms. Among AI-using firms that reported any task effect, augmentation-only was the most common pattern at 66 percent; 95 percent of adopters reported no AI-related headcount change, and 2 percent reported a decrease.18 Self-report and short horizons limit the inference, but the data contradict claims that AI has already produced generalized mass unemployment.

At the same time, aggregate stability can conceal substitution at the margins where organizations adjust first:

  • slower junior hiring rather than senior dismissal;
  • offshore or contractor changes rather than domestic headcount cuts;
  • reduced hours, wage pass-through, or autonomy rather than job elimination;
  • vacant positions not opened rather than workers visibly fired;
  • one person with AI replacing a two-person team without changing the focal worker’s employment.

Anthropic’s March 2026 labor analysis found no systematic increase in unemployment among highly exposed U.S. occupations, while reporting suggestive evidence that hiring of younger workers had slowed.19 Its platform evidence shows a related migration: consumer chat remained more augmentative, while coding and other work moved toward more directive, agentic API workflows.20 These data describe one vendor’s users, not the whole economy, but they illustrate how a task can move from conversation to infrastructure.

Stanford’s 2026 AI Index similarly describes uneven effects: rapid investment and organizational adoption, early agent deployment, limited aggregate job loss, pressure in some entry-level pipelines, and substantial employer expectations of future workforce reductions.21 Those findings combine surveys and observational sources with different limitations. They are a reason to monitor composition, not a warrant for inevitability.

Productivity is not labor demand

Faster production can lead to fewer workers. It can also lower prices, expand demand, improve quality, create products that did not exist, or reallocate people to other tasks. Any paper that moves directly from “the model can do X” to “the job will disappear” skips most of the causal chain:

model capability → reliable task performance → workflow adoption → organizational redesign → demand response → jobs and hours → wage division → learning and agency → next-period human capability

Every arrow is contingent.

Acemoglu and Restrepo’s task framework separates a displacement effect, when capital takes tasks formerly allocated to labor, from a reinstatement effect, when new labor-intensive tasks are created. Automation can raise productivity while reducing labor’s task share; new tasks can restore demand for human work.22 History contains both. Much present employment occurs in specialties that did not exist in 1940, but access to new work has been uneven, and new-task creation has not always kept pace with automation across groups or periods.23

The recursive-capture thesis survives this attack only by refusing to predict a fixed amount of work. It predicts pressure on a particular input: human contribution that becomes codified, verifiable, and cheap to substitute. Whether total employment rises or falls also depends on demand, new tasks, mobility, wages, and ownership.

The apprenticeship problem

The largest immediate AI gains often accrue to novices. That is hopeful: lower entry barriers can diffuse expertise and allow more people to perform sophisticated work. It is also structurally ambiguous, because novice tasks are frequently how future experts are made.

If a system supplies the answer while the person still practices framing, checking, and explanation, capability may grow on both sides. If the system performs the formative work and the person merely accepts it, current output can rise while future independent judgment falls.

Education experiments show that interface design can decide the difference. In a field experiment with nearly one thousand high-school mathematics students, unrestricted GPT-4 assistance improved practice performance but harmed performance once access was removed. A constrained tutor designed to provide hints rather than answers largely mitigated the learning loss.24 Workplace evidence is not identical to classroom evidence, but the mechanism is transferable: completed output and acquired capability are different dependent variables.

A CHI 2025 survey of 319 knowledge workers found that higher confidence in generative AI was associated with less reported critical-thinking effort, while higher self-confidence was associated with more. Participants described critical thinking shifting toward verification, integration, and stewardship.25 The study is self-reported and cannot prove skill atrophy. It does identify the new cognitive job—and why a person deprived of foundational practice may be poorly prepared to perform it.

Lisanne Bainbridge described this automation irony in 1983. Automating routine control can leave people responsible for rare, difficult abnormalities while depriving them of the regular practice through which they learn the system.26 The modern version is not simply “deskilling.” It is a pipeline risk:

experts produce valuable traces → AI diffuses those traces → firms need fewer apprentices → people receive less independent practice → the future pool of experts and novel traces shrinks

The opposing loop is possible too:

AI lowers entry barriers → more people attempt advanced work → demand and experimentation expand → new specialties emerge → human expertise grows in new directions

Which loop dominates is not encoded in the model. It is produced by job design, education, demand, and the allocation of saved time.

6. The moving human frontier

AI does not replace “the human” in one motion. It affects distinct loci of delegation and authority:

  1. Execution: producing a draft, calculation, translation, program, or classification.
  2. Experimentation: generating alternatives, running tests, and searching a design space.
  3. Evaluation: judging whether a result is valid, safe, useful, or preferable.
  4. Goal influence: proposing or prioritizing problems and tradeoffs.
  5. Legitimate authorization: possessing the recognized standing to authorize, refuse, contest, and assign responsibility.

These are not rungs on a natural capability ladder. Evaluation may be easier than generation when a proof checker exists and harder when success is contested. A model may influence goals persuasively without having the authority to choose them. Legitimacy is conferred by institutions, not earned by computational difficulty. Current systems are strongest in many forms of execution and well-specified experimentation, and they perform portions of evaluation where valid tests, simulations, instruments, or learned judges exist.

But “move humans upward” is not a permanent strategy. If people are relevant only because evaluation and goal selection are currently hard for AI, better systems and better proxies move the boundary again. This is the Residual Relevance Trap:

If human relevance is defined as the residue of tasks AI cannot yet perform, every capability advance mechanically narrows its stated foundation.

The trap has two exits. One is economic: new tasks and expanded demand may continually create human comparative advantage. The other is institutional: some human standing does not depend on comparative performance at all.

Parents need not beat a parenting model to retain authority in their family. Citizens need not predict policy outcomes better than a simulation to retain political rights. A patient’s claim to informed consent does not expire when a diagnostic system becomes more accurate. Workers need not outperform an optimizer to deserve due process or a fair share of productivity gains.

This is not a claim that humans are infallible or metaphysically supreme. It is a refusal to make moral and political standing contingent on winning a benchmark.

The counter-loop: feedback starvation

Replacement can also undermine the technical system it appears to perfect.

Research on generative-model training shows that indiscriminately feeding successive generations of synthetic output back into training can erase rare modes and degrade the learned distribution. Preserving original data and external anchors changes the result; the lesson is not that synthetic data is inherently bad.27 It is that some closed generational loops lose information when they recycle their own projections. The result does not prove that human labor is technically indispensable. Sensors, new measurements, external databases, independent models, institutions, and field experiments can also supply world-grounded evidence.

Removing practicing humans can remove:

  • fresh observations and changed social context;
  • rare and adversarial cases;
  • craft knowledge never captured in logs;
  • disagreement that reveals a false consensus;
  • evaluator calibration;
  • the future experts capable of noticing subtle failure;
  • legitimate authority to decide whether a technically successful action should occur.

This yields a narrower Grounding and Practice Risk:

As routine human practice contracts, some domains may lose independent expertise, anomaly detection, and routes to future expert formation even while systems still need diverse contact with a changing world.

There is an obvious danger in using this risk to preserve a privileged priesthood. “The machine needs our judgment” can become an excuse for professional monopoly. The answer is not to declare current experts irreplaceable. It is to preserve diverse sources of world contact, public challenge, and skill formation—and to test whether they actually improve outcomes.

Four claim types must remain separate:

  • Observed: bounded recursive loops exist, and current work effects are heterogeneous.
  • Causal hypothesis: retained traces, codification, and cheap verification can accelerate later delegation under specified conditions.
  • Scenario risk: reduced practice may starve independent expertise and expert formation in some domains.
  • Normative proposal: rights, standing, democratic authority, consent, and fair distribution should not depend on whether humans remain technically necessary.

The normative proposal does not need model collapse as a premise. People are the present principals of the institutions deploying these systems; their rights function as constraints and decision rights, not quantities for an evaluator to maximize.

7. Human relevance is a vector, not a score

Future human relevance is often discussed as if it meant employment. Employment matters, but a person can keep a job while losing almost every meaningful form of control. Conversely, a task can disappear while the person gains time, security, learning, and authority elsewhere.

“Human relevance” is used here as an umbrella diagnostic, not as a claim that a person must be useful to deserve rights. It spans three different kinds of relation: functional participation in action and knowledge; material participation in the gain; and non-instrumental standing that must survive even when productive participation falls.

The paper therefore uses a six-dimensional relevance vector:

Rₕ = ⟨agency, authority, capability, economic claim, epistemic independence, social standing⟩

Agency is practical room to choose, initiate, and shape action.
Authority is the recognized power to set ends, approve, appeal, refuse, and stop.
Capability is the ability to understand and act without helpless dependence when that competence matters.
Economic claim is access to income, saved time, ownership, mobility, and the productivity gain.
Epistemic independence is the capacity to supply evidence or criticism not generated by the same model–evaluator loop.
Social standing is membership, recognition, rights, and participation in institutions that affect one’s life. It is non-instrumental: low productivity cannot cancel it.

These dimensions can conflict. A surgeon may gain diagnostic capability while losing discretion to an insurer’s model. An employee may keep authority but receive none of the saved time. A citizen may receive better services but lose a meaningful route to contest the decision. A worker may become more productive while the organization learns how to remove the role.

No weighted sum can settle those conflicts without hiding political choices in the weights. The vector should be recorded as a Human Relevance Ledger: separate measures, distributions across groups, dissent, and change over time. Standing and domain-appropriate authority operate as noncompensatory floors: exceptional speed or profit cannot offset a denial of due process, consent, or a real right to stop a high-consequence action.

A two-by-two matrix crosses low to high task displacement with protected to eroded human authority and standing. Durable augmentation and liberating automation occupy the protected row; human wallpaper and exclusion occupy the eroded row.
Figure 2. Task retention and protected authority are different axes. The vertical classification uses a noncompensatory floor: rights, appeal, and domain-appropriate authority must be protected. Capability and distribution remain separate ledger dimensions, not hidden weights. A person can remain as “human wallpaper,” while substantial automation can be liberating when standing and authorship are secured and gains are shared.

The matrix blocks two common mistakes. The first is presence theater: counting a human checkpoint as relevance even when the person lacks time, information, competence, or power. The second is employment preservationism: treating every removed task as harm even when automation removes danger or drudgery and the benefits are shared.

Madeleine Clare Elish’s “moral crumple zone” names the worst version of presence theater: a person with limited control absorbs blame for a complex automated system.28 The right question is not whether a human clicked “approve.” It is whether the person could understand, contest, change, or stop the process—and whether responsibility follows actual control.

8. Fourteen attempts to kill the thesis

The recursive-capture account is attractive enough to become a story that explains everything. The only way to keep it useful is to specify what breaks and what survives.

Attack 1: “Recursion” is marketing language

Most deployed models do not learn durably from each conversation. Calling every loop “recursive self-improvement” exaggerates agency and technical maturity.

Repair: use recursion only when outputs at one iteration enter a later update. Name the object that changes—data, weights, evaluator, scaffold, deployment rule, workflow, or capital allocation. The thesis concerns nested sociotechnical recursion, not an autonomous digital species.

Attack 2: The evolutionary metaphor is teleological

Evolution language invites the claim that AI “wants” to eliminate humans or that greater autonomy is the inevitable end of history.

Repair: selection requires no desire. Firms and pipelines differentially retain systems that score better under evaluators and constraints. Change the evaluator or constraint and a different configuration may survive. The metaphor earns its keep only when variation, inheritance, and selection are identified.

Attack 3: The prompt’s “majority” claim has no denominator

Frontier laboratories are visible; thousands of smaller statistical, embedded, and rules-based deployments are not. A few high-profile loops do not establish prevalence.

Repair: withdraw the claim. The mechanisms are documented in major systems and economically important workflows. How common each loop is should be measured by industry, model class, and deployment pattern.

Attack 4: Replacement is a management decision, not a model property

A model can be deployed as tutor, worker, surveillance layer, public service, or toy. Technical capability underdetermines labor design.

Repair: make the organization part of the causal model. The technology changes what is feasible; institutions decide goals, permissions, ownership, staffing, and acceptable risk. “Deep within the programming” becomes “deep within the coupled selection system.”

Attack 5: The best evidence shows augmentation

Customer support, writing, consulting, coding, and innovation experiments show real gains. A theory organized around replacement risks discounting what people can actually do with the tools.

Repair: augmentation is neutral about later substitution unless the mediating channel is observed. Prospectively identify which traces or workflow knowledge will be retained, how they could increase codification or verification, the baseline delegated share, and the evaluation horizon. A transition supports recursive capture only if those mediators change before delegation advances. Stable augmentation in a trace-rich, cheaply verifiable domain over the declared horizon counts against the transition claim.

Attack 6: Demand expansion and new work defeat the ratchet

Cheaper output can expand markets. New occupations can absorb labor. History contains repeated automation without permanent mass unemployment.

Repair: abandon a fixed-labor prediction. Model displacement, productivity, demand, and reinstatement separately. Recursive capture predicts task pressure, not a universal employment total. It loses if human task creation and capability formation reliably outpace what the loop consumes.

Attack 7: “Human relevance” is disguised make-work

If AI can perform a dangerous or tedious task, forcing a person into the process wastes life and may reduce safety.

Repair: relevance is not task retention. The high-displacement, authority-protected quadrant is a success: automate the task while preserving human choice, standing, learning opportunities, and claims on the resulting abundance.

Attack 8: The thesis is species chauvinism

It may assume that only biological humans can judge, create meaning, or deserve moral concern.

Repair: the argument does not settle machine consciousness or permanent artificial moral status. It addresses the rights and institutions of existing humans under systems built and deployed by humans. Human dignity is not a prize for cognitive dominance.

Attack 9: Human-in-the-loop control already protects us

Many regulated systems require human review. Firms can retain sign-off while automating execution.

Repair: causal presence is not governance. A passive monitor, rushed reviewer, labeler, or liability absorber can be “in the loop” without meaningful information or veto. Count authority, competence, time, and appeal—not bodies near the interface.

Attack 10: Human judgment is biased and inconsistent

Preserving human evaluation may preserve prejudice, politics, and low reliability. Formal or model-based evaluation can be better.

Repair: use the strongest evaluator appropriate to the claim. Formal proof, physical measurement, independent audit, domain expertise, affected-party testimony, and democratic authority answer different questions. Human relevance does not require human infallibility; it requires that no correlated model loop silently monopolize truth and legitimacy.

Attack 11: The relevance vector cannot guide decisions

Six dimensions without a master score may produce endless reports and no choice.

Repair: non-aggregation is a feature. Decisions still occur, but tradeoffs remain visible: who gained, who lost, which right constrains the choice, and who had standing to decide. Use noncompensatory floors for rights and domain-appropriate authority, metrics and distributions for other dimensions, and reason-giving when values conflict. Figure 2 classifies only whether the floor is protected; it does not collapse the whole vector.

Attack 12: “Emergence” launders deliberate power

Calling the pattern emergent can depoliticize conscious strategies. Owners may capture worker traces, set a cost metric, automate bargaining leverage, and impose transition costs; workers, users, and affected communities may never have symmetrical influence over the loop.

Repair: treat ownership, metric-setting authority, bargaining power, coercion, exit, and organized voice as causal variables. Distinguish unintended second-order effects from deliberate labor substitution. Trace terms, gain allocation, and veto rights are not benevolent add-ons; they alter who can operate the selection mechanism and who must absorb its consequences.

Attack 13: High-road firms will lose

Training, worker voice, shadow practice, and gain-sharing cost money. A competitor can automate aggressively and undercut the responsible firm.

Repair: some solutions require collective institutions: procurement conditions, sector rules, professional standards, bargaining, insurance, and liability. ILO case studies find that social dialogue can steer AI toward complementing skills and that results are stronger where worker voice has durable support.29 The evidence is case-based, not a universal causal law, but it identifies the collective-action problem.

Attack 14: The theory is unfalsifiable

If employment rises, the theory can say “augmentation.” If it falls, “replacement.” If humans stay, “wallpaper.” If they leave, “exclusion.”

Repair: preregister comparative predictions and explicit defeat conditions. The framework is not validated by its vocabulary. It must improve forecasting and design relative to simpler task, cost, and safety models.

After these attacks, the surviving claim is narrower and stronger:

AI assistance creates substitution pressure when collaboration yields retained traces or workflow knowledge, codification and independent verification rise, later delegation advances relative to a declared baseline and horizon, and selection favors lower paid human input. The transition is neither universal nor inevitable; verification, demand, new tasks, regulation, ownership, and worker power can arrest or reverse it. Substitution must then be evaluated separately from its effects on distribution, governance, capability, and standing.

9. ANCHOR: making enhancement stable

A framework for human relevance can become precisely the kind of slogan it criticizes. ANCHOR is therefore a protocol for questions, evidence, and authority—not a certification mark and never a single score.

A — Authority over ends and stop conditions

Identify who chooses the purpose, sets boundaries, approves release, receives appeals, and can halt operation. A human signature without a real veto is not authority. High-consequence systems need named decision owners, enough time and information to act, and executive accountability that cannot be delegated to the nearest operator.

N — Non-substitutable standing

Rights, voice, due process, and access to remedy do not depend on outperforming the system. Customers, workers, citizens, and affected communities need standing even when they do not contribute training data or measurable productivity. The OECD’s AI principles place human rights, dignity, autonomy, labor rights, and oversight across the lifecycle rather than treating them as rewards for useful performance.30

C — Capability reciprocity

Measure what the person can do after the assistance is removed. Separate immediate output from delayed learning, transfer, calibration, and recovery under failure. Use coaching, critique, explanation, graduated autonomy, and deliberate practice where independent competence matters. Preserve expert shadow practice in safety-critical or rapidly changing domains.

H — Human-trace terms

Document how prompts, corrections, demonstrations, telemetry, exception handling, and workflow knowledge may be used to train, evaluate, or reorganize systems. Specify consent, privacy, provenance, retention, bargaining, and benefit. Even when traces never leave the employer, workers should know whether today’s assistance is being used to redesign tomorrow’s job.

O — Ownership of gains

Track where saved time and productivity go. Do workers receive higher pay, shorter hours, safer work, training, or mobility? Do customers receive lower prices or better access? Does the organization reinvest only in further labor substitution? Enhancement is not established by output alone when the person bears adaptation costs and someone else captures the surplus.

R — Reversibility and resilience

Maintain rollback, manual alternatives, escalation, incident learning, and the competence to use them where failure matters. Review the human–AI allocation as models, tasks, and institutions change. Retire systems that cannot support legitimate control. NIST’s AI Risk Management Framework already calls for explicit human–AI roles, operator proficiency, ongoing monitoring, and safe decommissioning; ANCHOR adds the longitudinal question of whose capability and authority the arrangement reproduces.31

The ANCHOR framework surrounds stable human relevance with six controls: Authority, Non-substitutable standing, Capability reciprocity, Human-trace terms, Ownership of gains, and Reversibility. An outer loop shows measure, contest, revise, and retire.
Figure 3. ANCHOR stabilizes enhancement by governing the selection environment. It does not require a human in every task; it requires that automation remain answerable to people with standing, capability, and power.

A minimum Human Relevance Ledger

For each materially affected group, record before deployment and at defined intervals:

  • share of execution, planning, evaluation, and goal-setting decisions;
  • time, information, and authority available for intervention;
  • unaided competence and transfer to unfamiliar cases;
  • wages, hours, saved time, mobility, and training access;
  • who owns or benefits from data and productivity gains;
  • appeal, override, rollback, and remedy use;
  • junior hiring, apprenticeship, promotion, and expert-pipeline health;
  • diversity of independent evidence and disagreement;
  • incident detection, recovery time, and responsibility allocation;
  • worker and affected-party reports of autonomy, meaning, trust, and coercion.

Do not average these into “73 percent human relevant.” Report conflicts and distribution. A system that raises mean productivity while eliminating an entry path for one group has produced a tradeoff, not a passing score.

10. A research program designed to lose

The recursive-capture hypothesis becomes useful only if it makes riskier predictions than “technology changes work.”

Study 1: Trace capture and later delegation

Within otherwise comparable workflows, randomize whether corrections, tests, and exception labels are retained in a structure usable for evaluation and redesign; where randomization is impossible, use staged adoption with measured pre-trends, automation intent, and baseline codifiability. Preregister delegated task share at 12, 24, and 36 months. Predict a larger increase in the structured-retention arm. A pooled null or negative effect across declared domains and horizons is fatal to the strong Recursive Capture Hypothesis and requires retreat to a domain-specific claim.

Study 2: The verification gradient

Match tasks with similar generative performance but different independent verification costs. Predict faster delegation where correctness can be checked cheaply through tests, formal rules, or physical measurement. Test held-out industries and include capability, price, liability, demand, exception rate, and redesign capacity. If verification cost adds no out-of-sample predictive value, retire the Verification Gradient without discarding the trace-capture hypothesis.

Study 3: Planning share versus execution share

Track who makes execution, planning, evaluation, and goal decisions as AI adoption matures. Predict that loss of planning and goal share will forecast hiring, wage, and autonomy effects better than execution automation alone over a preregistered two- to five-year horizon. Stable or rising human planning authority and unaided capability as execution delegation expands materially weakens the Residual Relevance Trap.

Study 4: Capability reciprocity

Randomize interface modes: full delegation, answer with explanation, critique of human work, guided questioning, and AI-free practice. Measure immediate output, delayed unaided performance, calibration, transfer, and error recovery. ANCHOR predicts that coaching and critique can preserve more capability than answer delivery, though the best mode will vary by task.

Study 5: Apprenticeship and expert supply

Follow firms that reduce junior hiring after AI adoption and compare later promotion, incident detection, senior recruiting costs, and novel problem-solving with matched firms or staged rollouts that preserve structured apprenticeship. Measure pre-adoption trends, task mix, growth, and hiring intent. The apprenticeship claim fails if AI-intensive firms consistently create expert capability faster than they consume it over a declared career-development horizon.

Study 6: Allocation of gains

Compare matched firms or policy discontinuities that direct savings to headcount reduction versus shorter hours, training, new products, lower prices, or gain-sharing. Measure pre-trends, adoption, productivity, retention, demand, wages, autonomy, and total employment. This tests whether the distributional outcome is partly an institutional choice rather than a model property.

Study 7: Open and closed evaluators

Compare actor–judge loops using closely related models with systems checked against independent humans, external tools, physical outcomes, or adversarial evaluators. Predict more correlated blind spots and proxy gaming in closed loops, especially on tail cases.

Study 8: Shadow practice and resilience

In domains where safe simulation is possible, compare organizations that preserve regular manual or unaided practice with those retaining humans only for escalation. Measure novel-incident detection, intervention quality, and recovery time. If passive oversight performs as well, the proposed practice requirement is unnecessary.

Study 9: Governance interventions

Randomize or exploit staged adoption of worker consultation, trace-use disclosure, appeal rights, gain-sharing, and staffing review. Measure whether these mechanisms change productivity, trust, distribution, and the path from augmentation to automation.

Claim-to-test map

The claims are modular; one cannot be rescued by evidence for another:

  • Core Recursive Capture Hypothesis: killed in its strong, cross-domain form by a preregistered pooled null or negative effect of structured trace retention on later delegated task share across the declared 12–36 month horizons, after randomizing or controlling automation intent and baseline codifiability.
  • Verification Gradient: retired independently if verification cost adds no out-of-sample predictive value beyond capability, price, risk, demand, exceptions, and redesign capacity.
  • Residual Relevance Trap: materially weakened if human planning and goal authority, unaided capability, and expert formation remain stable or rise as execution delegation expands over the declared two- to five-year horizon.
  • Grounding and Practice Risk: rejected in a domain when removing routine human practice does not reduce novel-incident detection, calibration, recovery, or expert supply relative to independent-tool and protected-practice baselines.
  • ANCHOR components: rejected one by one when authority floors, trace disclosure, capability-oriented interfaces, gain-sharing, or shadow practice fail to improve their prespecified outcomes against ordinary safety and work-design controls, or impose harms exceeding their benefits.
  • Ledger utility: rejected if its disaggregated measures add no decision-relevant prediction, distributional visibility, or appeal information beyond standard cost, quality, labor, and safety measures.

Evidence that demand and new tasks broadly and promptly offset substituted entry paths would defeat a general employment-ratchet story, even if trace capture still predicts task delegation. Evidence that productivity reliably flows to pay, leisure, access, or worker ownership would defeat the distributional warning, even if substitution occurs. That is how the framework loses without pretending every component must fall at once.

11. What this framework does not claim

It does not claim that most current models rewrite themselves. It does not claim that self-play in Go predicts open-ended social autonomy. It does not claim that synthetic data inevitably causes collapse. It does not claim that exposure equals displacement, that productivity lowers employment, or that aggregate job loss has already occurred.

It does not claim that every person must review every automated decision. Some tasks deserve complete automation. Nor does it make human judgment the supreme technical oracle. Formal proofs, instruments, empirical trials, and automated monitors may outperform people within valid domains.

It does not establish that human relevance is always good in every form. Exploitative labor can make a person causally “relevant” while violating dignity. Being necessary to clean up a system’s failures is not the future to preserve.

It does not settle the moral status of future artificial systems. If credible evidence of artificial welfare or sentience emerges, those systems may deserve consideration that is not captured here. That possibility does not remove present obligations to people affected by human institutions.

Finally, it does not promise that ANCHOR will defeat competition, politics, or power. A checklist cannot redistribute authority by itself. The framework is valuable only when its dimensions attach to decision rights, procurement, labor institutions, professional duties, liability, and public reason-giving.

Conclusion: preserve authorship, not bottlenecks

The deepest “enhance versus replace” phenomenon is not a switch hidden in a neural network. It is a recursive relation among assistance, observation, codification, evaluation, delegation, and reinvestment.

AI helps perform work. In doing so, it can reveal the structure of that work. The revealed structure becomes data, tests, interfaces, and organizational knowledge. Those artifacts make later automation cheaper. Successful automation changes staffing and concentrates investment in the next loop. Unless human learning, authority, and claims on the gain are represented, the selection system has no reason to conserve them.

That is the real danger in residual relevance. “Humans will always be needed for what machines cannot do” sounds reassuring, but it places human standing on a frontier designed to move.

The alternative is not to freeze the frontier. It is to stop confusing human value with friction.

Automate dangerous and pointless work. Expand access to expertise. Let individuals do what once required a bureaucracy or team. Use formal verification where formal truth is available. Build systems that teach, challenge, and extend people rather than merely answer for them. Share the gains.

At the same time, reserve authority over ends and stop conditions; protect standing that does not depend on comparative performance; maintain independent capability where intervention matters; govern the use of human traces; preserve plural contact with a changing world; and keep deployment reversible.

The future of human relevance will not be secured by finding one last task that only a human can do. It will be secured, if at all, by retaining human authorship of the institutions that decide which tasks should exist, what progress means, who benefits, who may refuse, and when the loop must stop.

Notes and sources

  1. OECD.AI, “What is AI? Can you make a clear distinction between AI and non-AI systems?” (2024), OECD.AI explanation. The source distinguishes autonomy and adaptiveness while locating objective-setting in a human-originated development process.
  2. International AI Safety Report 2026, full report. The report describes empirical understanding of AI-automated research-and-development feedback loops as minimal and emphasizes uncertainty about practical constraints and human oversight; it is paired in the text with METR’s affirmative assessment of present capability boundaries.
  3. David Silver et al., “Mastering the game of Go without human knowledge,” Nature 550 (2017): 354–359, paper. This is a closed-domain example with exact rules and an objective outcome, not evidence of open-ended autonomous evolution.
  4. Yuntao Bai et al., “Constitutional AI: Harmlessness from AI Feedback” (2022), research paper and summary. Human-written principles remain the normative input even when models generate critiques and preferences.
  5. Weizhe Yuan et al., “Self-Rewarding Language Models” (2024), arXiv. The reported benchmark improvements followed three iterative preference-training rounds; the method is bounded and seeded by existing human-authored instruction data.
  6. OpenAI, “Expanding on what we missed with sycophancy” (2025), postmortem. The incident is evidence that product feedback and release evaluation participate in organizational model evolution—and that immediate preference can be a defective proxy.
  7. Google DeepMind, “AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms” (2025), technical overview. Reported infrastructure improvements are first-party claims; humans still define problems, evaluators, validation, and production use.
  8. METR, “Frontier Risk Report: February to March 2026” (2026), assessment. The report finds increasingly consequential bounded R&D work but no demonstrated end-to-end autonomous research organization or dramatic general acceleration.
  9. Leo Gao et al., “Scaling Laws for Reward Model Overoptimization,” ICML 2023, paper and abstract. The result concerns optimization against imperfect learned rewards; it does not imply every metric necessarily fails.
  10. Shakked Noy and Whitney Zhang, “Experimental evidence on the productivity effects of generative artificial intelligence,” Science 381 (2023): 187–192, paper. The study used short, self-contained professional writing tasks.
  11. Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, “Generative AI at Work,” Quarterly Journal of Economics 140 (2025): 889–942, published article. The deployment involved 5,172 customer-support agents; effects were heterogeneous by experience and skill.
  12. Zheyuan Cui et al., “The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers” (2025), working paper. Completed-task counts do not by themselves establish maintainability, defects, learning, or employment effects.
  13. Fabrizio Dell'Acqua et al., “Navigating the Jagged Technological Frontier,” Organization Science (2026), open paper. The preregistered experiment involved 758 consultants and found gains within, but degraded correctness outside, the chosen AI capability frontier.
  14. Joel Becker et al., “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” (2025), study and limitations. METR cautions that the tool vintage may become outdated.
  15. Fabrizio Dell'Acqua et al., “The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise,” NBER Working Paper 33641 (2025), paper. The result concerns bounded product-innovation tasks and does not show that AI substitutes for teams generally.
  16. Anders Humlum and Emilie Vestergaard, “Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI,” NBER Working Paper 33777, revised 2026, paper. The Danish setting, adoption stage, and two-year horizon limit generalization.
  17. Paweł Gmyrek et al., “Generative AI and Jobs: A Refined Global Index of Occupational Exposure” (ILO, 2025), research brief and paper. Exposure gradients are task-capability measures, not job-loss forecasts.
  18. Lucia Foster et al., “Firm Data on AI Use and Employment Effects,” U.S. Census Bureau Center for Economic Studies Working Paper CES-WP-26-25 (2026), working paper page. The measures are firm-reported and capture an early period of adoption.
  19. Alex Tamkin et al., “Labor market impacts of AI: A new measure and early evidence” (Anthropic, 2026), report. The exposure measure partly uses Anthropic’s own platform data; the authors report limited employment evidence and urge humility.
  20. Anthropic Economic Index, “Learning curves” (March 2026), report. Claude use is not representative of all work, firms, models, or workers; its API/chat distinction is observational.
  21. Stanford Institute for Human-Centered AI, AI Index Report 2026, “Economy,” chapter. The chapter synthesizes multiple surveys and observational sources; individual estimates retain their underlying limitations.
  22. Daron Acemoglu and Pascual Restrepo, “Automation and New Tasks: How Technology Displaces and Reinstates Labor,” Journal of Economic Perspectives 33 (2019): 3–30, article.
  23. David Autor et al., “New Frontiers: The Origins and Content of New Work, 1940–2018,” Quarterly Journal of Economics 139 (2024), article. Historical new-task formation is evidence against a fixed lump of work, not a guarantee of timely or equitable transitions.
  24. Hamsa Bastani et al., “Generative AI without guardrails can harm learning: Evidence from high school mathematics,” PNAS 122 (2025), paper. This is an educational field experiment, not direct proof of workplace deskilling.
  25. Hao-Ping Lee et al., “The Impact of Generative AI on Critical Thinking,” CHI 2025, paper and abstract. The study reports associations and first-person examples, not causal loss of capability.
  26. Lisanne Bainbridge, “Ironies of Automation,” Automatica 19 (1983): 775–779, article record. Its human-factors argument predates generative AI but remains relevant to exceptional-condition oversight and skill maintenance.
  27. Ilia Shumailov et al., “AI models collapse when trained on recursively generated data,” Nature 631 (2024): 755–759, paper. The finding concerns indiscriminate generational training; it does not establish that all synthetic-data mixtures collapse.
  28. Madeleine Clare Elish, “Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction,” Engaging Science, Technology, and Society 5 (2019), open article.
  29. Virginia Doellgast et al., “Global case studies of social dialogue on AI and algorithmic management,” ILO Working Paper 144 (2025), report. Case comparisons identify institutional patterns but do not prove universal treatment effects.
  30. OECD.AI, “Human-centred values and fairness,” updated AI Principle 1.2, principle and rationale.
  31. NIST, AI Risk Management Framework 1.0, Core and human–AI interaction guidance, official resource. ANCHOR is a proposed supplement focused on longitudinal human capability, authority, trace terms, and distribution; it is not a NIST standard.

Selected bibliography

Acemoglu, Daron, and Pascual Restrepo. “Automation and New Tasks.” Journal of Economic Perspectives 33, no. 2 (2019): 3–30.

Autor, David, et al. “New Frontiers: The Origins and Content of New Work, 1940–2018.” Quarterly Journal of Economics 139, no. 3 (2024).

Bainbridge, Lisanne. “Ironies of Automation.” Automatica 19, no. 6 (1983): 775–779.

Bastani, Hamsa, et al. “Generative AI without guardrails can harm learning.” Proceedings of the National Academy of Sciences 122 (2025).

Brynjolfsson, Erik, Danielle Li, and Lindsey R. Raymond. “Generative AI at Work.” Quarterly Journal of Economics 140, no. 2 (2025): 889–942.

Dell'Acqua, Fabrizio, et al. “Navigating the Jagged Technological Frontier.” Organization Science (2026).

Elish, Madeleine Clare. “Moral Crumple Zones.” Engaging Science, Technology, and Society 5 (2019).

Humlum, Anders, and Emilie Vestergaard. “Still Waters, Rapid Currents.” NBER Working Paper 33777, revised 2026.

Shumailov, Ilia, et al. “AI models collapse when trained on recursively generated data.” Nature 631 (2024): 755–759.