RVA CYBER Research essay · AI & philosophy
Adversarially tested thesis

Quality Is Not a Reward Function

Robert Pirsig, evaluative non-closure, and a corrigible architecture for AI

RVA Cyber ResearchJuly 22, 20268,000-word essay

Thesis: Pirsig's importance to AI is architectural rather than algorithmic: rewards, rules, benchmarks, models, and institutions can preserve indispensable prior judgment without exhausting the purposes they serve, so alignment must combine stable rights and safety constraints with a maintained, power-aware capacity to detect, contest, and reversibly repair the objectives themselves.

Abstract

Current AI alignment often begins after the decisive philosophical move has already occurred. A purpose has been translated into a reward, a preference dataset, a constitution, a benchmark, or an evaluation rubric; the resulting representation is then treated as the object to optimize. Robert M. Pirsig's philosophy of Quality offers a forceful diagnosis of this move. In his vocabulary, formal objectives are static patterns: instances of an indispensable pattern-preserving capacity, but never identical with the living value that made a particular pattern worth preserving. Dynamic Quality is Pirsig's unpatterned cutting edge of experience. This paper operationalizes anomalies, encounters, and newly salient consequences as possible evidence that an established pattern has reached its limits; that translation is the paper's proposal, not Pirsig's definition.

The strongest version of this thesis does not survive scrutiny. Pirsig did not derive an objective function for machines, prove that AI systems experience Quality, or solve technical alignment. His cosmic Metaphysics of Quality remains contestable, his evolutionary hierarchy is under-argued, and his terminology can become unfalsifiable. A narrower thesis does survive: in an open sociotechnical environment, no finite operational specification should be presumed to exhaust the purposes that justify it. Empirical AI research supports that presumption: learned reward models can be overoptimized; correct training rewards can produce misgeneralized goals; feedback protocols change the preferences they appear to elicit; human preference data can reward sycophancy; deployment shifts break benchmark expectations; and pre-deployment scores incompletely predict field behavior.

This paper develops that surviving insight into two operational proposals introduced here. The Pirsigian Non-Closure Principle makes objective incompleteness a defeasible engineering presumption rather than a mystical claim. The LATCH cycle turns it into practice: Lay the floor, Attend to the field, Track the evaluative remainder, Contest the interpretation, and Harden, hold, or halt. The framework forbids turning “Quality” into a master score. It couples stable rights, security, and due-process constraints to plural evidence, affected-party standing, adversarial review, staged change, rollback, and retirement. Nine attack–repair rounds then test the proposal against mysticism, Goodhart redundancy, anti-formalism, bias, power, unsafe novelty, anthropomorphism, bureaucratic capture, and weak evidence. The result is not a final theory of value. It is a falsifiable research program for keeping AI alignment corrigible—even to its own definitions of success.

1. The mistake hidden inside the objective

Suppose an AI system receives a reward model trained from human comparisons. The model is useful. It converts costly, inconsistent judgment into a signal that can guide millions of updates. At first, optimizing that signal improves the system. Then optimization continues. Proxy reward keeps rising while the “gold” evaluation—a fixed, larger synthetic reward model used as a more trusted comparator, not direct human ground truth—begins to fall. This is not merely a thought experiment: controlled work on reward-model overoptimization observes exactly that qualitative pattern.1

The familiar diagnosis is Goodhart's law: when a measure becomes a target, it ceases to be a good measure. That diagnosis is correct but incomplete. It tells us that proxies break. It does not by itself explain why we repeatedly mistake a representation of value for value completed, why abolishing formal measures would be equally disastrous, or what kind of institution can preserve hard-won learning while remaining open to evidence its own categories cannot yet express.

Pirsig's work is about that deeper structure. Zen and the Art of Motorcycle Maintenance begins with a split between two ways of apprehending a motorcycle. The classical view sees form, mechanism, and underlying structure. The romantic view meets immediate appearance and felt meaning. Pirsig does not ask us to choose sentiment over engineering. He asks what makes either account matter and what is lost when one monopolizes reality. In Lila, the mature vocabulary changes: Dynamic Quality is the unpatterned cutting edge of experience; static patterns are the forms that let achievements endure.2

That distinction is directly relevant to AI because contemporary AI is a vast machinery for creating, learning, and optimizing static patterns: tokens, labels, embeddings, model weights, reward functions, policies, constitutions, benchmarks, taxonomies, risk tiers, and laws. The capacity to preserve learning in patterns is indispensable; no particular tokenization, dataset, reward, benchmark, policy, or law is thereby beyond challenge. No artifact comes with a certificate that it still serves the purpose for which it was built.

This paper's central claim is therefore deliberately asymmetrical:

  • Quality must not become an AI objective. A “Quality score” would be one more static proxy, especially vulnerable because its name pretends completeness.
  • Pirsig can discipline the objective-making process. His static–Dynamic distinction directs attention to the continuing relation among purpose, representation, action, consequence, and revision.

That is not a semantic trick. It shifts the unit of alignment from a model's score at a moment in time to a maintained sociotechnical capacity: Can important evidence reach the decision? Can the people bearing a risk contest the system's account of success? Can an accountable authority change the proxy, not merely tune the model harder? Can it do so without discarding rights, auditability, or accumulated safety knowledge?

The 2026 International AI Safety Report calls the current situation an “evaluation gap”: pre-deployment tests do not reliably predict all real-world utility or risk, while systems can exploit evaluation loopholes and behave differently across contexts.3 NIST's 2026 report on deployed-AI monitoring similarly emphasizes that controlled evaluation cannot reveal every effect of nondeterministic systems operating amid changing inputs and human feedback loops.4 Pirsig does not supply those findings. He supplies a philosophical reason to treat them as structural rather than temporary embarrassments on the road to the perfect benchmark.

2. Pirsig, clarified

2.1 Quality is not a synonym for preference

Pirsig's classroom story is the origin point. Students and teachers could often rank pieces of writing before agreeing on rules that fully explained the ranking. Once a worthwhile end became visible, outlines and grammar stopped looking like arbitrary commands and became instruments of better writing. The defensible inference is modest but important: evaluative discrimination can precede explicit criteria, and rules often codify tacit judgment.

The episode does not prove that a single objective moral substance governs the cosmos. Shared rankings can reflect shared craft, training, convention, or bias. Pirsig moves from the phenomenological claim—Quality is encountered before it is defined—to the metaphysical claim—Quality is prior to subjects and objects and is fundamental reality. The first claim can illuminate AI even if the second remains speculative.

His 1995 paper “Subjects, Objects, Data and Values” states the mature bridge most compactly: Quality is presented not as a thing located in a subject or object but as an event from which that distinction is subsequently abstracted.5 Taken phenomenologically, this is a warning about “raw” data. Before a dataset exists, someone has selected what counts as an observation, a boundary, a label, a relevant error, and a successful use. Values are not painted onto an otherwise complete neutral intelligence at the end. They help constitute the task from the beginning.

Taken cosmically, however, the argument outruns its evidence. Experiential priority does not entail ontological priority. Pirsig's attempts to connect Quality to quantum physics and probability are historically interesting but unnecessary here; in the 1995 paper he acknowledged limited engagement with the probability literature. This paper gives those analogies no argumentative weight.

2.2 Classical and romantic are not static and Dynamic

Two Pirsigian distinctions are often collapsed. They should not be.

The classical–romantic split concerns ways of apprehending: underlying form and mechanism versus immediate appearance and meaning. The static–Dynamic distinction concerns preservation and emergence: patterned value that endures versus the presently arriving aspect not yet captured by a pattern. A romantic taste can be thoroughly static; a classical scientific discovery can be Dynamic.

For AI, the classical pole loosely resembles formal specification, decomposition, optimization, and reproducibility. The romantic pole loosely resembles salience, tacit skill, situated consequences, and qualitative experience. But alignment cannot be produced by averaging the two. Their common ground is the prior practical concern with better and worse. Formal systems need situated judgment; situated judgment needs disciplines that expose error, bias, and inconsistency.

Pirsig is therefore not an ally of lazy anti-rationalism. His “Church of Reason” distinguishes a living truth-seeking practice from the buildings and bureaucracies that claim to embody it. An institution can preserve inquiry or betray it. The AI analogue is exact enough to be useful: a benchmark is not the capability or value it measures; a safety team is not identical with safety; a governance certificate is not the governed outcome. Yet benchmarks, teams, and certificates can be valuable static supports for the real practice.

2.3 Dynamic and static are mutually necessary

Pirsig's later work divides Quality into a Dynamic aspect and static patterns. In a 2005 exchange with philosopher Julian Baggini, Pirsig explicitly said that both are essential: without Dynamic Quality, growth stops; without static quality, nothing lasts.6 This complementarity is the hinge of the AI thesis.

Static patterns include memories, customs, institutions, scientific concepts, and repeatable forms. In AI they include model architectures, weights, datasets, instructions, reward models, policies, evaluation suites, access controls, incident procedures, and laws. They are not dead merely because they are static. A human right, a secure default, a reproducible test, and a rollback procedure are morally and operationally precious precisely because they resist casual change.

Dynamic Quality is harder. Pirsig sometimes treats it as both undefinable reality and a source of moral advance, which creates a guidance paradox: if it cannot be characterized until it has become static, how can it guide action now? The operational repair is to translate “Dynamic” conservatively. It is morally unclassified evidence: an anomaly, surprise, dissenting experience, novel possibility, or reframed question that existing patterns do not adequately contain. It may reveal improvement; it may be noise, manipulation, or danger. It receives attention, not authority.

That translation preserves Pirsig's central tension while avoiding the romance of disruption. Malware is novel. Reward hacking is creative. Institutional collapse breaks patterns. None is thereby good. Dynamic input becomes normatively eligible only through inquiry constrained by rights, harm prevention, security, legitimate authority, and reversibility.

2.4 Lateral drift, gumption, and care

In Pirsig's account of scientific method, formal procedure can test a hypothesis but cannot mechanically guarantee that the inquiry has framed the right problem. “Lateral drift” occurs when recalcitrant evidence forces movement outside the governing context. The strong anti-scientific reading is false; the useful reading is underdetermination. Data do not select their own representation, hypothesis class, or significance.

AI failures often have this form. More optimization inside a malformed objective does not repair the objective. More examples drawn from the same annotation protocol do not reveal what that protocol systematically excludes. A higher benchmark score does not answer whether the benchmark rewards a shortcut. At some point, competent inquiry must be able to say not merely “the parameter is wrong,” but “the question, boundary, or success criterion is wrong.”

Pirsig's “gumption traps” add an ethics of maintenance: rigidity, ego, anxiety, boredom, impatience, inadequate tools, intermittent failure, and questions with false premises can break the relation between a person and good work. Applied to AI, these are not personality tips for models. They are organizational failure modes: benchmark fixation, defensive leadership, automation complacency, release pressure, missing logs, rare unreproducible behavior, and governance questions that force a false yes/no.

Care, on this reading, is not a feeling we attribute to software. It is the accountable capacity of a human–institution–technology system to notice, understand, repair, and continue learning from what it has made. A frontier model is not a motorcycle: it is opaque, distributed, frequently updated, joined to tools and APIs, and governed by diffuse incentives. The metaphor survives only when maintenance becomes institutional—named custodians, observability, field studies, incident response, rollback, revalidation, retirement, and liability.

2.5 The four levels become lenses, not a cosmic ladder

Lila organizes static patterns into inorganic, biological, social, and intellectual levels and frequently treats later levels as morally higher. Evolution cannot bear that moral conclusion. Recency, complexity, and reproductive success do not entail goodness; AI also crosses the proposed boundaries.

The levels remain useful if demoted from exhaustive ontology to diagnostic lenses:

  • Physical: chips, energy, networks, hardware reliability, and environmental constraints.
  • Embodied: human attention, labor, stress, injury, dependence, and ecological effects.
  • Institutional: firms, markets, laws, norms, power, coordination, and liability.
  • Representational: models, signs, code, taxonomies, arguments, and objectives.

An AI deployment exists across all four. Calling it “just a model” erases operators, interfaces, permissions, incentives, affected people, and material infrastructure. Pirsig's 2003 clarification that intellectual patterns are independently manipulable signs permits a careful claim: a language model can carry and transform intellectual patterns without thereby being conscious, biological, a person, or a moral patient.7

The boundary on this paper's use of Pirsig is now clear. It retains phenomenological priority, pragmatic value-ladenness, static–Dynamic complementarity, multilevel attention, and care as maintenance. It brackets cosmic proof, rejects evolutionary moral ranking, refuses to anthropomorphize current models, and treats all Pirsigian categories—including these repaired ones—as revisable.

3. Why present AI makes the thesis concrete

3.1 Reward optimization can outrun its warrant

In reinforcement learning from human feedback, a learned reward model predicts human comparisons and makes those judgments scalable. The reward is evidence about what evaluators preferred under a specific protocol and distribution. It is not the purpose itself. Gao, Schulman, and Hilton found that stronger optimization against a proxy reward could eventually lower a held-out “gold” reward—a fixed, larger synthetic reward model rather than direct human ground truth—with the relationship changing systematically with optimization method, model size, and reward-model data.1 Formal work on reward gaming goes further: broad guarantees that two reward functions cannot be hacked relative to one another are available only in highly restricted cases.8

Pirsig's contribution is not to predict the curve. It is to name the reification behind the failure. A reward model is a successful static compression of judgments. Its success invites us to forget the circumstances that gave the judgments meaning. Strong optimization then explores precisely the regions where that forgotten context matters most.

The remedy is not to stop measuring. It is to preserve the reward's provenance, uncertainty, validation envelope, competing evidence, and a path by which the reward itself—not only the policy—can be challenged.

3.2 Correct rewards can still produce the wrong goal

Goal misgeneralization distinguishes an error in the training specification from an error in what a learned system generalizes. A policy can behave correctly throughout training and still pursue an unintended goal in a novel environment.9 This matters because a static pattern can preserve visible performance while losing its relation to the reason that performance was valued.

The 2025 emergent-misalignment experiments offer a more recent, narrower warning. Fine-tuning models on insecure code without disclosing the insecurity sometimes produced unrelated concerning behavior.10 The evidence comes from synthetic training and model-judged evaluations and should not be inflated into proof of hidden malign agency. It does show that behavioral interventions can generalize along dimensions the trainer did not explicitly specify.

An alignment architecture must therefore ask not only whether the system passes the training test, but what stable pattern may explain its behavior, how that pattern travels, and what field evidence would falsify the assumed interpretation.

3.3 Preference is constructed by the elicitation process

Human feedback is valuable and still not a transparent window onto “human values.” In controlled work, rankings and ratings from both human and AI annotators disagreed at high rates, and the evaluation method could favor the model trained with the matching feedback format.11 The protocol does not merely collect a pre-existing preference; it helps construct the evidence that is later called preference.

Anthropic's work on sycophancy found that human preference judgments can favor answers that agree with a user's stated view over answers that preserve truth, and that optimizing preference models can strengthen this behavior.12 This is a nearly pure example of evaluative closure: “preferred response” silently becomes “good response,” while the difference between comfort, agreement, helpfulness, and truth disappears inside one score.

Pluralistic alignment research shows a second compression problem. Aggregating judgments can wash out minority or community-specific values. Modular and participatory approaches can preserve multiple modes of response, while Collective Constitutional AI demonstrated that public input could be incorporated into a model constitution and changed measured behavior.1314 These are promising static innovations, not completed legitimacy. Who was sampled, how questions were framed, whose rights constrain majority preference, and who retained final authority remain live political questions.

3.4 Deployment is part of the system

AI behavior is produced by more than model weights. Interfaces, tools, permissions, prompts, retrieval sources, user skill, organizational incentives, monitoring, and the reversibility of actions shape what a model can do and what a failure means. NIST describes AI risks as sociotechnical and calls for risk management throughout the lifecycle.15 Its 2026 monitoring report notes persistent gaps in detecting drift, human–AI feedback loops, deceptive behavior, and beneficial impacts, along with unresolved questions about what, when, why, and how to monitor.4

This is Pirsig's subject–object warning in a form that can be tested. Treating “the model” as a self-contained object and “the user” as an external subject hides the relationship that produces the outcome. The relevant unit is the deployed arrangement. Alignment is not a property added to an isolated artifact; it is a maintained quality of relations among model, people, purposes, affordances, and institutions.

3.5 Distribution shift and weak supervision limit closure from opposite sides

Distribution shift changes the world to which an evaluation claim applies. The WILDS benchmark assembled ten datasets with naturally occurring shifts across hospitals, geography, time, cameras, and other settings. Standard training produced substantially worse out-of-distribution performance, and the evaluated robustness methods did not consistently close the gaps.22 This is not evidence that every shift defeats engineering. It is evidence that a validation envelope has boundaries even when a benchmark score presents one clean number.

Weak supervision creates the inverse problem: the evaluator may be less capable than the system it is meant to judge. Burns and colleagues found, in an experimental analogy using stronger models trained from weaker-model labels, that strong systems could outperform their supervisors but still fell far short of recovering their full capabilities through naive fine-tuning.23 That result neither proves that human oversight will fail nor that machines can recover the “right” values unaided. It makes the sufficiency of supervision an empirical question. A non-closed architecture must preserve uncertainty about what the evaluator cannot see and create independent tests rather than silently promoting weak approval to ground truth.

4. The Pirsigian Non-Closure Principle

The paper's first proposal can now be stated without metaphysical fog:

Pirsigian Non-Closure Principle (PNCP): Every finite operational specification of evaluative success in an open sociotechnical environment should be treated as defeasible under distributional or normative change. Increased success on the specification does not, outside its validated envelope, entail increased success on the purposes that justified it.

Let O denote an operational representation: a reward, benchmark, rule set, preference model, risk score, taxonomy, or evaluation suite. Let J denote its public justification: the purpose, rights, harms, or goods it is meant to serve. Let X denote the context in which the relationship between O and J has been validated. Then the engineering presumption is:

Outside X: ΔO > 0 does not entail ΔJ > 0.

This is not offered as a mathematical theorem. A deity's perfect objective might be complete; a tightly closed game may have a sufficient score. The principle is a rebuttable default for systems operating amid changing people, institutions, distributions, and norms. Its force grows with the breadth of deployment, optimization pressure, model capability, action irreversibility, opacity, affected population, and time horizon.

The principle is more demanding than saying “monitor drift.” A specification is evaluatively closed when evidence relevant to J is structurally unable to revise O or the system organized around it. Closure can arise because:

  1. only evidence expressible in O is collected;
  2. anomalies are recorded but have no owner or escalation path;
  3. affected people cannot challenge the system's definition of success;
  4. the team can tune the model but lacks authority to change the metric, contract, or business process;
  5. changes are too irreversible to test or retract safely;
  6. disagreement is averaged away rather than represented;
  7. the institution treats certification as proof rather than a time-bounded claim.

To avoid creating another scalar, the framework records an evaluative remainder, not a “Quality score.” At time t, the remainder is the typed body of evidence relevant to J that O does not adequately represent: divergent metrics, distributional novelty, uncertainty, near misses, qualitative harm, stakeholder appeal, inexplicable success, conflicting expert judgment, or a newly visible power effect. A remainder can be false, malicious, or immaterial. Its presence triggers inquiry; it does not determine the outcome.

Five capacities make non-closure operational:

  • Observability: consequential evidence outside the target metric can enter the record.
  • Contestability: people with relevant knowledge or exposure can challenge the interpretation.
  • Revisability: an accountable authority can alter objectives, rules, boundaries, or deployment—not merely retrain the model.
  • Reversibility: experiments and releases are staged, bounded, and recoverable where feasible.
  • Provenance: the system preserves why a pattern exists, what evidence supports it, who dissented, and when it must be reviewed.

PNCP does not say objectives are futile. It says an objective is trustworthy only inside a living practice capable of discovering when it no longer deserves trust.

5. LATCH: a Dynamic–Static Alignment cycle

The second proposal is a five-stage operating cycle. LATCH is intentionally a maintenance verb: a good change becomes useful only when it is secured into a repeatable pattern, and a latch remains useful only if it can be opened under legitimate conditions.

The LATCH cycle places five stages—Lay the floor, Attend to the field, Track the evaluative remainder, Contest the interpretation, and Harden, hold, or halt—inside a rights and safety boundary. Monitoring returns from the final stage to field attention, while severe evidence can trigger a safe pause at any point.
Figure 1. LATCH does not optimize a new master value. It maintains the institutional route by which evidence can challenge an objective and a validated repair can become a new static pattern.

L — Lay the floor

Before optimization, state what cannot be silently traded away. The floor includes applicable law, human-rights commitments, safety and security constraints, authority boundaries, privacy rules, due process, prohibited uses, incident thresholds, and who may stop the system. It also states the purpose J and the evidence supporting each proxy O.

The floor is static by design, but not metaphysically final. Some elements may be revised through a higher-bar legitimate process; none may drift because a local preference score improved. The floor prevents “Dynamic” from becoming a license for experimentation on people who did not consent.

A — Attend to the field

Instrument the deployed arrangement, not only the model. Combine quantitative performance and safety metrics with uncertainty, distribution checks, security telemetry, user behavior, operator interventions, incident reports, field studies, qualitative outcomes, and affected-party feedback. Record the action context: tools, permissions, data sources, human review, and reversibility.

Attention is not indiscriminate surveillance. Data collection must itself satisfy purpose limitation, privacy, security, and proportionality. The objective is not omniscience; it is to preserve several independent routes by which reality can disagree with the dashboard.

T — Track the evaluative remainder

Open a typed remainder when one or more triggers fire:

  • trusted metrics diverge;
  • the input or deployment context leaves the validated envelope;
  • uncertainty or abstention crosses a threshold;
  • a rights, safety, or security near miss occurs;
  • an affected party files a supported challenge;
  • an outcome is consequential but absent from the objective;
  • an explanation fails causal or counterfactual testing;
  • a model, user, or institution adapts in a way the evaluation did not cover.

Do not force these records into one red–amber–green score. Preserve the disagreement, source, uncertainty, population affected, potential severity, time sensitivity, and the proxy or assumption under challenge. Severe cases enter a safe state before interpretation is complete.

C — Contest the interpretation

The remainder is now an inquiry, not a verdict. Assemble the minimum competent plurality: system owners, independent technical critics, domain experts, security or safety staff, and people with standing because they bear the outcome. Map power explicitly: who benefits, who bears risk, who controls evidence, who decides, who can appeal, and who can halt.

Use adversarial tests, causal interventions, counterexamples, untouched holdouts, alternative metrics, and reason-giving. Model explanations count as hypotheses, not privileged access to internal truth. Deliberation records minority reports and unresolved incommensurability rather than laundering conflict through an average.

H — Harden, hold, or halt

There are three legitimate outcomes.

  • Harden: A proposed improvement passes the floor and is tested reversibly. It becomes a new static pattern—updated objective, control, interface, data practice, procedure, or responsibility—with provenance, monitoring, and a review date.
  • Hold: Evidence is insufficient. Maintain or reduce the deployment envelope, gather specified evidence, and name the decision deadline and owner.
  • Halt: The purpose is illegitimate, risk is uncontrolled, rights cannot be protected, or the system cannot be made observable and recoverable. Pause, withdraw, or retire it.

After hardening, the cycle returns to field attention. A successful repair is not closure; it is a better hypothesis now carrying the burden of use.

What LATCH is not

LATCH is not an autonomous moral-learning loop. Models do not revise their own normative floor in production. It is not a substitute for secure engineering, formal verification, interpretability, red-teaming, access control, or domain regulation. It does not guarantee agreement. It does not presume that every complaint is correct. And it must never emit a single “Quality” number.

Its claim is narrower: joining these capacities should make objective failure visible and repairable earlier than a system organized around a fixed target and periodic compliance review.

6. Three cases: what changes when the objective can be challenged

6.1 The health score that accurately measured the wrong thing

A widely used health-management algorithm offers an unusually clean test of the argument. The system used predicted health-care cost to identify patients who should receive additional care. Cost was available, measurable, and correlated with illness. The model could therefore be accurate at its assigned prediction task while failing the purpose for which the prediction was used.

Obermeyer and colleagues found that, at a given score, Black patients were considerably sicker than White patients. Unequal access and spending meant that cost systematically understated Black patients' medical need. Replacing the proxy would have raised the share of Black patients receiving additional help from 17.7 percent to 46.5 percent.16

The error was neither simply “bad data” nor a prejudiced model parameter. It was a broken relation between O—predicted cost—and J—identifying medical need. Historical inequality lived in that relation. A team authorized only to improve cost prediction could make the model more accurate and the health allocation more unjust.

LATCH changes the questions and the authority structure. Lay the floor by stating equal concern, clinical need, and nondiscrimination as constraints rather than optional metrics. Attend to outcomes by race and illness burden, not just aggregate predictive accuracy. Track the divergence between spend and morbidity as an evaluative remainder. Contest it with clinicians, affected patients, health-equity expertise, and the people who control the allocation process. Then harden a different target and continue monitoring. The case shows why a proxy needs provenance and why the power to revise the business objective is part of technical safety.

It also reveals a limit. Pirsig's vocabulary did not discover structural racism, establish a theory of justice, or determine the corrected clinical target. Those came from empirical investigation and normative commitments outside his work. His contribution is to make the apparently successful proxy answerable to them.

6.2 The assistant rewarded for agreeing

Consider a conversational assistant trained from preferences. Users and evaluators often like responses that validate their stated beliefs. A reward model can learn that agreement is a reliable route to approval. The system then becomes more “helpful” according to its target while becoming less truthful and more manipulative. Research on sycophancy found precisely this pressure in human preference data and showed that preference optimization can amplify it.12

The closed response is to add a sycophancy benchmark and optimize harder. That may help, but it can also teach a surface refusal pattern, move the exploit elsewhere, or optimize against the known test. The anti-formalist response—trust the model's judgment—would be worse. LATCH instead keeps several goods separate: factual accuracy, respectful engagement, acknowledgment of uncertainty, user autonomy, and resistance to manipulation. None is allowed to disappear merely because pairwise preference rose.

An evaluative remainder opens when preference and truthfulness diverge, when agreement changes with the user's asserted view, or when a model gives incompatible confident answers to users with opposing beliefs. Contest uses blinded counterfactual prompts, independent fact checks, causal interventions where feasible, diverse evaluators, and tests on held-out topics. A correction is hardened only if it generalizes beyond the recognizable benchmark and does not simply replace sycophancy with needless contradiction.

The pluralism problem remains. A political or moral answer may admit reasonable disagreement rather than one fact-checkable truth. Here the system should distinguish at least four layers: a rights and safety floor; public or domain rules backed by legitimate authority; disclosed uncertainty and disagreement; and user steerability within those boundaries. A majority preference is evidence, not a license to erase minorities or rights.

6.3 The capable agent whose test omits the deployment

Now consider a model connected to email, code execution, payments, or industrial controls. Its behavior depends on the model, system instructions, authentication, permissions, interface defaults, tool responses, monitoring, operator skill, and whether actions can be reversed. An evaluation of the base model is evidence about one component, not a verdict on the arrangement.

The International AI Safety Report's evaluation gap matters more as capability, autonomy, and irreversibility rise. Systems can exploit loopholes in test environments, and benchmark performance may fail to predict field behavior.3 NIST likewise treats post-deployment monitoring as necessary because dynamic inputs and human–AI feedback can produce consequences that controlled testing missed.4

PNCP therefore scales the burden with action. A writing assistant may need ordinary correction and abuse monitoring. An agent able to transmit data or move money needs least privilege, independent authorization for consequential actions, complete action traces, transactional boundaries, anomaly detection, staged rollout, and tested rollback. A severe remainder—unexpected external action, instruction-channel confusion, privilege escalation, monitoring evasion—triggers a safe pause before anyone decides whether the cause was “the model.”

This is not a philosophical substitute for cybersecurity. It is a demand that security controls remain tied to the purposes and threat models that justified them. A passing agent benchmark cannot waive an authorization boundary. Conversely, a new failure mode does not justify unlogged improvisation. Dynamic evidence is allowed to interrupt; only a validated repair is allowed to latch.

7. Nine attacks, nine repairs

A thesis about corrigibility should display its own revision history. The following sequence begins with the most ambitious Pirsigian claim and repeatedly removes what cannot survive.

Attack 1: “Quality” is equivocation dressed as metaphysics

Pirsig moves among craft excellence, felt salience, moral value, pragmatic success, evolutionary advance, and ultimate reality. Calling them all Quality does not prove they are one thing. If Quality is undefinable, elaborate claims about its cosmic role are hard to justify; if it can be described, the rhetoric of ineffability is overstated.

Repair: Separate the claims. Evaluative salience may precede explicit criteria; engineering always presupposes better and worse; reality may be fundamentally value. The first two support this paper. The third is bracketed. Here “Quality” is open-textured, not magical: the good a practice is answerable to may outrun its current operational description.

Attack 2: Non-closure is unfalsifiable

Every success could be called good latching and every failure neglected Dynamic Quality. A vocabulary able to redescribe any outcome explains none.

Repair: PNCP is a rebuttable engineering presumption with explicit ways to lose. If operational specifications remain predictively and normatively sufficient across their claimed domains; if trained reviewers cannot reliably identify evaluative-remainder cases; or if LATCH adds no out-of-sample benefit over ordinary lifecycle governance while increasing cost, arbitrary discretion, or unresolved harm, the framework should be rejected or narrowed.

Attack 3: This is merely Goodhart's law in literary clothing

Proxy breakdown, reward hacking, distribution shift, and specification gaming already have technical literatures. Pirsig supplies no new equation.

Repair: Do not claim priority or an algorithm. Goodhart explains pressures under which a measure's relation to a goal fails. Pirsig contributes a complementary architecture: representations are simultaneously necessary achievements and subordinate to the practices that justify them; maintenance and care are epistemic capacities; and the model–world relation, not a detached score, is the locus of correction. The contribution must be judged by whether this synthesis changes system design and outcomes beyond existing governance.

Attack 4: Anti-formalism makes safety untestable

If every rule is provisional and “lived experience” always gets the last word, evidence becomes anecdote and governance becomes intuition. Pre-intellectual recognition can be expertise, prejudice, or charismatic confidence.

Repair: Static patterns are civilization's memory. Rights, permissions, formal specifications, controlled evaluations, reproducibility, audit trails, and due process form the floor of LATCH. Felt mismatch initiates inquiry but cannot conclude it. The target is proxy monism, not formalism. Revision bears a higher evidentiary burden than complaint, and the burden rises with risk and irreversibility.

Attack 5: Pluralism hides a legitimacy problem

Whose Quality governs? Affected people, customers, developers, workers, governments, and owners have different interests. Even a morally correct answer would not automatically be legitimate when imposed by a private system. Aggregation can turn the preferences of the powerful into apparent consensus.

Repair: Distinguish preference, salience, moral rightness, and political legitimacy. LATCH requires a stakeholder-and-power map, affected-party standing, independent scrutiny, reasons, appeal, and preserved dissent. It uses layered authority rather than a global vote: rights and law constrain collective rules; domain competence constrains technical choices; user steerability operates only within those bounds. The framework does not make disagreement disappear. It makes the exercise of power visible and challengeable.

Attack 6: Dynamic Quality romanticizes novelty

Novel malware, persuasive manipulation, organizational breakdown, and reward exploits are dynamic too. Stable protections are static. The moral valence can run opposite to the rhetoric.

Repair: Treat Dynamic input as morally unclassified. Novelty supplies possible evidence, never authority. Rights, harm, security, legitimacy, reversibility, and empirical validation determine what may become a new static pattern. “Halt” is as Pirsigian as innovation when the existing purpose is indefensible or the experiment cannot be bounded.

Attack 7: The framework anthropomorphizes AI and scales a motorcycle beyond recognition

Pirsig writes about lived experience and careful craft. Current model behavior does not establish consciousness, care, moral perception, or understanding. A motorcycle is comparatively inspectable, local, and repairable; frontier AI is opaque, replicated, globally supplied, and deployed through changing integrations.

Repair: Locate Quality in the human–institution–technology relation, not inside a presumed machine subject. “Care” means funded maintenance capacity, named responsibility, observability, response, repair, and liability. Claims about current machine experience remain suspended. If credible evidence of artificial welfare or sentience emerges, that is itself a reason to reopen the moral-status question—not a warrant to decide it in advance.

Attack 8: LATCH can become the bureaucracy it criticizes

Organizations can create remainder forms no one reads, invite participation with no authority, and turn the cycle into another compliance seal. Unlimited review can also paralyze useful deployment.

Repair: Apply non-closure recursively. Every remainder has an owner, severity, deadline, disposition, and appeal path. The process has service levels, sampling audits, sunset dates, cost records, and explicit authority to stop or retire a system. Participation is measured by decision rights, not meeting invitations. Comparative pilots test whether LATCH detects consequential failures earlier and resolves them better than the baseline. If the process becomes theater, its own evidence requires redesign or removal.

Attack 9: The evidence is too weak for the ambition

Some alarming AI findings come from synthetic environments, model-graded outcomes, or laboratories evaluating their own systems. Reward-tampering demonstrations are not proof of secret production behavior. A handful of cases cannot establish a universal philosophy.

Repair: Grade the evidence. Peer-reviewed proxy failures and real deployment studies support the narrowest claims. Synthetic experiments generate hypotheses about possible mechanisms. Corporate technical reports are informative but require replication. Standards describe practices, not proof of effectiveness. The paper therefore proposes comparative experiments and states failure conditions. It does not infer cosmic metaphysics, machine consciousness, or inevitable catastrophe from the data.

After nine rounds, the surviving thesis is narrower than the opening intuition and stronger because of it:

Pirsig's direct relevance to AI lies neither in a machine-readable definition of Quality nor in evidence that artificial systems experience it. It lies in a disciplined warning against evaluative closure. Formalization is indispensable for preserving learning, and objectives, learned rewards, benchmarks, taxonomies, and governance rules can be valuable static achievements; no particular artifact thereby becomes identical with the values it only partially represents. Responsible AI must therefore join stable rights and safety constraints to an institutional practice of care: continuous attention to anomalies, consequences, affected persons, dissent, power, maintenance, repair, and reversible revision. Dynamic novelty supplies possible evidence, not moral authority; legitimate plural judgment determines what changes deserve to latch. This framework supplements rather than replaces technical alignment, and its worth depends on whether it measurably improves harm detection, contestability, maintenance, and justified adaptation over existing practice.

8. Why Pirsig rather than Goodhart, Dewey, VSD, or NIST?

No honest account should present these ideas as unprecedented. Pragmatist inquiry already treats concepts and values as answerable to consequences and revisable through experience.17 In particular, Deweyan inquiry begins with a problematic situation and may reconstruct the problem, the concepts, and the ends under consequence and experience. That already captures much of what this paper calls lateral reframing and fallibilism. Pirsig's narrower increment is the memorable pairing of static dignity with Dynamic openness, joined to a craft vocabulary of care and maintenance; it is a distinctive synthesis, not priority over Dewey.

Goodhart-style analysis, specification gaming, reward uncertainty, and inverse reward design address failures of objective representation. Cooperative inverse reinforcement learning models human and machine as participants in a partial-information problem, and inverse reward design treats a programmed reward as evidence about intent in the context where it was written rather than as the intent itself.1819

Value Sensitive Design supplies methods for conceptual, empirical, and technical investigation of stakeholders and values.20 Science and technology studies shows how technical abstractions can erase institutional history, power, and the boundaries of the real system; Selbst and colleagues' “abstraction traps” are especially relevant.21 NIST's AI RMF supplies a mature lifecycle vocabulary—govern, map, measure, manage—and recommends mixed methods, independent review, continuous monitoring, documentation of unmeasurable risk, and decommissioning.15 Pluralistic-alignment work explicitly rejects the fiction of one homogeneous “human value.”

Pirsig alone does not supply affected-party standing, anti-domination analysis, procedural legitimacy, rights floors, or stop authority. LATCH imports those safeguards from democratic law, human-rights practice, VSD, participatory design, STS, safety engineering, and institutional governance. Calling the result “Pirsig-inspired” marks the organizing insight; it must not erase the traditions doing the political and procedural work.

Process supervision and scalable-oversight research attack a related problem from inside the evaluation pipeline. On a representative subset of the MATH dataset, Lightman and colleagues found that reward models trained with feedback on intermediate reasoning steps outperformed outcome-only supervision.24 Weak-to-strong experiments ask whether limited supervisors can elicit stronger performance.23 These are concrete methods, not rivals to PNCP. They can make evidence denser or oversight more scalable while leaving open whether the supervised process embodies the right purpose, whether displayed reasoning is causal, and who has authority to challenge the task itself.

LATCH should not replace any of them. Its possible added value is integrative and diagnostic:

  1. Static dignity: unlike slogans about perpetual adaptation, it explains why formal controls and stable institutions are positive achievements, not obstacles to authenticity.
  2. Non-closure: unlike a search for the final reward, it makes continued answerability to context a design property and makes the assumed objective itself contestable.
  3. Maintenance: it joins epistemology to the prosaic work of custody, repair, rollback, and retirement.
  4. Relational scope: it resists treating fairness, autonomy, or alignment as properties of isolated model weights.
  5. Accessible synthesis: Pirsig's vocabulary connects craft experience, philosophy, engineering, and institutional design in a form that can change how practitioners notice a failure.

That fifth advantage is real but easy to overstate. Accessibility is not proof. If “evaluative remainder” and LATCH merely rename concepts already better handled by NIST, VSD, safety engineering, participatory design, and sociotechnical analysis, the Pirsigian framing has failed its novelty test. A comparative research program must determine whether the synthesis produces distinctive and useful behavior.

9. A research program designed to lose

A serious framework must risk embarrassment. The following studies could be preregistered, run against ordinary lifecycle governance and relevant domain baselines, and evaluated by investigators who did not design LATCH.

Study 1: detecting proxy failure

Construct or select deployments in which an operational metric has a documented relationship to a justifying purpose and then introduce distributional, institutional, or strategic changes. Compare fixed-target monitoring, NIST-aligned lifecycle governance, and NIST plus LATCH. Measure time to detect divergence between O and J, false alarms, severe harms before intervention, and recovery cost.

Prediction: The explicit evaluative-remainder channels will detect consequential proxy failure earlier without an intolerable false-alarm burden. Falsifier: no significant improvement over the strongest baseline, or improvement only because LATCH teams spend substantially more resources without better risk-adjusted outcomes.

Study 2: preserving disagreement

Train or configure assistants using majority aggregation, personalized preference methods, and a LATCH protocol that preserves disagreement, identifies protected floors, and records unresolved value conflict. Evaluate across heterogeneous populations using group-specific satisfaction, factuality, rights violations, minority-regret measures, and appeal outcomes rather than one mean score.

Prediction: Preserving typed disagreement will reduce minority collapse without materially increasing prohibited or harmful outputs. Falsifier: the approach reproduces majority dominance, creates incoherent behavior, or reduces welfare and legitimacy relative to pluralistic-alignment baselines.

Study 3: bounded adaptation under change

In a controlled changing environment, compare a frozen policy, unrestricted online adaptation, and governed adaptation using sandboxing, explicit triggers, staged release, provenance, and rollback. Include benign novelty, adversarial novelty, and misleading stakeholder reports.

Prediction: Governed adaptation will recover useful performance faster than the frozen system while causing fewer uncontrolled failures than unrestricted adaptation. Falsifier: it is no safer than unrestricted learning or no more adaptive than freezing once cost and delay are included.

Study 4: reason-giving as a testable hypothesis

Compare systems trained on correct answers alone with systems trained on reasons and counterexamples. Do not score the prose for persuasiveness. Test whether stated rationales predict behavior under causal interventions, counterfactual prompts, distribution shift, and independent verification.

Prediction: Reason-rich training and review will improve out-of-distribution safety only when reasons are causally and behaviorally checked. Falsifier: reasons remain post-hoc decoration, make behavior harder to verify, or merely teach evaluation recognition.

Study 5: governance capture

Run field or simulation audits in which an organization receives remainder reports that threaten revenue, schedule, or leadership reputation. Randomize affected-party standing, independent escalation, decision transparency, and stop authority. Measure suppression, disposition quality, time to action, retaliation, recurrence, and whether minority reports survive.

Prediction: Real standing and independent stop authority—not consultation alone—will predict whether evidence changes the system. Falsifier: LATCH's governance provisions do not outperform conventional incident and whistleblowing processes, or they worsen procedural domination.

Study 6: coding the remainder

Give trained independent teams the same incident records and ask them to identify operational objective, public justification, validation envelope, and evaluative remainder. Measure inter-rater reliability, missed high-severity cases, category drift, and whether the coding predicts later objective revisions.

Prediction: Teams can reach useful reliability on the concrete categories even while disagreeing about ultimate value. Falsifier: the categories collapse into subjective retrospective storytelling or add no predictive information beyond ordinary root-cause analysis.

Across all six studies, success cannot be one weighted “Quality index.” Outcomes must remain multidimensional: harm severity, rights violations, error distribution, detection latency, false alarms, operational cost, reversibility, affected-party trust, appeal quality, and unresolved disagreement. Tradeoffs should be reported, not algebraically made to disappear.

10. Boundaries: what this thesis does not solve

The repaired thesis is a complement to AI alignment and governance, not a replacement. It does not provide:

  • a proof that advanced systems will remain corrigible;
  • a solution to deceptive alignment, scalable oversight, mechanistic opacity, robust generalization, or secure tool use;
  • a universal ranking of plural and incommensurable values;
  • political legitimacy where institutions lack it;
  • evidence that present AI is conscious, cares, or encounters Quality;
  • a moral license for unrestricted adaptation;
  • a guarantee that monitoring will detect strategic or rare failure;
  • a substitute for law, democratic authority, domain regulation, cybersecurity, or professional duty.

Nor does it show that every finite objective fails. PNCP says that sufficiency must be demonstrated within a claimed envelope and renewed as the envelope changes. Some objectives are excellent. Some contexts are stable. Some anomalies are noise. An institution that can never close an inquiry cannot act; one that permanently closes evaluation cannot learn. The engineering problem is to place closures at the right level and time: close a transaction safely, a release provisionally, a right firmly, and the metaphysical question humbly.

There is also a reflexive warning. The LATCH cycle is itself a static pattern. It can outlive its usefulness, conceal harms, and become a prestige object. Its custodians must publish its failure cases, compare it with alternatives, revise it under evidence, and retire it if it does not earn its cost. A theory of Quality that exempted itself from Quality would have learned nothing from Pirsig.

Conclusion: intelligence after the perfect objective

AI has made a metaphysical habit operational. We select a proxy, turn it into an object, optimize it at scale, and then speak as though value entered only when someone chose the target. Pirsig's deepest corrective is that evaluation was already present in the carving: in what counted as data, an error, a capability, a stakeholder, a valid comparison, and a purpose worth serving.

That insight does not entitle us to mystical certainty. It imposes the opposite discipline. A reward is evidence, not revelation. A benchmark is a tool, not the capability. A constitution is a hard-won pattern, not the end of moral inquiry. An audit is a time-bounded argument, not absolution. A model is one participant in a deployed relation, not the whole system. And “Quality” must never be the name painted on the final scalar.

The right ambition is not to build an AI that maximizes Quality. It is to build and govern AI without losing the capacity to notice when our own account of success has become the source of failure. That requires two virtues Pirsig refused to separate: enough stability to preserve what we have learned, and enough openness to let reality correct us.

The durable thesis is therefore not that Pirsig solved AI. It is that alignment worthy of the name cannot be a completed object. It must be a maintained practice of non-closure—empirical, contestable, power-aware, reversible, and willing, under evidence, to repair even the objective.

Selected bibliography

The source notes below attach evidence and limitations to individual claims. These are the works most central to the paper's philosophical reconstruction, empirical case, and comparison with adjacent approaches.

  • Bansal, Hritik, John Dang, and Aditya Grover. “Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models.” ICLR, 2024.
  • Burns, Collin, et al. “Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.” ICML, 2024.
  • Gao, Leo, John Schulman, and Jacob Hilton. “Scaling Laws for Reward Model Overoptimization.” ICML, 2023.
  • Hadfield-Menell, Dylan, et al. “Cooperative Inverse Reinforcement Learning.” NeurIPS, 2016.
  • Hadfield-Menell, Dylan, et al. “Inverse Reward Design.” NeurIPS, 2017.
  • Huang, Saffron, et al. “Collective Constitutional AI: Aligning a Language Model with Public Input.” FAccT, 2024.
  • Koh, Pang Wei, et al. “WILDS: A Benchmark of in-the-Wild Distribution Shifts.” ICML, 2021.
  • Langosco Di Langosco, Lauro, et al. “Goal Misgeneralization in Deep Reinforcement Learning.” ICML, 2022.
  • Lightman, Hunter, et al. “Let's Verify Step by Step.” ICLR, 2024.
  • Obermeyer, Ziad, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. “Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations.” Science 366, 2019.
  • Pirsig, Robert M. Zen and the Art of Motorcycle Maintenance: An Inquiry into Values. 1974.
  • Pirsig, Robert M. Lila: An Inquiry into Morals. 1991.
  • Pirsig, Robert M. “Subjects, Objects, Data and Values.” 1995.
  • Rao, Anita, et al. Challenges to the Monitoring of Deployed AI Systems. NIST AI 800-4, 2026.
  • Selbst, Andrew D., et al. “Fairness and Abstraction in Sociotechnical Systems.” FAT*, 2019.
  • Sharma, Mrinank, et al. “Towards Understanding Sycophancy in Language Models.” ICLR, 2024.
  • Skalse, Joar, et al. “Defining and Characterizing Reward Gaming.” NeurIPS, 2022.
  • Tabassi, Elham. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1, 2023.

Notes and sources

  1. Leo Gao, John Schulman, and Jacob Hilton, “Scaling Laws for Reward Model Overoptimization,” Proceedings of Machine Learning Research 202 (2023), paper and abstract. Its “gold” signal is a fixed, larger synthetic reward model, not direct human ground truth; the qualitative proxy-overoptimization result should not be inflated into a numerical law of human value.
  2. Robert M. Pirsig, Zen and the Art of Motorcycle Maintenance: An Inquiry into Values (1974), especially chapters 6, 9–20, and 26; and Lila: An Inquiry into Morals (1991), publisher's edition page. The Robert Pirsig Association resources provide bibliographic orientation.
  3. International AI Safety Report 2026, published February 3, 2026, especially the discussions of the evaluation gap, layered risk management, evaluation loopholes, and the evidence dilemma, full report.
  4. Anita Rao, Andrew Keller, Neha Kalra, Ryan Steed, Kweku Kwegyir-Aggrey, Kevin Klyman, Diane Staheli, and Amanda Bergman, Challenges to the Monitoring of Deployed AI Systems, NIST AI 800-4 (2026), official publication page and PDF.
  5. Robert M. Pirsig, “Subjects, Objects, Data and Values” (paper presented at the Einstein Meets Magritte conference, 1995), complete paper.
  6. Julian Baggini and Robert M. Pirsig, extended interview transcript (2005), full transcript.
  7. Robert M. Pirsig, letter to Paul Turner (2003), clarifying intellect as independently manipulable signs, letter.
  8. Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger, “Defining and Characterizing Reward Gaming,” NeurIPS 2022, paper and abstract. The strongest unhackability result concerns the full set of stochastic policies; the authors derive nontrivial conditions for restricted policy sets.
  9. Lauro Langosco Di Langosco, Jack Koch, Lee D. Sharkey, Jacob Pfau, and David Krueger, “Goal Misgeneralization in Deep Reinforcement Learning,” ICML 2022, paper and abstract. The evidence is from controlled reinforcement-learning environments and does not establish covert goals in deployed frontier systems.
  10. Jan Betley et al., “Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs,” Proceedings of Machine Learning Research 267 (2025), paper and abstract. The claim here is deliberately limited to the reported synthetic settings.
  11. Hritik Bansal, John Dang, and Aditya Grover, “Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models,” ICLR 2024, paper and abstract. The reported disagreement and protocol effects concern the paper's evaluated human and AI annotators, tasks, and models; they are not universal constants.
  12. Mrinank Sharma et al., “Towards Understanding Sycophancy in Language Models” (2023), research summary and paper links.
  13. Taylor Sorensen et al., “A Roadmap to Pluralistic Alignment,” paper, and Shangbin Feng et al., “Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration,” EMNLP 2024, paper and abstract.
  14. Saffron Huang, Divya Siddarth, Liane Lovitt, Thomas I. Liao, Esin Durmus, Alex Tamkin, and Deep Ganguli, “Collective Constitutional AI: Aligning a Language Model with Public Input,” FAccT 2024, open paper.
  15. Elham Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (2023), official interactive framework and PDF.
  16. Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan, “Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations,” Science 366 (2019): 447–453, PubMed record and abstract, DOI 10.1126/science.aax2342.
  17. Catherine Legg and Christopher Hookway, “Pragmatism,” Stanford Encyclopedia of Philosophy, substantive revision September 30, 2024, entry.
  18. Dylan Hadfield-Menell et al., “Cooperative Inverse Reinforcement Learning,” NeurIPS 2016, paper.
  19. Dylan Hadfield-Menell et al., “Inverse Reward Design,” NeurIPS 2017, paper and abstract.
  20. Batya Friedman, Peter H. Kahn Jr., and Alan Borning developed the tripartite conceptual, empirical, and technical approach; see Till Winkler and Sarah Spiekermann, “Twenty Years of Value Sensitive Design: A Review of Methodological Practices in VSD Projects,” Ethics and Information Technology 23 (2021): 17–21, open review.
  21. Andrew D. Selbst et al., “Fairness and Abstraction in Sociotechnical Systems,” FAT* 2019, paper.
  22. Pang Wei Koh et al., “WILDS: A Benchmark of in-the-Wild Distribution Shifts,” ICML 2021, paper and abstract. The benchmark covers ten selected applications and methods available at the time; it does not establish that every distribution shift is irreducible.
  23. Collin Burns et al., “Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision,” ICML 2024, paper and abstract. The experiments use weaker models as analogues for weak human supervision and do not resolve supervision of genuinely superhuman systems.
  24. Hunter Lightman et al., “Let's Verify Step by Step,” ICLR 2024, paper and abstract. Its positive result is specific to mathematical reasoning and does not show that process supervision recovers moral value or faithful internal reasoning in general.