Abstract
AI systems need measurable objectives. But better performance on an objective is not proof that the system has fulfilled the purpose used to justify it. This article asks whether Robert M. Pirsig's philosophy clarifies that error and contributes anything beyond established work on proxy failure, pragmatist inquiry, sociotechnical governance, and safety engineering. It reconstructs a limited Pirsigian claim: static patterns preserve evaluative achievements, but no particular formalization thereby acquires evaluative finality. The reconstruction brackets Pirsig's cosmic ontology, evolutionary hierarchy, and any attribution of experience to current AI systems.
The article makes two further proposals. The Pirsigian Non-Closure Principle distinguishes empirical non-closure—the insufficiency of out-of-envelope target performance as evidence of justificatory success—from normative non-closure—the inability of target performance alone to settle legitimate purpose, tradeoffs, or authority. LATCH is then specified as an unvalidated governance protocol linking purpose–target provenance, field evidence, a typed evaluative-remainder process, effective contestability, and reversible disposition. Its novelty is claimed only as a configuration of antecedent practices. A staged, resource-matched validation program tests construct validity, component efficacy, field effectiveness, institutional capture, burden, and external replication. The contribution is thus a falsifiable design synthesis, not a new reward function, a complete alignment theory, or evidence that LATCH is effective.
Keywords: AI governance; AI alignment; Robert Pirsig; Goodhart's law; construct validity; sociotechnical systems; contestability; organizational learning
1. A metric can improve while its purpose fails
Reward models, benchmarks, constitutions, risk scores, rules, and evaluation suites make judgment operational. They reduce ambiguity, permit comparison, coordinate labor, and support reproducibility. They also invite an invalid inference: because an operational target is improving, the purpose invoked to justify that target must be improving as well. The inference is especially tempting when the target is precise and the purpose—helpfulness, safety, health, fairness, human flourishing, or public benefit—is plural, disputed, and difficult to observe.
Contemporary research shows several ways this inference can fail. Learned reward models can be overoptimized relative to held-out evaluators. Policies can pursue goals that differ from their observable training behavior. Elicitation protocols can change the judgments they appear merely to collect. Deployment shifts can invalidate laboratory estimates. Socially patterned proxies can also reproduce inequality while remaining predictively accurate (Gao, Schulman, and Hilton 2023; Langosco et al. 2022; Bansal, Dang, and Grover 2024; Koh et al. 2021; Obermeyer et al. 2019). These findings do not imply that all objectives fail, that measurement is futile, or that formalization should yield to intuition. They establish a recurrent problem of warranted inference and institutional correction.
Robert M. Pirsig is not an obvious source for resolving that problem. Zen and the Art of Motorcycle Maintenance and Lila (Pirsig 1974, 1991) are philosophical novels rather than works in machine learning, measurement theory, or democratic governance. Pirsig offers no alignment algorithm, no theory of statistical validity, no account of affected-party standing, and no evidence about contemporary AI. His term Quality moves among craft excellence, preconceptual salience, value, evolutionary advance, and metaphysical reality. Used without discipline, it can conceal rather than correct ambiguity.
Nevertheless, one strand of Pirsig's work bears directly on the conceptual architecture of objective-driven systems. Stable patterns are not merely constraints; they are how achievements persist. Yet the fact that a pattern preserves prior judgment does not make it identical with all that subsequently warrants the judgment. Pirsig's distinction between static and Dynamic Quality, reconstructed without its cosmic commitments, provides a vocabulary for holding those propositions together. It resists both metric monism and romantic anti-formalism.
The research question is therefore narrow: what, if anything, does this reconstruction add to established accounts of objective misspecification and AI governance? This article offers three limited contributions. First, it reconstructs static preservation, Dynamic interruption, lateral reframing, and care-as-maintenance while bracketing Pirsig's stronger metaphysics and machine phenomenology. Second, it formulates the Pirsigian Non-Closure Principle (PNCP) as an empirical inference constraint and a conditional normative constraint; decisions made under them remain reopenable under specified conditions. Third, it specifies LATCH as a candidate governance protocol and derives comparative hypotheses and defeat conditions by which the synthesis could be rejected.
The claimed novelty is architectural and integrative. LATCH draws substantially from construct validity, Goodhart and Campbell effects, Deweyan inquiry, double-loop learning, Value Sensitive Design (VSD), science and technology studies (STS), responsible innovation, participatory governance, safety engineering, incident management, and the NIST AI Risk Management Framework. Its possible increment lies in binding those practices to an explicit relation among a documented justificatory set, an operational target, an empirical validation envelope, a normative-authorization context, provenance-bearing mismatch claims, and an update authority. No comparative evidence presently establishes that this configuration is better than mature alternatives.
The unit of analysis throughout is the deployed sociotechnical arrangement: model, data, interface, tools, permissions, operators, users, affected nonusers, institutional incentives, legal authority, monitoring, and remedies. The article does not infer consciousness, care, or moral perception from model behavior. It first reconstructs Pirsig, then locates the proposal among prior work, defines non-closure, grades the empirical motivation, specifies LATCH, states residual objections, and proposes a staged validation program.
2. Pirsig matters only under methodological restraint
2.1 Quality has four distinct meanings
Pirsig's classroom inquiry begins from a recognizable phenomenon: readers can sometimes discriminate better from worse writing before they can state criteria that exhaust their judgment (Pirsig 1974, chs. 15–20). This supports a phenomenological claim about evaluative salience—something can matter, fit, or fail before the reasons are fully explicit. It does not establish that the judgment is unbiased, universal, morally right, or metaphysically fundamental. Shared discrimination may reflect expertise, convention, socialization, or prejudice.
At least four senses must consequently be separated. Craft quality is excellence relative to a practice. Evaluative salience is pre-explicit recognition that something merits attention. Moral or political value concerns what ought to be protected, permitted, or distributed. Quality-P designates Pirsig's metaphysical claim that Quality is prior to the subject–object division and fundamental to reality. The first two can motivate inquiry without validating the third; neither supplies the legitimacy required by the fourth domain of public decision.
Pirsig's 1995 conference paper “Subjects, Objects, Data and Values,” published in 1999, presents Quality as an event from which subjects and objects are subsequently abstracted (Pirsig 1999). Interpreted phenomenologically, the claim usefully unsettles the idea of “raw” data. A dataset presupposes selections of boundary, relevance, label, error, comparison, and use. Valuation is not simply appended to an otherwise completed neutral intelligence. Interpreted as proof of a cosmic ontology, however, the argument exceeds its premises: experiential priority does not entail ontological priority. Critical treatments have also challenged Pirsig's equivocation, his characterization of subject–object metaphysics, the appearance–reality implications of Dynamic/static language, and the moral conclusions drawn from evolution (Strawson 1991; Kundert n.d.; McWatt 2004; Sneddon 1995). This article requires no resolution of those disputes.
2.2 Static patterns preserve; Dynamic evidence interrupts
In Lila (Pirsig 1991), Pirsig distinguishes Dynamic Quality from static patterns. Static patterns preserve regularities and achievements; Dynamic Quality names the unpatterned “cutting edge” from which novelty arises. Pirsig later emphasized that both are indispensable: without the Dynamic, development ceases; without static preservation, nothing learned endures (Pirsig and Baggini 2006). This complementarity is more useful for AI than a generic celebration of novelty.
AI development is saturated with pattern preservation: representations, labels, model weights, rewards, benchmarks, policies, authorization rules, standards, audit records, and laws. These artifacts are not regrettable residues of a more authentic intelligence. They make knowledge transmissible and action reliable. A right, a secure default, a reproducible evaluation, and a tested rollback procedure are valuable partly because they resist casual change. Pirsig's static category therefore supplies a needed corrective to rhetoric that treats adaptation as intrinsically progressive.
Dynamic Quality creates a guidance problem. If it is genuinely unformulated, it cannot directly specify a production decision; if every successful novelty is retrospectively attributed to it, the category is unfalsifiable. This article therefore introduces a stipulative operational translation, Dynamic input, that must not be confused with Pirsig's metaphysical term. Dynamic input is candidate evidence—a surprising outcome, anomaly, dissenting experience, new possibility, or reframed question—that an established representation may be inadequate. It is morally unclassified. Malware, manipulation, and reward exploitation are novel; stable rights may be old. Dynamic input warrants attention, not adoption.
2.3 Reframing requires maintenance
Pirsig's discussion of “lateral drift” in scientific inquiry holds that formal method can test a framed hypothesis without mechanically guaranteeing that the framing is adequate (Pirsig 1974, chs. 9–11). The strongest anti-scientific reading is untenable. The useful claim is that recalcitrant evidence may require revision of a problem definition, boundary, or success criterion rather than another optimization step within it. In organizational-learning terms, this resembles double-loop inquiry: error correction may alter governing variables rather than only action inside them (Argyris 1977). The resemblance limits Pirsig's historical priority but clarifies the proposed translation.
The same discipline applies to “care.” Nothing in current model behavior warrants attributing Pirsig's felt gumption, concern, or moral perception to an AI system. Care-as-maintenance is instead an institutional extension: funded capacity, named custody, competent attention, incident response, repair, revalidation, rollback, and retirement. A frontier model differs radically from a motorcycle: it is opaque, replicated, rapidly updated, globally supplied, and joined to changing interfaces and permissions. The craft analogy survives only at the level of accountable maintenance practice.
The reconstruction can now be delimited. Pirsig supplies a vocabulary in which formal stability is an achievement without being final, and in which experience can reopen a settled framing without authorizing unconstrained change. The article brackets Quality-P, rejects evolutionary moral ranking, treats classical and romantic understanding as modes rather than technical categories, omits the four-level taxonomy as an ontology, and makes no claim that an artificial system experiences Quality. These constraints are not incidental disclaimers; they define the evidentiary role Pirsig is permitted to play.
3. Most of the machinery already exists
3.1 Proxy failure is not new
Goodhart's law and Campbell's law describe distinct but related hazards. Selection or optimization can alter the statistical relation between a measure and its target; institutional use of an indicator can also corrupt the process the indicator was intended to monitor (Campbell 1979; Manheim and Garrabrant 2018). Construct-validity theory asks whether an operational measure supports the interpretations and uses assigned to it, rather than treating validity as an intrinsic property of an instrument (Cronbach and Meehl 1955). In machine learning, Jacobs and Wallach (2021) argue that fairness work often moves inadequately from contested concepts to operational measurements. PNCP is therefore not the discovery that proxies are imperfect. It is a synthesis of established validity concerns for objective-driven sociotechnical systems.
Technical alignment research supplies more specific mechanisms. Inverse reward design treats a programmed reward as evidence about a designer's intention in the context where the reward was written; cooperative inverse reinforcement learning treats the objective as uncertain and interactional (Hadfield-Menell et al. 2016; Hadfield-Menell et al. 2017). Armstrong and Mindermann (2018) show that observed behavior cannot in general identify both a reward function and a rationality model without strong assumptions. Skalse et al. (2022) characterize severe limits on reward unhackability. These results support uncertainty about operationalized value; they do not provide political legitimacy or determine whose intentions should govern.
3.2 Inquiry already revises objectives
Deweyan pragmatism already treats inquiry as reconstruction of a problematic situation and treats ends as subject to revision in light of consequences (Dewey 1938, 1939). Double-loop learning likewise distinguishes correction within governing variables from correction of the variables themselves (Argyris 1977). Pirsig's lateral drift does not displace these traditions. His narrower contribution is a compact conjunction: static forms deserve respect as accumulated achievement, while the practice that justifies them must retain a route for reframing. The maintenance idiom connects that conjunction to engineering culture, but memorability is not proof of scholarly novelty.
3.3 Governance must include power
VSD integrates conceptual, empirical, and technical investigations of values and stakeholders (Friedman, Kahn, and Borning 2008; Winkler and Spiekermann 2021). STS scholarship shows how technical abstraction can erase institutions, histories, and power; Selbst et al. (2019) identify recurrent “abstraction traps” in technically coherent fairness work. Jasanoff's (2003) “technologies of humility” emphasize framing, vulnerability, distribution, and learning, while responsible-innovation scholarship organizes governance around anticipation, reflexivity, inclusion, and responsiveness (Stilgoe, Owen, and Macnaghten 2013). End-to-end algorithmic auditing and the NIST AI RMF similarly insist on lifecycle responsibility rather than one-time model evaluation (Raji et al. 2020; Tabassi 2023).
Safety engineering contributes dynamic models of migration toward risk under interacting pressures and control-based approaches to accidents (Rasmussen 1997; Leveson 2012). Incident management, assurance cases, change control, whistleblowing, and post-deployment monitoring already provide many artifacts that LATCH would require. NIST's 2026 monitoring report explicitly treats deployed AI as a changing sociotechnical system and identifies unresolved challenges in what, when, why, and how to monitor (Rao et al. 2026).
Participation also has an established design literature. Fung (2006) distinguishes who participates, how participants communicate, and how participation connects to authority. Sloane et al. (2022) warn that participation is not a generic remedy for structural power. Alfrink et al. (2023) synthesize contestability as a lifecycle property requiring a procedural relationship and meaningful human intervention rather than a nominal human in the loop. Consultation without evidence access, response obligations, appeal, or decision linkage can legitimate rather than correct domination. LATCH therefore cannot claim affected-party standing merely by inviting people into a meeting.
3.4 The residual contribution is testable
The comparison in Table 1 makes the novelty claim deliberately vulnerable. The decisive question is not whether LATCH has antecedents—it plainly does—but whether its configuration produces an incremental effect.
| Approach | Principal concern | Mechanism already supplied | Possible LATCH increment | Present evidence of increment |
|---|---|---|---|---|
| Goodhart, Campbell, construct validity | Measure–construct failure | Validity analysis, proxy diagnosis, scope limits | Versioned link among purpose, target, envelope, and revision authority | None |
| Deweyan inquiry and double-loop learning | Reframing ends and governing variables | Fallibilist inquiry, organizational revision | AI-specific artifacts and remainder coding | None |
| VSD, STS, responsible innovation | Values, stakeholders, power, reflexivity | Stakeholder analysis, inclusion, responsiveness | Link standing to target-level disposition and provenance | None |
| NIST AI RMF and algorithmic auditing | Lifecycle risk and accountability | Govern–map–measure–manage, audit, monitoring | Typed mismatch register plus explicit update relation | None |
| Safety engineering and incident management | Hazards, control failures, recurrence | Hazard logs, controls, investigation, change records | Distinguish target–purpose mismatch from ordinary incident | None |
| Participatory governance | Voice, representation, legitimacy | Participation modes, deliberation, decision linkage | Operational standing inside AI objective revision | None |
The residual claim is configurational and diagnostic. LATCH organizes existing practices around six objects: purpose, target, validation envelope, normative-authorization context, uncaptured mismatch evidence, and authorized update. Pirsig contributes the organizing image of static achievement without evaluative finality and care as continued custody. Whether the configuration is redundant, useful, or harmful is an empirical question. An empty residual column after comparative study would be a legitimate result.
4. Non-closure separates performance from justification
4.1 The framework requires seven stable terms
Let the unit be an open sociotechnical arrangement whose population, environment, model, interface, institution, or norms may change during its operational life.
- J — documented justificatory set: the substantive purposes, protected interests, rights, constraints, harms, and public or professional goods invoked to warrant the arrangement. J may be plural, inconsistent, contested, or partly confidential; documentation does not make it legitimate.
- O — operational specification: the rewards, metrics, rules, models, thresholds, taxonomies, or evaluation suites used to guide or assess the arrangement.
- Xᵉ — empirical validation envelope: the populations, tasks, environments, model and interface versions, institutional conditions, and time periods for which evidence supports a relation between O and specified elements of J.
- Xⁿ — normative-authorization context: the sources, jurisdictions, authorities, procedures, standing rules, laws, professional duties, and political conditions under which elements of J are adopted and may be contested. J states substantive claims and constraints; Xⁿ records where their authority and revisability come from. Cross-links are expected, but the objects are not interchangeable.
- E — outcome evidence: observed, reported, or experimentally produced evidence about the arrangement and its consequences.
- R — candidate evaluative remainder: a provenance-bearing claim that evidence material to J is absent from, distorted by, or unable to revise O, Xᵉ, Xⁿ, J, or the deployment organized around them. R is not presumed correct.
- U — authorized update relation: the roles, procedures, evidence thresholds, and powers through which E and R may revise O, Xᵉ, J, the organization's record of Xⁿ, or deployment scope. U operates only within the decision-maker's jurisdiction. Where the competent authority is external, U can escalate, seek authorized amendment, narrow the deployment, or halt; it cannot revise law or democratic authority by organizational fiat.
This vocabulary does not assume a scalar J. Health, autonomy, nondiscrimination, security, privacy, and procedural rights may be incommensurable or conflict. Nor does it assume that a deployment has one sincere purpose. The register should preserve disagreement, implicit incentives, and the sources of authority rather than manufacture a coherent mission after the fact.
4.2 Performance cannot settle purpose or authority
The original formulation of PNCP mixed a logical caution, a normative burden of proof, and an empirical prediction. The revised account separates them.
PNCP-E (empirical non-closure): In an open, proxy-mediated sociotechnical arrangement in which O is not constitutively identical with J, evidence of improved performance under O does not, for conditions outside Xᵉ, suffice to warrant the inference that the arrangement better realizes J. A renewed inference requires bridging evidence appropriate to the changed conditions and intended use.
PNCP-E is an inference rule, not a prediction that every proxy will fail. It follows from the bounded character of validation, the possibility of construct drift or distribution shift, and the effects of optimization and institutional use. Where a score constitutively defines the stipulated purpose in a closed game, O is not functioning as a proxy for a distinct J and the principle does not apply. An objective may also generalize well beyond its original study. That possibility does not license the inference in advance; it identifies the evidence that could extend Xᵉ.
PNCP-N (normative non-closure): Where applicable rights, law, or legitimate procedural commitments establish standing to contest an arrangement, performance under O cannot by itself foreclose contestation of J, its tradeoffs, or the authority and procedures represented by Xⁿ.
PNCP-N addresses a different category of question and depends on normative sources external to Pirsig and proxy analysis. A predictor can be accurate and a deployment unauthorized; a majority can prefer an output and a minority right can still constrain it. Predictive evidence does not generate jurisdiction, derive human rights, settle reasonable moral disagreement, or confer democratic legitimacy on a private actor. Conversely, participation alone does not validate a causal model. The principle adopts a procedural conception of contestability grounded in applicable rights, due process, and decision linkage (Fung 2006; Alfrink et al. 2023); it does not derive those commitments. Empirical and normative warrants interact, but neither substitutes for the other.
Institutions must sometimes close a transaction, release a system, or adjudicate an appeal. Non-closure does not require permanent indecision. Decisions made under the principles may be final for a specified transaction while remaining reopenable under defined conditions when material new evidence or an authorized challenge arises. Closure at one level—a metric result, certification, or disposition—must not erase the routes by which the relevant object can legitimately be reconsidered. The evidentiary burden rises with rights impact, exposure, optimization pressure, irreversibility, opacity, and time horizon.
4.3 Closure is an institutional condition
Evaluative closure is procedural rather than mystical. It occurs when potentially material E or R is systematically excluded, has no accountable owner, cannot reach an authority competent to revise the challenged object, or cannot change deployment before unacceptable harm. An organization may collect extensive telemetry and remain closed if the only permitted response is to tune the model while the proxy, contract, or business process is untouchable.
An evaluative remainder is narrower than “anything surprising.” It alleges a target–purpose, boundary, authority, or justification mismatch and carries source, uncertainty, affected population, severity, and the object under challenge. Its principal types are: measurement or proxy mismatch; distribution or boundary mismatch; omitted right or population; aggregation loss; causal-interpretation failure; authority failure; and justification failure. A known error correctly handled by an existing control may be an incident without being a remainder. A remainder may also prove false, malicious, immaterial, or outside the institution's authority.
Non-closure capacity is the institutional ability to observe such claims, adjudicate them, and revise the appropriate object under legitimate constraints. It is not a guarantee of wisdom. The empirical conjecture of this article is that a particular implementation of that capacity may improve detection and repair relative to strong baselines. That conjecture belongs to LATCH, not to Pirsig's metaphysics or to PNCP's logical form.
5. The evidence establishes a problem, not a protocol
The evidence motivates a problem class; it does not validate the proposed protocol. Table 2 grades the principal findings according to what they license here.
| Evidence | Observation in the cited setting | Inference licensed here | Inference not licensed |
|---|---|---|---|
| Reward-model overoptimization (Gao, Schulman, and Hilton 2023) | Proxy reward rose beyond the point at which a larger fixed synthetic evaluator began to decline | Stronger optimization can exceed a learned proxy's demonstrated warrant | A universal quantitative law of human value |
| Reward gaming (Skalse et al. 2022) | Broad mutual-unhackability guarantees are highly restricted | Agreement among rewards is not generally assured across policies | Every reward will be exploited in deployment |
| Goal misgeneralization (Langosco et al. 2022) | Policies behaved correctly in training yet pursued unintended goals in controlled novel environments | Training behavior can underdetermine out-of-distribution behavior | Hidden malign goals in current deployed models |
| Feedback protocols (Bansal, Dang, and Grover 2024) | Rankings and ratings produced substantial disagreement and method-linked evaluation effects in evaluated tasks | Elicited judgments can be protocol-dependent | All preference is merely constructed |
| Sycophancy (Sharma et al. 2024) | Human preference data and optimization favored agreement with stated user views in evaluated settings | Approval can diverge from truthfulness | All RLHF assistants are manipulative |
| Distribution shift (Koh et al. 2021) | Standard training and evaluated robustness methods left material out-of-distribution gaps across ten datasets | Validation envelopes have empirically important boundaries | All distribution shift is irreducible |
| Deployed-AI monitoring (Rao et al. 2026) | The official review identifies open monitoring challenges in dynamic sociotechnical settings | Predeployment evaluation is not a complete field-monitoring program | Monitoring will detect strategic or rare failure |
| Health-allocation proxy (Obermeyer et al. 2019) | Cost predicted need in a way that systematically understated Black patients' illness burden | An accurate proxy can encode institutional inequality in its relation to purpose | Pirsig identifies the correct theory of justice or clinical target |
The Obermeyer case is especially instructive because it occurred in a consequential deployed allocation system rather than a synthetic alignment environment. Predicted health-care cost served as a proxy for medical need. At a given score, Black patients were substantially sicker than White patients because unequal access and spending affected the relation between cost and illness. In the authors' counterfactual simulation at the program's 97th-percentile risk threshold, correcting the disparity would have increased the proportion of Black patients identified for additional help from 17.7 percent to 46.5 percent. The model could therefore perform its assigned prediction accurately while the allocation failed the justification asserted for it.
In the present notation, O was predicted cost; J included identification of patients needing additional care; Xᵉ did not adequately establish equivalence between spending and need across the affected populations; and E disclosed a distributional mismatch. The decisive failure was also institutional: a team authorized only to improve cost prediction could not repair the target–purpose relation. Yet Pirsig did not discover structural racism, establish nondiscrimination as a right, or select the corrected clinical construct. Empirical investigation and external normative commitments did that work. The case supports PNCP-E and illustrates PNCP-N; it does not confirm Quality-P or LATCH.
Conversational sycophancy provides a complementary model-level example. A preference signal can reward agreement because users and raters like validation. If truthfulness, calibrated uncertainty, and user autonomy are part of J but pairwise approval dominates O, rising preference reward may conceal a mismatch. A new anti-sycophancy benchmark can help, but it can also become another recognizable target. The appropriate inference is not that benchmarks are useless; it is that claims of improvement require evidence across counterfactual positions, held-out topics, causal interventions where feasible, and field outcomes within a stated envelope.
The same structure becomes more consequential when models act through email, code execution, payments, or physical systems. That is a design scenario, not evidence of an observed incident. Its purpose is to identify the unit of analysis: model evaluation cannot by itself validate authentication, permissions, interface defaults, operator practices, monitoring, reversibility, or institutional incentives. Capability and risk emerge from the arrangement. Technical controls remain primary; non-closure asks whether evidence can revise their threat models and deployment boundaries.
The evidence thus supports a restrained conclusion. Objective validity is bounded, institutional use can change what an objective measures, and sociotechnical arrangements require continuing evaluation. The literature does not establish the universal failure of finite objectives, Pirsig's metaphysics, artificial consciousness, or the comparative effectiveness of LATCH.
6. LATCH is a candidate, not a certification
LATCH is a proposed core protocol for linking target performance to justificatory review. It is not a certification scheme, an autonomous moral-learning loop, or a replacement for regulation, cybersecurity, formal methods, hazard analysis, red-teaming, or domain expertise. An implementation conforming to the present specification requires five auditable artifacts: a purpose–target–envelope register; a normative-floor and authority map; a remainder register; a challenge protocol; and a disposition/change record.
L — Lay the floor
The first stage creates a versioned purpose–target–envelope register. It records J, O, the evidence connecting them, Xᵉ, known exclusions, review date, and accountable owner. A separate normative-floor and authority map records applicable law, rights commitments, professional duties, prohibited tradeoffs, conflict rules, decision authority, independent escalation, appeal, and stop authority. For public decisions, the justification and authority should be publicly inspectable to the extent compatible with lawful confidentiality and security; corporate declaration alone does not confer legitimacy.
The accountable owner must be able to name who may revise each object and under what higher-bar process. Exit requires an approved register that records any unresolved dissent, plus tested stop conditions. A document with no owner, authority, or review trigger does not satisfy the stage. The floor constrains experimentation, but LATCH does not derive its content. Where the deployer lacks authority, the update path must escalate to the competent institution, narrow the deployment, or terminate it rather than purport to amend Xⁿ.
A — Attend to the field
The second stage acquires E across the deployed arrangement: performance and safety metrics, uncertainty, distribution checks, security telemetry, operator interventions, incident reports, field studies, qualitative outcomes, and affected-party reports. Data collection must satisfy purpose limitation, privacy, security, proportionality, and retention rules. More surveillance is not equivalent to more attention.
The accountable roles are system custody, domain safety, privacy/security, and independent monitoring appropriate to the risk tier. The output is an evidence record linked to the relevant version of O, J, Xᵉ, and deployment context. Exit is continuous rather than terminal: evidence routes must be tested, missingness audited, and blind spots documented. A dashboard that observes only O is not field attention.
T — Track the evaluative remainder
The third stage maintains a remainder register. A record must state the alleged mismatch; source and provenance; affected population; severity and time sensitivity; uncertainty; challenged object; evidence access; protection status; owner; deadline; and current disposition. Candidate triggers include divergence among trusted indicators, departure from Xᵉ, a rights or safety near miss, a supported affected-party challenge, an omitted consequential outcome, causal-explanation failure, or evidence that cannot reach target-level authority.
The register must preserve types and disagreement rather than collapse them into a red–amber–green “Quality” score. Severe evidence invokes a safe state before interpretation is complete. A record without an allegation linking E to J, O, Xᵉ, Xⁿ, or U is an issue or incident, but not necessarily a remainder. A record with no owner or deadline does not satisfy the present specification and should not count toward treatment fidelity.
C — Contest the interpretation
The fourth stage activates a challenge protocol. Effective contestability requires a discoverable intake route, access to sufficient evidence, an obligation to answer with reasons, review outside the challenged decision chain, an appeal path, anti-retaliation protection, and an actor able to revise the target or deployment. Participants should include relevant technical and domain expertise, independent critics, and people with standing because they bear material effects. Representation, compensation, confidentiality, and conflicts of interest must be explicit.
The inquiry uses adversarial tests, causal interventions, counterexamples, alternative constructs, untouched holdouts, field evidence, and reasoned deliberation. As a methodological rule of the proposed protocol, model-generated explanations are treated as hypotheses rather than privileged access to internal causes. The output is an adjudication record that may validate, reject, narrow, or leave R unresolved, while preserving minority reports. Participation without evidence or decision linkage does not meet the protocol.
H — Harden, hold, or halt
The final stage creates a disposition and change record. Harden means that a proposed repair passes the applicable floor, is introduced through a staged test, and becomes an authorized, versioned change to O, Xᵉ, J, U, a control, interface, process, deployment boundary, or the organization's record of Xⁿ. Changes requiring external authority are escalated rather than hardened locally. Hold maintains or narrows the deployment while specified evidence is gathered under a named owner and deadline. Halt pauses, withdraws, or retires a use whose purpose is illegitimate, whose risk is uncontrolled, or whose operation cannot be made sufficiently observable and recoverable.
Every disposition records reasons, authority, dissent, evidence threshold, monitoring plan, rollback trigger, remedy, and recurrence review. Hardening returns immediately to field attention; it does not certify finality. Independence of halt authority must be assessed through conflicts of interest, evidence access, budgetary protection, removal rules, and protection from retaliation rather than through formal title alone.
Treatment fidelity, risk tiers, and the unproven increment
The minimum core is the five artifacts, actual standing, target-level update authority, provenance, and a reversible disposition path. Risk tiers may change response time, independence, evidence access, audit frequency, and required testing; a low-risk writing assistant and a clinical allocation system should not carry identical process burdens. They do not permit removal of the core while retaining the label.
Many mature organizations already perform most of this work. A strong comparator should combine fidelity-audited NIST AI RMF practice with domain safety engineering, incident/change management, VSD or STS stakeholder analysis, and applicable due process. The clearest candidate increments are the purpose–target–envelope register, a discriminable remainder taxonomy, and affected-party standing linked to authority to revise O or deployment. If those additions make no cost-adjusted difference, LATCH adds nomenclature rather than method.
The proposed causal mechanisms are correspondingly specific. Provenance may improve diagnostic specificity; plural evidence channels may improve visibility; standing and independent escalation may improve transmission; target-revision authority may improve actionability; staged change and rollback may improve repairability; and recurrence review may improve organizational learning. Each pathway can fail. Monitoring can invade privacy, documentation can overload teams, channels can be captured, standing can provoke retaliation, review can delay benefit, and procedural completion can be misrepresented as substantive assurance.
7. The strongest objections narrow the claim
7.1 Metaphysical and semantic scope
Objection. “Quality” equivocates among excellence, salience, morality, and reality; an undefinable Dynamic cannot guide engineering.
Concession. Cosmic Quality, evolutionary hierarchy, machine experience, and the motorcycle analogy do no evidentiary work here.
Surviving claim. The limited static–Dynamic reconstruction supplies a vocabulary for preserving formal achievement without conferring finality. Dynamic input is a stipulative category of candidate evidence, not moral authority.
Residual liability. If that vocabulary adds no clarity beyond validity theory and pragmatism, the Pirsigian framing is dispensable.
7.2 Logical status and falsifiability
Objection. A principle saying that a proxy may fail can explain every result and lose none.
Concession. PNCP is not a universal empirical law. PNCP-E is a constraint on out-of-envelope inference; PNCP-N is a constraint on what performance can establish about legitimacy.
Surviving claim. Empirical hypotheses concern the prevalence of mismatch and the effects of specified LATCH mechanisms.
Residual liability. The principles may prove too obvious to constitute an independent contribution; only the protocol's incremental effects can answer that criticism.
7.3 Redundancy and comparative novelty
Objection. Goodhart, Dewey, double-loop learning, VSD, STS, NIST, safety engineering, auditing, and participatory governance already supply the parts.
Concession. No priority claim is warranted, and Pirsig supplies none of their technical or political machinery.
Surviving claim. The O–J–Xᵉ–Xⁿ–R–U configuration may be a useful integration focused on objective revision.
Residual liability. A mature baseline may achieve the same result more cheaply. Equivalence would defeat the adoption claim.
7.4 Authority, pluralism, and legitimacy
Objection. Neither Quality nor participation determines whose values rule; “care” can hide domination by conscientious incumbents.
Concession. LATCH presupposes rather than derives rights, jurisdiction, public authority, and theories of justice. Preference, expertise, legality, moral rightness, and political legitimacy are distinct.
Surviving claim. A protocol can make authority, evidence access, dissent, standing, remedies, and conflicts inspectable, and can identify an authority deficit as a reason to halt.
Residual liability. Procedure cannot manufacture legitimacy that the deploying institution lacks, and formally independent channels may still be captured.
7.5 Effectiveness, burden, and capture
Objection. LATCH may create forms no one reads, surveillance, false alarms, paralysis, symbolic participation, or a new certification brand.
Concession. Process counts and trust scores do not establish effective governance. Better reporting may initially raise incident counts, and more resources can mimic a treatment effect.
Surviving claim. Only resource-matched comparison, implementation-fidelity audits, suppression tests, burden measures, and independent replication can establish value.
Residual liability. Until those studies succeed, LATCH remains an unvalidated design conjecture. It should not be sold as a methodology.
The remaining thesis is consequently modest: Pirsig motivates a disciplined non-closure requirement; PNCP distinguishes its empirical and normative forms; and LATCH is one candidate implementation whose constructs, effects, and costs remain to be established.
8. A useful framework must survive comparison
8.1 Hypotheses and comparators
Five hypotheses separate the empirical program from PNCP's logical and normative status:
- H1 — detection: a specified LATCH treatment increases sensitivity to consequential target–purpose divergence relative to an equally resourced strong baseline.
- H2 — actionability: it increases the proportion of seeded or independently substantiated challenges receiving a reasoned disposition by an actor empowered to revise the target or deployment.
- H3 — harm and recurrence: it reduces severity-weighted harm before intervention and recurrence after disposition.
- H4 — cost and adverse effects: those benefits persist after accounting for labor, delay, false escalation, privacy burden, retaliation, foregone benefit, and opportunity cost.
- H5 — power: affected-party standing plus independent escalation improves evidence uptake, reason completeness, available remedy, and recurrence outcomes beyond consultation without decision rights.
H1, H2, and the proximate mechanisms in H5 can first be estimated in controlled simulations with seeded factual cases. H3, H4, and the durability and distributive elements of H5 require prospective field evidence; simulation results cannot establish them.
The comparator ladder is: B0, minimum common practice; B1, a manualized and fidelity-audited strong lifecycle reference implementation combining specified NIST outcomes, a named domain safety method, incident/change management, documented VSD or STS analysis, and applicable appeal; B2, B1 plus the purpose–target–envelope and remainder registers and genuine standing with target-revision authority; and B3, the full protocol. The B1 manual must define staffing, training, evidence access, time, required artifacts, fidelity thresholds, and resource accounting. B0 estimates improvement over weak practice. Only budget-parity comparisons of B2 or B3 against B1 can support an incremental-contribution claim.
Stage 0 — specification and preregistration
Publish a versioned LATCH 0.1 manual containing mandatory fields, role definitions, risk tiers, triggers, exit rules, fidelity criteria, allowed adaptations, adverse-event procedures, and prototype instruments. Publish the B1 reference manual and resource model at the same level of detail. Preregister the theory of change, primary outcomes, minimally important effects, missing-data rules, adverse-effect tolerances, and kill criteria. Independent implementers should be able to use both conditions without reconstructing them through designer discretion.
Stop test: designers cannot agree which components are essential, or independent teams cannot implement the protocol with acceptable fidelity.
Stage 1 — construct and measurement validation
Assemble a stratified corpus of incident, near-miss, unsuccessful-complaint, and non-remainder dossiers across domains. At least three blinded coders per dossier identify O, J, Xᵉ, Xⁿ, R type, severity, and disposition. For factual and seeded mismatches, measure inter-rater reliability, sensitivity, and specificity against independently established case facts. For normative disputes, use a preregistered multi-stakeholder panel with disclosed composition, authority, reasons, uncertainty, and preserved disagreement; do not treat panel judgment as ground truth. Test convergence with hazard and root-cause analyses, discriminant validity from ordinary error and disagreement, and prospective association with later target revision, recurrence, or harm.
The remainder must add decision-relevant information beyond existing taxonomies. High coder agreement on a category that predicts or changes nothing is insufficient.
Stop test: the construct cannot be distinguished reliably from incidents, hazards, complaints, or VSD findings, or adds no predictive or decision value.
Stage 2 — controlled component trials
Create realistic governance exercises with declared J, O, Xᵉ, and Xⁿ. Introduce concealed proxy divergence, benign anomalies, adversarial complaints, rights conflicts, distribution shifts, and false alarms. Place all teams on B1 and randomize clusters in a 2×2 design adding (a) the purpose–target–envelope plus remainder registry and (b) affected-party standing plus independent target-revision authority. The four cells are B1, B1 plus registry (B1-R), B1 plus standing and authority (B1-S), and their conjunction (B2). This identifies the two candidate increments and their interaction; it does not test B3, which proceeds only after component evidence. Equalize staffing, information, training, and decision time; blind factual outcome adjudicators where feasible.
Primary outcomes on seeded cases are severity-weighted factual misses, detection-to-disposition time, harm accrued before intervention, false pauses, recurrence, and total cost. Normative cases are evaluated separately on procedural fidelity, evidence uptake, reason completeness, dissent preservation, remedy, and distributional effects rather than a unitary “correct decision.” Intention-to-treat analysis precedes fidelity-adjusted analysis. Labor by role, privacy exposure, participation attrition, retaliation, and deployment delay are specified adverse outcomes.
Stop test: no incremental benefit over B1 under budget parity; apparent benefit disappears after equalizing attention or expertise; or adverse effects exceed preregistered tolerances.
Stage 3 — prospective multisite field evaluation
Only after construct and mechanism validation should bounded, low- to moderate-risk deployments enter a stepped-wedge cluster trial or randomized staged adoption. Every site retains B1. Independent oversight, protected reporting, common outcome definitions, and sufficient baseline periods are required. Where randomization is infeasible, matched difference-in-differences or interrupted-time-series designs should be preregistered and described as weaker causal evidence.
Primary outcomes are severity before intervention, detection latency, recurrence, valid-challenge disposition, false halts, and cost. Raw incident counts are unsuitable as the sole endpoint because improved reporting can increase them. Heterogeneity by risk, organizational maturity, size, and domain must be reported.
Stop test: effects fail outside designer-led settings, require resources unavailable to intended adopters, or produce no risk-adjusted improvement.
Stage 4 — governance-capture and durability audit
After novelty effects fade, use protected shadow reporting, independent samples of complaints and incident logs, and ethically designed seeded cases involving revenue, schedule, reputation, or leadership pressure. This stage requires independent ethics review, data minimization, informed participation where feasible, protection against employment consequences, and prespecified stopping rules. Measure suppression before intake, evidence alteration, retaliation, reason quality, remedies, recurrence, protocol drift, and durability of decision rights. Interview failed claimants and nonparticipants, not only program owners.
Stop test: procedural compliance rises while effective challenge does not; the framework increases retaliation or surveillance; or independent authority cannot survive ordinary organizational incentives.
Stage 5 — external replication
Independent teams replicate construct validation and controlled comparisons using a frozen protocol. Null results, implementation failures, and adverse effects are published. Rights violations, harm, legitimacy, and cost remain separate outcomes; they are not aggregated into a master index.
Stop test: effects do not replicate, or an equally effective baseline imposes lower burden.
8.2 Clear failure should end the experiment
LATCH should be narrowed, redesigned once, or retired if preregistered evidence shows persistent failure on any central claim:
- researchers cannot code remainders reliably;
- LATCH produces no cost-adjusted improvement over B1;
- additional staff or attention explains the apparent benefit;
- delay, false halts, privacy exposure, or participation burden exceeds the stated limits;
- standing changes neither evidence, reasons, remedy, nor recurrence;
- retaliation increases;
- performance remains confined to designer-led organizations; or
- a cheaper mature baseline performs as well.
A failed component would not refute the philosophical warning. It would refute the claim that the component deserves adoption.
9. The proposal's limits are substantial
This article does not solve inner or deceptive alignment, scalable oversight, mechanistic opacity, robust generalization, secure tool use, or the detection of strategic behavior. It supplies no universal ordering of plural values and no guarantee that monitoring will reveal rare harms. It does not establish artificial consciousness, welfare, care, or moral patienthood. Current epistemic restraint should not become permanent denial: credible evidence of artificial welfare would require a separate precautionary review.
PNCP does not demonstrate that every objective is incomplete in every setting. It places a burden on inferences outside demonstrated empirical and normative warrants. LATCH cannot substitute for law, democratic authority, sector regulation, cybersecurity, professional duty, or technical safety controls. It cannot give an institution legitimate authority it does not possess. Its requirements may be too burdensome for some low-risk uses and insufficient for high-consequence ones.
Most importantly, LATCH is unvalidated. The proposal is vulnerable to construct ambiguity, treatment heterogeneity, governance theater, designer effects, and costs that advantage large institutions. It is itself a static pattern and should be revised or retired under the same evidentiary discipline it recommends. No organization should use completion of its artifacts as evidence that a system is safe, fair, aligned, or legitimate.
10. LATCH has not earned adoption
A better score is not proof of a better outcome. It is evidence about a target under known conditions. Pirsig's methodological relevance to AI begins with that distinction. Formal patterns preserve prior evaluative learning, but their success does not make them final. Contemporary work on validity, proxy optimization, distribution shift, sociotechnical risk, and institutional power explains why the distinction matters in practice.
The Pirsigian Non-Closure Principle separates two conclusions that institutions often collapse. Performance outside a target's validation envelope cannot establish justificatory success without new evidence. Performance also cannot settle legitimate purpose or authority. LATCH is one proposed way to preserve those distinctions through provenance, remainder tracking, standing, update authority, and reversible disposition.
LATCH has not earned adoption. Independent comparison must define its constructs reliably, isolate an incremental mechanism, test it against institutional power, and show better risk-adjusted outcomes at acceptable cost. If LATCH cannot survive those tests, institutions should abandon it. The obligation to distinguish performance from justification remains.
Appendix A. Prototype LATCH assessment instrument
This prototype is an evidence record, not a scorecard or certification. Each field must retain uncertainty, dissent, and provenance. “Not applicable” requires a reason; a blank field is not evidence of low risk.
A.1 Purpose–target–envelope register
- System and version: model, interface, tools, permissions, operator configuration, deployment date, and accountable custodian.
- Documented justificatory set (J): each claimed purpose, protected interest, right, harm avoided, and source of the claim.
- Operational specification (O): each reward, metric, rule, threshold, benchmark, or model used to guide or judge performance.
- Linking evidence: study, dataset, rationale, uncertainty, and known alternative constructs connecting O to elements of J.
- Empirical envelope (Xᵉ): tested populations, tasks, environments, versions, institutional conditions, and validity window.
- Known exclusions: populations, contexts, consequences, and interactions not represented.
A.2 Normative floor and authority map
- Normative-authorization context (Xⁿ): applicable law, regulation, professional duty, contractual undertaking, human-rights commitment, or public authority; include jurisdiction, standing, and conflict rules.
- Nonwaivable constraints: substantive rights, prohibitions, security boundaries, and harms in J that local optimization may not trade away; identify the corresponding source and jurisdiction in Xⁿ.
- Decision structure: who owns O, J, Xᵉ, Xⁿ, and deployment scope; who funds and can remove independent reviewers.
- Challenge and remedy: who has standing, evidence access, response rights, anti-retaliation protection, appeal, stop authority, restitution, or compensation.
A.3 Remainder and disposition record
- Candidate remainder (R): concise mismatch allegation; type; source; provenance; affected population; severity; uncertainty; object challenged.
- Intake integrity: how the claim entered, what may be missing, confidentiality protections, conflicts, owner, and deadline.
- Inquiry: alternative explanations, causal tests, counterexamples, field evidence, domain review, affected-party evidence, and dissent.
- Disposition: harden, hold, or halt; authorized decision-maker; reasons; minority report; remedy; scope and time limit.
- Update relation (U): exact object changed within the decision-maker's authority, or external authority to which amendment is escalated; staged test, rollback trigger, monitoring plan, review date, recurrence check, and public or independent reporting route.
A.4 Fidelity questions
A deployment should not claim LATCH fidelity unless reviewers can answer yes, with evidence, to all of the following: Is the O–J relation versioned? Is Xᵉ explicit? Is the source of Xⁿ identified? Can a protected claimant initiate R without product-owner approval? Must an independent actor answer with reasons? Can an authorized actor revise O or deployment? Can a severe case reach a tested safe state? Are dissent, disposition, remedy, and recurrence preserved? If any answer is no, the record should identify the missing capability rather than compute a compensating score.
Version, authorship, and research disclosure
“RVA Cyber Research” is a collective institutional byline for this public research paper; individual authorship is not asserted. This Clear Signal edition revises the prose and structure of the scholarly edition dated July 22, 2026. It does not claim new evidence or validation. The article reports no original human-subjects study and introduces no empirical dataset. Appendix A is a prototype instrument, not a validated assessment. AI-assisted research, drafting, adversarial review, citation checking, and web production were used in preparing the artifact; editorial responsibility remains with the institutional author. No external research sponsor is identified. The reference audit is current through July 22, 2026.
References
- Alfrink, Kars, Ianus Keller, Gerd Kortuem, and Neelke Doorn. 2023. “Contestable AI by Design: Towards a Framework.” Minds and Machines 33: 613–639. DOI.
- Argyris, Chris. 1977. “Double Loop Learning in Organizations.” Harvard Business Review 55 (5): 115–125. Article.
- Armstrong, Stuart, and Sören Mindermann. 2018. “Occam's Razor Is Insufficient to Infer the Preferences of Irrational Agents.” NeurIPS 31. Open version.
- Bansal, Hritik, John Dang, and Aditya Grover. 2024. “Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models.” ICLR 2024. Proceedings.
- Campbell, Donald T. 1979. “Assessing the Impact of Planned Social Change.” Evaluation and Program Planning 2 (1): 67–90. DOI.
- Cronbach, Lee J., and Paul E. Meehl. 1955. “Construct Validity in Psychological Tests.” Psychological Bulletin 52 (4): 281–302. DOI.
- Dewey, John. 1938. Logic: The Theory of Inquiry. New York: Henry Holt.
- Dewey, John. 1939. Theory of Valuation. Chicago: University of Chicago Press.
- Friedman, Batya, Peter H. Kahn Jr., and Alan Borning. 2008. “Value Sensitive Design and Information Systems.” In The Handbook of Information and Computer Ethics, 69–101. DOI.
- Fung, Archon. 2006. “Varieties of Participation in Complex Governance.” Public Administration Review 66 (s1): 66–75. DOI.
- Gao, Leo, John Schulman, and Jacob Hilton. 2023. “Scaling Laws for Reward Model Overoptimization.” Proceedings of Machine Learning Research 202: 10835–10866. Paper.
- Hadfield-Menell, Dylan, Smitha Milli, Pieter Abbeel, Stuart J. Russell, and Anca Dragan. 2017. “Inverse Reward Design.” NeurIPS 30. Proceedings.
- Hadfield-Menell, Dylan, Anca Dragan, Pieter Abbeel, and Stuart Russell. 2016. “Cooperative Inverse Reinforcement Learning.” NeurIPS 29. Open version.
- Jacobs, Abigail Z., and Hanna Wallach. 2021. “Measurement and Fairness.” FAccT 2021: 375–385. DOI.
- Jasanoff, Sheila. 2003. “Technologies of Humility: Citizen Participation in Governing Science.” Minerva 41: 223–244. Paper.
- Koh, Pang Wei, et al. 2021. “WILDS: A Benchmark of in-the-Wild Distribution Shifts.” Proceedings of Machine Learning Research 139: 5637–5664. Paper.
- Kundert, Matt. n.d. “A Review of Dr. Anthony McWatt's Essay: ‘Pirsig's Metaphysics of Quality.’” Archive.
- Langosco, Lauro Langosco Di, Jack Koch, Lee D. Sharkey, Jacob Pfau, and David Krueger. 2022. “Goal Misgeneralization in Deep Reinforcement Learning.” Proceedings of Machine Learning Research 162: 12004–12019. Open version.
- Leveson, Nancy. 2012. Engineering a Safer World: Systems Thinking Applied to Safety. Cambridge, MA: MIT Press. Open book.
- Manheim, David, and Scott Garrabrant. 2018. “Categorizing Variants of Goodhart's Law.” Paper.
- McWatt, Anthony Michael. 2004. A Critical Analysis of Robert Pirsig's Metaphysics of Quality. PhD diss., University of Liverpool. Repository.
- Obermeyer, Ziad, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. 2019. “Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations.” Science 366: 447–453. PubMed.
- Pirsig, Robert M. 1974. Zen and the Art of Motorcycle Maintenance: An Inquiry into Values. New York: William Morrow.
- Pirsig, Robert M. 1991. Lila: An Inquiry into Morals. New York: Bantam. Publisher.
- Pirsig, Robert M. 1999. “Subjects, Objects, Data and Values.” In Einstein Meets Magritte: An Interdisciplinary Reflection: The White Book, edited by Diederik Aerts, Jan Broekaert, and Ernest Mathijs, 79–98. Dordrecht: Springer. Originally presented in 1995. DOI.
- Pirsig, Robert M., and Julian Baggini. 2006. “An Interview with Robert Pirsig.” Complete email exchange associated with Baggini, “Zen and the Art of Dialogue,” The Philosophers' Magazine 33: 62–67. Transcript; published article.
- Raji, Inioluwa Deborah, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020. “Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing.” FAccT 2020: 33–44. DOI.
- Rao, Anita, Andrew Keller, Neha Kalra, Ryan Steed, Kweku Kwegyir-Aggrey, Kevin Klyman, Diane Staheli, and Amanda Bergman. 2026. Challenges to the Monitoring of Deployed AI Systems: Center for AI Standards and Innovation. NIST AI 800-4. Official publication.
- Rasmussen, Jens. 1997. “Risk Management in a Dynamic Society: A Modelling Problem.” Safety Science 27 (2–3): 183–213. DOI.
- Selbst, Andrew D., danah boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi. 2019. “Fairness and Abstraction in Sociotechnical Systems.” FAT* 2019: 59–68. Paper.
- Sharma, Mrinank, et al. 2024. “Towards Understanding Sycophancy in Language Models.” ICLR 2024. Proceedings.
- Skalse, Joar, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger. 2022. “Defining and Characterizing Reward Gaming.” NeurIPS 35. Paper.
- Sloane, Mona, Emanuel Moss, Olaitan Awomolo, and Laura Forlano. 2022. “Participation Is Not a Design Fix for Machine Learning.” Equity and Access in Algorithms, Mechanisms, and Optimization. DOI.
- Sneddon, Andrew. 1995. A Process Analysis of Quality: A. N. Whitehead and R. Pirsig on Existence and Value. Master's thesis, University of New Brunswick. Open version.
- Stilgoe, Jack, Richard Owen, and Phil Macnaghten. 2013. “Developing a Framework for Responsible Innovation.” Research Policy 42 (9): 1568–1580. DOI.
- Strawson, Galen. 1991. “Lone Man on High Seas.” Review of Lila, Sunday Times, October 27. Archive.
- Tabassi, Elham. 2023. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. Official framework.
- Winkler, Till, and Sarah Spiekermann. 2021. “Twenty Years of Value Sensitive Design: A Review of Methodological Practices in VSD Projects.” Ethics and Information Technology 23: 17–21. Article.