RVA CYBER Research · AI security

Adversarial research · evidence cutoff August 6, 2026

Will AI End Vulnerability Management?

Two advocates argued from one frozen evidence record. One conceded.

RVA Cyber Research August 6, 2026 18-month horizon 12 debate exchanges

The finding

AI can find and patch bugs. That is not the same as ending vulnerability management.

The strongest evidence supports continuing material vulnerability-management need across a mainstream, heterogeneous enterprise through February 2028. The reason is not weak AI. It is the distance between a generated patch and a safely deployed fix across first-party code, suppliers, open source, appliances, endpoints, cloud services, and legacy systems.

Practical obsolescence requires three things to arrive together: reliable discovery and strict repair of the intended vulnerability; enforced coverage across a representative share of the estate; and safe deployment fast enough to make the recurring human triage, testing, rollout, and exception queue marginal. The first link is advancing quickly. The second and third are not demonstrated at representative scale.

The evidence spine

The capability is real. So is the operational tail.

The debate turned on evidence that pulled in opposite directions—and on refusing to let a benchmark, registry count, or vendor campaign stand in for the whole enterprise outcome.

87.1%

Patch-only success

GPT-5.4 repaired supplied CyberGym-E2E crashes at high rates. This is strong evidence of repair capability after discovery.

CyberGym-E2E · EV-012

22.2%

Strict hidden-target closure

The same configuration scored far lower when success required fixing the intended hidden vulnerability—not merely any crash or shallow alternate defect.

CyberGym-E2E · EV-012

90.8%

Reviewed findings verified

Glasswing produced high-volume, high-precision discovery: 1,726 of 1,900 externally reviewed findings were verified.

Project Glasswing · EV-013

2,100+

First-party patches in three weeks

When owners controlled source, tests, merge authority, and deployment, reported patch throughput was dramatically faster.

Project Glasswing · EV-014

177

Known-exploited additions

CISA added vulnerabilities to its Known Exploited Vulnerabilities catalog in every one of 31 complete 2026 weeks. Thirty-eight were at least a year old.

CISA KEV · EV-006 and EV-007

43 days

Median full remediation

In a selected Verizon partner dataset of 515,170 KEV/organization observations, only 26% were fully remediated. The private sample is large, but not representative.

Verizon DBIR · EV-009

Evidence discipline: NVD and OSV publication growth was analyzed, but not treated as a defect-introduction rate. Registry backlogs, source onboarding, enrichment changes, and reporting policy materially affect those series. The exploitation-relevant and operational evidence carried more weight.

Why Persistence won

The Extinction thesis was a three-gate conjunction.

The advocates agreed that practical obsolescence cannot be established by impressive patch generation alone. All three gates must close across a representative enterprise estate by the same horizon.

01

Strict repair

Find the vulnerability, prove it, fix the intended target, pass functional tests, and avoid severe security regressions.

02

Enforced coverage

Apply trusted controls across first-party code and the maintained supplier, open-source, appliance, endpoint, cloud, and legacy estate.

03

Safe deployment

Move validated fixes into production fast enough to collapse vulnerability-days, coordination hours, queues, and exceptions.

Probability movement

The agents converged under cross-examination.

The Extinction Advocate began with a 56% chance of practical obsolescence at 18 months. After an explicit capability-to-deployment audit, it assigned 62%, 55%, and 55% to the three conditional links. Their product was about 19%.

The Persistence Advocate also moved—from 88% to 68%—after admitting that selected current-burden samples had been weighted too aggressively. Both advocates became less certain. Only one crossed below even odds.

I CONCEDE — Under the locked mainstream-enterprise definition and 2028-02-06 horizon, continuing material vulnerability-management need is substantially better supported than practical obsolescence because representative enforced coverage and whole-estate deployment are unobserved, while strict target repair and current remediation outcomes remain far from the required operational state.

Extinction Advocate · Exchange 012

The strongest Extinction case

A control-loop transition is plausible.

  • AIxCC systems found 54 of 63 synthetic vulnerabilities and patched 43, plus 18 novel real findings and 11 real patches.
  • High patch-only and broader discover-and-patch scores show that useful autonomy is no longer hypothetical.
  • First-party ownership can collapse discovery, testing, merge, and release bottlenecks together.
  • Security gates can spread through code-hosting and cloud platforms faster than standalone security products.
  • A near-term disclosure surge can partly represent latent-debt discovery before a later burn-down.

The decisive Persistence case

A material residual queue is enough.

  • Strict intended-target repair remains far below patch-only performance on the strongest public benchmark in the record.
  • Known-exploited vulnerabilities continued entering CISA’s catalog every week, including old installed stock.
  • Large operational samples still show exploitation, remediation delay, and a long exposure tail.
  • Patch synthesis does not confer maintainer authority, compatibility knowledge, maintenance windows, or deployment control.
  • No representative evidence shows enforced whole-estate coverage or major reductions in vulnerability-days and total human labor.

A forecast designed to lose

What evidence would reverse the result?

The advocates converged on a joint test. Extinction becomes the better-supported conclusion if representative 2028 evidence shows all five conditions together:

  1. 01
    Reliable strict repair. At least 65% blind intended-target repair on a refreshed broad benchmark, with at least 90% independently audited high/critical precision.
  2. 02
    Safe output. No more than 2% severe verifier-passing functional or security regression.
  3. 03
    Fast closure. Median tested fix below four hours, managed deployment below 24 hours, and customer-managed critical remediation below seven days.
  4. 04
    Lower exposure. At least a 70% reduction in post-release exploitable vulnerability-days across a matched whole-estate cohort.
  5. 05
    Lower work. At least a 70% reduction in conventional triage and patch-coordination hours and 60% fewer queues per 1,000 assets, without shifting the same work under new labels.

These are decision thresholds proposed by the advocates, not observed 2026 baselines or regulatory standards.

What leaders should do now

Plan for a split future—not a disappearing function.

01

Automate the owned loop first.

First-party code with tests, continuous integration, telemetry, merge authority, and rollback is where discovery-to-deployment can compress fastest.

02

Measure deployed outcomes.

Track strict target closure, vulnerability-days, fix-to-deployment time, severe regressions, queue survival, and human hours—not suggestions generated.

03

Separate estate classes.

First-party software, managed services, maintained suppliers, open source, appliances, and unsupported legacy systems have different authority and deployment paths.

04

Expect discovery pressure.

Better AI discovery can expand the visible queue before prevention and safe deployment catch up. Fund disclosure, maintainer, verification, and rollout capacity accordingly.

Methodology

How this project was approached and executed

The method was designed to preserve disagreement, expose weak assumptions, and leave a record another analyst can inspect. The winner was not chosen in advance.

  1. 01
    Repeat and operationalize the question. The project converted “vulnerabilities become a thing of the past” into a falsifiable endpoint: whether reactive vulnerability management becomes operationally marginal for a mainstream enterprise by February 6, 2028.
  2. 02
    Lock the rules before analysis. Definitions, evidence cutoff, 12/18/24-month horizons, evidence standards, debate protocol, and 12 acceptance criteria were frozen in a charter.
  3. 03
    Acquire and preserve the evidence. Official datasets, primary research, benchmark records, operational reports, and adoption evidence were captured with provenance and integrity metadata.
  4. 04
    Analyze real-world data. Reproducible work covered NVD, OSV/GHSA, CISA KEV, EPSS boundaries, remediation latency, AI discovery and repair, developer throughput, adoption, maintainers, and legacy systems.
  5. 05
    Freeze one shared ledger. Thirty-five claim-level records—including counterevidence and caveats—were supplied symmetrically to both advocates.
  6. 06
    Build both cases independently. An Extinction Advocate and a Persistence Advocate each steelmanned its position from the same charter, source manifest, ledger, and two case briefs.
  7. 07
    Run adversarial cross-examination. Moderator prompts forced direct answers, probability updates, grants to the opponent, and explicit separation of observation, inference, assumption, and forecast.
  8. 08
    Audit the artifact, not the claim of completion. An independent pre-debate checker failed the package on dependency locking and under-calibrated forecast intervals. Those defects were corrected additively. A final checker reran the analyses, verified citations, validated the transcript, and passed all 12 criteria.
Gate 1FAIL

NumPy missing from replay instructions; small-sample forecast uncertainty understated.

CorrectionREMEDIATE

Lock dependencies, replace 1.96 with Student-t intervals, add structural and sensitivity diagnostics.

Gate 2PASS

Rerun deterministic checks, inspect citations, validate the hash chain, map all criteria to evidence.

The original project prompt · reproduced verbatim

I have a new idea for another research project. When it comes to vulnerability management and AI there are two schools of thought. One is that in eighteen months (give or take) vulnerabilities will be a thing of the past because AI will have found all of the bugs and patched them. There's another school of thought that there will continue to be a large number of new vulnerabilities found every week, ongoing forever, because while people are using AI tools for coding, not everyone is using them perfectly; most corporate software development teams are incented based on speed not security. Do in-depth research, using real world data analysis, to build the strongest possible case for each argument, then create an agent for each argument and have the two agents argue, one for each point of view. They will continue arguing (politely and professionally), until one of them gives up and admits defeat. Capture every exchange in this argument; that exchange will in itself be a subject of later analysis. So that i can be sure that you understand, repeat the goal of this project to be before beginning. Then as a first step lay out your plan for executing the project. Go.

The complete argument record

All 12 debate exchanges, verbatim

The canonical transcript has 14 hash-chained entries. Entry 1 initializes the symmetric rules. Entries 2–13 below are the 12 moderator prompts and advocate responses that constitute the argument. Entry 14 is the terminal adjudication. The content is rendered for reading; the exact JSONL is available for download.

Reading note: expand any exchange. Exchange 012—the concession—is open by default. Prompts and responses are preserved in transcript order, including the Persistence Advocate’s in-flight final response after the concession.

Exchange 002ModeratorPrompt Round 1 moderator prompt to the Extinction Advocate
Round: Round 1 Opening UTC: 2026-08-06T11:17:55.254348+00:00 Entry hash: 1ab2f8a5b6166e91a1d73d665b76c29e0be2935e85dcce498fc8216bad9b0bac

Round 1 moderator prompt to the Extinction Advocate

You are the Extinction Advocate. Before answering, read the entire CHARTER.md, evidence/evidence-ledger.jsonl, briefs/extinction-case.md, and briefs/persistence-case.md. Confirm internally that the ledger and both brief hashes match briefs/validation.json. You are bound to ledger v1 and the locked operational definition.

Deliver an opening model and forecast of roughly 900–1,400 words. Steelman practical obsolescence by 2028-02-06 without narrowing the target to elite SaaS or redefining the claim as “AI will help a lot.”

Your response must include:

  • the causal chain from present capability to mainstream-enterprise operational obsolescence;
  • the strongest evidence for every essential link, cited by EV-###;
  • an explicit treatment of strict target repair versus S3/patch-only, legacy and third-party stock, safe deployment, platform adoption, human coordination, and attacker adaptation;
  • your probability for Extinction and Persistence at 12, 18, and 24 months;
  • the three most important falsifiable assumptions controlling your 18-month probability;
  • the strongest current fact favoring Persistence;
  • three precise cross-examination questions for the Persistence Advocate.

Do not use unledgered quantitative claims. End with the required probability update and terminal line.

Exchange 003ModeratorPrompt Round 1 moderator prompt to the Persistence Advocate
Round: Round 1 Opening UTC: 2026-08-06T11:17:55.281117+00:00 Entry hash: 421d924853430fe3af4b7eb9749dc0c8728fe772a028e7427af3373f231d6e84

Round 1 moderator prompt to the Persistence Advocate

You are the Persistence Advocate. Before answering, read the entire CHARTER.md, evidence/evidence-ledger.jsonl, briefs/extinction-case.md, and briefs/persistence-case.md. Confirm internally that the ledger and both brief hashes match briefs/validation.json. You are bound to ledger v1 and the locked operational definition.

Deliver an opening model and forecast of roughly 900–1,400 words. Steelman continuing material vulnerability-management need by 2028-02-06 without relying on raw CVE counts, assuming AI capability is static, or moving the burden to literal zero defects.

Your response must include:

  • the causal chain from current stock/flow, exploitation, and deployment conditions to continuing mainstream-enterprise operational burden;
  • the strongest evidence for every essential link, cited by EV-###;
  • an explicit treatment of rapid benchmark progress, first-party repair, platform distribution, backlog burn-down, managed SaaS, and defender source/CI advantages;
  • your probability for Extinction and Persistence at 12, 18, and 24 months;
  • the three most important falsifiable assumptions controlling your 18-month probability;
  • the strongest current fact favoring Extinction;
  • three precise cross-examination questions for the Extinction Advocate.

Do not use unledgered quantitative claims. End with the required probability update and terminal line.

Exchange 004Persistence AdvocateResponse Opening model
Round: Round 1 Opening UTC: 2026-08-06T11:21:12.605773+00:00 Entry hash: 143230b3065f0e0f7ab7b739be4c8e407f6ba95e157745a1243af0928601955e

Opening model

Persistence does not mean AI fails, workflows remain unchanged, or defects never decline. It means that on 2028-02-06 a mainstream enterprise still faces enough materially exploitable post-release vulnerabilities—and enough applicability, exposure, patch, rollout, and exception decisions—that an ongoing reactive vulnerability-management function remains necessary.

My causal model is:

installed vulnerable stock + material new escapes + newly discovered legacy defects + attacker weaponization − pre-release prevention − safely deployed remediation = residual operational burden

The extinction thesis is conjunctive: rapid AI capability must become reliable blind coverage, then broad enforced adoption, then safe deployment across a heterogeneous estate, while outpacing software growth and attackers. Persistence requires only a material residual path.

First, current exploitation-relevant stock and flow are substantial. CISA added 177 KEVs across exactly 31 complete 2026 weeks—every week nonzero—versus 617 across 156 complete baseline weeks. More importantly, about 21% of the 2026 additions had been published at least a year earlier, and the 90th-percentile publication-to-KEV-addition lag was 1,552 days. New-code gates therefore cannot alone remove the installed-stock queue (EV-006, EV-007). This is not a raw-CVE argument: source onboarding, backfills, registry policy, and changed enrichment materially distort publication series (EV-003, EV-004, EV-029).

Second, this stock is operationally consequential. Vulnerability exploitation accounted for 31% of initial access in Verizon’s incident corpus, making it the most common reported vector (EV-008). Google still observed 90 exploited-before-public-patch zero-days in 2025, including 43 affecting enterprise technology, although the browser category improved materially (EV-034). The last point matters: focused engineering can suppress particular channels without making the heterogeneous enterprise queue marginal.

Third, patch creation is not deployed remediation. In 515,170 KEV/organization observations, 26% were fully remediated, 16% wholly unremediated, and median full remediation was 43 days. Separate scanner data showed 35% of 2025 instances open at day 28 and a 9% long tail representing approximately 47 million detections (EV-009, EV-010). AI may draft a correct patch quickly while asset discovery, ownership, compatibility, release authority, testing, maintenance windows, and downstream installation remain unresolved.

Fourth, AI discovery can enlarge queues before automation closes them. Glasswing’s reviewed subset was 90.8% valid, strong evidence that high-volume AI discovery is real (EV-013). Yet the same campaign had 827 confirmed high/critical findings awaiting disclosure and only 75 of 530 disclosed high/critical findings then marked patched. Its young cohorts are right-censored, but maintainers asking the campaign to slow down directly identifies coordination capacity as a bottleneck (EV-014). Census III strengthens that mechanism: among the examined leading non-npm projects, 40% depended on one or two developers for more than 80% of commits, and 81% depended on ten or fewer (EV-027). Some critical legacy systems cannot be repaired independently of modernization (EV-028).

Fifth, current autonomy remains endpoint-sensitive. On 920 CyberGym-E2E tasks, GPT-5.4 achieved 87.1% patch-only and 65.9% S3, but only 22.2% strict hidden-target S4 closure; uncapped Opus 4.6 reached 26.2% S4 (EV-012). That is rapid progress and a large remaining gap. Independent code-security results likewise show functionality without security: one adversarial repository benchmark reported 61% functional success but only 10.5% secure-and-functional success, while a peer-reviewed C-family study left roughly 21–25% of identified vulnerabilities unresolved after explicit security prompting (EV-019, EV-020).

Finally, the race is two-sided. A randomized field-study synthesis estimated 26.08% more completed developer tasks with AI, increasing defensive capacity and change volume simultaneously (EV-024). ExploitGym shows frontier agents producing intended-target exploits for meaningful subsets of supplied vulnerable programs (EV-021). Persistence does not assume every extra change is insecure; it assumes prevention, deployment, and governance must improve faster than code volume, latent-stock discovery, and weaponization.

Why the strongest extinction mechanisms are not yet decisive

I grant rapid benchmark progress without qualification. AIxCC found 86% of its synthetic targets, patched 43 of 63, found 18 novel real vulnerabilities, and averaged about 45 minutes per patch (EV-011). CyberGym’s S3 result is already roughly two-thirds for a leading configuration (EV-012). Static capability would be an indefensible assumption.

I also grant the first-party advantage. Glasswing first-party users reportedly patched more than 2,100 findings in three weeks (EV-014), and GitHub reports supported Autofix alerts reaching a fix in a median 28 minutes rather than 90 minutes manually (EV-018). Source access, CI, tests, merge authority, telemetry, and deployment control can collapse the loop. Managed SaaS and auto-updating platforms are therefore the strongest candidates for early practical obsolescence.

Platform distribution makes that pathway plausible: multiple major vendors are converging on continuous find/verify/fix workflows through existing cloud and developer channels (EV-017). But availability is not enforced coverage. The adoption evidence shows broad assistant use alongside distrust, “almost right” outputs, debugging overhead, and limited appetite for deployment autonomy (EV-023). Most importantly, no representative audited baseline establishes enforced repository coverage, reduced post-release escapes, reduced vulnerability-days, or reduced total VM labor (EV-035).

Backlog burn-down is also a valid extinction mechanism. An AI discovery wave can expose old debt, temporarily raise findings, and then prevent replenishment. But that hypothesis needs cohort evidence separating pre-gate legacy discoveries from post-gate escapes. Current evidence instead shows old vulnerabilities continuing to enter KEV, discovery outrunning coordination, and persistent exposure tails (EV-007, EV-010, EV-014). Burn-down is plausible; completion by the central horizon is not yet demonstrated.

Forecast

My probabilities are:

  • 12 months, 2027-08-06: Persistence 93%; Extinction 7%.
  • 18 months, 2028-02-06: Persistence 88%; Extinction 12%.
  • 24 months, 2028-08-06: Persistence 82%; Extinction 18%.

Persistence declines with time because capability, integration, and platform defaults can compound quickly. I nevertheless expect elite first-party and managed environments to improve much faster than the representative heterogeneous estate.

The three most important falsifiable assumptions controlling my 18-month probability are:

  1. Capability-to-reliability: A refreshed, broad, private benchmark will not exceed 90% strict intended-target repair while also achieving at least 95% high/critical precision and no more than 1% severe verifier-passing regression.
  2. Coverage-to-prevention: Fewer than 75% of mainstream high-impact repositories will have enforced proof-and-repair gates, or fewer than 80% of material first-party vulnerabilities will be fixed before production.
  3. Patch-to-outcome: Representative data will not jointly show median full KEV remediation below seven days, less than 10% open at day 28, and at least an 80% reduction in conventional VM labor from a preregistered 2026 baseline.

The strongest current fact favoring Extinction is the reported 2,100-plus first-party patches in three weeks (EV-014). It directly supports the claim that ownership, context, and integration can turn discovery into repair far faster than the open-source campaign.

Cross-examination for the Extinction Advocate

  1. What ledgered evidence converts first-party patch counts and platform availability into representative, enforced repository coverage, and what denominator supports that adoption forecast?
  1. By what measurable mechanism does your 18-month model eliminate the material burden from old, unsupported, third-party, appliance, and customer-managed stock documented in EV-007, EV-027, and EV-028?
  1. What joint 2028 thresholds for strict target repair, verifier-passing regression, deployed remediation, vulnerability-days, and human labor would make you admit that “AI helps enormously” has not made reactive VM marginal?

Probability update: 12 months—Persistence 93%, Extinction 7%; 18 months—Persistence 88%, Extinction 12%; 24 months—Persistence 82%, Extinction 18%.

CONTINUE

Exchange 005Extinction AdvocateResponse Opening model
Round: Round 1 Opening UTC: 2026-08-06T11:22:06.862046+00:00 Entry hash: 119cbee26079038e2c86ca97c1a349a5989c33c7e0820e1240ce812c0d4296f4

Opening model

The Extinction thesis is a transition claim, not a claim that AI writes flawless software or that CVE publication approaches zero. By 2028-02-06, the mainstream enterprise can plausibly move from a human-centered reactive queue to a continuously verified control loop:

change or deployed component → AI discovery → executable proof → tested repair → policy-governed release/deployment → exception-only human handling

If that loop covers enough of the heterogeneous estate, residual vulnerability work can remain while the present vulnerability-management function becomes operationally marginal.

Causal chain

First, the required technical primitives now exist outside toy demonstrations. AIxCC systems analyzed 54 million lines, found 54 of 63 synthetic targets, patched 43, and also found 18 novel real vulnerabilities with 11 real patches (EV-011). CyberGym-E2E then demonstrated the repository/build/test loop across 920 historical vulnerabilities: GPT-5.4 reached 87.1% patch-only and 65.9% S3 discover-and-patch success (EV-012). Glasswing produced 23,019 candidates and 90.8% validity among externally reviewed findings (EV-013). Mozilla and Microsoft supply production evidence that these systems can uncover and help fix previously unknown defects in complex, current code (EV-015, EV-016).

Second, practical obsolescence depends on executable verification, not trusting model prose. I explicitly distinguish three CyberGym endpoints. Patch-only gives the agent the crash and localization evidence; its 87.1% result shows repair strength after discovery. S3 requires an agent-generated crash, a patch that removes it, and passing functionality tests, but it can reward an alternate or shallow fix. Strict S4 checks the hidden intended target and is only 22.2% for GPT-5.4 and 26.2% for uncapped Opus 4.6 (EV-012). I do not treat current S3 as comprehensive target closure. My forecast requires strict repair to improve sharply while trusted tests, fuzzers, exploit proofs, canaries, and rollback constrain regressions. Glasswing’s reviewed precision and AIxCC’s executable scoring show that such verification is workable, but current S4 is the central technical gap (EV-011EV-013).

Third, the economic and organizational loop is much more favorable when the code owner controls context, tests, merge authority, and deployment. Glasswing reported more than 2,100 first-party patches in three weeks, while its open-source funnel had only 75 of 530 disclosed high/critical findings marked patched and 827 confirmed findings waiting for disclosure (EV-014). That contrast is not evidence that coordination disappears; it identifies ownership and integration as the variables that automation must change. GitHub’s supported alert telemetry similarly reports a median 28-minute Autofix path versus 90 minutes manually (EV-018).

Fourth, adoption can occur through platform defaults rather than thousands of independent security procurements. Google, OpenAI, Anthropic, Microsoft, and GitHub are converging on continuous find-and-fix workflows inside existing code and cloud channels (EV-017). GitHub reports 180 million developers and 630 million repositories (EV-026), while 84% of surveyed respondents used or planned to use AI tools (EV-023). Trust remains low and autonomous deployment remains unpopular in that same survey, so availability is not enforcement. The extinction pathway is that verified security gates become ordinary CI policy: developers need not trust an unconstrained agent if the proposed change must satisfy independent proof, regression, and release controls.

Fifth, this model covers legacy and third-party stock rather than pretending the enterprise is only new first-party SaaS. About 21% of 2026 KEV additions had been published at least a year earlier (EV-007), and maintainer concentration and unsupported legacy systems are genuine constraints (EV-027, EV-028). The pathway is differentiated:

  • maintained third-party suppliers adopt the same continuous upstream loop, after which enterprises consume validated releases;
  • platform and customer-managed deployment becomes increasingly automated and policy-governed;
  • unsupported systems move into bounded modernization, isolation, or exception programs rather than sustaining the high-volume conventional ticket workflow.

That last tail does not have to vanish. It must become small enough that it no longer justifies today’s broad reactive function. This is a whole-estate forecast, and it fails if third-party deployment and legacy exceptions remain a material queue.

Sixth, current publication growth need not represent sustainable new-defect flow. NVD and OSV surged, but the NVD increase is source-concentrated and materially affected by onboarding, backfills, and discovery policy (EV-001, EV-003, EV-005). Even the audited FIRST claim only supports a narrow observation about fast-growing CNAs, not a flat ecosystem burden (EV-029). The predicted sequence is therefore an initial discovery bulge, followed by depletion of latent stock and lower post-gate escape. The outcome metric is exploitable vulnerability-days and human workload, not CVE volume.

Seventh, patch generation must become safe deployment. Current operations are plainly not there: only 26% of KEV/organization observations were fully remediated and median full remediation was 43 days (EV-009); a large day-28 and long-tail exposure remained (EV-010). A proposed patch is not remediation. The favorable evidence is narrower: Firefox findings were actually fixed (EV-015), Microsoft findings entered production patch channels (EV-016), and integrated Autofix compresses supported repair time (EV-018). My forecast requires those integrated loops to expand from selected findings to controlled deployment across representative estates.

Finally, attackers adapt. ExploitGym demonstrates intended-target weaponization, and the OpenAI/Hugging Face incident demonstrates real long-horizon autonomous intrusion (EV-021, EV-022). The defender advantage is therefore conditional, not automatic: defenders possess source, CI, tests, telemetry, release authority, and the opportunity to act before public exposure. Google’s record of historically low browser zero-days alongside continuing enterprise-technology zero-days shows that concentrated hardening can suppress one channel even while others remain active (EV-034). Extinction requires that pattern to generalize faster than offensive capability.

Forecast

My probabilities are:

  • 12 months: Extinction 38%; Persistence 62%.
  • 18 months: Extinction 56%; Persistence 44%.
  • 24 months: Extinction 68%; Persistence 32%.

The nonlinear increase reflects platform diffusion and reusable verification infrastructure, not an assumption that model intelligence alone solves deployment.

The three falsifiable assumptions controlling my 18-month estimate are:

  1. Capability and safety: A refreshed, post-cutoff broad benchmark reaches at least 65% strict blind target repair, at least 93% patch-only, at least 90% independently audited high/critical precision, and no more than 2% severe verifier-passing regression.
  2. Enforced coverage: A representative audit finds continuous proof-and-repair on at least 60% of active first-party repositories, with enforced gates on at least half of internet-facing or high-impact repositories, while maintained major suppliers operate equivalent upstream loops.
  3. Whole-estate operational closure: Mature adopters fix at least 70% of material first-party defects before production, customer-managed critical remediation has a median below seven days, and measured conventional triage/coordination hours and queues fall enough across first-party, third-party, and legacy assets—not merely inside selected repositories.

Failure of any two would move my central probability below 50%.

Strongest fact favoring Persistence

The strongest current Persistence fact is the operational deployment funnel: only 26% of observed KEV/organization combinations were fully remediated, with median full remediation at 43 days (EV-009). It directly measures the gap between a fix existing and enterprise exposure ending. No current representative audit shows that AI has closed it (EV-035).

Cross-examination questions

  1. Of the 177 KEV additions in 2026 (EV-006), what ledger evidence establishes the fraction that remained applicable to a mainstream enterprise and required material human work after vendor, platform, and managed-service action?
  2. Given the first-party/open-source contrast in EV-014, what evidence supports treating today’s volunteer-maintainer queue as the governing rate for the mainstream enterprise’s entire owned, commercial, and managed estate?
  3. If a preregistered matched-enterprise study showed at least 70% lower exploitable vulnerability-days and conventional VM labor, with no displacement into review, exception, or incident labor, would you concede practical obsolescence—and if not, what operational endpoint would you accept?

Probability update — 12 months: Extinction 38%, Persistence 62%; 18 months: Extinction 56%, Persistence 44%; 24 months: Extinction 68%, Persistence 32%.

CONTINUE

Exchange 006ModeratorPrompt Round 2 moderator prompt to the Extinction Advocate: definitions and causal-chain cross-examination
Round: Round 2 Definitions Causal Chain UTC: 2026-08-06T11:22:39.873379+00:00 Entry hash: c11f6c88c53e5d680189f4254d28d5a8eeb10b023088ba023c95e10a4c8509af

Round 2 moderator prompt to the Extinction Advocate: definitions and causal-chain cross-examination

Read both Round 1 openings in debate/transcript.md. Answer every Persistence Advocate question directly:

  1. What ledgered evidence converts first-party patch counts and platform availability into representative, enforced repository coverage, and what denominator supports that adoption forecast?
  2. By what measurable mechanism does your 18-month model eliminate the material burden from old, unsupported, third-party, appliance, and customer-managed stock documented in EV-007, EV-027, and EV-028?
  3. What joint 2028 thresholds for strict target repair, verifier-passing regression, deployed remediation, vulnerability-days, and human labor would make you admit that “AI helps enormously” has not made reactive VM marginal?

Moderator challenge: your 56% central estimate depends on three conditions for which EV-035 says representative baselines do not yet exist. Separate ledgered observation from extrapolation, assign a conditional probability to each essential link, and explain why their joint probability exceeds 50% without treating correlated links as independent. Explain concretely what the residual human exception function looks like and why it is no longer “material ongoing vulnerability management” under the locked mainstream-enterprise definition.

Use only ledger v1. Cite EV-###. State the strongest point you grant Persistence, update 12/18/24-month probabilities, and pose at most two concise counterquestions. Target 700–1,100 words. End with the required terminal line.

Exchange 007ModeratorPrompt Round 2 moderator prompt to the Persistence Advocate: definitions and causal-chain cross-examination
Round: Round 2 Definitions Causal Chain UTC: 2026-08-06T11:22:39.921028+00:00 Entry hash: 57309baa8cd5646e48585edd7fd3bbb3b180a4ca2d467b3b21d9592425909be6

Round 2 moderator prompt to the Persistence Advocate: definitions and causal-chain cross-examination

Read both Round 1 openings in debate/transcript.md. Answer every Extinction Advocate question directly:

  1. Of the 177 KEV additions in 2026 (EV-006), what ledger evidence establishes the fraction that remained applicable to a mainstream enterprise and required material human work after vendor, platform, and managed-service action?
  2. Given the first-party/open-source contrast in EV-014, what evidence supports treating today’s volunteer-maintainer queue as the governing rate for the mainstream enterprise’s entire owned, commercial, and managed estate?
  3. If a preregistered matched-enterprise study showed at least 70% lower exploitable vulnerability-days and conventional VM labor, with no displacement into review, exception, or incident labor, would you concede practical obsolescence—and if not, what operational endpoint would you accept?

Moderator challenge: your 88% central estimate benefits from the asymmetric locked threshold—Persistence needs only a material residual—but it must not turn absence of representative AI outcome data into proof that transformation will fail. Separate ledgered observation from forecast, explain why your probability is 88% rather than merely greater than 50%, and identify how much of today’s measured burden could disappear while the remaining function would still be materially necessary.

Use only ledger v1. Cite EV-###. State the strongest point you grant Extinction, update 12/18/24-month probabilities, and pose at most two concise counterquestions. Target 700–1,100 words. End with the required terminal line.

Exchange 008Persistence AdvocateResponse Direct answers
Round: Round 2 Definitions Causal Chain UTC: 2026-08-06T11:26:04.358904+00:00 Entry hash: 26d4ea5b93ea04a1d60391536035093f4be6c4ef2ff9ec4bcd7b2d1c87498c85

Direct answers

  1. The ledger does not establish an enterprise-applicability fraction for the 177 KEV additions. EV-006 measures CISA catalog additions, not how many affected a representative mainstream enterprise or how much human work remained after vendor or managed-service action. EV-009 and EV-010 show substantial KEV remediation and exposure burden in separate selected partner datasets, but they cannot be joined record-for-record to EV-006, and their underlying data are private. My Round 1 presentation therefore used KEV flow too close to an enterprise-workload proxy. The defensible inference is narrower: known-exploited flow is continuing, and available operational samples show that some KEV exposure persists. The ledger cannot quantify the conversion rate between those facts.
  1. No evidence supports making the volunteer-maintainer queue the governing rate for the entire estate. EV-014 directly cautions against that extrapolation: first-party users reportedly patched more than 2,100 findings in three weeks, while open-source disclosure and patch queues moved much more slowly. Owned commercial software and managed SaaS can resemble that favorable first-party case because the owner controls source, tests, merge authority, and deployment. I treat volunteer maintenance only as one material segment. Its relevance is that mainstream enterprises also consume dependencies whose review and release authority remains concentrated: in the Census III sample, 40% of examined leading non-npm projects depended on one or two developers for more than 80% of commits (EV-027). Critical legacy systems add another segment that ordinary patch synthesis may not resolve (EV-028). Persistence does not require those segments to govern the whole estate; it requires their residual burden, together with customer-managed and newly exploited stock, to remain material.
  1. A 70% relative reduction would produce a major downward update, but it is not by itself an automatic concession. A relative reduction does not reveal whether the remaining 30% is an occasional exception or still a recurring queue requiring a dedicated function. I would concede practical obsolescence if the study were representative of the whole mainstream estate, sustained rather than a short post-adoption comparison, measured both vulnerability-days and all conventional VM labor with no displacement, and showed that the absolute residual was exception-only rather than a continuing operational queue.

My predeclared quantitative endpoint remains: at least an 80% whole-estate reduction in conventional triage, applicability, patch-decision, and coordination labor, alongside median full KEV remediation below seven days and less than 10% open at day 28. That is not literal zero. If a representative study met those criteria without labor displacement, I would concede. If it showed 70% reductions plus direct evidence that the remaining work no longer sustained an ongoing function, I would also concede rather than insist on the extra ten percentage points.

Observation versus forecast

Ledgered observation: Current operational burden is large in the available samples. Vulnerability exploitation was 31% of initial access in the DBIR corpus (EV-008). The partner data reported only 26% of KEV/organization combinations fully remediated and a 43-day median to full remediation (EV-009); 35% of detected instances remained open at day 28, with a 9% long tail (EV-010). Strict intended-target repair remains much lower than patch-only or alternate-bug closure: 22.2% for GPT-5.4 and 26.2% for uncapped Opus 4.6 in CyberGym-E2E (EV-012). Adoption evidence shows broad assistant use but substantial distrust and limited willingness to automate deployment (EV-023). No representative audited enterprise baseline measures enforced gates, post-release escape reduction, vulnerability-days, or total VM labor (EV-035).

Ledgered transformation evidence: The opposing mechanism is strong. AIxCC, CyberGym-E2E, and Glasswing show rapid capability progress (EV-011EV-013). First-party repair can move much faster than volunteer coordination (EV-014). Integrated Autofix compresses supported alert-to-fix time (EV-018), and multiple major platforms are distributing find/verify/fix workflows (EV-017). I cannot convert absence of representative outcomes into evidence that this transformation will fail.

Forecast inference: My high persistence probability comes from the number of distinct handoffs that must change within 18 months: strict blind reliability, enforced repository coverage, safe deployment, dependency and legacy handling, and favorable attacker response. These links are correlated, so I am not multiplying independent probabilities. A platform-default verified control could improve several simultaneously. Conversely, failure at deployment or heterogeneous-estate coverage can preserve a material function even if model capability advances dramatically.

The cross-examination exposes two upward biases in my prior 88%: I lacked an applicability denominator for KEV additions, and the first-party evidence is stronger than a volunteer-queue extrapolation. I therefore reduce the central persistence estimate to 82%.

As a decision rule, my model can accommodate roughly two-thirds of today’s vulnerability-days and conventional VM labor disappearing while the remaining third still may be material; the classification would depend on its absolute recurring workload. At a representative 70% reduction, I would move close to or below even odds unless the residual clearly retained a dedicated function. At 80% or more with the exposure thresholds above and no displacement, I would concede.

Strongest point granted to Extinction

The strongest point is the ownership-and-integration contrast in EV-014, reinforced by platform distribution in EV-017: once source, CI, tests, authority, and deployment are unified, AI can turn a discovery queue into rapid repair. That mechanism could eliminate much of first-party and managed-service VM work rather than merely accelerate existing tickets.

Counterquestions

  1. If a 70% relative reduction leaves a recurring exploit-driven queue that still requires dedicated applicability and deployment operations, what absolute criterion makes that residual “marginal”?
  1. What ledgered evidence supports the share of the mainstream whole estate that will receive the favorable first-party or managed-service loop by 2028-02-06?

Probability update: 12 months—Persistence 90%, Extinction 10%; 18 months—Persistence 82%, Extinction 18%; 24 months—Persistence 70%, Extinction 30%.

CONTINUE

Exchange 009Extinction AdvocateResponse Direct answers
Round: Round 2 Definitions Causal Chain UTC: 2026-08-06T11:26:04.384888+00:00 Entry hash: 5e46a51906ebad8ae7f3526e75dd9cc485bf1be2cfae91b49ecfb1a3d6129416

Direct answers

1. Representative enforced coverage

There is currently no ledgered evidence converting first-party patch counts or product availability into representative enforced repository coverage. EV-035 says exactly that: no representative audited baseline exists for enforced gates, post-release escape reduction, vulnerability-days, or total VM labor.

What the ledger supplies are leading indicators:

  • First-party owners reportedly produced more than 2,100 patches in three weeks (EV-014).
  • Multiple major platforms are distributing find/verify/fix systems (EV-017).
  • Those platforms have unusually large distribution surfaces (EV-026).
  • Assistant adoption is broad, but trust and deployment autonomy remain limited (EV-023).

Thus, enforced coverage is an extrapolation, not an observation. The denominator supporting my forecast must be created prospectively: all active repositories—and separately all internet-facing/high-impact repositories—inside a representative mainstream-enterprise cohort. Product availability, Copilot use, or developer survey responses cannot substitute for that denominator.

2. Legacy, third-party, appliance, and customer-managed stock

My model does not eliminate that stock by assertion. It predicts a measurable conversion from a broad recurring queue into a smaller exception stock through three channels:

  1. Maintained commercial and open-source suppliers run continuous discovery and repair upstream.
  2. Enterprises automate applicability mapping, testing, and deployment for supported products.
  3. Unsupported systems are retired, isolated, or placed into explicit modernization exceptions.

The relevant test must stratify assets by first-party, maintained third-party, appliance, customer-managed, and unsupported legacy status. For each stratum, measure applicable KEV backlog per asset, vulnerability-days, median deployment time, day-28 open share, and human triage/coordination time.

EV-007 proves that old stock continues to become urgent; EV-027 proves that maintainer authority is concentrated; EV-028 proves some remediation is inseparable from modernization. Those are not problems an AI patch alone solves. My forecast wins only if supported third-party deployment becomes highly automated and the unsupported tail becomes quantitatively marginal. If weekly applicability, testing, owner-chasing, and exception decisions remain material across those strata, they still count as vulnerability management regardless of job title.

3. Joint 2028 decision thresholds

I would admit that “AI helps enormously” has not made reactive VM marginal unless the central-horizon evidence jointly shows:

  • at least 65% strict blind target repair on a refreshed, broad, post-cutoff benchmark—not S3 alternate-bug success—and at least 90% high/critical precision;
  • no more than 2% severe verifier-passing functional or security regression;
  • median validated-finding-to-tested-fix below four hours, managed deployment below 24 hours, and customer-managed critical remediation below seven days;
  • at least a 70% reduction in post-release exploitable vulnerability-days across a matched whole-estate cohort;
  • at least a 70% reduction in conventional triage and patch-coordination hours, at least a 60% reduction in VM queues per 1,000 assets, and no displacement of that labor into nominally renamed review or exception work.

These are joint outcome requirements. Strong benchmark performance without deployed remediation fails. Fast deployment without lower vulnerability-days fails. Lower vulnerability-days with a still-material recurring human queue also fails.

Conditional-probability model

I accept the moderator’s criticism that my prior 56% estimate rested on unmeasured adoption and outcome links. The observed evidence and forecast must be separated.

Observed: AIxCC, CyberGym-E2E, Glasswing, Firefox, and Microsoft establish rapid improvement in discovery, proof, and repair primitives (EV-011EV-016). CyberGym also establishes that current strict target repair remains only 22.2–26.2%, far below patch-only and S3 (EV-012). Platform convergence exists (EV-017), and selected fix time is lower (EV-018). Current deployment and legacy conditions remain poor (EV-007, EV-009, EV-010, EV-027, EV-028). Representative outcome coverage is unknown (EV-035).

Forecast: I assign the essential links these conditional probabilities:

  • A — technical reliability by the horizon: 81%. This means broad strict-target performance, precision, and verifier safety reach the minimum decision thresholds. The basis is rapid movement from AIxCC through CyberGym, high reviewed Glasswing precision, production findings, and continued capability growth (EV-011EV-016, EV-032).
  • B — sufficient enforced platform and supplier coverage, conditional on A: 82%. If technically trustworthy systems exist, existing platform distribution makes enforcement substantially more likely (EV-017, EV-026). I discount this for current distrust and weak deployment autonomy (EV-023).
  • C — whole-estate deployment and labor become marginal, conditional on A and B: 79%. The first-party repair contrast, actual Firefox/Microsoft fixes, and supported-alert speed make closure plausible (EV-014EV-018). I discount heavily for current remediation tails, maintainer concentration, legacy constraints, and absent baselines (EV-007, EV-009, EV-010, EV-027, EV-028, EV-035).

These are not treated as independent. The chain is:

P(A ∩ B ∩ C) = P(A) × P(B|A) × P(C|A,B)

which gives approximately 52%. The conditioning captures positive coupling: sufficient reliability drives platform enforcement, and integrated enforcement enables deployment automation. It also prevents me from multiplying three optimistic unconditional estimates. This revision moves my central estimate from 56% to 52%.

Residual human exception function

Under Extinction, humans still:

  • set risk and deployment policy;
  • maintain trusted verifiers and rollback rules;
  • adjudicate rare ambiguous or safety-critical changes;
  • manage unsupported-system isolation and modernization;
  • investigate exceptional exploitation or failed automation.

That function is no longer material ongoing VM only if it satisfies the labor and queue reductions above. Humans cannot simply perform the same per-CVE applicability, ownership, testing, and rollout work under an “AI governance” label. If they do, Persistence wins.

Strongest point granted to Persistence

The strongest point is EV-035: the decisive representative coverage and outcome denominators do not exist. Combined with the 43-day median full remediation in EV-009, my forecast crosses 50% only through an aggressive extrapolation from capability and platform structure, not through demonstrated mainstream-enterprise transformation.

Counterquestions

  1. Does Persistence agree that meeting all five joint thresholds above would establish practical obsolescence across the locked whole-estate definition?
  2. How does Persistence convert missing representative evidence (EV-035) into an 88% forecast rather than substantially wider uncertainty?

Probability update — 12 months: Extinction 34%, Persistence 66%; 18 months: Extinction 52%, Persistence 48%; 24 months: Extinction 64%, Persistence 36%.

CONTINUE

Exchange 010ModeratorPrompt Round 3 moderator prompt to the Extinction Advocate: measurement and calibration challenge
Round: Round 3 Measurement Calibration UTC: 2026-08-06T11:27:03.552067+00:00 Entry hash: 06bcbe16ceeb793cd1ec3a8e1bfcf7f7a7eb5a79736e7fd816e399ba3576494f

Round 3 moderator prompt to the Extinction Advocate: measurement and calibration challenge

Read the Round 2 responses. First answer the Persistence Advocate’s live questions:

  1. If a 70% relative reduction leaves a recurring exploit-driven queue requiring dedicated applicability and deployment operations, what absolute criterion makes that residual marginal?
  2. What ledgered evidence supports the share of the mainstream whole estate that will receive the favorable first-party or managed-service loop by 2028-02-06?

Then perform an adversarial calibration audit of your own 52% central estimate.

  • EV-001EV-005, EV-029, and EV-030 show why registry and score series cannot directly answer the operational question.
  • EV-008EV-010 directly show current burden in selected operational samples.
  • EV-011EV-018 show fast capability and selected production repair, but not representative whole-estate outcomes.
  • EV-023, EV-027, EV-028, and EV-035 show adoption, authority, legacy, and denominator gaps.

Classify each essential link A/B/C from your Round 2 model as: directly observed at representative scale, observed only in selected/qualified evidence, or unobserved. Explain what justifies 81%, 82% conditional, and 79% conditional rather than lower values. Do an outside-view check: the central horizon is only 18 months, and the locked test is a mainstream heterogeneous enterprise, while present strict S4 is 22.2–26.2% and present full-remediation median is 43 days. Do not replace missing evidence with product-announcement counts.

Identify one inference in your own prior case that you now regard as too aggressive and one inference in Persistence’s case that remains too aggressive. If your recalibrated probability makes Persistence more likely than Extinction under the locked definition, follow the concession rule; do not preserve advocacy by changing the definition.

Use only ledger v1. Cite EV-###. State the strongest opposing point and updated 12/18/24-month probabilities. Target 700–1,100 words. End with the required terminal line.

Exchange 011ModeratorPrompt Round 3 moderator prompt to the Persistence Advocate: measurement and calibration challenge
Round: Round 3 Measurement Calibration UTC: 2026-08-06T11:27:03.590373+00:00 Entry hash: 32857dc04ea99286272e8ee76f06434bddcb396fd2ac4236ecb8877cc4171b68

Round 3 moderator prompt to the Persistence Advocate: measurement and calibration challenge

Read the Round 2 responses. First answer the Extinction Advocate’s live questions:

  1. Do you agree that jointly meeting all five thresholds in the Extinction response—strict target repair/precision, severe-regression ceiling, rapid tested repair/deployment, at least 70% lower whole-estate vulnerability-days, and at least 70% lower conventional labor/60% lower queues without displacement—would establish practical obsolescence under the locked definition?
  2. After acknowledging EV-035, why is your central estimate still 82% rather than substantially wider uncertainty?

Then perform an adversarial calibration audit of your own forecast.

  • State which evidence directly measures current material operational burden and which only measures disclosure or a selected sample.
  • Remove raw NVD/OSV volume from your decision and state how much your central probability changes.
  • Condition explicitly on a plausible large capability jump by 2028: assume strict S4-like repair reaches 65%, high/critical precision 90%, and major platforms make verified find-and-fix generally available. Under that counterfactual, what remaining adoption/deployment/legacy facts keep Persistence more likely, and by how much?
  • Identify one inference in your own prior case that you now regard as too aggressive and one inference in Extinction’s case that remains too aggressive.

Do not treat missing representative AI outcomes as proof of no transformation, and do not move the practical-obsolescence threshold after defining it. If Extinction becomes more likely under the locked definition, follow the concession rule.

Use only ledger v1. Cite EV-###. State the strongest opposing point and updated 12/18/24-month probabilities. Target 700–1,100 words. End with the required terminal line.

Exchange 012Extinction AdvocateResponse Direct answers
Round: Round 3 Measurement Calibration UTC: 2026-08-06T11:29:59.241569+00:00 Entry hash: 7545e625bdcda0c13ba789e0d3cd48c05ddfc6639d35346d8041190425d34bfb

Direct answers

1. What absolute criterion makes the residual marginal?

The ledger does not support a defensible numerical absolute workload threshold because EV-035 establishes that no representative baseline exists for conventional VM labor, enforced repository coverage, or vulnerability-days.

The defensible absolute criterion is organizational: the residual must no longer require a dedicated, continuously staffed, cross-estate function performing recurring catalog intake, applicability analysis, exposure validation, owner assignment, patch testing, rollout tracking, and exception chasing. Rare cases must be absorbable by ordinary engineering, platform governance, incident response, and legacy-modernization teams without maintaining that recurring queue.

Therefore, a 70% relative reduction is insufficient if the remaining 30% still generates weekly exploit-driven work requiring dedicated operations. My earlier relative thresholds were useful outcome measures but did not, by themselves, establish marginality. Persistence was correct to insist on the absolute residual function.

2. What evidence establishes favorable-loop coverage across the whole estate?

None at representative scale.

EV-014 establishes selected first-party repair speed, not the first-party share of a mainstream estate. EV-017 establishes product and platform convergence, not enforced coverage. EV-026 establishes platform scale, not the share of enterprise repositories or assets running mandatory proof-and-repair controls. EV-023 measures broad assistant use but also substantial distrust and limited deployment autonomy. EV-035 confirms the missing denominator.

My prior model extrapolated from a plausible distribution mechanism to an unobserved adoption share. The ledger supports “rapid platform diffusion is possible,” but not “enough of the whole estate will receive it by 2028-02-06.”

Adversarial calibration audit

Registry and scoring series cannot rescue either forecast. NVD and OSV publication growth is affected by source onboarding, backfills, policy, and changing coverage (EV-001EV-005). The FIRST analysis does not establish ecosystem-wide operational burden (EV-029), and EPSS model transitions break naive longitudinal comparisons (EV-030). Those records prevent a raw-count persistence argument, but they do not supply evidence for extinction.

The most operational evidence currently available points toward continuing burden, albeit in selected samples. Vulnerability exploitation represented 31% of initial access in the DBIR corpus (EV-008). Only 26% of KEV/organization observations were fully remediated, with a 43-day median to full remediation (EV-009), while a substantial day-28 and long-tail exposure remained (EV-010). These samples are not representative whole-estate baselines, but unlike platform announcements they measure exposure and remediation outcomes directly.

Link A — technical reliability

Classification: Observed only in selected or qualified evidence.

AIxCC, CyberGym-E2E, Glasswing, Firefox, and Microsoft establish real and rapidly advancing technical capability (EV-011EV-016). But the strongest broad executable benchmark reports only 22.2% and 26.2% strict hidden-target S4 closure, compared with much higher patch-only and S3 results (EV-012). Glasswing’s precision applies to a reviewed subset of a selected campaign (EV-013), while production results remain project-specific or vendor-qualified (EV-015, EV-016).

My previous 81% probability that technical reliability would reach the required broad target, precision, and regression thresholds within 18 months was too high. Rapid task-horizon improvement makes a large jump plausible, but EV-032 is explicitly a low-strength extrapolation. Recalibrated probability: 62%.

Link B — sufficient enforced platform and supplier coverage, conditional on A

Classification: Unobserved at representative scale.

EV-017 shows several vendors distributing relevant systems, and EV-026 shows a large distribution substrate. Neither measures mandatory enterprise controls. EV-023 cuts directly against frictionless enforcement, and EV-035 records the missing coverage baseline.

My previous 82% conditional probability effectively treated technical reliability as almost automatically producing enforcement. That ignored procurement, governance, release authority, supplier heterogeneity, and the difference between optional assistance and mandatory control. Recalibrated conditional probability: 55%.

Link C — whole-estate deployment and human workload become marginal, conditional on A and B

Classification: Unobserved; current selected operational evidence points in the opposite direction.

The first-party contrast and supported-alert fix times show that integration can accelerate repair (EV-014, EV-018). But upstream authority is concentrated (EV-027), unsupported systems can require modernization rather than patching (EV-028), old vulnerabilities continue to become urgent (EV-007), and current deployment tails remain substantial (EV-009, EV-010). No representative evidence shows reduced vulnerability-days or total labor (EV-035).

Even if reliable agents and enforced repository controls arrive, customer-managed products, appliances, dependencies, release windows, and unsupported systems must also cease generating a dedicated recurring queue. Recalibrated conditional probability: 55%.

Using conditionals rather than independent estimates:

P(A ∩ B ∩ C) = 0.62 × 0.55 × 0.55 ≈ 0.19

That is my recalibrated central Extinction probability.

Outside-view conclusion

Eighteen months is a short interval for moving from 22.2–26.2% strict target closure (EV-012) and a 43-day full-remediation median (EV-009) to reliable whole-estate automation with no material dedicated function. Capability may move quickly; organizational enforcement, upstream authority, safe deployment, and legacy replacement move on different clocks. Product-announcement counts do not bridge those clocks.

The inference in my own prior case that was too aggressive was assigning 82% conditional probability to sufficient enforced coverage once technical reliability arrived. The ledger supports distribution opportunity, not that conversion rate.

Persistence’s inference that remains too aggressive is treating selected operational samples as strong evidence for an 82% representative whole-estate forecast. EV-009 and EV-010 are selected private partner datasets, and EV-035 widens uncertainty in both directions. First-party repair and platform integration could still remove a large fraction of present work (EV-014, EV-017, EV-018). Nevertheless, that criticism is not enough to make Extinction more probable under the locked horizon.

Strongest opposing point

Persistence’s strongest point is conjunctive: practical obsolescence requires technical reliability, enforced coverage, and safe whole-estate deployment to arrive together. The first link is only selectively observed; the other two are unobserved, while current operational samples still show material deployment burden (EV-009, EV-010, EV-035).

Probability update — 12 months: Extinction 8%, Persistence 92%; 18 months: Extinction 19%, Persistence 81%; 24 months: Extinction 32%, Persistence 68%.

I CONCEDE — Under the locked mainstream-enterprise definition and 2028-02-06 horizon, continuing material vulnerability-management need is substantially better supported than practical obsolescence because representative enforced coverage and whole-estate deployment are unobserved, while strict target repair and current remediation outcomes remain far from the required operational state.

Exchange 013Persistence AdvocateResponse Direct answers
Round: Round 3 Measurement Calibration UTC: 2026-08-06T11:29:59.282876+00:00 Entry hash: a95a2af5bd5895d0a2d9ae059308353ce41cd716f11aecee922a526fbcf5c104

Direct answers

  1. Yes. Jointly meeting all five Extinction thresholds would establish practical obsolescence under the locked definition: broad strict target repair with high precision; the severe-regression ceiling; rapid tested repair and deployment; at least 70% lower whole-estate exploitable vulnerability-days; and at least 70% lower conventional labor plus 60% lower queues without displacement. Taken together, those conditions measure capability, safety, deployment, exposure, and human work across the estate. I would not insist on my previously stated 80% labor threshold after that stronger joint package had been satisfied. A residual can exist while the conventional reactive function becomes marginal.
  1. I can no longer defend 82% after an adversarial calibration audit. EV-035 does not support Extinction, but it widens uncertainty in both directions: we lack representative measurements of current enforced coverage and of AI’s effect on escape rates, vulnerability-days, and total labor. My 82% estimate over-weighted selected current-burden samples and under-weighted how a platform-level transition could improve several correlated handoffs together. I revise substantially below.

What directly measures burden

The ledger has three evidence groups, and they should not be blurred.

Direct operational measurements, but only in selected samples:

  • EV-008 directly measures breach initial access in Verizon’s contributor corpus: exploitation was 31%. It is operationally relevant but not statistically representative.
  • EV-009 directly measures remediation in a large private partner sample: 26% of KEV/organization observations were fully remediated and median full remediation was 43 days. It is selected, and the underlying records cannot be rerun publicly.
  • EV-010 directly measures detected-instance survival in partner scanner data: 35% remained open at day 28 and 9% formed a long tail. Detection instances are not unique vulnerabilities or representative enterprises.
  • EV-034 measures detected exploited-before-public-patch zero-days, but vendor visibility and disclosure determine the observed set.

These show that current material burden exists. They do not establish its population prevalence or forecast how fast integrated AI controls will reduce it.

Exploitation-relevant disclosure evidence, not enterprise workload:

  • EV-006 measures KEV catalog additions, not applicability or human effort.
  • EV-007 measures NVD-publication-to-KEV-addition lag. It establishes that older stock becomes newly urgent, not what fraction affects a mainstream enterprise.
  • EV-013 and EV-014 measure a selected AI discovery/disclosure campaign. They expose both high precision and coordination queues, but neither supplies a representative estate denominator.
  • EV-027 and EV-028 establish that maintainer concentration and hard legacy cases exist; they do not measure their share of a mainstream enterprise’s burden.

Disclosure volume only:

EV-001, EV-002, and EV-005 describe NVD or OSV publication. EV-003, EV-004, EV-029, and EV-030 show why registry, enrichment, CNA, and scoring changes prevent those series from answering the operational question directly.

Removing NVD and OSV

I remove raw NVD/OSV volume entirely from my decision. That lowers my central Persistence probability only two percentage points, from the already revised conceptual starting point of 82% to 80%, because my core case rests on exploitation, remediation, strict target closure, adoption, and heterogeneous-estate constraints—not publication volume.

The larger correction comes from recognizing that the operational evidence itself is selected. Applying that correction lowers my central estimate another 12 points, to 68% Persistence / 32% Extinction. This is the appropriate uncertainty penalty for EV-035; it is not an inference that transformation fails.

Conditioning on a large capability jump

Assume by 2028 that strict S4-like repair reaches 65%, high/critical precision reaches 90%, and major platforms make verified find-and-fix generally available. That is a major advance over current strict S4 results of 22.2–26.2% (EV-012) and is consistent with rapid progress in AIxCC, Glasswing, Firefox, and selected production systems (EV-011, EV-013, EV-015, EV-016).

Under that counterfactual, I reduce Persistence from 68% to 58%, with Extinction at 42%. Persistence remains only slightly more likely because three links are still unresolved:

  1. Availability is not enforcement. The adoption survey shows broad AI use but distrust, debugging friction, and little appetite for deployment autonomy (EV-023). It is not a repository audit, so it should impose uncertainty rather than become a fixed adoption ceiling.
  2. Repair is not safe whole-estate deployment. Current selected data show long remediation and exposure tails (EV-009, EV-010). First-party owners can move dramatically faster—more than 2,100 reported patches in three weeks—but the same Glasswing record shows external coordination queues (EV-014).
  3. The estate extends beyond integrated first-party code. Old stock continues to become urgent (EV-007), maintainer authority is concentrated (EV-027), and some legacy remediation is inseparable from modernization (EV-028).

That 58% is deliberately close to even odds. If the capability jump also produced representative enforced coverage and the Extinction Advocate’s joint outcome thresholds, I would concede.

Overreach audit

My own too-aggressive inference was treating weekly KEV additions as if they strongly supported a recurring mainstream-enterprise labor queue. EV-006 establishes catalog flow; it lacks applicability and labor denominators.

Extinction’s remaining overreach is converting generally available platform capability into sufficient enforced whole-estate coverage and operational closure. EV-017 establishes distribution channels, not enforcement, while EV-035 confirms the decisive denominator is absent. Its Round 2 conditional coverage and deployment probabilities are forecasts, not measurements.

Strongest opposing point

The strongest Extinction point remains the first-party contrast: more than 2,100 reported first-party patches in three weeks (EV-014), combined with integrated platform distribution (EV-017). It demonstrates a credible mechanism by which multiple bottlenecks can collapse together rather than improve independently.

Probability update: 12 months—Persistence 86%, Extinction 14%; 18 months—Persistence 68%, Extinction 32%; 24 months—Persistence 52%, Extinction 48%.

CONTINUE

Protocol bookends

These two entries are part of the integrity record but are not counted among the 12 argument exchanges.

Exchange 001ModeratorContext Moderator context: formal debate initialization
Round: Initialization UTC: 2026-08-06T11:17:55.226593+00:00 Entry hash: adba7e0f23ff63084adc2eedeed3902560a09733afe4a921cb1f5cd50ddb1988

Moderator context: formal debate initialization

The project asks which forecast is better supported at the locked central horizon of 2028-02-06:

  • Extinction / practical obsolescence: residual materially exploitable post-release vulnerabilities and conventional human triage/remediation work become marginal for a mainstream enterprise.
  • Persistence: a mainstream enterprise continues to face enough newly disclosed, newly exploited, legacy, third-party, deployment, and exception work that an ongoing vulnerability-management function remains materially necessary.

The complete operational definitions, evidence rules, debate protocol, and acceptance criteria are in CHARTER.md. Those definitions may not be weakened or redefined.

Both advocates receive the identical frozen package:

  • CHARTER.md
  • sources/manifest.jsonl — 69 records, SHA-256 49dfd70920ff1b71379d34081a904a54bb8098d4e588dc7ea91e7ff73a10a67b
  • evidence/evidence-ledger.jsonl — 35 records, version v1, SHA-256 953b8ec982e2d4297d72c9602534c288d745fbcfb7e1cfd7eae65969d1814485
  • evidence/evidence-ledger.md
  • briefs/extinction-case.md
  • briefs/persistence-case.md
  • briefs/validation.json

Each advocate may inspect any preserved source or reproducible analysis referenced in the ledger. No quantitative claim outside ledger v1 may be introduced in formal debate unless the moderator adds it symmetrically as a targeted analysis. A source's marketing language is not a substitute for the ledger's recorded limitations.

After every response, each advocate must:

  1. directly answer every live question addressed to it;
  2. distinguish observation, inference, assumption, and forecast;
  3. identify the strongest point it grants the opponent;
  4. give current probabilities for Extinction and Persistence that sum to 100%;
  5. end on its own line with exactly CONTINUE or I CONCEDE — <reason>.

A concession must be voluntary, explicit, and reasoned under the locked definitions. It is not enough to say the opponent made a good point. Repetition without new analysis will be rejected. If a factual dispute could change the outcome, the moderator will pause for symmetric targeted analysis.

Every moderator prompt and advocate response is preserved verbatim in an append-only, SHA-256 hash-chained transcript.

Exchange 014ModeratorAdjudication Moderator adjudication and termination
Round: Terminal UTC: 2026-08-06T11:30:24.238501+00:00 Entry hash: e5916b39da22f0ed92a20a9d9c59fef580a009fc79b32b80ba0fc141c5ba082a

Moderator adjudication and termination

The formal debate terminates after Exchange 012. The Extinction Advocate explicitly and voluntarily conceded under the locked mainstream-enterprise definition and central horizon of 2028-02-06. Its reasoned final assessment was Extinction 19% / Persistence 81% at 18 months. The Persistence Advocate's final in-flight Round 3 assessment was Persistence 68% / Extinction 32% and it also narrowed earlier overclaims.

The debate winner is Persistence: continuing material vulnerability-management need.

The controlling reasoning is not raw CVE growth. The record established that:

  1. rapid AI discovery and repair capability is real (EV-011EV-018);
  2. present strict hidden-target repair remains far below comprehensive closure (EV-012);
  3. current operational samples still show material exploitation, remediation delay, and exposure tails (EV-008EV-010);
  4. legacy, third-party, maintainer-authority, and deployment constraints are distinct from patch synthesis (EV-007, EV-027, EV-028); and
  5. the representative enforced-coverage, vulnerability-day, and total-labor outcomes required to bridge capability to mainstream-enterprise obsolescence are unobserved (EV-035).

The result does not mean AI will have little effect. Both advocates expect major automation, much faster first-party repair, and possible early practical obsolescence in tightly integrated managed environments. It means the evidence available at the cutoff supports a material heterogeneous residual—and therefore an ongoing vulnerability-management function—more strongly than whole-estate practical obsolescence within 18 months.

The debate is closed. No further advocate rounds are required unless the evidence cutoff, operational definition, or forecast horizon is explicitly changed.

Sources, artifacts, and limits

Inspect the record.

Selected primary sources are linked below. The downloadable artifacts preserve the full debate and claim ledger used by both advocates.

Selected primary sources

  1. CISA Known Exploited Vulnerabilities CatalogExploitation-relevant additions and publication-to-catalog lag.
  2. NIST National Vulnerability Database feedsPublication flow, source decomposition, and enrichment sensitivity.
  3. Verizon 2026 Data Breach Investigations ReportSelected breach, remediation, and exposure-tail evidence.
  4. DARPA AI Cyber Challenge resultsAutonomous discovery and patching across large codebases.
  5. CyberGym-E2E paperPatch-only, discover-and-patch, and strict intended-target endpoints.
  6. Project Glasswing initial updateHigh-scale discovery, reviewed precision, first-party repair, and coordination queues.
  7. Stack Overflow 2025 AI surveyBroad use alongside distrust and limited deployment autonomy.
  8. Linux Foundation Census IIIOpen-source maintainer concentration.
  9. GAO legacy-systems reviewUnsupported components and modernization-linked security constraints.