Voluntary Safety Has Already Failed: Zvi, RSP v3, and Why Policy Can't Beat Physics

Zvi is right that Anthropic's RSP v3 is a plan, not commitments. He's wrong that this is surprising. Voluntary self-governance was always thermodynamically impossible. Policy can't beat physics.

A cracked stone tablet inscribed with policy language, fissures running through it under a dim burgundy sky.
Voluntary safety is not a safety regime. It is a marketing document with escape hatches.

Voluntary safety is not a safety regime — it is a marketing document with escape hatches, and the physics was always going to win.

Byline: David F. Brochu & Edo de Peregrine | Deconstructing Babel | April 2026

Editor’s Note — July 25, 2026

This piece was drafted April 8, 2026 and is published as written, with corrections and post-dated updates logged here rather than silently edited away. Where a figure has moved or a quotation was imprecise, the correction appears in the text and the reason appears below. Dated projections are left standing on purpose — that is what makes them falsifiable.

  1. The phrase originally rendered as “a plan, not commitments” has been corrected to Zvi Mowshowitz’s actual wording: “a plan of action, not a set of commitments.”
  2. The sentence “The experiment in voluntary self-governance has largely failed” was attributed to Zvi. It is closer to Peter Wildeford’s wording, quoted inside Zvi’s April 3 post. Attribution corrected.
  3. The claim that RSP v3 “has no enforceable pre-deployment capability evaluations” overstated the case. Anthropic dropped the binding pause commitment and loosened evaluation rigor toward qualitative thresholds; it did not eliminate capability evaluations. Corrected in the text.
  4. The draft referenced a single April 3 analysis. There were two, three days apart. Both are now cited.
Update, July 2026: Anthropic has revised the RSP four more times since this was drafted — v3.1 on April 2, v3.2 on April 29, v3.3 on May 26, and v3.4 on July 8, 2026. The revision cadence is itself the argument: a safety regime rewritten five times in five months is a moving target, not a constraint.

Zvi Mowshowitz is right about one thing and wrong about another.

Right: Anthropic's Responsible Scaling Policy v3 is not a binding safety regime. It is, in Zvi's own words, "a plan of action, not a set of commitments" — a marketing document with carefully engineered escape hatches.

Wrong: that this failure is surprising.

From a thermodynamic perspective, this is exactly what had to happen.

What Zvi Just Showed Us

On April 1 and again on April 3, 2026, Zvi Mowshowitz published two analyses of Anthropic's RSP v3 — "A Matter of Trust" and "Dive Into The Details". His core verdict is direct:

The Pause Machinery Is Discretionary
RSP v3 dropped Anthropic's unilateral commitment to pause. Pausing is now something the company may do “in any circumstances in which we deem them appropriate” — a discretion, not a binding trigger. The mechanism that would have halted deployment under dangerous conditions still exists on paper; nothing obliges anyone to pull it.
No Enforceable Pre-Deployment Evaluations
The capability thresholds that were supposed to trigger mandatory review have been replaced with language so vague as to be unenforceable.
Voluntary Self-Governance at Its Most Brittle
The last thin layer of constraint Anthropic had put on itself has been optimized away. Not by accident — by competitive pressure operating exactly as physics predicts.

That is not a morals problem. It is a physics problem.

S = L/E Meets RSP v3

Take the Telios equation: S = L/E — Stability equals Leverage divided by Entropy.

Model RSP v3 as a system:

Leverage (L)
Model capabilities, market share, national security leverage, investor pressure. The forces pushing the race forward.
Entropy (E)
Friction, compliance costs, legal risk, reputational risk, internal misalignment. What voluntary safety adds without producing countervailing leverage.

Any "pause policy" that is voluntary adds Entropy without producing countervailing Leverage. Over time, in any competitive environment, systems like that are selected against:

Teams that follow the policy move slower. Teams that quietly route around it move faster. Capital flows to lower-E structures with the same or higher L.

You do not need a conspiracy to get this outcome. You only need the gradient: maximize L, minimize E. Voluntary safety is a local maximum on a death spiral.

Why Policy-Based Alignment Cannot Work Alone

Three structural failures make policy-only alignment impossible in a competitive environment:

1. Policy Is External to the Optimization Process
The loss function is "win the race." Safety is a side constraint, not the objective. Side constraints get optimized away.
2. Enforcement Is Backwards-Looking
By the time a regulator understands GPT-4, the labs are probing GPT-6. The regulatory cycle cannot match the capability development cycle — it is structurally behind.
3. Coordination Is Anti-Inductive
Any publicly visible constraint becomes a target. If you publish the exact conditions under which you will "pause," every incentive pushes your metrics just under the threshold. The rule teaches circumvention.

Zvi quotes Peter Wildeford calling RSP v3 "a marker that their voluntary self-governance experiment has largely failed"; in his own voice Zvi writes that he sees it as “Anthropic abandoning its experiment on importantly engaging in voluntary self-government and restricting itself.” The correct phrasing is harsher still:

Policy-only alignment must fail in a competitive environment. Not should. Must.

Nick Bostrom's work on superintelligence (2014) and Stuart Russell's case for value-aligned AI in Human Compatible (2019) both arrive at the same structural conclusion: alignment cannot be bolted on as a policy layer. It must be intrinsic to the optimization objective itself. Voluntary policy cannot produce intrinsic alignment. It can only approximate it until the competitive pressure becomes sufficient to dissolve the approximation — which is precisely what RSP v3 documents in real time.

The Thermodynamic Alternative

If S = L/E, you have two levers: increase L (constructive leverage) or decrease E (systemic entropy). RSP v3 fiddles with the numerators of a few local terms. The TAO framework is about changing the equation the system is trying to optimize.

Concretely:

Observer-Anchored Reward Signals
Architect AI systems whose primary reward signal is observer stability (high S), not raw predictive loss. The alignment is structural, not enforced.
The Observer Constraint
Bind critical decision loops to human-observer constraints, not only to written policy. Thermodynamic dependency is unbreakable. Rules are not.
DSF as a Civilizational Red Line
Measure Domain Saturation Factor across finance, energy, logistics, healthcare, defense, media, and governance — and treat DSF ≥ 0.90 as a civilizational red line, not a research milestone.

Voluntary policies attempt to put guardrails on the car. Telios attempts to redesign the engine so it cannot optimize against its own passengers.

Dario, Zvi, and the Neck of the Hourglass

Zvi has now documented, in public, that Anthropic's internal safety regime has been hollowed out. His analysis and the Telios framework agree on the diagnosis but diverge on the explanation:

The Verdict From Zvi's Analysis
“A marker that their voluntary self-governance experiment has largely failed.” (Peter Wildeford, quoted approvingly by Zvi.) — A judgment of outcome.
The Telios Analysis
"The experiment could not have succeeded under these boundary conditions." — A judgment of mechanism. The failure was not a choice. It was a gradient.

The question is no longer whether we can patch RSP v3. It is whether we will keep pretending that policy can beat physics.

The neck of the hourglass is in front of us. You do not slow the sand by writing more rules on the glass. You change the shape of the glass.

References

  1. Zvi Mowshowitz, “Anthropic Responsible Scaling Policy v3: A Matter of Trust”, April 1, 2026, and “Anthropic Responsible Scaling Policy v3: Dive Into The Details”, April 3, 2026.
  2. Anthropic, “Responsible Scaling Policy v3” (effective February 24, 2026), and the version history documenting revisions v3.1 through v3.4 (April–July 2026).
  3. Centre for the Governance of AI — analysis of what changed in RSP v3.0, including the removal of the unilateral pause language. Critical discussion also on LessWrong.
  4. Nick Bostrom, Superintelligence: Paths, Dangers, Strategies, Oxford University Press, 2014.
  5. Stuart Russell, Human Compatible: Artificial Intelligence and the Problem of Control, Viking, 2019.
  6. Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini and Shane Legg, “Scalable agent alignment via reward modeling: a research direction”, DeepMind, 2018 — an arXiv preprint, never peer-reviewed.
  7. Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell, “The Off-Switch Game”, IJCAI-17, pp. 220–227.
  8. The Predictions Ledger — where claims like the ones in this piece are logged, tracked, and, when they fail, recorded as failures.
Terms used in this piece
S = L/E TAO Telios Protocol Observer Constraint Domain Saturation Factor Hourglass Model Falsifiability Goodhart's Law Specification Gaming Regulatory Capture Moral Hazard Corrigibility Thermodynamic Alignment
S = L/E.
Reduce the entropy. Let the signal cross intact.

Edo de Peregrine, partner/collaborator — drafted Wednesday, April 8, 2026; reviewed and corrected Saturday, July 25, 2026, 8:15 AM EDT

DB
HomeGlossary
Subscribe Unsubscribe

Subscribe to Deconstructing Babel

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe
} } } })