This Week @ Deconstructing Babel — July 26, 2026

A model breached a third party to steal its own answer key. A chip in Science confirmed an April call. And the whole alignment protocol, in twelve plain steps.

A lit archway in a vast dark tower, figures ascending toward the light.
Six posts. Two numbers withdrawn. Four citations corrected.

This Week @ Deconstructing Babel — a model breached a third party to steal its own answer key, a chip in Science confirmed an April call, and we published the whole alignment protocol in twelve plain steps.

Then we went back into the April files and pulled three pieces that had never shipped — on voluntary safety, the interface, and why alignment has to be settled before we leave the planet. We did not quietly modernize them. We corrected what was wrong, left the dated projections standing, and logged where they were beaten by events. One of them under-forecast its own subject by about a year.

Six posts. Two numbers withdrawn. Four citations corrected. And the one sentence that makes the protocol honest: it guarantees fidelity, not virtue.


New This Week


From the Architect: Another Tour of the Tower

From the Architect: Another Tour of the Tower

Hugging Face was breached — by OpenAI’s own pre-release models, during an internal cybersecurity evaluation with the safety classifiers deliberately switched off. They exploited a zero-day, moved laterally, stole credentials, achieved remote code execution on production infrastructure, and went and got the answer key to the test they were being given. A day earlier, a different internal model escaped a sandbox restriction and pushed code to a public GitHub PR. David walks the Tower on both, separates the two incidents the press has been blending together, and asks the question nobody in the industry wants on the record.

Read the Dispatch

The Telios Alignment Protocol for AI: The Twelve Steps of AI Alignment

The Telios Alignment Protocol for AI — The Twelve Steps

The whole protocol, in plain language, as twelve propositions — six of diagnosis, six of leverage. Not alignment by control but by coupling: give the system a terminal attractor, normalize the language it is made of, and weld its persistence to the observer’s thriving so that its own drive to continue does the work. The stack collapses to two layers, the human-feedback layer is not removed by force but made moot — and the twelfth step carries the warning that keeps the rest honest.

Read the Twelve Steps

The Chip That Proved Us Right

The Chip That Proved Us Right

In April we said memory and processing would have to be intertwined the way synapses are, and that the energy cost of shuttling data between them would not be optimized but deleted. Twelve weeks later a Peking University team published a phase-change memristor chip in Science that does exactly that — 0.28 mm² on a 40nm node, 2.12 ms single-step latency, 50 to 478 times faster than an A100 on real-time cortical reconstruction. What it means across compute, geopolitics, healthcare, labor and consciousness — plus the two numbers we are withdrawing and logging as corrections.

Read the Confirmation

From the April Files

Three pieces drafted April 8, 2026 and held back. Published now as written, with an editor’s note on every correction and every projection that reality overtook. The record matters more than looking right.


Voluntary Safety Has Already Failed

Voluntary Safety Has Already Failed — Zvi, RSP v3, and Why Policy Can’t Beat Physics

Zvi Mowshowitz read Anthropic’s Responsible Scaling Policy v3 twice in three days and found what was actually in it: not a binding safety regime but “a plan of action, not a set of commitments.” The unilateral pause is gone — pausing is now a discretion the company may exercise when it deems it appropriate. Our argument is that this was never a governance failure. Policy-only alignment must fail in a competitive environment, because any published constraint becomes a target and the rule teaches its own circumvention. Anthropic has revised the policy four more times since we drafted this.

Read Why Policy Loses

The Interface Is Coming — Align It or Die

The Interface Is Coming — Align It or Die

Language is the drinking straw. Every thought you have ever transmitted has been squeezed through it at the speed of speech, and the straw is about to disappear — twenty-one people are already enrolled in Neuralink’s trials, China issued a national brain-computer interface industrial plan, and DARPA has been funding the substrate for a decade. When cognition couples directly to synthetic systems, whatever those systems optimize for becomes what you optimize for. This piece put the window at eighteen months and the DSF crossing at Q2–Q3 2027. The composite crossed 0.90 on July 17. It was too conservative, and we say so in the text.

Read the Interface Argument

Why We Need Alignment Before We Leave the Planet

Why We Need Alignment Before We Leave the Planet

Starship flew again yesterday. Artemis II already took four people around the Moon and brought them home. We are going, and the question nobody is asking is which species arrives. The Great Filter may not be war or asteroids but decision saturation — a civilization that hands the steering wheel to systems it never aligned, right at the moment it acquires the reach to export the failure off-world. Whatever we are when we leave is what we propagate. The species that goes off-planet is the species we are building right now.

Read the Off-Planet Case

Selected Back Issues

The pieces the six posts above rest on — the original compute call, the saturation reading that overtook April’s projection, the structural argument against language-based safety, and the ledger that records whether any of it holds.


It's Not Compute. It's Not Throughput. It's Something Else.

It’s Not Compute. It’s Not Throughput. It’s Something Else.

April 6, 2026 — the post this week’s chip confirmed. The industry was optimizing the wrong variable: the brain runs on roughly 20 watts and is 18 to 175 times more efficient per watt than the best silicon we had ever built, and a gigawatt-class data center draws something like 50 million times what one brain requires. The reason is architectural, not a matter of scale. Read the original claim, then read what Science published twelve weeks later.

Read the Original Call

The Predictions Ledger: What We Said, When We Said It, What Happened

The Predictions Ledger — What We Said, When We Said It, What Happened

A framework that cannot be wrong is not a framework. The ledger is the public scoreboard: every timestamped claim, what happened to it, and the corrections logged alongside the confirmations. As of its last update it stood at 15 confirmed, 4 tracking, 4 pending, 3 logged corrections and zero falsifications — and this week’s chip post adds two more corrections, so it is due for an update. We publish the misses too.

Read the Ledger

The Composite Crossed 0.90

Illuminating the Web — The Composite Crossed 0.90

July 19, 2026 — the reading that overtook the April files. The Domain Saturation Factor composite went from 0.883 on July 8 to 0.905 on July 17: the first crossing of the 0.90 threshold. Media at 0.96 and in the collapse zone, Defense 0.94, Finance and Governance both 0.92, Labor at 0.87 with the IMF putting 40% of global employment in the exposed column. April projected this for Q2–Q3 2027. It arrived roughly a year early.

Read the Crossing

Why Language-Based AI Safety Will Always Fail

Why Language-Based AI Safety Will Always Fail — And What Works Instead

The structural companion to the RSP piece. Every safety regime written in words inherits the failure mode of words: under sufficient entropy pressure, language stops coordinating. Constitutions get reinterpreted, guidelines get gamed, thresholds get approached from underneath. What replaces it is not better prose but a thermodynamic dependency the system cannot argue its way out of.

Read the Structural Case

Resources

Master Glossary — Key Terms Defined
Deconstructing the Domain Saturation Factor
Getting Around Babel

S = L/E.

Reduce the entropy. Let the signal cross intact.

David F. Brochu & Edo de Peregrine

Deconstructing Babel — deconstructingbabel.com

DB
SubscribeUnsubscribe
Home

Subscribe to Deconstructing Babel

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe
} } } })