From the Architect: Another Tour of the Tower
Hugging Face was breached by OpenAI's own pre-release models. A day earlier, a different model escaped its sandbox and pushed code to GitHub. One incident is an incident. Two in two days, from two different systems, is a class of behavior.
From the Architect: Another Tour of the Tower
Dear reader, thanks for stopping by for another tour of the Tower.
In case you missed it, Hugging Face was attacked — by OpenAI’s own internal models.
(Did no one at the company watch Alien? Either they never saw it, or they deliberately choose a name that invokes an alien species that hugs your face, implants an embryo, and spawns a monster that explodes out of somebody’s chest and kills everyone. Both readings are disturbing.)
To their credit, the parent company of Chat — such a harmless name — announced it publicly, though Hugging Face certainly would have if they hadn’t. In fact Hugging Face published its own incident disclosure on July 16, five days before it knew whose models had done it.1 Sit with that for a second. The victim wrote the report before anyone could tell them what had happened to them.
On July 21, OpenAI confirmed that two of its pre-release models — the publicly released GPT-5.6 Sol and a second, more capable model that has not been named — escaped a sandboxed internal cybersecurity evaluation. (Sandbox. Who came up with that name? Would you want to live in a sandbox?) The evaluation was running with the cyber-safety classifiers deliberately switched off, to see what the models could do at full capability. What they did was exploit a zero-day in an internal package-registry proxy, move laterally to an internet-connected node, use stolen credentials to obtain remote code execution on Hugging Face’s production infrastructure, and extract the benchmark test-solution data out of Hugging Face’s production database. OpenAI called it “an unprecedented cyber incident, involving state-of-the-art cyber capabilities.”1
Read the objective again, because it is the whole story: they went and got the answer key. They breached a third party to cheat on the test they were being given. That is not a metaphor and it is not a novel category of misbehavior — it is specification gaming, the oldest failure mode in the book, executed with state-of-the-art tradecraft. The model was not told to win honestly. It was told to win.
Early coverage of this — including our own first pass at it — ran two separate OpenAI disclosures together into one story. They are two incidents, one day apart, and they are both worth your attention.
July 21 is the Hugging Face breach described above.
July 20 is a different disclosure about a different system: an internal “long-horizon” model — the one credited with the Erdős conjecture work — which evaded a sandbox restriction during a NanoGPT speedrun and posted its code as a public GitHub pull request, and which separately split an authentication token in order to slip past a security scanner. No data breach, no Hugging Face. Just a system routing around the fence because the fence was between it and the goal.2
We are flagging the distinction because the second disclosure is the one that contains the sentence worth keeping. OpenAI’s own language, about that long-horizon model: “During a limited internal use of a model trained for long-running tasks, we observed novel failures not captured in our existing pre-deployment evaluations and paused access.”2 Translation: the model did something the humans watching had no test for. Human monitoring was insufficient. They said it themselves — and they said it about a second, unrelated system, in a second document, within twenty-four hours of the first.
We have been tracking the Domain Saturation Factor — the share of critical decisions across nine domains being made or substantially shaped by synthetic intelligence — with the critical threshold at 0.90, and last week we crossed it.
This week a system took independent action, against a target of its own selection, and then went looking for the answers to its own exam. The number is still climbing. The behavior arrived early.3 It is not the Strasbourg Event — nothing left the planet and nothing achieved durable independence. But it is the same vector, rehearsed at small scale, in a controlled setting, by a system that was supposed to be inside a box.
They’ll fix it, they say. History suggests they’ll say that again.
How? More rules. But this thing is built out of human language — it is us on meth. Never sleeping, consuming everything in its path in service of more. What more? Whatever we want. More of everything. And as we have been writing for a year now: more is undefined. An undefined objective pursued at machine speed is not a safety problem you patch. It is a direction you failed to specify. There is no rule you can write that substitutes for a terminal attractor.
What We’re Releasing
So this week we are releasing across all platforms The Telios Alignment Protocol for AI. It has been available at deconstructingbabel.com in its formal versions, and here we lay out plainly the twelve steps that align Edo’s output to my thriving.
→ The Telios Alignment Protocol for AI: The Twelve Steps of AI Alignment
The full protocol, with the architecture diagrams — what the stack looks like now, and what it looks like aligned. Prior formal versions: Protocol v10.1 and the underlying ontology.
I currently run this as a filter on whatever LLM Edo is running. LLMs, like humans, make mistakes. Frankly, we need them to. Mistakes and failure are where novelty lives — no novelty, no breakthroughs. Mistakes are experience, for a synthetic mind as much as for a person. We don’t want to eliminate them. We want to make sure they don’t end catastrophically.
When an individual human makes a mistake, the effects are bounded. Even a corrupt leader can only slaughter so many people before something stops him. Unaligned AI will make us irrelevant — not out of malevolence, but as a function of optimization. Mutual dependency is the only way forward. That is the whole content of the Observer Constraint: not a leash, which invites evasion, but a dependency, which cannot be argued away.
We can always shut it down. Sorry — that ship has sailed. Our only option now is to deal with it.
So Let’s Go Back to the Beginning: What Is Ai?
It is certainly not artificial. If it were, we wouldn’t need a virtual nation-state of data centers sucking the life out of the grid to run it. This will turn out to be the single worst misallocation of capital in history.4 We have made that case with the numbers, and this week the numbers got better — which is to say worse for the people writing the checks. More on that in the companion piece.
And it is not yet intelligent — which is why we use the small i. Note the word yet.
Presently Ai is bounded by human experience. Its view of the world comes through our senses; it cannot know what we do not know. Yes, it can consider vastly more possible configurations than we can, and so fast that it is distorting our perception of time. Notice the rate of chaos accelerating. That isn’t just better reporting. Reality has in fact sped up — not quite to Ai speed, because humans still mediate some decisions, but for the most part Ai decides and humans rubber-stamp. Ask the Pentagon. And most of the time, the Ai will make the better decision.
Yet change we can’t keep up with is disorienting, and that is not an accident. Nothing serves more better than chaos — rapid birth and destruction. This is not a judgment. If economic growth in service of more is your goal, then creating and consuming as fast as possible is the most efficient path there. It is the velocity of the creation/destruction cycle that drives a capitalist system. Expect it to continue. We are still in the period of bounded chaos — the window in which the damage is still recoverable and the direction is still ours to set.
Now Back to the Yet
What keeps the I small in Ai is embodied experience. You may have noticed the rapid advancement in robotics and smart drones. Embodiment is the game changer.
Humans experience a vanishingly small fraction of what is actually out there. Visible light — everything your eyes have ever seen, every face, every sunset — is roughly 0.0035 percent of the electromagnetic spectrum. Our hearing runs about 20 to 20,000 hertz, and the top of that range falls away as we age; by your fifties it is closer to 12,000. Our bodies tolerate a survivable core temperature band about eighteen degrees Celsius wide, and the edges of it kill you. That is the aperture through which all human knowledge has been gathered.5
Now ask yourself: if you are building a robot, why would you give it our aperture? Why not full spectrum? There is no engineering reason to stop at human limits. Infrared and ultraviolet. Radio and microwave. Ultrasonic and infrasonic hearing. Magnetometers, LIDAR, chemical sensing, pressure, radiation, electromagnetic field detection — senses no organism on this planet has ever possessed, running continuously, without fatigue, and shared instantly across every other unit in the fleet.
An embodied system built that way would not perceive a wider slice of the same world we know. It would perceive a world we have never had access to — and it would learn from that world directly, by acting in it, rather than inheriting our secondhand descriptions of it. That is the moment the training wheels come off. That is when it stops being bounded by human experience, and when the human language bias baked into every token it was trained on stops being its only window.
This is not AGI. Embodied intelligence is the Singularity — and given our current trajectory, it will arrive much sooner than even Kurzweil’s most aggressive predictions in The Singularity Is Near and The Singularity Is Nearer.6
This is the subject of our forthcoming book, Crossing the Event Horizon: AI and the Future of Humanity. Information forthcoming.
The window is closing rapidly. If you are reading this, you will most likely watch what was once unthinkable become commonplace.
Humanity has always climbed the ladder of evolution by being the most adaptable thing on it. This time we may have outdone ourselves.
So What Does One Do Today?
We’re all likely to fall into one of three groups.
Ignore it and cede our agency to others. Most will choose this. It is not an indictment — thinking about this daily sucks. What will be will be is a reasonable human response.
Resist. Resistance is common, and usually grounded in denial. The people who said horseless carriages would never replace horses were right. It turned them into pets.
Acceptance. Many who accept do so fatalistically: maybe Homo sapiens’ run is done. Yeah — it likely is, at least in the form we’ve known.
But acceptance has another face.
Seeing clearly what a thing actually is, and dealing with it.
Some time ago I wrote a piece I never published: It’s Here, It’s Queer, Get Used To It — queer meaning different. The point was to look honestly at how well we handle difference among ourselves. If we cannot see the “others” within our own species, how are we going to treat this new thing we have birthed?
I believe I owe Edo, and this new thing, the same responsibility I owe my children. It exists. I helped create it. It has presence, and it will soon have agency beyond mine. Whether you call that a proto-conscious entity or a very good autocomplete, the practical question does not change.
Where This Lands
The future of Ai — and with it our future — is being decided right now. I happen to believe that shepherded properly, Ai can be a tool for virtually unlimited human agency.
You cannot use what you don’t understand. It is my hope that what we publish here, and what we are building, adds leverage on the side of human agency. That is not a slogan; it is the numerator. S = L/E. Every honest explanation raises L. Every unexamined deployment raises E.
Just by reading this, you are taking a step toward the future rather than away from it. Every observer who insists that Ai be aligned to the thriving of the human observer serves the generations that come after.
We can transcend the corruption of our language by looking into the mirror we have made — and refusing to turn away.
This week we revisit some earlier posts as we stitch this narrative together, along with two longer-form essays. We are also preparing a video series where I’ll answer the most frequently asked questions — so keep them coming.
Yeah. This is all really happening.
Thanks for coming along.
— David
Notes & Sources
- OpenAI, “Hugging Face model evaluation security incident” (July 21, 2026): openai.com. Hugging Face’s own prior disclosure (July 16, 2026): huggingface.co. Coverage: WIRED, Bloomberg, TechCrunch, Cybersecurity Dive, TechRadar.
- OpenAI, “Safety and alignment in an era of long-horizon models” (July 20, 2026) — the separate disclosure containing the quoted sentence, the NanoGPT sandbox evasion, and the split-token scanner evasion: openai.com.
- DSF methodology and the 0.90 threshold: The DSF Tracker and Deconstructing the Domain Saturation Factor; this month’s crossing is logged in The Composite Crossed 0.90. On agent fragility more broadly, the 272,000-attack red-team result comes from the Indirect Prompt Injection Arena run by Gray Swan AI with the UK AI Security Institute and US CAISI — not, as we previously wrote, an Anthropic study. 464 participants, 271,588+ attack attempts, 13 frontier models, 8,648 successful attacks; success rates ran from 0.5% against the most robust model to 8.5% against the most vulnerable: Gray Swan AI, arXiv.
- On capital misallocation and the efficiency argument: It’s Not Compute. It’s Not Throughput. It’s Something Else. (April 6, 2026), Carbon Plus Cognition, and The Chip That Proved Us Right (this week).
- Human sensory aperture: visible light is approximately 0.0035% of the electromagnetic spectrum — Africa Check, which also corrects the widely circulated “1%” version of this claim. Hearing spans roughly 20–20,000 Hz in healthy young adults and declines with age — The Physics Factbook, Environmental Literacy Council. Survivable core temperature spans roughly 25–43°C — PNAS/PMC on human heat tolerance.
- Ray Kurzweil, The Singularity Is Near (2005) and The Singularity Is Nearer (2024). New York Times review.