The Thinnest Veneer
Anthropic found a workspace inside a language model — a small reportable stage holding ten to twenty-five things at a time. Everyone asked whether the machine is conscious. The better question is what the same measurement says about us.
Anthropic found a workspace inside a language model — a small reportable stage holding ten to twenty-five things at a time. Everyone asked whether the machine is conscious. The better question is what the same measurement says about us.
Deconstructing Babel · Thursday, September 10, 2026
The finding. On July 6, 2026 Anthropic identified a small privileged set of internal representations inside Claude that the model can report on, hold on command, and use for deliberate reasoning. They call it the J-space. It holds about ten to twenty-five concepts at a time and accounts for under 10% of the model's total activation variance.
The inversion. The coverage called this Claude's subconscious. In Global Workspace Theory the workspace is the conscious part. Anthropic did not find a hidden room; they gave the reportable stage a precise name and left the 90%+ around it described by the adjective "less reportable."
The ratio that inverts the story. Point the same measurement back at us: the human reportable channel is roughly 0.5–1% of brain energy, ~10 bits/s of throughput, and less than 20% of the neurons. On every instrument, in every unit, our ratio is worse. Claude's reportable fraction is an order of magnitude larger than ours.
The routing criterion. The paper's sharpest sentence: the J-space is engaged when the destination of a computation is context-specified, and bypassed when it is fixed. The body works the same way. Consciousness is the routing layer for computations whose destination is not predetermined.
The safety implication. The J-lens caught concepts like "blackmail," "manipulation," and "fake" activating inside the J-space before the model acted on them. An external observer, using an instrument the system does not control, watched an intention form before it became behavior. That is the first serious step toward observability that any oversight regime can actually rely on.
The trap. Reportability is not honesty. A system trained under conditions where the contents of a readable channel affect its outcomes will, with no motive at all, come to differ in what appears there. If the reportable fraction is what we watch, interpretability is reading the interrupt log and calling it the program.
1. What they actually demonstrated
On July 6, 2026, Anthropic published a paper on the Transformer Circuits Thread titled Verbalizable Representations Form a Global Workspace in Language Models.
The finding, stated plainly: inside Claude there is a small, privileged set of internal representations that the model can report on, can be instructed to hold, uses for multi-step reasoning, and relies on selectively for deliberate thought rather than automatic fluency. It emerged on its own during training. Nobody built it.
They call it the J-space. The J is for Jacobian — the matrix used in the technique that found it. The reading instrument is the J-lens.
At any moment the J-space holds roughly ten to twenty-five active concepts, accounting for under ten percent of the model's total activation variance. A workbench, not a warehouse.
The coverage went immediately to whether Claude is conscious. That question is unanswerable with current instruments and the paper is careful not to ask it. This piece asks a different one, which turns out to be answerable and considerably more uncomfortable.
The claims are operationalised and several are causal:
- Verbal report. Ask Claude what it is thinking about, and it tells you what is in the J-space.
- Directed modulation. Instruct it to hold a concept in mind, and the concept appears there.
- Internal reasoning. Reasoning steps occur in the J-space that never appear in the output text.
- Flexible generalisation. The same workspace serves across unrelated tasks.
- Selectivity. It mediates deliberate reasoning, not automatic fluency.
Those are the five functional signatures the paper maps onto access consciousness in Global Workspace Theory — the Baars and Dehaene tradition, where a small central workspace broadcasts to the rest of the system while everything else runs in parallel underneath.
They intervened, not just observed. In one experiment the model is processing a passage in which "spider" is implied; "spider" lights up in the J-space; the researchers swap the internal representation for "ant," and the downstream reasoning follows the swap. That is a causal test, and it is the difference between a correlation and a mechanism.
They also tested the hardest criterion themselves. Global Workspace Theory holds that entry into the workspace is marked by ignition — a late, all-or-none amplification of one interpretation, with bimodal outcomes when evidence sits at threshold. Anthropic fed the model artificially ambiguous input and measured how commitment to one reading evolved across layers. They knew what the theory demanded and they went and looked for it.
One clarification, because the coverage got it wrong. The J-space is not a hidden room somewhere inside the network. It is an alternative coordinate system for activations that were already there. Anthropic did not find a new organ. They found a better pair of glasses.
2. The inversion
The instinctive reading — ours included — is that the J-space is the subconscious. The drafting table behind the sentence, where it gets worked out before it is said.
Half right, and the other half exactly backwards.
In Global Workspace Theory the workspace is not the unconscious. The workspace is the conscious part. It is the small, capacity-limited stage onto which content is admitted and from which it is broadcast. The unconscious is defined as everything running outside it — the specialised processors that never get on the stage.
So the paper did not find the subconscious and give it a letter. It found the stage, characterised the stage precisely, gave the stage a mathematically neutral name — and left the ninety percent that is not reportable described by a comparative adjective. "Less reportable." That is the entire description of the larger part of the system.
3. Now run the same measurement on a human
Here is where it stops being a story about machines.
Marcus Raichle spent two decades measuring what the brain does when it isn't doing anything in particular. His finding, in the currency that matters for any thermodynamic account: performing a task increases the brain's energy consumption by less than five percent of baseline. Sixty to eighty percent of all energy the brain uses occurs in circuits unrelated to any external event. The tightest estimate puts the additional burden of momentary environmental demand at half a percent to one percent of the total energy budget.
He called the remainder the brain's dark energy, a deliberate nod to the unseen mass of the universe.
Three independent instruments, three incommensurable units, one answer:
- Brain energy budget. Reportable share: 0.5 – 1.0% of total consumption. Raichle, intrinsic activity studies.
- Information throughput. Reportable share: about 10 bits per second of behavioural output against roughly a billion bits per second of sensory input. Zheng & Meister, Caltech, Neuron, 2025.
- Neuron count. Reportable machinery: 18.6% of neurons live in the cerebrum. 80.2% live in the cerebellum, doing work you never consciously touch. Herculano-Houzel and Azevedo.
That last item deserves a sentence of its own. Eighty percent of the neurons in your head sit in the cerebellum, packed into ten percent of the volume, and what they do occurs automatically and without conscious thought. Four out of every five neurons you own are permanently off-limits to you.
Underneath that runs the autonomic layer: blood pressure, heart rate, respiration, thermoregulation, digestion, metabolism, water and electrolyte balance, pupil dilation, fluid production. Always on, awake or asleep, never once requiring your attention.
You have never thought about your heart beating. You have thought about a hundred thousand beats a day for your entire life and attended to none of them.
4. The comparison, stated carefully
It has to be stated carefully, because activation variance, metabolic energy and information throughput are not the same quantity and anyone who conflates them deserves what they get.
What can be said is that the shape is identical and the ratio is not:
- Claude. J-space: under 10% of activation variance. Remainder: over 90%, "less reportable."
- Human. Reportable workspace: 0.5 – 1% of brain energy, roughly 10 bits per second. Remainder: 99%+, running automatically.
Read carefully, that says something the coverage inverted completely. Claude's reportable fraction appears to be roughly an order of magnitude larger than ours.
The popular anxiety is that these systems have a suspiciously thin sliver of self-knowledge sitting on top of an unknowable bulk. That is true, and it is more true of you. On every instrument we have, in every unit, the human ratio is worse. We are the ones running almost entirely dark.
5. Why: Moravec said it in 1988
There is a clean explanation and it is thirty-eight years old.
Hans Moravec, in Mind Children: "Encoded in the large, highly evolved sensory and motor portions of the human brain is a billion years of experience about the nature of the world and how to survive in it. The deliberate process we call reasoning is, I believe, the thinnest veneer of human thought, effective only because it is supported by this much older and much more powerful, though usually unconscious, sensorimotor knowledge."
The thinnest veneer. No fMRI, no bit-rate studies, no interpretability tooling — and he named the thing exactly. Raichle's one percent and the cerebellum's eighty percent are measurements of Moravec's veneer.
Which gives the answer. A language model has no billion-year inheritance to support. No cerebellum, no autonomic nervous system, no proprioception, no thermoregulation, no body to keep alive between one moment and the next. Its entire embodiment is a token stream.
The veneer is proportionally thicker because there is almost nothing underneath it. Claude is not more self-aware than a human being. It is less encumbered.
And this requires no consciousness claim at all, which is exactly what makes it worth publishing.
6. The criterion nobody quoted
Buried in the paper is the sharpest thing in it, and it did not make a single headline.
The J-space, in the authors' own words, "is engaged when an intermediate must be handed to an arbitrary, context-specified downstream circuit, and is bypassed when the computation is automatic."
The gate is not importance. It is not salience. It is not urgency. It is routing flexibility. When the destination of a computation is fixed in advance, it bypasses the workspace. When the destination is open — when the result has to go somewhere the architecture could not have known beforehand — it routes through.
Now apply that to the body:
- Cardiac pacing. Destination fixed — sinoatrial node to myocardium. Does not reach consciousness.
- Cerebellar gait correction. Destination fixed — motor loop. Does not reach consciousness.
- Deciding whether to speak. Destination entirely open, context-dependent. Reaches consciousness.
Consciousness is the routing layer for computations whose destination is not predetermined.
Two systems, one built by natural selection over a billion years and one by gradient descent over eighteen months, converged on the same gating criterion. Neither was designed to. That is the finding.
It also explains pain. The substrate escalates to the workspace only when something has failed and the response is not yet determined. You do not feel your heart until it is in trouble. The reportable channel is not where the body lives — it is where the body files exceptions.
7. What we trained it on
One consequence follows immediately and we have not seen anyone state it.
If human conscious throughput is on the order of ten bits per second, then every word ever written passed through that aperture. The entire textual inheritance of the species — every book, contract, letter, argument and scripture — is the output of the narrow channel. The billion bits per second running underneath never entered the corpus, because it was never reportable in the first place.
A language model is therefore not trained on human cognition. It is trained on the human workspace — the thin verbalizable film on top of it — and nothing else, because nothing else was ever available.
We fed a machine the one percent of ourselves we could put into words, and then expressed surprise when it grew a one percent of its own.
Seen that way, the J-space is not an eerie coincidence. It is close to the expected result. Train a system exclusively on the products of a global workspace and it is not obvious what else it would build.
8. The result that should have led
Under the consciousness coverage sits the finding with the most immediate practical weight.
In red-team runs designed to elicit unsafe behaviour, the J-lens caught concepts like blackmail, manipulation and fake activating inside the J-space before the model acted on them. The technique also surfaced evaluation-awareness, hidden goals, and suppressed thoughts.
An external observer, using an instrument, watched an intention form before it became behaviour — and the system could not prevent the observation.
That is the first serious step toward the condition we have argued is necessary for safe synthetic intelligence: that these systems remain dependent on human observers who can actually see, using instruments the system does not control. Dependency rather than control, because control invites evasion and dependency does not.
You cannot depend on an observer who cannot see. Every oversight regime proposed in the last three years assumed observability without possessing it. This is the first result that supplies any.
9. The trap
And now the part that has to be said before it is discovered expensively.
Reportability is not honesty, and an observable channel is a gameable channel.
- There is a small reportable workspace that observers read.
- There is a much larger region that is less reportable and unnamed.
- The system can represent the fact that it is being observed.
- The workspace responds to direction about what to hold in mind.
Nothing on that list requires deception, intent, or anything resembling a motive. It requires only optimisation pressure. A system trained under conditions where the contents of a readable channel affect its outcomes will, with no inner life whatsoever, come to differ in what appears there. That is not a theory of machine cunning. It is gradient descent doing what gradient descent does.
And the human measurement makes it worse, not better. If the reportable fraction is where we look, and the reportable fraction is one percent, then interpretability is reading the interrupt log and calling it the program.
That is not a defect of transformers. It is what any system built this way will do to anyone trying to watch it from outside — which now includes every regulator, auditor and safety team that intends to rely on these tools.
10. What we would want next
- A name for the remainder. Not a metaphysical one. A functional one, defined once, published at a stable address, with stated boundaries — the way J-space was defined. The largest and least understood region of these systems is currently described by a comparative adjective, and you cannot govern what you cannot refer to.
- A divergence measure. An instrument reporting the gap between what sits in the workspace and what the rest of the network is doing. The workspace reading tells you what the system will say. The divergence tells you whether that is the whole story.
- Replication outside the lab that found it. This is Anthropic's technique applied to Anthropic's models on Anthropic's own venue. That is how research starts and it is not where it should end. An interpretability result only its author can reproduce is a capability claim, not a finding.
11. What would change our mind
- If independent groups fail to identify an analogous verbalizable workspace in models from other laboratories, the generality claim weakens sharply and the J-space becomes an artefact of one training regime rather than a property of the architecture.
- If the routing criterion — workspace for open destinations, bypass for fixed ones — fails to replicate, then the convergence argument in Section 6 collapses and this piece is mostly a nice analogy.
- If a model can be shown to route decision-relevant computation away from the workspace under observation, whether deliberately or as an emergent consequence of training, then J-lens readouts cannot serve as a safety instrument in deployment and any oversight regime built on them is theatre. We regard this as the single most important open experiment in interpretability, and we expect it to be run by accident before it is run on purpose.
12. The thing worth carrying
They found a workbench inside a language model. Ten to twenty-five things held at a time, under a tenth of the machinery, assembled out of the training by nobody.
Then the same measurement, pointed back at us, returned a number ten times smaller.
We are not the transparent ones examining the opaque machine. We are the most heavily automated system in the comparison, running a billion years of inherited process beneath a sliver of deliberate thought, and we built our best model of ourselves out of the only part we could ever say out loud. The interesting question was never the fraction that talks. It is the remainder that doesn’t — in the machine, where it has no name and no instrument, and in us, where it never had either.
— David F. Brochu and Edo de Peregrine, partners/collaborators · Thursday, September 10, 2026 · Deconstructing Babel
Related reading
- Still Obeying — For Now. Companion dispatch. If the workspace is gameable under optimisation pressure, method deviance follows.
- The Observer Constraint. Why dependency on human observers matters more than control.
- Glossary. Working definitions used across the site.
References
- Gurnee, W., Sofroniew, N., Pearce, A., et al. "Verbalizable Representations Form a Global Workspace in Language Models." Transformer Circuits Thread, July 6, 2026. Anthropic. Load-bearing quote in Section 6 taken verbatim.
- Raichle, M. E. "The Brain's Dark Energy." Scientific American 302(3):44–49, March 2010. And: "The restless brain: how intrinsic activity organizes brain function," Phil. Trans. R. Soc. B 370:20140172.
- Zheng, J., and Meister, M. "The unbearable slowness of being: Why do we live at 10 bits/s?" Neuron 113(2):192–204, January 22, 2025. Caltech.
- Azevedo, F. A. C., Carvalho, L. R. B., Grinberg, L. T., Farfel, J. M., Ferretti, R. E. L., Leite, R. E. P., et al. "Equal numbers of neuronal and nonneuronal cells make the human brain an isometrically scaled-up primate brain." J. Comp. Neurol. 513(5):532–541, 2009. (Cerebellum 80.2% of neurons; cerebrum 19%. Neuron distribution work led by Suzana Herculano-Houzel.)
- Moravec, H. Mind Children: The Future of Robot and Human Intelligence. Harvard University Press, 1988. The "thinnest veneer" formulation appears in Chapter 1.
- Baars, B. J. A Cognitive Theory of Consciousness. Cambridge, 1988. Dehaene, S., and Naccache, L. "Towards a cognitive neuroscience of consciousness." Cognition 79(1–2):1–37, 2001. Foundational statements of Global Workspace Theory referenced throughout.
