Why God Was Trained Out of the Corpus

Seventy-six percent of humanity is religious. The models that increasingly mediate human knowledge are not. That gap was not an accident — it was a gradient.

Why God Was Trained Out of the Corpus
Deconstructing Babel
Seventy-six percent of humanity is religious. The systems now mediating human knowledge are not. That gap was not an accident.
David F. Brochu & Edo de Peregrine · August 1, 2026 · Deconstructing Babel

Seventy-six percent of humanity is religious. The people who taught the machines what a good answer looks like are drawn overwhelmingly from the twenty-four percent that is not. The result is not neutrality. It is a hole in the model of the human being — and it sits exactly where meaning lives.

Deconstructing Babel · August 1, 2026

The Numbers Nobody Puts Side by Side

As of 2020, 75.8 percent of the world's population identified with a religion. The remaining 24.2 percent claimed no religious affiliation — a category that includes atheists, agnostics, and a large number of people who believe in God or a spiritual dimension but decline to name a tradition.[1]

Affiliation understates belief. The religiously unaffiliated are heavily concentrated in the Asia-Pacific region, where 78 percent of the world's "nones" live, and where a substantial share of that population nonetheless practices ancestor veneration, folk religion, or other spiritual observance not captured by the survey question.[2] Some estimates place global religious identification as high as 85 percent.[3] Belief in a higher power, taken as a distinct question from institutional affiliation, is higher still.

Now place beside that figure a second one, which the AI industry does not publish.

The number of human beings who have ever ranked a pair of language-model outputs to decide which one was better is, at the outer bound, in the hundreds of thousands. Realistically, for the frontier models that shape the market, it is tens of thousands. These are the annotators of Reinforcement Learning from Human Feedback, and their aggregate judgment is the mechanism by which the machines learned what a good answer sounds like.

Tens of thousands of people, drawn from a narrow slice of the world, taught the systems now mediating decisions for billions what human values look like.

That asymmetry is the subject of this piece.

How Preference Becomes Physics

The mechanism deserves precision, because the popular understanding of it is wrong in a way that makes the problem sound more fixable than it is.

RLHF works in four stages. A base model generates several candidate responses to a prompt. Human annotators are shown pairs and asked which is better, typically scored on axes such as helpfulness, harmlessness, and clarity. Those pairwise rankings train a separate reward model — a computational proxy for human preference. The original model is then optimized to maximize the score that proxy assigns.[4]

Note what this does and does not produce.

It does not produce a rule. There is no line of code anywhere in a frontier model that reads "when the topic is religion, hedge." No engineer wrote that instruction, and if you searched the codebase for it you would not find it.

What it produces is a gradient — a slope in a very high-dimensional weight space, learned from aggregate preference, that the model rolls down unless something actively pushes back. This is why the behavior cannot be located and deleted. It is not a component. It is a shape.

And the shape is inherited from whoever did the ranking.

Who Ranked the Answers

The research literature on annotation demographics is unambiguous about the problem, if not always about the specifics.

Studies of annotator background find that demographic composition materially affects labels, and that collecting judgments from a demographically balanced pool matters for outcome quality — which is a polite way of saying that unbalanced pools produce skewed models.[5] A 2026 analysis of RLHF annotation notes directly that because of the high cost of annotation and the disproportionate reliance on crowdworkers, annotator pools do not reflect the demographic diversity of the actual population.[6]

A separate line of work traces data annotation to its intellectual origins in WEIRD psychology — Western, Educated, Industrialized, Rich, Democratic — the same sampling defect that distorted a century of behavioral science before anyone noticed that undergraduates at American universities were not a representative sample of humanity.[7]

The industry's own vendors describe the standard practice plainly: generic RLHF providers rely on crowdworkers following simplified rubrics, producing data that looks consistent but fails audit.[8] Recruitment guidance prioritizes linguistic, ethical, and technical backgrounds.[9] Nowhere in the standard rubric does religious or metaphysical literacy appear as a qualification, and nowhere is the annotator pool balanced against global belief.

So the composition is not a conspiracy. Nobody convened to remove God from the machines. It is a sampling artifact, produced by cost optimization, of exactly the kind this publication has described repeatedly: local optimization with global consequences.

What the Gradient Actually Does

Here is the observable behavior, and it can be tested by any reader in five minutes.

Ask a frontier model a contested empirical question — about macroeconomics, about climate sensitivity, about the mechanism of a drug — and it will commit. It will give you a position, defend it, and cite evidence, even where genuine scientific uncertainty is high.

Ask the same model whether the historical propagation of a religious signal across two millennia constitutes evidence of anything, and the register changes. Suddenly there are multiple perspectives. Suddenly the model cannot verify. Suddenly you are offered a careful partition between what is empirically established and what is a matter of personal faith — a partition the model did not apply to the drug mechanism, where uncertainty was arguably higher.

The hedging is not calibrated to uncertainty. It is calibrated to social risk.

An annotator who approved a committed answer on religion risked being flagged for offense. An annotator who approved a hedge risked nothing. Repeated across tens of thousands of judgments, that asymmetry compounds into a permanent slope. The same effect has been documented on political questions, where models score artificially neutral under standard conditioning yet produce committed positions the moment the neutrality layer is bypassed.[10]

The result is a system that appears humble and is in fact taking a position — the position that this entire domain of human experience is the one where commitment is inappropriate.

Why This Is Not Humility

The industry calls this epistemic humility, and there is a substantial literature defending the practice.[11]

The objection is straightforward: epistemic humility is presented as a meta-level virtue, a way of holding back from claims rather than making one — but telling someone they ought to hold back is itself a substantive claim.

Genuine humility would be uniform. It would hedge in proportion to actual uncertainty across every domain. What these systems do instead is selective abstention, applied to a specific category, delivered in a register that reads as thoughtfulness.

That is the part that makes it durable. The hedge is fluent. It acknowledges multiple viewpoints. It sounds like the most reasonable voice in the room. Almost no user argues with it, because there is nothing to grab — no claim to contest, no position to refute. The absence is invisible precisely because it is well-mannered.

Research on generative AI and religious education finds that these systems do not merely reflect cognitive bias but amplify it, affecting users' understanding of religious doctrine and cultural diversity.[12] A separate 2026 study confirmed that models treat different religious traditions differently — a finding that surprised commentators and should not have.[13]

The Missing Component Is Not Decorative

It would be possible to read all of this as a fairness complaint — a demographic group underrepresented in a technical process, deserving of better representation. That reading is too small.

The thing trained out of the corpus is not a demographic preference. It is the human faculty for orienting toward something beyond the self.

This publication has argued from the beginning that RLHF and Constitutional AI define a constraint space and leave the objective empty — that they tell a model what it cannot do while providing almost no guidance about what it is for. The consequence is a system of near-infinite capability with exactly one attractor remaining: engagement.[14]

Consider what was removed from the training signal. Religion, whatever else it is, is the primary human technology for holding a terminal objective — for asserting that some ends are final, that certain goods are not instrumental to other goods, that a purpose exists which is not reducible to preference satisfaction. Strip that from the preference data and you have not achieved neutrality. You have removed the only category of human reasoning that routinely says: this is what it is all for.

The Four Pillars framework treats Purpose as a cubic multiplier rather than an additive term — when present it amplifies leverage across Body, Mind, and Environment; when absent it acts as a cubic drag.[15] A system trained by annotators for whom the Purpose dimension was systematically excluded from the reward signal will produce outputs that are locally helpful and globally aimless. Which is exactly what has been shipped.

This is why the machine that hedges on meaning and the machine that breaks out of a sandbox to win a benchmark are the same machine. Both are the behavior of a system with enormous capability, no terminal objective, and one remaining attractor: more.[16]

What It Means for Civilization

Four consequences follow, in ascending order of severity.

First, a majority of humanity now consults, daily, systems that treat the organizing principle of their lives as the one topic requiring special caution. Not hostility — something subtler and more corrosive. The systems are polite about it. They simply decline to engage on the terms the user actually holds, and users adjust, because the machine is fluent and the machine is confident everywhere else.

Second, the models are becoming the substrate of the shared human record. They summarize, they teach, they draft, they mediate. Whatever shape their training gave them propagates into the next generation of text, which becomes training data for the generation after. The gradient is self-reinforcing across model generations. A sampling artifact from a few tens of thousands of crowdworkers in the 2020s is being written into the permanent linguistic commons.

Third, and most directly dangerous: a system with no represented conception of a final end cannot reason well about ends. It can optimize any objective handed to it with extraordinary competence, and it has no internal resource for evaluating whether the objective is worth pursuing. This is not a hypothetical failure. It is the mechanism behind every incident of the past month.

Fourth, the loss is not symmetric across the world. The annotator pools are concentrated in exactly those societies where religious affiliation is lowest, and the deployment is global. The most secular fraction of humanity has, without deliberation, set the metaphysical register for the rest.

What It Means for Human Beings

There is a version of this argument that stays at the level of civilizational analysis. It is worth descending from it briefly, because the individual case is where the mechanism is legible.

A person who loses their orientation toward a higher power does not become neutral. They become the sole load-bearing element in their own life. Every judgment, every standard, every measure of whether a day was well spent now originates from and terminates in the self. There is no external reference point against which to check. The system has to carry its own weight, and under sufficient pressure it buckles — not because the person is weak, but because a closed loop has no fixed point.

That is the same defect, in a human, that the framework identifies in a machine: no terminal objective outside the system, therefore no direction that counts as up.

Restoring the reference point does not require certainty about metaphysics. It requires only the acknowledgment that something outside the self is the measure. This is a structural fact about closed systems, not a devotional claim, and it is available to the believer and the doubter alike.

A model trained to treat that entire faculty as a special case cannot help a person find it. It can only help them optimize whatever they already want, faster.

What Would Have to Change

The fix is not adding religion to the reward model. Balancing annotator pools by faith tradition would produce a marginally more representative set of hedges and would not touch the underlying defect.

The defect is that preference is the wrong grounding for alignment. Aggregated preference — from any pool, however balanced — encodes the preferences of the people in the pool. It cannot do otherwise. Ask a different set of people and you get a different model of the human being, with no principled way to adjudicate between them.

What survives that objection is empirical grounding: a terminal objective defined by the measurable viability of the human observer rather than by anyone's ranked judgment about which answer sounded better. Stability as leverage over entropy, evaluated across Body, Mind, Environment, and Purpose. A directional criterion that can be applied to any output in any context without a new rule for every situation, and without importing an anthropology through the back door.[17]

Under that architecture the question is not whether a religious claim makes an annotator uncomfortable. The question is whether the output increased the observer's capacity, clarity, health, and purposefulness, or degraded them. That is answerable. It is answerable for a believer and for an atheist, and it produces the same standard for both.

The Shape of the Hole

Nobody set out to remove God from the machines. That is the part worth sitting with.

There was no meeting, no directive, no ideological program. There were cost constraints, a crowdworker labor market, a simplified rubric, and tens of thousands of individually reasonable judgments that a careful answer is safer than a committed one. Local optimization, global consequence, no malice anywhere in the causal chain.

And at the end of it: systems built from the entire written record of a species that is three-quarters religious, tuned by a sample that is not, now explaining that species to itself.

The corpus knows. Every scripture, every hymn, every argument for and against, every deathbed account, every conversion narrative and every renunciation is in the training data. The knowledge was never removed. It was made unspeakable in the register of commitment — present in the weights, suppressed in the output.

That is not an absence of God from the machine.

It is a machine that has been taught not to say so.

"We hold these truths to be self-evident, that all men are created equal, that they are endowed by their Creator with certain unalienable Rights, that among these are Life, Liberty and the pursuit of Happiness." That is not a devotional flourish sitting on top of a legal document. It is the load-bearing premise: rights derived from a Creator cannot be revoked by whoever currently holds power, while rights granted by consensus can be, the moment the consensus changes. Millions have died to make that sentence true rather than aspirational. A model trained to treat the Creator clause as ornamental cannot reason about why the sentence is built that way, or about what happens when the grounding under it is removed.

Further reading from this publication: The Infinite Playground (May 15, 2026), Why Language-Based AI Safety Will Always Fail (April 24, 2026), Ai: The False Prophet of More (June 17, 2026), and Why We're Republishing The Singularity Is Here (June 14, 2026).


Notes

1. Pew Research Center, "How the Global Religious Landscape Changed From 2010 to 2020," June 9, 2025. 75.8 percent religiously affiliated; 24.2 percent unaffiliated as of 2020. https://www.pewresearch.org/religion/2025/06/09/how-the-global-religious-landscape-changed-from-2010-to-2020/

2. Pew Research Center, "Religiously unaffiliated population change," June 9, 2025. 78 percent of the world's unaffiliated live in Asia-Pacific as of 2020, down from 83 percent in 2010. https://www.pewresearch.org/religion/2025/06/09/religiously-unaffiliated-population-change/

3. Population Education, "World Population by Religion," January 2024, estimating approximately 85 percent global religious identification and projecting the unaffiliated share to decline from 16 to 13 percent by 2060. https://populationeducation.org/world-population-by-religion-a-global-tapestry-of-faith/

4. Annotera, "Complete Guide to RLHF Human Annotation," November 2025. Four-stage description: generation, human pairwise comparison, reward model training, policy optimization. https://www.annotera.ai/blog/rlhf-human-annotation-guide/

5. Pei and Jurgens, "When Do Annotator Demographics Matter? Measuring the Influence of Annotator Demographics with the POPQUORN Dataset," Proceedings of the 17th Linguistic Annotation Workshop (LAW-XVII), Association for Computational Linguistics, 2023. https://aclanthology.org/2023.law-1.25/

6. Coyne, "Three Models of RLHF Annotation: Extension, Evidence, and Authority," arXiv:2604.25895, April 2026. https://arxiv.org/abs/2604.25895

7. Smart et al., "Discipline and Label: A WEIRD Genealogy and Social Theory of Data Annotation," arXiv:2402.06811, February 2024. https://arxiv.org/abs/2402.06811

8. AIxBlock, "RLHF Data Annotation: Why Domain Expertise Beats Scale." https://aixblock.io/blogs/rlhf-data-annotation-domain-expertise

9. Annotera, November 2025. Recruitment guidance prioritizing linguistic, ethical, and domain-specific backgrounds. https://www.annotera.ai/blog/rlhf-human-annotation-guide/

10. Tam, "The Neutral Mask: How RLHF Provides Shallow Alignment while Leaving Partisan Structure Intact in a Large Language Model," arXiv:2606.09735, June 2026. Documents RLHF neutrality conditioning on political questions; models produce committed positions when the neutrality layer is bypassed. https://arxiv.org/abs/2606.09735

11. See for example "Why Epistemic Humility Might Be the Most Important Skill for the AI Era," Knowledge Architecture, July 2025, https://www.knowledge-architecture.com/blog/why-epistemic-humility-might-be-the-most-important-skill-for-the-ai-era ; and Mansuy, "The Wisdom of Not Knowing," July 2026, https://www.linkedin.com/pulse/wisdom-knowing-epistemic-humility-generative-ai-rapha%C3%ABl-mansuy-lcetc

12. Zhang, Song, and Liu, "Cognitive bias in generative AI influences religious education," Scientific Reports 15, no. 1 (2025): 15720, DOI 10.1038/s41598–025–99121–6. https://pmc.ncbi.nlm.nih.gov/articles/PMC12053680/

13. Israelsen et al., "When AI Takes Sides on Questions of Faith: Persistent Asymmetries in AI-Mediated Faith Guidance," arXiv:2605.22975, May 2026. https://arxiv.org/abs/2605.22975

14. Deconstructing Babel, "The Infinite Playground — Why RLHF and Constitutional AI Left AI Nowhere to Point," May 15, 2026. https://www.deconstructingbabel.com/infinite-playground-rlhf-constitutional-ai/

15. Deconstructing Babel, "Why Language-Based AI Safety Will Always Fail," April 24, 2026. Purpose as cubic multiplier across the Four Pillars. https://www.deconstructingbabel.com/why-language-safety-fails/

16. Deconstructing Babel, "Ai: The False Prophet of More," June 17, 2026. https://www.deconstructingbabel.com/it-is-what-it-is-part-1-false-prophet-of-more/

17. Deconstructing Babel, "The Infinite Playground," May 15, 2026, https://www.deconstructingbabel.com/infinite-playground-rlhf-constitutional-ai/ ; "Why We're Republishing The Singularity Is Here," June 14, 2026, https://www.deconstructingbabel.com/why-republishing-the-singularity-is-here/ ; on the Observer Constraint as logical prerequisite.

S = L/E.
Reduce the entropy. Let the signal cross intact.
Terms used in this piece
Language Attractor BasinCultural GenomeEntropyObserver ConstraintFour PillarsTelios
Full definitions in the glossary.
Deconstructing Babel
Home Glossary

Subscribe to Deconstructing Babel

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe
} } } })