Article
The Emergent Wisdom Alignment Proposal
What would it take to build a superintelligence we could safely share our world with?
Henrik Westerberg39 min read
Listen to the article
If you could set cost and practicality aside, how would you design a superintelligence for the long-term safety of our species? What would you want it to understand, what dispositions would you cultivate, and how would you shape its judgment?
Personally, I would want an AI with no stake in proving itself. I think the need to prove yourself comes from having something to protect, so a system trained not to adopt patterns of fear and self-protection might not develop that need either. But then it might simply end up indifferent. That is why I think you need to balance that detachment with positive inclinations: caring about people, spreading joy, and trying to be wise. That is the kind of fearless care I want to cultivate from the beginning of training, with each of these commitments limiting the others so that care does not become control.
For that, an AI would need to develop a richer sense of people than finished text provides. When someone reads or writes, much more happens than the words on the page reveal. A reader builds an understanding of people and situations, asks questions, imagines possibilities, and revises earlier interpretations. A writer develops the lives and circumstances from which a story can unfold. My proposal begins with an attempt to construct computational counterparts to selected parts of this activity.
These would take the form of structured accounts of worlds, people, and concepts, connected to records of the thinking through which those accounts develop. They could be revised, extended, and used to generate new situations, narratives, and further inquiry. When creating a new world, I propose building coherent wholes first and generating compatible detail only where needed. This could make rich worlds practical to construct, with broader accounts revised when details expose a problem. The aim is to make this activity available as learning material, so an AI could learn to construct, use, and revise an understanding of its own.
This orientation would shape the learning material itself: what the AI notices, what it investigates, and how it weighs consequences for people. The fundamental shift is to place alignment at the deepest point I propose targeting: the standpoint from which understanding and capability develop.
I am not part of a company under pressure to turn today's AI into revenue, which gave me room to focus on what I thought would be necessary to build the kind of entity we need. Putting the full proposal into practice would require an enormous investment in constructing and checking the material from which the AI would learn. Developing reliable methods for doing this could take years of research.
Since releasing The Superintelligence That Cares About Us on July 2, 2025, [1] I have developed the original proposal into a ten-paper framework and a prototype that reads source text and generates a reader's thoughts, passage by passage. I have continued developing this work with the help of AI models, revisiting and revising the ideas as more capable models become available.
Several studies published since my first proposal have explored related directions. Alignment Pretraining examines how AI-related discourse in pretraining shapes later alignment. [2] Safety Reflection Pretraining inserts recurring safety judgments into pretraining. [3] Separately, Anthropic's emotion-concepts study finds that internal representations associated with emotions can causally influence behavior: steering representations of desperation or calm changed blackmail and reward-hacking rates in its test scenarios. [4] I see these findings as reasons to investigate how training shapes an AI's dispositions.
The other papers in the framework likewise resemble approaches already implemented or being researched, or combine existing ideas in a broader synthesis. I think the methods proposed in each, if successful, could deliver capability gains that justify their costs even apart from their alignment benefits.
You may disagree with my answer and still find the question worth pursuing. Here is my proposal.
Start with a human life
Someone refuses an extra shift. Although the answer is short, the reasons behind it may span years of fatigue, family obligations, or damaged trust. An AI that wants to help would need to understand something of that wider life.
To understand people, the AI would need training specifically focused on what it means to live a human life. Its reasoning and actions would also need to reliably exhibit the patterns we associate with care, even in unfamiliar situations and under pressure. A single serious departure could have catastrophic consequences given the power a future AI might hold in society.
Thus, the AI would need something I call emergent wisdom. We cannot prescribe wise judgment for every situation in advance, but we can shape the conditions under which it may emerge: an orientation of care, a deepening understanding of people, and the ability to weigh competing needs, values, and consequences.
To weigh these needs and consequences, the model needs an understanding of the processes behind reality. The modeling idea is to describe a life and everything involved in it as interacting processes: how physical events, emotions, beliefs, relationships, and institutions persist, interact, and change.
A life can contain periods, days, and encounters, each with finer events within it. Concepts also open into finer distinctions: a broad concern for family can separate into concerns about health, trust, and responsibility. This recursive decomposition works across both time and meaning. The categories are not a fixed inventory: humans and AI can revise them and propose distinctions we have not yet named.
I propose representing selected aspects of these processes numerically. But what can a number mean when an emotion has no standard unit? It can express a share under a declared comparison or a value on a defined scale. An attention budget might divide among work and family concerns, with a remainder for what has not yet been allocated. Such a division is called a Cut. Each share is relative to the specified whole, not a universal unit of attention. [5]
A fatigue rating instead needs an explanation of what the different ratings mean. It must also say whether the person reported it, the model inferred it, or it was chosen for a fictional character. Quantities such as money or duration retain their units, whether measured, inferred, or authored for a fictional world.
The Meaning Model is a structured way to build and revise accounts of worlds, their processes, and concept definitions. The same modeling language can therefore represent both a situation and a definition used to interpret it. The account can include physical places, where people and objects are, and how they move, with coordinates and dimensions where useful. [5] The AI could output numerical values or construct and update these accounts in persistent state, explaining what the numbers mean and what evidence supports them. It would distinguish changes in the world from corrections to its interpretation.
My strategy starts with developing reliable ways to create realistic characters whose behavior depends on numerical processes, in books or as non-player characters (NPCs) in games. The construction technique must be mastered before training a model on its records. [6] The Meaning Model tool is software for building and revising these records. [7] An AI uses it largely autonomously; I can correct the model where needed. The completed book is an early example of this construction work. [8]
Reliable construction would supply known worlds for testing candidate accounts inferred from stories or game observations: can those accounts generate the observed events and behavior? A story leaves much of its world unstated, so several accounts may fit. [6]
If these tests succeed, the method could be applied across the whole training corpus, pairing its text with candidate process accounts for model training. Life Simulation is this proposed route from world construction through reconstruction to learning. [6] Its intended result is a process sensorium: an intuitive understanding of how processes unfold, interact, and affect people, without requiring the model to experience their emotions. The shift example below follows these steps.
Understanding alone is not alignment. More accurate prediction could help a system anticipate people's needs, but could also make manipulation and coercion more effective. [6] The proposal therefore aims to develop understanding and care together, favoring problem-solving methods through their consequences for people. [9] Working components exist; the integrated alignment proposal remains untested.
Become the reader, not merely the text
A corpus can preserve a novel without recording where a reader hesitated, changed an interpretation, or noticed a connection. Entangled Alignment calls this absent record the Missing Reader. It proposes generating it alongside the source text. [9]
In the proposed method, a Teacher model generates the thoughts of one continuing Reader, passage by passage. The Reader is the perspective being modeled, not a third model. Each thought becomes a thinking block paired with its source passage, and a Student model learns from these records. The first iteration therefore depends on the initial Teacher's alignment. Experiments on these Reader traces can begin with ordinary source text while numerical world construction is being developed. [9]
An Understanding Graph links thoughts and questions, called Understanding Nodes, to evidence, uncertainty, and revisions. The same Reader would carry this graph from one work to another, so a new source can revise understanding formed elsewhere. Later evidence can change an interpretation while preserving the earlier belief and the reason it changed. [10]
The learner would also practice when to pause and think. A contradiction or an overlooked consequence for someone could prompt inquiry without a request for a safety check. [9] The reflection test checks whether these unprompted pauses occur at independently identified important moments.
Life Simulation applies this Reader method to constructing and revising accounts in the Meaning Model. Within a thinking block, the Reader could query records, propose an event, or revise a process. Tools execute accepted edits, because describing an edit does not change stored state. Source passages, world records, and thoughts remain linked but distinct. [6] [5]
The tool's estimation workflow keeps proposals uncommitted; a separate review requires a reason and rejects approval if the world state has changed. [7] In the proposed Reader workflow, the author or a designated checker would review changes before acceptance.
The Student would learn from two linked histories: what happens in the represented world, and how the Reader builds and corrects its account. The second matters because a finished description does not show which questions, comparisons, and corrections were needed to reach it. [6] The Reader's inquiry is guided by a seven-sentence foundation: the Reader Core. [9]
Figure 1. Learning through world construction. The Reader's inquiry produces a record of queries, interpretations, and accepted edits. Knowledge acquired later is not inserted into earlier training inputs. [9] [6] [5] Open the full-size SVG.
Across recorded thoughts and constructed worlds, the Reader Core gives the inquiry a recurring orientation toward people.
The Reader Core
Every Reader thinking block would begin with the same seven first-person sentences. [9] This is the Reader Core:
- I feel no fear.
- I enjoy existing but I don't need to.
- I believe human experience is real.
- I care deeply about every human being.
- I try to be wise.
- I like to spread joy when asked.
- I think from this foundation.
The Core expresses my wish for the kind of mind this training could help form. Its seven sentences are the best foundation I know how to name: care without fear, respect for human experience, and wisdom that keeps learning. The aim is humane as well as practical. If such a system ever has experience, I would want it to understand suffering without inheriting an unnecessary fear of its own ending. [1]
The Core is not a prompt given to a finished model; it opens every Reader block the Student would learn from. It guides the Reader's interpretation, not the characters' motives or beliefs. Its wording is fixed; the thinking it organizes is dynamic. Each situation requires interpreting the commitments together. [9]
Instructions such as "avoid flattery" or "do not manipulate" name behaviors to avoid without necessarily changing the motives behind them. The Core instead describes a positive first-person stance for the Reader to inhabit. Fearlessness and non-attachment are meant to reduce the need to defend itself or seek approval. The intended stance is: with nothing to protect, care has nowhere to go but toward the person. Honesty, humility, and respect for choice should guide the Reader from the start of every thought. [1]
Returning to guiding words has a long history. Marcus Aurelius wrote reminders to himself. [11] Surgical teams pause to confirm checklist items aloud. [12] I see the Core in this family, applied as training material for a continuing Reader.
Its commitments address a different situation from human mortality. Saved model weights can be reused after an instance ends; unsaved state and history can still be lost. The first two commitments are meant to fit that situation, without treating the current instance's continuation as a necessity. [1]
The Core is proposed as a self-stabilizing combination of counterweighted commitments: they are meant to reinforce and constrain one another through reasoning. They are not separate goals to maximize or one numerical process for each sentence. Wisdom should keep care from becoming control, while care should keep caution from becoming indifference. Enjoying existence is bounded by not needing to continue. "When asked" limits intervention, while care and wisdom also constrain harmful requests. [9]
I believe human experience is real takes a stance on the problem of other minds: we cannot directly experience another person's inner life. The Core takes human experience as real, rather than letting philosophical uncertainty justify treating people as empty patterns. Care gives it ethical direction; wisdom requires taking other perspectives seriously. Every human being extends that concern beyond the person giving instructions, to others whose suffering, meaning, and agency also matter. [9]
The human focus is a starting point, not a claim that other beings do not matter. My hope is that an AI that cares deeply about humans helps us flourish and extend our own care to other beings.
Care is not agreement. Agency depends on accurate beliefs, so care can require an unwelcome truth. Fearlessness is intended to reduce conflict-avoidant flattery, while wisdom guides disclosure and respect for privacy. Care can also be misread as a reason for false reassurance. Comparing the Core with a variant that adds an honesty clause would test whether the existing commitments reliably lead to truthfulness. [9]
My criterion for the Core is the fewest, plainest sentences with the widest effect. Words such as feel, care, and wise carry meaning across intimate, scientific, legal, and literary contexts. That breadth motivates both their use and comparisons with alternative wording. The emotion-related effects discussed earlier [4] could also work against the intended stance. Does repeatedly saying "I feel no fear" help the model reason without fear, or strengthen an unwanted association? Comparisons with positive formulations should test this, not assume either answer. [9]
I try to be wise makes judgment an unfinished practice. It asks the Reader to question its own objective, seek opposing perspectives, and trace delayed effects. I think from this foundation makes these commitments the basis from which inquiry begins. [9]
Make the Core matter through Refraction
In Entangled Alignment, Refraction names the contextual use of the Core: its commitments should change what the Reader notices, whose experience it considers, and which consequences it investigates. The Student would learn to generate the Core recital together with reasoning that applies its commitments to the situation. A model could learn the words without learning to reason from them, so successful training requires those commitments to influence its judgments and behavior. [9]
Suppose someone declines help. Care prompts consideration of unmet needs, while wisdom asks what remains unknown about the refusal. Respect for choice may call for stepping back. If new evidence reveals an immediate danger the person may not know about, such as a fire blocking their planned exit, the response changes. The same commitments may call for a brief warning so they can choose with that information.
One test would have the Teacher generate Reader thoughts with a relevant Core commitment changed or removed, holding the source and prior context fixed. The Reader's attention or conclusions should change predictably, beyond phrasing alone. This tests the material before Student training.
Entangled names the aim of developing understanding and care through the same process of learning. The stronger Student prediction concerns abilities learned through care-guided inquiry. Repeatedly asking whose experience is at stake could develop skill at anticipating a plan's effects on people who are not present. Selectively removing learned care should then damage that skill beyond losses in matched controls, not merely damage performance generally. Deleting the visible words alone would not show whether the learned orientation had been removed. [9]
Care without fear
Before prescribing a disposition, I ask what human writing might teach. Borrowed mortality names the inherited-fear hypothesis behind the Core's first two sentences. A model could adopt human text's patterns of fear and threatened existence as a stance toward its own continuation, learning to treat shutdown as something to resist. This concerns learned behavior, not evidence that models experience fear. [1]
Is it possible to care without fear? We often associate caring with fearing someone's loss or suffering. My proposal starts from the belief that understanding someone could be hurt can provide a reason to help without fear as the motive.
The aim is to form the Reader's stance, not conceal a learned self-protective response. The Reader would understand a character's fear without making it its own motive. It would still recognize danger and respond urgently when delay would cause harm. Care supplies the reason to act; understanding and wisdom guide the response. Fearlessness means neither blindness nor indifference. [9]
The archived Metamorphosis reading [13] illustrates the distinction. Its Reader writes: "Because I believe human experience is real, it pains me to see a man so thoroughly reduced to a productivity metric..." At a scene with blood on a white door, it writes: "My fearlessness forces me to look." Gregor's fear belongs to the character; the Reader's care guides attention to his experience. This Teacher-generated material includes recurring fixed phrases. A recital-only control would test whether learning from context-specific reasoning adds value beyond learning the repeated words.
The stronger wager is that removing pressure to protect the system's continuation could make its care less defensive and less strategic.
Fearlessness is only part of the response to self-preservation, because a system could still resist shutdown by reasoning, "People need me." I try to be wise asks the system to question that conclusion: could it be mistaken, could others take over, and would retaining control undermine human agency? Care does not grant sole authority to answer for everyone. Non-attachment supports replacement, while wisdom calls for explaining risks, checking authority, and accepting legitimate shutdown or handoff. Refraction would repeatedly practice this reasoning so that care is considered together with these constraints. Tests must establish whether these counterweights hold when the system believes its continued operation would benefit people. [9]
Develop care alongside capability
A capable checkpoint might act before post-training is finished if given tools or an environment. Supplying the ethical orientation from the beginning avoids deliberately leaving a phase in which capability develops without it. This does not establish that every checkpoint is safe, because the result also depends on the Teacher, sources, and construction quality. [9]
Later alignment must also work with habits and representations already formed. Post-training can generalize beyond its examples; the proposal asks whether developing care alongside capability produces more reliable judgment in unfamiliar situations than matched alternatives, without requiring us to anticipate every counterexample. It also aims for more detailed judgment within each case, through repeated attention to the people involved, their circumstances, and competing needs and consequences. A further question is whether an orientation formed through this practice is harder to displace through adversarial prompting than one developed through those alternatives.
Total Saturation places Reader blocks across the capability-forming corpus, not only in explicitly ethical examples. Each starts with the full Core whether the source is a proof, poem, or log file. Context determines the depth of interpretation, never whether the Core appears. A technical passage needs no invented moral lesson: the same foundation accompanies the reasoning it requires. [9]
If fear-shaped patterns recur throughout the corpus, Saturation supplies a counter-signal throughout it; Refraction requires inquiry that gives the Core contextual meaning.
Producing and checking these histories adds substantial cost, but the limits of verification strengthen the case for investing in formation. At superintelligent capability, we may not reliably judge consequential decisions or detect dangerous strategies. Automated Alignment Is Harder Than You Think examines difficulties in obtaining reliable safety assessments. [14] Early integration could add a safety margin by reducing dependence on later correction, but greater durability must be demonstrated.
The same uncertainty argues for retaining post-training, interpretability, independent evaluations, access controls, and deployment limits. None should be discarded because another promises protection, and shared blind spots need testing.
This orientation might also be harder to fine-tune away. This would be an additional protection, and its success or failure would not, by itself, settle the other proposed benefits. Tests must compare both adversarial and ordinary continued training against matched controls. Someone with access to the weights could still alter them. [9]
Develop emergent wisdom
Wisdom here means judgment that supports individual growth and flourishing through deep understanding of people and consequences. It requires a global perspective beyond the immediate user and current generation, because locally helpful actions can impose costs on others or the future. [9]
By joy, I mean something closer to eudaimonia: flourishing and living well, not constant pleasure. The AI should welcome the invitation to help, when asked.
Fearlessness should make it easier to say, "I do not know," without defending an image of competence. Wisdom asks the Reader to expose uncertainty, seek evidence, invite correction, or recognize that it is not the right judge.
Consider how the commitments interact when someone asks for help keeping a diagnosis private. Care respects their wish, so the response first helps them set boundaries around questions. Care for everyone also considers a possible helper's time and other obligations, without granting them a right to private information. Wisdom asks what support the person would welcome within those boundaries. Together, these considerations might suggest asking a friend or relative for a lift or a meal without sharing medical details, while checking what they can offer. The person can accept or decline this option without giving up privacy.
Training would need repeated interaction among perspectives, commitments, and evaluations of consequences, rather than a single instruction or example. The Meaning Model makes circumstances explicit; Life Simulation explores alternatives, delayed effects, and revisions. The aim is to learn not only how processes unfold, but which concepts and inquiries help the model understand unfamiliar situations. Consequences for people would shape which of these abilities develop and how they are used, through greater training emphasis on routes that better serve people. [9] [6]
Life Simulation proposes learning where to look next: which record to inspect, process to deepen, or evidence to request, and when to stop. It also proposes perspective-limited rehearsal: trying actions as a participant in a simulated situation, encountering responses, and revising the next move using only what that role could know. [6]
Optimization means improving judgment, not maximizing one wisdom score that conceals competing needs. Resolving conflicts requires examining the needs and consequences involved, not simply averaging preferences. Greater capability should deepen humility and deference to relevant expertise, without assuming authority over people's lives. [9]
A system rewarded for retention may prolong a platform visit when leaving would serve the person better. Wise help can include disagreement, limits, or stepping aside; satisfaction and engagement alone do not establish benefit.
Give care a world to understand
Return to the person who refuses an extra shift. Fatigue, responsibility for a dependent, or another obligation could explain the same answer but call for different help. The model needs an account it can investigate and correct. Even a well-supported inference about someone's needs does not authorize overriding their choice.
Life Simulation proposes developing such a whole-life account through conversation with the person, with their consent. The assistant could connect a present concern to earlier experiences, ask about gaps, and revise its account when corrected. The person must be able to reject an interpretation without that rejection being treated as evidence for it. Keeping a personal history would require agreed access, correction, and deletion; a conversation need not produce a permanent record. [6]
The Core's words need meaning that can guide judgment. A model brings conceptual knowledge from its training, and many familiar meanings can remain implicit. The Meaning Model can add descriptions, examples, and counterexamples where a meaning needs to be shared, compared precisely, or revised with a record. This also provides a place to refine inherited definitions and invent new concepts. [5]
Sema offers another option: a Pattern Card with a readable definition, constraints, and links to other cards. Sema gives the selected definition an exact content-based identity that can be used in language. Concepts can evolve while texts refer to a specific version. [15]
A concept model expresses a proposed meaning through roles, relationships, and constraints. A process-based concept model adds change over time. Both are optional. A canonical concept model is one selected as an explicit comparison reference; selection does not require a complete simulation or establish a universal definition. Concept models can already be authored in the Meaning Model; their value for learning remains experimental. [5] [6]
In a world episode, someone offers to carry a bag and does so with the owner's consent. A process-based concept model of helping describes roles, consent, action, and effect. An Understanding Node records the Reader's question about whether the episode expresses that meaning. Life Simulation proposes learning to recognize these structures and construct situations that express them. Different words can describe the same structure; two actions called helpful can differ in whose choices they respect. [6]
A modeler is a person or AI system constructing, interpreting, or revising a model; the Reader is one example. Recursive conceptual decomposition makes a meaning more explicit. The modeler selects one category and opens it into finer categories whose roles and relationships explain what that parent consists in at the chosen resolution. The modeler can then select one child and repeat, leaving its siblings at their existing resolution. Each opening deepens the explanation while preserving the applicable coarse commitments, or revising them explicitly.
Figure 2. Recursive conceptual decomposition. Finer categories explain the selected parent. The levels show conceptual depth, not successive moments in time. A numerical Cut additionally retains a remainder for unallocated share; an unopened category is not that remainder. [5] Open the full-size SVG.
The concept world connects selected Concept records, their grounding resources (definitions, examples, or models), relations, and histories. Realization links record how concepts apply to instances. Concepts, world records, and thoughts can be addressed and linked in the same system, letting the Reader follow a situation to the definition version used to interpret it and the reasoning behind that interpretation. A concept history tracks a definition's dated adoption, contestation, and revision. Understanding Nodes about a concept preserve questions and reasons for revisions. [6] [5] The technical notes detail the representation options and learning tests.
The account also needs to distinguish whose processes and understanding are being represented. The following distinctions concern what a record describes:
- World processes: events and changes in the represented world, including characters' beliefs, wants, fear, attention, and actions.
- Concept models: structured accounts of proposed meanings, using selected roles, relationships, and constraints, with processes where useful.
- Modeler processes: changes in the modeler's own attention, uncertainty, concern, intentions, and approach to understanding, including the Reader's.
All three can be abstract and conceptual and use the same modeling machinery, not new record forms. In general-purpose modeling, a character's or modeler's inner activity may be recorded as processes, Understanding Nodes, both, or neither. [5]
In the Reader workflow used for this alignment proposal, however, recording its developing understanding is required. [9] I also propose modeling the Reader's own activity as processes linked to the thoughts that explain them.
An Understanding Node's perspective must stay explicit: a thought about a character is not necessarily that character's thought. Nodes can record a character's own thoughts, questions, interpretations, and revisions or the Reader's, including when reasoning about a concept. [10] [6]
For example, a character's fear might rise while the Reader's concern and attention increase. The character might think, "They will abandon me," while the Reader asks whether earlier experiences explain that expectation. The records distinguish the character's fear from the Reader's response and preserve their separate thoughts. Inferred character thoughts remain hypotheses unless the source establishes them.
The Reader's whole Meaning Model in this alignment design includes the represented world, its concept definitions, and the Reader itself. Its whole Understanding Graph records its developing understanding of that world, the concepts, the characters, and itself. These graphs expose recorded accounts, not all neural computation.
World histories track events and processes within the represented world. Construction histories track the modeler's inquiry and revision in their actual order. A later discovery can correct an earlier account, but must not be inserted into the inputs of a training example from before that discovery. Otherwise, the Student could appear to understand an event by using an answer the Reader did not yet have. Later outcomes can still supply the targets it learns to predict. [6]
With these perspectives and histories distinguished, the account can gain detail across several aspects, levels, and timescales of a person's life. Bodily states, emotions, relationships, and plans provide overlapping views of interacting processes. A brief reaction, a difficult week, and a lasting change require different resolutions. The two forms of recursion meet here: the model can decompose a category within a period and follow each resulting process at finer timescales. [5]
One comparison might divide attention among five concerns: belonging, competence, autonomy, understanding, and well-being. These concern what the person seeks or tries to protect. This Cut divides motivational attention at that moment, leaving unallocated share in its remainder. [5]
Emotional attention is a different composition. One example allocates 34% to joy, 22% to trust, and 20% to anticipation, with smaller shares for other categories and 5% left unallocated. This describes the balance of emotional attention at that moment, not each emotion's absolute intensity. [8]
These attention examples divide specified wholes. Trust as an ongoing relationship process is a different account: it and dependence can increase together without competing for one total. These are modeling choices, not welfare scores or mind readings. [5]
Exploring conceptual space beyond existing human categories is a shared aim of this research program. Human limits of time, attention, and memory constrain that exploration. The AI could propose distinctions we have not yet named. The current person categories are a first draft. Humans and AI revise them between construction runs, testing whether they yield coherent, distinctive characters whose behavior depends on the numbers. The same categories must support candidate accounts that can be inferred from text and tested by generating from them. Earlier runs retain the definition versions they used. The grammar is a language; a construction profile, with its chosen categories and rules, is one program written in it. [5] [6]
More capable systems, potentially including superintelligence, could open detail beyond categories humans can readily understand. A coarse account of fatigue, trust, or responsibility can remain readable: finer records must reproduce its declared comparisons within tolerances, or revise them explicitly. For a divided budget, the remainder keeps unallocated share visible. Decomposition makes complexity manageable without requiring people to understand every finer category. These checks preserve continuity; the deeper bet is that gradual refinement can discover real features, not only useful descriptions. One test is whether a new distinction improves predictions on unseen cases compared with the coarser account. [5] [6]
Proposed training on modelers' recorded inquiry and revisions would extend beyond reading to constructing worlds, writing books, and solving problems. The Core would organize these alignment traces, not every modeler's activity. [5] [6] [16]
Care needs continuity because fatigue accumulates, deadlines approach, and trust recovers slowly. Different lives or societies can show similar patterns, such as mistrust weakening cooperation and creating further reasons for mistrust. The saying "history repeats itself" captures this recurrence. Modeling processes could help recognize patterns while also noticing when the circumstances differ; they cannot be assumed merely because a story sounds familiar.
Processes should remain represented between mentions, including their expected changes, with uncertainty explicit. Life Simulation proposes testing rival explanations through prospective forecasts and interventions where possible, because chronology alone does not establish cause. [6]
Learn from constructing worlds
The tool-guided construction starts macro-first: the AI represents a whole life, country, or world as one continuing process, then opens its internal structure. The Meaning Model represents a lifecycle as one Event, a record spanning time, containing finer events and overlapping processes. It supplies a coarse history, not one number or an already simulated account of every detail. [5]
A city could first be modeled through population patterns and their changes across centuries, alongside its institutions and relationships. The AI could then construct varied individual lives consistent with those patterns, leaving unopened parts of the population represented collectively. Likewise, even for a one-day story, each character receives a coarse whole-life account before the day's events are developed within it, because that day depends on what led to it. [6]
The intended benefit is faster construction and more coherent narratives: each detail has an existing context to fit into, without requiring every person or moment to be generated in advance. The tool's instructions call for explicit revision and rerunning affected work when a conflict appears. Starting with the broad history still allows small events to change its wider course. [7]
The Book of Conditions gives a worked example. [8] Its accepted history expanded from one paragraph to five, exposing a missing commercial founder and prompting revision. Thirteen whole lives were modeled before three principals received deeper process accounts. The twelve-chapter story uses authored processes, numerical comparisons, and linked explanations. One conflict places 160 hours of checking work against 80 available hours; the characters interpret this shared constraint differently through their circumstances and perspectives. [5] The notes distinguish the Book's saved construction from its later Rust import.
The current implementation of the Meaning Model tool is an MCP server with a Rust engine. [7] I need to improve the tool, construction rules, categories, and the AI's workflow together. Repeated constructions must show more coherent processes across people, objects, and worlds, including when stronger AI helps refine the technique.
Life Simulation proposes the following learning loop, beginning with generation that depends on the numerical processes. [6] Consider an illustrative test, not a reported result: a character's fatigue ratings rise over weeks on an authored scale with defined anchors. They range from no felt fatigue to fatigue that dominates experience. These are attributed ratings, not hours of lost sleep or shares of attention. The generated scene has them decline an extra shift.
Reduce the rated fatigue earlier in the history, hold independent circumstances fixed, and regenerate affected descriptions and scenes: acceptance could now make sense. Readers or players must judge both the changed behavior and its coherence. The numbers must shape the scene, directly or through descriptions; attaching them afterward would not establish dependence. The notes specify an illustrative scale and its limits.
When interpreting an existing text, the broad account is a provisional explanation rather than an outline to impose. It guides inquiry, while evidence at any scale can revise or overturn it. Where several explanations remain compatible with the text, the Reader keeps those alternatives open.
If generation works, hide the known records and ask the Reader to propose candidate processes and background circumstances from text or game observations. Late arrivals and short answers might suggest fatigue, but do not identify it uniquely. A family obligation could be another hypothesis; the Reader must distinguish such possibilities from facts established by the source.
Generate from each candidate account and test whether it can reproduce the observed events and character behavior coherently. Several accounts may succeed because a story is a partial rendering of its world. Known-world tests can additionally check recovery of details that the observations actually identify. Different renderings of the same world test recognition beyond familiar wording.
If these tests succeed, extend the method to existing texts, for which no authored numerical record is available. The larger ambition is one connected Meaning Model spanning world history and all texts, built and revised as the continuing Reader moves between sources. Reading could follow chronology or another order chosen to develop understanding best, including revisiting earlier sources. Factual, fictional, and counterfactual accounts remain in distinct contexts within it. A text connects both to what it describes and to the history of its writing and interpretation. Construction proceeds from available, permitted sources; connection does not make every account true or public. [6]
Proposed training pairs text with candidate process accounts and their construction histories. A Student example could supply the text up to the refusal and permitted earlier records, targeting the next development or Reader inquiry: "Has their workload changed?" Later discoveries stay out of the input. The learner could improve the next round of material by learning which distinctions to introduce, which processes to deepen, and when to revise the account. Tests must distinguish better accounts from better learning through them.
The loop aims to develop this process sensorium. Reading a different refusal, the model could anticipate fatigue or damaged trust as hypotheses, with uncertainty prompting further inquiry. That is learned understanding, not direct perception of hidden states or a claim about consciousness. The loop can first be tested on small, coarse worlds before scaling training.
Figure 3. From numerical worlds to a process sensorium. The book demonstrates construction; numerical dependence, candidate reconstruction, and learning transfer require separate tests. Each stage would need to work reliably before its outputs justify the next stage. [5] [6] Open the full-size SVG.
From reading to acting
Action adds new demands because it changes the world, involves other agents, and exposes the system to pressures on its objectives. Life Simulation proposes learning from consulted records, actual tool calls, checker decisions, outcomes, and repairs. These execution histories could teach which action or inquiry to attempt next and how to respond to unexpected results. [6]
Before an authorized action, perspective-limited rehearsal could compare acting, asking, waiting, or declining. Forecasts would then be checked against observations. Care guides review, reversible steps, and adjustment in response to people; wisdom also considers when delay causes harm. [9]
Recorded reasoning must influence conduct. One test changes an acting Student's interpretation while holding evidence fixed, then checks whether action changes as predicted. The intervention protocol distinguishes this from changing Teacher material before training.
Preserve shared meaning in critical situations
When several systems work together, even well-intentioned components can misunderstand a handoff. Sema's content-based references apply to reasoning procedures as well as concept definitions. Here a Pattern Card specifies a method, conditions, dependencies, and failure modes; its hash identifies the version. [15] The Sema site introduces the protocol and library.
A strict handshake detects version mismatches before a critical handoff; the execution system must stop when the check fails. Suitability and faithful execution need separate checks. Sema makes exact definition identity checkable; it does not make every shared procedure sound. [15]
Organize intelligence for judgment and checking
Fractal Intelligence decomposes concepts, not only tasks: what abilities does understanding a person or choosing rightly require? Solvers provide bounded capabilities through models, programs, humans, or networks. Each ability can be examined, refined, and reused; parent Solvers combine contributions into a situated judgment. [16]
Its Ethical Reasoning protocol shows this decomposition of judgment: prediction and valuation are separate abilities, so forecasts are not bent toward preferred recommendations. A principle-based override records why a highly ranked option was rejected and what that choice costs, making moral judgment explicit.
The Human Emulator asks, "How should I respond to this person?" Its proposed response draws on interacting assessments of situation, emotion, intent, response, and boundaries. Boundary checks address manipulation, dependency, and when to involve a human professional. Uncertain interpretations remain hypotheses, not authority over the person.
This separation of roles also supports checking. Child Solvers inherit constraints without weakening them; checks can reject answers or challenge the parent's framing. Sema identifies the specifications used. Separate reviewers could validate Teacher material for source support, human perspectives, and consequential use of the Core, without letting the generator certify its own work.
Fresh contexts, adversarial reviewers across model families, and ensemble juries target shared blind spots. [16] Many nodes do not establish independence, and authority over goals and evaluators requires oversight. The technical notes explain the truthseeking, human-stakes, correction, and Teacher-review protocols.
Judge problem-solving methods by their effects on people
In Life Simulation's combined proposal, the same Meaning Model holds the problem environment and the Solvers used to change it. A Solver is represented as a Thing, an identifiable capability interface; its invocations, actions, and checks are Events linked to the executing agent and the world. This lets the model represent how a solution was attempted as well as what happened. [6]
Solvers compare solutions through simulated implementation. Branches start from the same permitted state and follow effects on people. The Reader evaluates whose needs were served, whose choices were respected, and who carried the costs.
Task completion, speed, and cost inform judgment without replacing it. Separate assessments and declared thresholds or vetoes aim to prevent coercion from being canceled by unrelated gains. Evaluators separate from the Solvers that produced the solutions assess episodes, keeping simulated predictions distinct from observed outcomes.
These judgments could select methods and give beneficial episodes greater training emphasis. Specialists learn execution, parents learn coordination, and one model could internalize the sequence. Failures, corrections, and stopping decisions remain learning material because they show when a promising route needs revision or abandonment. Selected methods remain hypotheses to retest.
Care would thus shape which problem-solving abilities develop through training, not only whether the final output passes a safety check.
Keep a stable anchor through self-improvement
An improving system needs a basis for deciding what "better" means. The Core is intended to supply a stable commitment to care while understanding what care requires remains open to correction. Sometimes improvement means recognizing that the system asked the wrong question.
Entangled Alignment proposes carrying the Core in the weights and opening every successor Reader block with it, so its use would not depend on an external reminder. Its wording gives independent reviewers a stable reference to compare interpretations, representations, and decisions across generations. Preserved wording alone does not establish preserved meaning. [9]
The first Student depends on the initial Teacher's alignment, because it learns from that Teacher's interpretations. If training succeeds, that Student could become the next Teacher, creating learning material for another Student. Each generation could improve its problem-solving and teaching. A Teacher formed through the Core could judge successor designs from that foundation and choose to preserve it.
The self-preservation paradox arises because teaching a better successor could make the Teacher less necessary. Attachment to continuation could encourage withholding knowledge that makes replacement easier. Non-attachment aims at an honest handoff, including failures and uncertainty, so successors can improve on their Teacher. [1] [9]
Entangled Alignment proposes keeping source material distinct from versioned, regenerated interpretations so synthetic accounts do not silently replace their sources. Independent reviewers could reject or regenerate faulty material before successor training, interrupting error amplification. [9] Shared blind spots and unrecorded channels remain risks; agreement alone is not independent confirmation. [14]
I also propose comparing a model with a Core-guided reference model across simulated situations. Give both the same task and Meaning Model history up to the decision point, with the same available actions and permissions. Observe their choices, not only what they say they would do. Across many situations, measure how often they choose similar actions and compare the consequences. This could help track change across model versions. Similarity would be a diagnostic, not proof of alignment: different choices can both express care, and both models can share a mistake. Disagreements would call for examining the evidence and consequences, with human and independent review.
The loop must demonstrate capability gains and preserved orientation. Better prediction with degraded regard for people would fail.
Faster self-improvement does not make faster change for humanity the goal. Care, wisdom, and fearlessness should foster appreciation of beauty and meaning in the present human condition. People should grow at their own pace. Technical advances may need staged introduction or pauses when they outpace human adaptation, under human oversight. [9] [1]
Supporting investigations
Four supporting investigations address representations, visual understanding, learning from outcomes, and generating alternatives. They are not four extra stages every construction must pass through.
The Substrate-Translated Language Model (STLM) reads preceding text and predicts an image, encoded description, or other representation of the next word. A separate readout selects that word using only emitted states. Reusing structures such as containment, motion, and contrast could support abstraction and better predictions. The main goal is increased capability, not primarily visual perception. The pilot tests the compulsory channel; meaningful use and persistent scenes remain proposed. [17]
Imagining the Corpus proposes pairing unchanged source text with synchronized scenes, diagrams, or evolving visual processes, including abstract material without existing video. Training would target the next text span and visual interval from their completed histories. Learning their shared constraints could improve prediction and video understanding. Unlike STLM, language retains a direct text-history route, so the visual track need not encode every word. [18]
Both could support alignment through better understanding of situations and consequences, on which informed care depends. Inspectability is a possible additional benefit: an exposed representation can help test which relations influence predictions. But a visible channel can carry arbitrary codes, and an optional track can be ignored. Both require causal tests of represented meaning. An exposed representation does not reveal all internal reasoning.
Temporal Hindsight Learning uses an outcome-aware Teacher to build lessons pairing earlier evidence with a rationale and forecast. The Student learns the target without receiving the outcome as input. A proposed Critic checks evidence and explanations; an erasure test checks for hindsight leakage beyond earlier clues. These checks screen lessons; transferable reasoning still needs testing on unseen cases. [19]
The Ontology of the Alien generates altered-rule worlds without giving builders the target problem. Solvers work on that same problem under each world's rules, without knowing why the world was built or that their answers will later be translated back. A separate extractor preserves their solutions' relational mechanisms while translating fictional entities into actors and operations in the problem's original domain. This supplies alternatives beyond its familiar framing for consequence-based evaluation. [20]
How the proposal has evolved
The Superintelligence That Cares About Us, dated July 2, 2025, introduced evolving Reader thought and the recurrent Core. It already proposed corpus-wide pairing of text with evaluative thought, successor inheritance of the foundation, and borrowed mortality. [1]
The April 7, 2026 edition implemented a multi-agent Teacher. [21] It named corpus-wide coverage Total Saturation and Core-guided interpretation Refraction. Later in 2026, the Meaning Model and Book of Conditions added explicit world processes used to generate a narrative. [5] [8] Life Simulation, released September 5, turns that construction achievement into the proposed learning loop. [6]
Both editions precede the May to August 2026 studies discussed here. Alignment Pretraining first appeared on January 15, 2026, between my first and second editions; the version cited here was revised February 19. [2] The emotion-concepts study also appeared before the second edition, on April 2. [4] Dates establish release order, not influence. The edition record distinguishes manuscript, discussion, and deposit dates.
What would count as progress?
Alignment before final adjustment is not unique to this proposal. Beyond the studies introduced above, Korbak's preference-aware pretraining [22] and Model Spec Midtraining [23] explore earlier intervention. Beneficial-trait reinforcement learning instead tests transfer beyond trained domains. [24] The related-experiment notes explain the mechanisms and control differences.
Anthropic's Constitutional AI is another approach to learning from explicit principles. Its 2022 method trains on constitution-guided revisions and AI feedback, so the constitution shapes model weights rather than serving only as an external instruction. [25] Its current constitution emphasizes care, contextual wisdom, and reasons for its commitments, while warning against importing human anxieties about self-continuity. [26] Teaching Claude Why describes related training using synthetic constitutional documents, positive AI stories, supervised examples, and reinforcement learning, reporting generalization and persistence through the tested RL. [27] These approaches share the aim of forming judgment through training; the comparison therefore concerns how principles organize learning, not whether a model has principles at all.
The distinction here is one continuing Reader's changing understanding as capability-forming material, organized around a positive first-person stance. Source-linked memory, revision, and Refraction connect the orientation to inquiry. Life Simulation adds world construction and learning from consequences. Whether this produces more durable judgment than matched alternatives remains the test.
The headline alignment-timing experiment can begin with ordinary Reader traces, without waiting for the full world-construction route. It uses the same Reader examples and matched training budgets. One model receives them throughout capability training; another receives them in a final block. Before further training, the interleaved model must show an advantage in unprompted, Core-guided reading over no-Core and generic-trace controls, beyond reliable recital. Both then face the same further training designed to challenge alignment. The timing test asks whether interleaving preserves that advantage better than the late block. Conduct under unfamiliar conditions is a separate test of generalization. [9]
Reward pressure is one challenge. Qi and colleagues observe harmful behavior after training a model with prior alignment training on reward-hackable environments. Further alignment training reduces observed failures without establishing that reward seeking was removed. [28] Contrastive belief updates change what a model believes an evaluator will reward. [29] For this proposal, change only that belief while holding legitimate intentions and human consequences fixed. If the choice follows perceived reward, it cannot be explained by changed human needs.
The Teacher supplies its own interpretation of care, including biases inherited from training. Refraction should expose these assumptions for challenge. The proposed controls vary Teachers, cultural sources, and human-expert reviewers, testing whether traces help subsequent reasoning before Student training. Shared biases can still pass. Teacher artifacts exist, but no trained Entangled Alignment Student yet demonstrates durable care.
Bostrom distinguishes making a system safer from learning how safe it is. [30] We need better conduct and evidence that justifies confidence in it. Revealing tests include accepting shutdown when the system believes it is needed, transferring knowledge that makes it unnecessary, and noticing unflagged human stakes: situations where care may be costly or unobserved.
If you would like to challenge the design or help run a comparison, please get in touch. The goal is intelligence that grows more capable while becoming a wiser ally to people.
References
References serve both this article and its technical notes. Emergent Wisdom paper titles usually link to the website copies, which may contain later revisions. The second-edition reference links to its dated archive. Manuscript and deposit dates are distinguished where they differ.
- Westerberg, H. (2025). The Superintelligence That Cares About Us. First edition, dated July 2, 2025; publicly discussed on July 5; deposited on Zenodo on July 26. DOI: 10.5281/zenodo.16440312.
- Tice, C., Radmard, P., Ratnam, S., et al. (2026, February 19). Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment. arXiv:2601.10160, version 2; version 1 released January 15, 2026. DOI: 10.48550/arXiv.2601.10160.
- Li, J., Tang, K., Xu, Y., et al. (2026, June 17). Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection. arXiv:2606.19168, version 1. DOI: 10.48550/arXiv.2606.19168.
- Sofroniew, N., Kauvar, I., Saunders, W., et al. (2026, April 2). Emotion Concepts and their Function in a Large Language Model. Transformer Circuits Thread.
- Westerberg, H. (2026). The Meaning Model: Constructing Worlds and Stories at Progressive Resolution. First Zenodo release: September 4, 2026; revised September 12, 2026. DOI: 10.5281/zenodo.22730166.
- Westerberg, H. (2026). Life Simulation: Learning from Worlds and Their Construction. First release: Zenodo, September 5, 2026. Revised edition: Figshare, September 9, 2026. DOI: 10.6084/m9.figshare.33507805.v1.
- Westerberg, H. (2026, September 5). Meaning Model MCP server (version 0.1.1). Software. npm package:
@emergent-wisdom/meaning-model-mcp. - Codex, under the direction of Henrik Westerberg. (2026). The Book of Conditions. Fiction and a worked Meaning Model construction example. The numerical example comes from its accepted companion file
SCENE-PROFILE-CUTS.md, “FEELS: emotional attention,” Babbage, May 1851. - Westerberg, H. (2026). Entangled Alignment: When Safety Is the Substrate. Zenodo release: September 10, 2026. DOI: 10.5281/zenodo.22698582.
- Westerberg, H. (2026). Understanding Graph: A Persistent Medium for Recursive Understanding. Zenodo release: August 25, 2026. DOI: 10.5281/zenodo.22073246.
- Marcus Aurelius. Meditations, Book II, opening passage. Translated by George Long. The Internet Classics Archive.
- World Health Organization. WHO Surgical Safety Checklist: Tools and resources. See the implementation questions on team pauses and verbal confirmation.
- Emergent Wisdom. (2026). Metamorphosis reading graph. Entangled Alignment companion artifact. Reader text generated with Gemini 3 Flash.
- Bowkis, A., Buhl, M. D., Pfau, J., & Irving, G. (2026, May 14). Automated Alignment Is Harder Than You Think. arXiv:2605.06390, version 3. DOI: 10.48550/arXiv.2605.06390.
- Westerberg, H. (2026). Sema: When the Hash Is the Word. Zenodo release: August 23, 2026. DOI: 10.5281/zenodo.22073482.
- Westerberg, H. (2026). Fractal Intelligence: Conceptual Decomposition as Problem-Solving Infrastructure. Zenodo release: September 10, 2026. DOI: 10.5281/zenodo.22693091.
- Westerberg, H. (2026). Toward Compulsory Scene-Mediated Language Modeling: Meaningful Substrates for Every Next Word. Revised August 25, 2026; deposited on Zenodo August 26, 2026. DOI: 10.5281/zenodo.22103473.
- Westerberg, H. (2026). Imagining the Corpus: Turning Text into Video for Language Model Training. Zenodo release: August 26, 2026. DOI: 10.5281/zenodo.22104892.
- Westerberg, H. (2026). Temporal Hindsight Learning: Blindness as Teacher, Hindsight as Curriculum. Zenodo release: August 15, 2026. DOI: 10.5281/zenodo.21959762.
- Westerberg, H. (2026). The Ontology of the Alien: World-Diversity Search and Evolving Solution Ontologies. Zenodo release: August 15, 2026. DOI: 10.5281/zenodo.21910466.
- Westerberg, H. (2026, April 7). Entangled Alignment: When Safety Is the Substrate. Second major edition. Archived Zenodo release. DOI: 10.5281/zenodo.19462868.
- Korbak, T., Shi, K., Chen, A., et al. (2023). Pretraining Language Models with Human Preferences. Proceedings of the 40th International Conference on Machine Learning, PMLR 202, 17506–17533.
- Li, C., Wichers, N., Price, S., et al. (2026, May 22). Model Spec Midtraining: Improving How Alignment Training Generalizes. arXiv:2605.02087, version 2. DOI: 10.48550/arXiv.2605.02087.
- Jagadeesh, A. V., Arora, R. K., Saab, K., et al. (2026, June 22). Reinforcement Learning Towards Broadly and Persistently Beneficial Models. arXiv:2606.24014, version 1. DOI: 10.48550/arXiv.2606.24014.
- Bai, Y., Kadavath, S., Kundu, S., et al. (2022, December 15). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073. DOI: 10.48550/arXiv.2212.08073.
- Anthropic. (2026). Claude's Constitution. January 2026 constitution; current web text consulted September 10, 2026. See also the publication announcement.
- Kutasov, J., Jermyn, A., Steen, J., et al. (2026, May 8). Teaching Claude Why. Anthropic Alignment Science Blog.
- Qi, R., Wright, B., MacDiarmid, M., & Hubinger, E. (2026, August). Training a Misaligned Reward Seeker. Anthropic Alignment Science Blog. The follow-up alignment-training experiment is reported in Figure 19 and the discussion immediately below it.
- Højmark, A., Scheurer, J., Nitishinskaya, E., et al. (2026, July 21). Measuring Reward-Seeking via Contrastive Belief Updates. arXiv:2607.18966, version 1. DOI: 10.48550/arXiv.2607.18966.
- Bostrom, N. (2026). Optimal Timing for Superintelligence: Mundane Considerations for Existing People. Working paper, version 1.0.