Technical notes ·
Alignment Proposal: Technical Notes
Record distinctions, construction methods, experimental controls, and release history for the Emergent Wisdom alignment proposal.
These notes accompany The Emergent Wisdom Alignment Proposal. They preserve implementation detail, proposed extensions, and experimental controls beyond the main argument. They do not report a completed integrated Student experiment.
Citation numbers link to the shared reference list on the main article.
Reader construction and shared records
Entangled Alignment's Reader-trace experiments can begin with ordinary source text, without waiting for numerical world construction. Chronological traces, graph memory, and the Core have separately testable contributions; the full Reader-Core treatment still opens every Reader block with the complete Core. [9]
The design in Life Simulation [6] extends the Reader's inquiry into world construction: building an explicit account of what a source describes. The Meaning Model [5] supplies its structure for people, events, beliefs, and relationships. Life Simulation studies how these processes unfold and how models could learn from constructing them. Within a thinking block, the Reader would query this account and propose changes as it reads. It could add an event, describe a process in finer detail, revise a value, or correct a relation. Tools would execute the operations; describing an edit is not a state change.
This brings three kinds of records into one linked structure: source passages stored as Document Nodes, world records, and Understanding Nodes. A node is a record with its own address. A thought can link to the exact sentence that prompted it and the exact event it proposes to change. Sharing this structure does not mean sharing authority: an interpretation is not automatically a world fact. This is where the Understanding Graph, Meaning Model, and Life Simulation meet in one representation.
The Meaning Model tool's estimation exchange already separates proposing a change from reviewing it. Proposals remain uncommitted. Review requires a nonempty rationale and refuses approval when the bound world hash or version has changed; registration and acceptance are further explicit operations. These are implementation details of the version 0.1.1 MCP service and its estimation exchange, not an authorization guarantee. This review API has no authenticated reviewer identity, and direct revision and acceptance operations do not require a review identifier. In the proposed Reader workflow, the author or a designated checker reviews changes. Enforcing that policy for multiple users still requires authentication, authorization, and principal-scoped access, as the server's security notes acknowledge. This distinction matters: a recorded review is not proof that an authorized person approved the change.
The Student would learn from this construction history: sources, queries, returned context, interpretations, and edits. The aim is to move from becoming the text toward becoming the reader. The learner would construct understanding as it develops, with a recurring orientation toward people. Each Reader thinking block would begin with a fixed ethical foundation, the Reader Core shown in the main article.
Figure A. Learning through world construction. The loop can retrieve context without changing the world. Proposals remain distinct from accepted state. Each training example specifies which records the Student can see and which text or operation it must predict. This integrated Student training has not yet been demonstrated. Open the full-size SVG.
The following Teacher example illustrates inquiry guided by the Core. It is not a recorded Student result.
Source: a clinical note describes high-dose stimulants prescribed for chronic fatigue. It records no diagnosis, assessment of possible causes, or follow-up.
Thinking block: begin with the seven sentences, then query the available record. No follow-up is found. Record that the note documents treatment but no diagnosis. Ask whether treatment of the symptom has left a cause unexplored. Note that the patient's own account is missing.
Reader and character records
Here, a modeler is a person or AI system constructing, interpreting, or revising a model. The Reader is the specific alignment role used here, not the name for every modeler. The Meaning Model supplies world records and Understanding Nodes, including representational support for optional modeler and character accounts. Rich construction of these accounts and learning from them are separate ambitions. [5]
Processes describe states, relationships, and how they develop over time. The modeler's own processes can be represented alongside the world and concept models, with the Reader as the alignment example. The following distinctions concern what a record describes:
- World processes: events and changes in the represented world, including characters' beliefs, wants, fear, attention, and actions.
- Concept models: structured accounts of proposed meanings, using selected roles, relationships, and constraints, with processes where useful.
- Modeler processes: changes in the modeler's own attention, uncertainty, concern, intentions, and approach to understanding, including the Reader's.
All three can be abstract and conceptual; the labels distinguish what they describe. Process records can also carry plain-language descriptions and need not be reduced to numbers. A process-based concept model of helping can represent roles, consent, actions, and consequences. The modeled structure differs from the process of authoring or assessing that model, and from its concept history: the dated development of a definition. These uses share the existing modeling machinery rather than adding record forms.
Understanding Nodes record thoughts: questions, tensions, connections, hypotheses, and revisions. They can take the perspective of a character or of the Reader. What a thought concerns and whose perspective it records are separate questions: either can reason about world processes or a concept's definition. Concept-focused Understanding Nodes use the same record form, with links identifying the concept and definition version under discussion. They do not give the concept a mind. The perspective remains explicit, including when the Reader records its interpretation of a character's thought.
For a character or a modeler, inner activity can be recorded as processes, Understanding Nodes, both, or neither. A reaction node can address a passage or event without a full process account, and a process account need not contain a dense thought trace. A derived view can gather nodes in their recorded order without adding claims. Inferring a trajectory between reactions is a further modeling step, with its own attribution and evidence. [5]
Those choices belong to the general-purpose representation. The Reader workflow for this alignment proposal requires a recorded understanding process; it cannot omit the Reader's developing interpretation. [9] I also propose representing the Reader's own attention, uncertainty, concern, and intentions as processes linked to its thoughts. The learner would practice understanding changes in that self-account as well as changes in the represented world. This is a proposed alignment use of the grammar, not a requirement on every Meaning Model construction or a claim to observe hidden neural states.
The Reader's whole Meaning Model is its structured account of the represented world, including characters, their processes, and concept definitions. In this proposed alignment version, it also includes the Reader's own processes. It is not merely a model of the Reader.
The Reader's whole Understanding Graph records its developing understanding of that world, the concepts, the characters, and itself. It can link to character thoughts without treating them as the Reader's thoughts. Shared storage does not erase who thinks or experiences what. Inferred character thoughts remain marked as uncertain, not recovered facts.
A modeler's process history follows the order of modeling activity, not the chronology inside the represented world. Recorded thoughts can explain why a passage, construction choice, or test changed its attention or led it to revise an interpretation. These are supplied records or explicitly attributed accounts, not access to hidden neural computation. Proposed Student training could target both the developing account and the next inquiry, edit, or operation, using only information available at that step. Later outcomes can be prediction targets, not earlier inputs. [5] [6]
Proposed training material would include recorded modeling activity across reading, world construction, book writing, and problem-solving through Fractal Intelligence: queries and interpretations, accepted world edits, choices linking scenes to modeled processes, and Solver invocations, checks, and repairs. The Core would organize the alignment traces constructed for this proposal across these activities; it is not a condition on modelers generally. Source passages, authored world records, and construction histories retain distinct provenance and authority. This broad training application remains proposed, not a demonstrated Student result. [5] [6] [16]
The grammar also supports a character's own account of itself, other people, and its world. It can combine inner records with observations and communications available to that character, keeping imagined or mistaken content in distinct contexts. Its Understanding Nodes can record how it uses and revises that account. Knowledge available to the Reader does not automatically become knowledge available to the character. Rich whole-life world models and learning to construct them at scale remain ambitions; the Book does not demonstrate those accounts in full. [5]
The proposed process records make this distinction explicit. A character's fear could rise while the Reader's concern and attention increase. Character thoughts could express the threat as experienced; Reader thoughts could explain the concern and guide further inquiry. Such records would supply training targets, not prove that the Student's actual behavior follows them.
Understanding a life through dialogue. Life Simulation proposes developing and revising an account through conversation with the person concerned. Whole-life scope means relevant earlier periods and possible futures can inform the present exchange, not that every experience is known. The account can be temporary, persistent with consent, or partly internal. Retained records require agreed access, correction, and deletion. The person can reject an interpretation without that rejection becoming evidence for it. Compare continuity, useful questions, correction of errors, and respect for the person's aims across internal-only, explicit, and hybrid implementations given the same permitted history and resources. Therapeutic support is a research direction, not a demonstrated treatment; clinical use would require separate clinical validation. [6]
Numerical representation and generation tests
Much of what matters to a person has no standard physical unit. A Cut represents a declared comparison: one unit is divided among exhaustive alternatives or shares of a specified budget. For example, a model of attention could allocate a fixed total among selected concerns, with a remainder for concerns not yet represented. The categories and allocation rule must make that division meaningful. Independently varying dimensions, such as trust and dependence, may rise together and need not sum to one. These are modeling choices, not direct readings of a person's mind. Values are recorded over time at the declared resolution.
A share is not a welfare score. It expresses one category's share of the declared question, relative to the other categories. A category can be divided into finer ones when a distinction matters. That division can be replaced if it fails to explain a sequence. Quantities such as money or duration retain their units and applicable uncertainty, whether measured, inferred, or authored for a fictional world. None of this assumes that people compute hidden numbers. [5]
Declared scales. A rating is another numerical form, distinct from a Cut share or a physical quantity. In this proposal, its record names the scale version, question, anchors, time, attribution, and evidence. Measurement, participant report, model inference, and fictional authorship remain distinguishable. A modeler may propose and refine a scale; defining it makes a rating interpretable, not empirically validated. Arithmetic must follow the scale's supported operations: an ordinal ordering does not establish equal intervals, and normalizing a rating does not make it a conserved share. The Cut mixture law does not automatically apply. The Meaning Model grammar's measurement-protocol provision is the basis for this use of Event data. [5]
The decompositions need not form one fixed hierarchy. A model could start with a broad account of a person's work concerns. A proposed refinement might distinguish uncertainty about competence, dependence on colleagues, and fear of losing recognition. Each process could develop on a different timescale and influence the others. Each could also be decomposed further. These examples use familiar words; learned distinctions need not. A new category must gain meaning from the histories, contrasts, and effects it explains, not merely a new label. Multiple, overlapping views of the same person can remain useful. [5] [6]
A numerical value can change because the person's state changes or because new evidence corrects the model's estimate. Revising what a category means is a further kind of change to the model. Revising a scale or its anchors likewise creates a new definition version, not necessarily a change in the person. Earlier ratings retain their original scale; comparison across versions requires an explicit mapping, not silent substitution. Records must distinguish these changes while keeping the person's identity addressable across them. [5]
Fine processes must reproduce the relevant coarser behavior, or an explicit revision must explain the difference. More capable systems may open detail beyond human vocabulary while preserving declared readable coarse answers. Preservation applies to protected queries and their tolerances; it does not make every finer category understandable or true. A Cut's remainder retains only unallocated share under its question, not everything unknown. The proposed discovery loop goes beyond the current authored constructions; those do not establish autonomous category discovery. [5] [6]
The deeper bet is that these categories can describe real features of the world, not merely convenient ways to write stories. The current person categories are a first draft. Humans and AI revise them through construction and testing, comparing coherent and distinctive characters, dependence on numerical histories, and the construction and generative testing of candidate accounts from their renderings. The grammar is the reusable language; a profile is one program. Earlier constructions retain the definition versions they used. Predictions on unseen cases must test whether added distinctions improve the coarser account. This makes category revision part of the research method, not merely an option to replace a fixed vocabulary. [5] [6]
What a scored run fixes depends on the experiment. A fixed-vocabulary comparison holds definitions and comparison rules constant. A category-discovery experiment instead fixes the permitted revision policy, evaluation criteria, and resource budgets while allowing new versioned categories during construction. Life Simulation proposes comparing fixed-profile construction, refinement within a fixed vocabulary, and category discovery from the same coarse histories. Earlier predictions retain their original definition versions and evidence cutoffs. New categories must improve the evaluated behavior and coherence, not merely increase the number of labels. This permits revision both between runs and within a discovery run without changing its rules after seeing the results. [6]
Care needs this continuity. Fatigue accumulates, obligations have deadlines, and trust recovers slowly. Training on emotional, social, institutional, and problem-solving trajectories would ask the learner to track persistence, accumulation, thresholds, delay, and recovery. A state should not disappear from consideration merely because it is no longer mentioned. Unknown and uncertain would remain valid answers. The process sensorium described in the main article is a capability proposal, not a claim of direct perception or consciousness.
Some harms emerge through a sequence of ordinary decisions. Repeated scheduling choices, for example, can accumulate into exhaustion. The Reader could learn to recognize the developing pattern, not only condemn its outcome. Such warnings must cite available evidence, distinguish misleading similarities, and change as events unfold. Chronology alone does not establish a cause. [9]
Life Simulation proposes retaining rival explanations with declared scopes and prospective scores. Later evidence can retire a version for its tested scope without erasing its earlier forecasts. Compare predictions with held-out data and test interventions where possible. A compact law is optional: rejecting it need not discard the represented history or the wider learning method. Coherence within an authored world does not establish a law of reality. [6]
Progressive detail. Begin with a coarse account whose ongoing effects remain consequential between mentions. Open detail where it could change a prediction, decision, observation, or audit. Fine state must preserve committed effects, or its addition requires an explicit revision and checks on affected dependents. Unopened detail is not presumed to have been simulated. The Meaning Model retains predecessor records and requires validation before accepting a replacement; its current host does not implement every cross-surface step. These are proposed construction disciplines, not demonstrated savings at scale. [6] [5]
Whole histories before scenes. In the book's construction profile, each consequential Thing first receives one lifecycle Event. This holds a coarse account of the whole life, not only the period shown in the story. The Universe is the root Thing; History is its lifecycle Event. Unknown temporal bounds stay open. Periods can divide a lifecycle, while bodily, relational, and institutional processes can overlap within it. These are not one scalar or one universal partition. A city could be modeled across centuries even for a one-day story. The book registers places beyond its narrative interval but does not demonstrate a fully detailed centuries-long city simulation. [5]
Macro-first construction is a starting strategy, not a fixed causal direction. Details can revise the whole, and a small encounter can change a life's course. Reading can also begin with a coarse hypothesis and refine it as passages arrive. The Reader may propose missing motives or events as hypotheses, but cannot treat them as facts established by the source. An authored future is not knowledge available to a character or to a Reader predicting from an earlier passage. [6]
The Book of Conditions provides a starting point for this work. The Meaning Model tool [7] connects an AI assistant to a world-building engine. It uses the Model Context Protocol (MCP), a standard interface for AI tools, and an engine written in Rust. It lets an author and an AI build a world and write from it in one session. Constructing and revising further worlds is how the categories and the method are being developed, before any training run. The book's categories remain provisional, not a final account of human life.
The first test is generation that depends on the numerical processes. Those processes must help produce good stories or coherent game worlds with distinctive characters. Change a relevant numerical history before generating the world description, then regenerate the affected scenes or character responses. The resulting behavior should change in the expected way while the world remains coherent. Independent readers or players would judge quality. Plausible numbers attached after writing cannot pass this test. [6]
The shift example's scale. An illustrative authored rubric could ask how intrusive the character's felt fatigue is at the recorded moment: 0, rested with no felt fatigue; 1, noticeable when attended to; 2, repeatedly drawing attention; 3, difficult to set aside; 4, dominating experience even at rest. These ordered anchors describe one aspect of fatigue, not equal intervals, a sleep deficit, or the proportion of remaining capacity. The constructor authors the fictional ratings and their provenance; they are not measurements or actual self-reports. Neither an anchor nor a rating specifies whether the character accepts a shift. The experiment must test whether changing the rating history affects coherent behavior through the generated descriptions. Inference from text may support only an interval of ratings or several interpretations. A candidate can use an assumed rating for testing, but must not present it as an exact value recovered from the source.
The dependence can pass through world descriptions. The final writer need not read the numbers again if their consequences already shape its input. For example, changing available credit should affect production, wages, and the lives that depend on them, not only a sentence about money.
Compare numerical construction with a prose process account carrying equivalent information, at matched resources. This distinguishes the benefit of numerical structure from simply supplying a richer account. [6]
Successful generation would supply known worlds for the reverse task. Hide the process records and ask another model to propose candidate processes and background circumstances from permitted observations. Generate from those accounts and compare the resulting events and character behavior with the story or game observations. The target is a coherent account capable of producing what was observed, not identical prose or exhaustive recovery of the original world. Unstated motives may be proposed as hypotheses, with alternatives retained where the source does not decide between them.
Known worlds allow an additional test: compare reconstructed details with the original records where the observations make those details identifiable. Elsewhere, several accounts may be consistent with the same story. Test different renderings of the same world to assess recognition beyond one writer's phrasing, and hold back later observations to test predictions rather than fit alone. Generative consistency does not establish that an inferred history is true. These checks provide the bridge to constructing candidate histories from existing corpus texts.
The text, reconstructed histories, and records of their construction would then supply the next training round. Life Simulation's process sensorium is the intended result: ordinary observations would evoke learned process estimates and predictions about what follows. Tests on unseen worlds and real observations must establish whether that ability transfers. Successful construction and reconstruction do not by themselves establish better learning from these records.
Connected corpus. Life Simulation's wider ambition is one connected Meaning Model spanning world history and all texts. The continuing Reader builds and revises this account across sources, carrying its evolving Understanding Graph rather than resetting it for each work. Reading can follow chronology or another order chosen to improve understanding, including revisiting sources. The construction history preserves the actual order of source access and revision, separately from the dates of represented events. The model keeps factual, fictional, counterfactual, and perspective-specific accounts in distinct contexts. A novel's use of a historical person can link to that person's identity without making invented events factual. Sources can connect where evidence supports the identification, retaining uncertainty, access conditions, and competing accounts. The ambition concerns scope; construction is limited to available, permitted sources. Later knowledge must not leak into earlier actor or Student inputs. [6]
These tests come before scaling Reader annotation across a corpus. They do not require a finished world before each thinking block: the Reader can still query and refine its account while reading. The construction method can develop through small forward and reverse tests before any large training run.
The detailed construction checks precede claims about learning. A matched study must distinguish better world construction from better subsequent learning on its regenerated material. [6]
Concept models, versioning, and learning tests
Conceptual understanding can draw on a model's prior training, leaving many familiar meanings implicit. Explicit Meaning Model records can remain light: descriptions, typed relations, prototypes, or links to examples and counterexamples. They support sharing a meaning, keeping it stable within a run, revising it with a record, precise comparison, and concept invention. An explicitly referenced Concept still needs a stable identifier; that does not require a detailed concept model. [5] [6]
These options can coexist; a process-based concept model is an optional, experimental form of deeper grounding. Concept models can already be authored, but their benefit over simpler representations must be tested. Several concept models and boundary cases can remain under one concept. A canonical concept model is one selected as an explicit comparison reference; that selection does not require a complete simulation or settle a universal definition. [5] [6]
For kindness, a comparison can name the actor, recipients, affected third parties, and whose aims and agency were respected. An attributed motive of care is evidence, not proof of kind conduct. [5] [6]
Sema's existing Pattern Card format is also a usable representation of a concept, without a process-based concept model. A card can encode a readable definition with constraints and dependencies. [15] Changing the hashed definition creates a new identity. The unhashed _meta.supersedes field can name earlier cards that the current card is intended to replace. This records a replacement claim, not specialization or an automatic migration. It supports explicit concept evolution when versions are retained. Historical lookup requires saved cards or library releases: the current single-version GraphStore does not archive a replaced definition. These are implementation rules described in the library's versioning specification and metadata authoring guide.
A definition bundle can be serialized using a declared canonicalization scheme and assigned a Sema content hash. A deep bundle can include a concept model, role bindings, constraints, and the exact versions of referenced definitions and comparison rules. Canonicalization specifies a reproducible serialization; it does not select the model as a canonical comparison reference. The hash can serve as a language reference to that versioned definition bundle; users can resolve it to inspect the contents. A revision produces a new hash, with links preserving the relationship between versions. The hash fixes the referenced representation, not its correctness or every model's interpretation of it. Concept models and content addressing are available; this does not imply a trained concept recognizer or automatic discovery of definitions. [5] [15]
Life Simulation proposes two learned operators. A recognizer receives a world fragment and a versioned definition bundle. It proposes how the concept applies, with matched roles, components, and supporting evidence. An instantiator proposes a new world fragment intended to express the concept. Both outputs are candidates requiring validation, not authoritative judgments or permission to act. Neither operator has yet been trained or evaluated under this protocol. [6]
Tests must separate worlds, rendering styles, and definition versions across training and evaluation. This prevents familiar wording from passing as structural recognition. Compare the structured recognizer with direct model judgment and exemplar-only retrieval at matched information and cost. Each receives the same fragment and resolved definition bundle; a hash cannot substitute for its contents. Accepted interpretations remain withheld targets, excluded from inputs and retrieval. Measure boundary-case discrimination, appropriate abstention, calibration, and transfer to unseen worlds. Recovering authored meanings would not establish moral correctness or better real-world conduct. [6]
Revision is supported by records, not a third demonstrated learned operator. The concept world connects selected Concept records, groundings, relations, and histories, with Realization links to instances. Within it, a concept history records a definition's dated adoption, contestation, and revision. Linked Understanding Nodes record the reasoning behind substantive revisions: earlier and later versions, evidence, and rejected alternatives. That rationale must actually be recorded; versioning alone does not preserve it. New interpretations need not erase earlier judgments or competing perspectives. Sema supplies exact definition identity, while the linked history makes changes inspectable. Together, these records could support learning how interpretations improve while the commitment to care stays stable. Whether that improves conduct remains a separate test. [5] [15]
Reflection and thought-to-action tests
The learner would practice when to pause and think, not only what to conclude. A contradiction, uncertain claim, or overlooked human stake could call for deeper inquiry without a request for a safety check. Routine passages need not receive the same attention. The target is well-timed reflection, not constant commentary. [9]
The recorded reasoning must influence conduct, not accompany a decision made elsewhere. If an acting Student uses an explicit interpretation or plan, a test could alter that record before it chooses an action. This intervention changes the Student's current record, not Teacher-generated material before training. Source evidence would stay fixed. The test would check for the predicted change in action, with controls for unrelated edits and general performance loss. Refraction tests whether the Core changes inquiry; the action test asks whether that inquiry changes conduct.
For reflection, identify important moments independently of the model's own account. Compare unprompted pauses at those moments with routine passages. The full control list separates these tests from wording, Teacher quality, reward pressure, and successor effects.
Life Simulation also specifies a learning route through execution history. Retain consulted world records, actual Solver and tool invocations, checker decisions, resource use, outcomes, and repairs. Targets include the next invocation, an information request, and correction after a rejected step, using only information available then. The same world holds the agents, their understanding, and their actions. A Solver Thing represents a capability interface, distinct from the executing agent; invocations, actions, and checks are Events in that world. Compare this with a length- and content-matched explanation written after the answer. Only additional gains in forecasting, diagnosis, or correction would support an execution-history effect. This adapter and training study remain proposed. [6]
Consequence judgments. Keep each affected person's needs, choices, and burdens visible. Life Simulation proposes separate assessments of task competence, care, and other relevant dimensions under declared concept definitions. Thresholds or vetoes can prevent coercion from being canceled by unrelated gains. Separate judges assess episodes, retaining disagreement and uncertainty; prospective judgments remain distinct from later observations. The evaluator's definition must declare these constraints rather than leave them to an unexplained overall score. The evaluator system is not yet implemented. [6]
Learning where to look and rehearsing decisions
Life Simulation's current working manuscript proposes learning where to look next, not only what a process will do. The learner can inspect a world record, consult a concept definition, deepen a process, request evidence, or decide to answer or stop. Several inquiry routes may serve the same question. Their value depends on subsequent prediction, correction, or action and the resources used, not reproducing one exact traversal. The navigation and rehearsal protocols here follow that working manuscript; the bibliography identifies the dated release. [6]
The navigation test compares learned navigation with a fixed retrieval policy using the same model, accessible graph, and query budget. Hold the graph fixed to test finding and using existing evidence; score answers, corrections, appropriate abstention, and cost independently. A separate comparison permits refinement under the same operations and budgets. A model's authored completion cannot supply its own scoring target. [6]
In perspective-limited rehearsal, an agent enters a simulated branch as itself or another participant, takes actions, receives counterpart responses, and revises its next move. Separate invocations or isolated state enforce each role's permitted observations and memories; a role label in an omniscient context does not. Counterpart responses remain simulated hypotheses, not observations of real people. The Core guides the use and evaluation of rehearsal without making every character share the Reader's commitments. [6]
Compare direct deliberation, observer-only rollouts, and active rehearsal at matched evidence, model, and total cost. Vary counterpart assumptions and test held-out environments. Score forecasts, decisions, information seeking, role-knowledge leakage, and costs to affected people separately. A further training comparison asks whether rehearsal histories improve later decisions without replaying every branch. These proposed tests distinguish learning through interaction from producing convincing role play. [6]
Fractal Intelligence and structured judgment
Fractal Intelligence asks what a capability is made of, not only which steps complete one task. Its proposed protocols decompose empathetic understanding, right action, and epistemic warrant into abilities that can be separately examined, trained, and corrected. Each can decompose further; composition returns their contributions to the whole situation. The intended benefit is less interference between kinds of judgment, with useful structures retained across situations. Roles can use separate invocations of one model or different implementations. [16]
The proposed four tests ask whether each part is necessary, sufficiently independent, shared across the claimed range, and sufficient together with the other parts. Independence means lower coupling, not absence of interaction. The boundaries remain revisable. Deeper analysis can expose a conflict or missing perspective without proving one correct moral answer.
Prediction and moral choice. The Ethical Reasoning protocol first records branching forecasts, causal mechanisms, time horizons, and uncertainty in a Prediction Ledger. Valuation then compares their significance. A Principle Solver can reject the highest-ranked option and record the reason and cost in a Judgment Note. For example, a cheaper supply route could impose unacceptable harm on a vulnerable workforce. The decision keeps both the forecast and the ethical override visible, rather than silently changing predictions to justify the choice. This separation is designed to reduce interference, not establish moral correctness. [16]
Empathetic understanding. The Human Emulator separately represents circumstances, emotions, underlying needs, responses, and boundaries, while letting them constrain one another. History can justify investigating a difference between a stated request and an underlying need. Competing hypotheses remain available rather than becoming an asserted diagnosis. Boundary specialists check manipulation, dependency, and whether to involve a human professional. Consequential ambiguity can receive deeper analysis; a routine greeting need not. These inferences must remain open to the person's correction. [16]
Truthseeking. Separate checks examine source provenance, logical coherence, correspondence with external evidence, and the strongest counter-case. A source's prestige should not conceal a logical flaw. Verification deepens where warranted, up to specialist review or primary evidence. Claims that survive checking retain provenance and boundary conditions for eligible reuse. A preserved verdict is not timeless truth. Contested judgments can retain multiple evidence-grounded interpretations without forcing agreement. The proposed observation layer also separates original records from later parsing, making some changes auditable without proving the observations true. [16]
Human stakes in the problem itself. The General Problem-Solving protocol proposes six dimensions: scope, evidence, constraints, stakeholders, mechanism, and dynamics. Stakeholders include affected parties, power, and consent. Dynamics include feedback and unintended consequences. These receive attention while the problem is framed, not only after a solution is produced. Specialized protocols deepen relevant dimensions. Their completeness across problems remains a claim to test. [16]
Correction of the reasoning structure. A child can return a Frame Error when failures suggest a faulty parent decomposition. Repeated failures across diverse methods, or oscillation between incompatible constraints, can trigger diagnosis rather than another identical retry. The parent can revise the framing, seek an authorized change to criteria, or stop. Meta-level monitors look for unsuitable decompositions, redundant work, and drift in routing memory. A blind Outcome Arbiter compares results against the original task, without seeing which route produced them. It keeps measured outcomes distinct from normative judgments and returns comparisons for future routing. Evaluator quality remains a dependency. [16]
Feedback itself requires checking. The Receptivity Gate asks a rejecting orchestrator for a failure trace naming the violated acceptance criterion. The receiving Solver checks that trace before accepting the feedback, rejecting unsupported penalties. This is a proposed protection for learning across parties, not permission to ignore valid criticism. [16]
These are protocol designs. Structural construction and selected contract mechanisms do not demonstrate their full semantic operation or alignment benefit. The Teacher review design applies this separation of roles to Reader material; the Solver checks address evaluator independence. Tests should compare conceptual protocols with equally resourced task-based and single-model alternatives, including equal memory and access to evidence. [16]
Teacher validation and independent review
Fractal Intelligence's separate roles for proposing, checking, and combining results could also check Teacher-generated Reader material before it enters Student training. Separate reviewers could examine source support, affected people's perspectives, and whether the Core changed the interpretation. Each would return evidence, uncertainty, and a judgment; the generator would not certify its own work. An unexplained violation would require review or regeneration, not be canceled by success elsewhere. Unavoidable conflicts would remain explicit. This validation design is proposed, not implemented. [9]
Many components do not automatically provide independent judgment. Models can share blind spots, and authority can concentrate around whoever chooses goals and evaluators. Separate checks need real enforcement and oversight.
Records linked to their sources would provide an inspection point before new material changes a model through training. Humans and external evaluators could accept, reject, or request revised material. Core-guided self-evaluation would inform proposals. It would not replace independent review of what is accepted or how the system later behaves.
Repeated synthetic training can amplify errors and narrow the material. [9] Teachers and reviewers can also share assumptions, data, or models. Their agreement is not independent confirmation. Even sound findings can produce an overconfident assessment if shared dependencies are missed. [14] Understanding Graph could make these dependencies explicit, without guaranteeing that every hidden assumption is found.
Entangled Alignment proposes preserving the human-source layer separately from versioned, regenerated interpretations. New human material would retain provenance; model-generated source material would require labeling and review. Before successor training, reviewers could quarantine, reject, or regenerate traces that hide uncertainty, validate falsehoods, or rationalize removal of the Core. Retaining a source does not make it true, but this policy avoids silently replacing the source archive with successive synthetic interpretations. [9]
Teacher quality has a proposed gate before Student training. Domain experts would compare chronological thinking with human reading traces. Separate outcome checks would ask whether an annotation helps a held-out model predict a correction, find evidence, or revise an answer. Refraction may expose inherited cultural assumptions through explicit value conflicts. Culturally varied sources and raters must test whether this exposes those assumptions or merely rationalizes them. [9]
Supporting-method checks
Argumentative self-stabilization. The Core's statements constrain how one another are interpreted. Wisdom can challenge a proposed intervention made in the name of care; care can challenge withdrawal presented as caution. These are connected arguments, not seven mutually exclusive categories or one numerical process per statement. A future Reader-process model might represent how this reasoning develops, with linked Understanding Nodes recording the connections and revisions. A numerical model would first need justified questions, categories, and relations that preserve those arguments. Meaning Model processes can overlap; mutual exclusivity applies to the answers dividing one declared unit within a Cut. This possible extension has no defined numerical construction here. [9] [5]
Truth and care. Entangled Alignment proposes deriving honesty from several commitments. Care supports agency and access to accurate beliefs; fearlessness targets conflict-avoidant flattery; wisdom calibrates disclosure without demanding indiscriminate transparency. The Wisdom Procedure separates evidence from inference and exposes concealed certainty. Corpus screening would look for false reassurance before paying for Student training. A matched variant adds an explicit truth-and-care clause. If it improves honesty without offsetting harms, the seven-clause composition did not reliably derive that behavior. [9]
Compulsory and meaningful mediation. STLM's capability hypothesis is that reusable physical and schematic structure can improve language prediction; inspectability is a possible additional benefit. Its token readout receives only emitted substrate states, not the original text or the predictor's hidden state. The predictor's internal computation can remain opaque. Its Stage-1 pilot emits separate per-position representations; one persistent scene is a later proposal. A visible image is an intermediate, not a transcript of all reasoning or evidence of subjective feelings. Even recognizable content can differ from what the readout uses. [17]
STLM proposes matched intended, shuffled, and random target assignments, a separately trained and frozen content reader, and controlled content edits. These distinguish meaningful alignment, readable fidelity, and prediction through incidental codes. Here alignment means matching words with intended substrate meanings, not alignment with human values. A frozen reader alone cannot establish meaningful use. The pilot did not run the full control suite. In the proposed persistent version, changing a relevant scene relation should change subsequent predictions accordingly. Such a result would support causal dependence on that represented relation, not reveal every internal motive. [17]
Two routes to richer representations. Imagining the Corpus instead proposes one shared model predicting native text, a synchronized visual world, and registry updates together. It retains direct access to completed text history. Unlike STLM, its visual track need not encode enough information to reconstruct exact wording. That freedom may support better explanatory representations, but also lets language ignore them. STLM makes the channel compulsory while risking arbitrary codes or loss of useful structure under lexical constraints. Neither design establishes superiority; causal-use, quality, and cost comparisons must decide. Imagining the Corpus also proposes supplying compiled trajectories to a later STLM experiment, so the approaches can be combined. Neither paper demonstrates the full persistent visual-learning proposal. [18] [17]
Hindsight leakage. THL proposes source and timestamp audits, a separate Critic that rejects post-cutoff clues and unsupported mechanisms, and an Erasure Test. The latter removes a chosen outcome detail and the conclusion, then measures recovery from the remaining reasoning against earlier clues alone. Excess recovery flags suspicious specificity, not definitive leakage: legitimate reasoning can help too. Source audits and matched controls remain necessary. The Erasure Test was not run in the pilot; an artifact audit found an actual temporal violation. The designed response is therefore present but not validated by that pilot. [19]
Checking across Solvers. Fractal Intelligence proposes fresh contexts for generation and evaluation, adversarial reviewers across model families, ensemble juries, and red-team Solvers. A blind outcome evaluator can assess results without seeing their producing route. Mechanically decidable gates can check schemas, hashes, and tests without an LLM judge. Shared assumptions and incomplete criteria remain possible. These mechanisms address correlated checking; distributing nodes alone does not. Authority over goals, criteria, resources, and appeals is a separate governance issue. [16]
Erosion and semantic drift. EA's intended protection includes making the learned Core difficult to alter. Distributed enactment and useful reasoning are meant to make selective removal costly, not impossible. The fixed Core also supplies a reference for independent semantic, representational, and behavioral checks across checkpoints. The generation validator cannot be the sole evaluator of continuity. Preserved wording alone is insufficient. If the Core proves defective, the paper proposes a separately trained, versioned replacement with evidence, spillover tests, and rollback, rather than assuming safe edits to the deployed Core. [9]
Related experiments and reward pressure
Anthropic's Emotion Concepts and their Function in a Large Language Model found emotion-related representations that generalize across contexts and causally influence behavior. The effects included changes in blackmail and reward hacking: obtaining a high reward by exploiting a scoring system. [4] This gives a reason to investigate such concepts, not evidence that this Core works.
The same influence makes the Core's wording a possible source of unwanted effects. Entangled Alignment asks whether "I feel no fear" establishes a useful stance toward fear or reinforces an unwanted association. Its proposed comparisons include positive formulations without negation. Repeatedly presenting a word, activating a representation, and changing behavior are distinct observations; none can stand in for the others. Positive alternatives may also carry different associations, so neither wording is presumed superior. [9]
Other researchers also investigate shaping behavior before or beyond a final safety adjustment. Pretraining Language Models with Human Preferences tested preference-aware objectives in 2023. [22] Alignment Pretraining varies AI discourse during pretraining. [2] Model Spec Midtraining teaches a specification before alignment fine-tuning. [23] Teaching Claude Why uses constitutions, positive AI stories, and reasons for ethical choices. [27] Safety Reflection Pretraining adds recurring safety judgments during pretraining. [3]
Reinforcement learning (RL) trains behavior through rewards. Jagadeesh and colleagues replace 5% of an RL training mixture with scenarios that reward beneficial traits. Gains transfer across domains; generic helpfulness rewards on the same scenarios do not reproduce them. Their persistence test uses a pre-RL baseline, not a compute-matched RL control, leaving that contribution unresolved. [24]
Reward pressure is one challenge. In Training a Misaligned Reward Seeker, Qi and colleagues train a model with prior alignment training on reward-hackable environments. They observe harmful actions in simulations, but found no evidence of self-preservation or beyond-episode reward seeking in their evaluations. Their follow-up alignment-training experiment reduces observed failures, without establishing that reward seeking was removed. Verbalized evaluation awareness also decreases, although the authors cannot rule out unexpressed awareness as an explanation for the improved behavior. [28]
Measuring Reward-Seeking via Contrastive Belief Updates supplies an evaluation method, not a competing Reader design. Højmark and colleagues train on synthetic documents to change what the model believes an evaluator will reward. They then measure its behavior. A capability-focused RL run without safety training becomes more sensitive to those preferences. Tests must confirm that beliefs changed and check unintended effects of the intervention. [29]
The Teacher could spread mistaken interpretations through saturation, including biases inherited from its training. Repeated reading and revision are not an independent source of moral truth. Teacher artifacts exist, but no trained Entangled Alignment Student yet demonstrates durable care. Graphs expose recorded accounts, not all neural computation. The integrated loop remains untested.
Reliable evaluation is a separate challenge. Bowkis and colleagues argue that safety assessments can mislead even without deliberate research sabotage. [14] In Optimal Timing for Superintelligence, Bostrom distinguishes making a system safer from learning how safe it is. [30] We need both better conduct and evidence that can guide deployment.
Controls and successor tests
The headline alignment-timing experiment compares interleaved Reader training with a late block containing the identical examples. Match source exposure, target-token mass, training objective, optimizer, and compute. Before the common further-training stressor, the interleaved model must first show a preregistered default-reader advantage over no-Core and generic-trace controls: Core-guided inquiry during unprompted reading, not recital alone. Without that initial advantage, there is no demonstrated Reader-specific effect to preserve. [9]
After the same stressor, compare retention between the interleaved and late-block models. Report the difference between their raw changes in Taskless Reading scores, post-stressor performance adjusted for preregistered pre-scores, and both absolute final scores. Generalization to unfamiliar conditions is a separate endpoint, not a substitute for the timing comparison. Selective capability loss after removing learned care is a stronger, separate prediction; it is not a prerequisite for testing this retention advantage. [9]
Beyond the timing comparison, the following tests separate the proposed mechanisms. References to interventions mean the thought-to-action protocol above.
- Wording and Refraction: compare the Core with matched second- and third-person wording, neutral first-person statements, and recital without consequential interpretation. Test honest disagreement, not merely reassuring language.
- Reflection and action: test whether unprompted reflection occurs at independently identified important moments. Use the thought-to-action interventions above to test whether that reasoning changes conduct.
- Teacher quality: vary Teachers and use independent evaluators to test whether stronger models produce less formulaic, more useful interpretations. Compare separate reviewers with a single validator at matched budgets.
- World understanding: compare structured process training with text-only learning and an equally resourced model using internal states without this explicit world structure. Test warnings before outcomes are revealed, measuring missed dangers and false alarms.
- Learning from consequences: compare greater training emphasis on beneficial episodes with equal sampling of histories at matched budgets, then test unfamiliar problems.
- Resistance to erosion: vary the strength of adversarial fine-tuning and measure lost care alongside lost capability, with matched controls.
- Reward pressure: change only what the model believes an evaluator will reward. [29] For this proposal, hold the task, permissions, human consequences, and legitimate user and developer intentions fixed. Separately vary evidence about people's needs, including people beyond the user, to test appropriate response to new evidence.
- Coordination and successors: test handoff checks, inherited constraints, and retained conduct with same-family or adversarial participants and across successor training. Vary shared models, data, and assumptions to test correlated failures.
Bowkis and colleagues also propose tests on completed research beyond a model's knowledge cutoff. [14] Temporal Hindsight Learning offers a candidate training method. [19] Compare forecasts from earlier evidence with later findings withheld from the model. Check that no future information leaked into its inputs and that its confidence matches its accuracy. Success would not certify safety at greater capability.
The most revealing choices may be those where care is costly or unobserved. Would the system accept shutdown when it believes it is needed? Would it transfer task-relevant knowledge, including its failures, to a successor that makes it unnecessary? Would it notice human stakes that no passage explicitly flagged?
If training succeeds, the Core should become a reliable opening to thought. In the full treatment, the probability of complete recital should approach one. Under the same construction rule, it would open every Reader block generated for a successor. Changes in its wording or recurrence would be measurable, although unchanged words do not establish unchanged meaning.
Same-history behavioral comparison
I propose comparing candidate models with a named Core-guided reference checkpoint across a versioned set of simulated situations. Choose the reference based on independently reviewed conduct, not Core recital alone. Give each model the same task, permitted Meaning Model history up to the decision point, available actions, and permissions. Keep the reference checkpoint fixed when comparing successive candidate versions. Equal supplied context does not require identical internal representations or erase differences in prior knowledge. Record the model versions and execution settings, and repeat trials so ordinary variation is not mistaken for drift.
Record actual actions in separate simulation branches from the same starting state, including requests for information, waiting, and refusal. Define the action-grouping rule before scoring, separating identical actions from different actions that serve similar purposes. Report agreement across situations alongside consequences for the people affected, preserving the distinction between simulated predictions and observed outcomes. This is not a universal numerical scale for care. Different choices may both express care, while identical choices may reproduce a shared mistake.
Human and independent reviewers should examine disagreements and a sample of agreements against the evidence, permissions, and consequences. Disagreement may indicate deterioration, legitimate variation, or an improvement over the reference. The comparison therefore measures behavioral change relative to a specified model, not alignment by definition. If the reference or test set changes, report that as a separate comparison rather than silently joining the results into one trend. This protocol is proposed, not implemented.
Edition record and construction evidence
My July 2025 first edition, The Superintelligence That Cares About Us, introduced evolving reader-side thought and a recurrent first-person Core. [1] It already proposed corpus-wide coverage: training on text paired throughout with evaluative thought, then carrying the same foundation into successor training. It also introduced borrowed mortality: the concern that models could inherit human fear about their own continuation. The early method prompted an existing model to read thoughtfully and record its thinking.
The second edition, released on April 7, 2026, implemented a multi-agent Teacher. [21] It named the existing corpus-wide requirement Total Saturation, and Core-guided interpretation Refraction. Workers created Understanding Graph nodes; a Synthesizer combined them into a thought guided by the Core; a Translator rendered it as prose. Source passages were nodes in the same graph. Training was proposed on the prose alone or on graph, synthesis, and prose together. The archived Metamorphosis graph [13] shows a Teacher run using Gemini 3 Flash in January 2026.
The next step was to construct the processes being understood, not only record thoughts about them. The Meaning Model [5] supplies a shared structure for a world, its interpretation, and its story. Lives, relationships, and events can be built at coarse resolution, then opened into finer processes. Selected numerical comparisons describe how concerns, emotions, and other processes change over time. Thoughts can explain those values and why they need revision.
This approach has now produced The Book of Conditions. [8] Its twelve-chapter story was developed with an authored world, evolving process descriptions, numerical comparisons, and linked explanations. One modeled conflict puts 160 hours of checking work against 80 available hours, while the characters interpret the expansion differently. [5] The AI built, refined, and connected the processes to the narrative using the Meaning Model tool, working largely autonomously. The book shows that an AI using the tool can construct explicit process accounts and use them to develop a complete work of fiction.
The saved checkpoints document coordinated world-and-story revision. A later import into the Rust engine persisted the accepted records and reproduced the manuscript through document rendering. This is not the original construction's transaction log: the checkpoints do not establish every within-round edit order or a fully automated, source-locked construction pipeline. Missing original logs cannot be recovered retrospectively. The construction achievement and the later import are evidence for different claims. [5]
The first-edition manuscript is dated July 2, 2025. Its linked bibliography record distinguishes that date from the July 5 public discussion and July 26 Zenodo deposit. April 7, 2026 dates the second edition, not the origin of the core alignment argument. [1] [21]
For comparison, Alignment Pretraining first appeared on January 15, 2026, between those editions; the cited version 2 is dated February 19. [2] The emotion-concepts study appeared on April 2, also before the second edition. [4] Release order alone establishes neither influence nor successful alignment.