https://chatgpt.com/share/6ab7f2ad-b7a0-83eb-9507-08b9864252b2
https://osf.io/y98bc/files/osfstorage/6ab7f247175aacf8ed3c3b23
Preregistered Study E4: Purpose Belt Ablation
Testing the Functional Irreducibility of Purpose Identity, Interpretation, Revision Attribution, and Hierarchical Latching
Study ID: WF-E4-PB-v1.0
Programme: The Science of World-Formation
Document Type: Confirmatory Preregistration
Version: 1.0 — 2026
Primary Target: Functional necessity and minimality of the Purpose Belt
Geometry: Explicitly out of scope
Abstract
This preregistered study tests whether an explicit Purpose architecture contributes behaviourally irreducible capabilities beyond those available to matched goal-directed, memory-bearing, and generic self-revising agents.
The study focuses on four candidate components: Purpose Identity, Purpose Interpretation, Revision Attribution, and Hierarchical Latching. These components are tested under long-horizon environments involving reinterpretation drift, ontology shift, factual surprise, misleading evidence, adversarial reframing, and heterogeneous causes of failure.
The central claim is deliberately narrow. The Purpose Belt is not assumed to make an agent generally more intelligent, more moral, or more capable on short tasks. Its proposed function is to maintain a persistent and auditable separation between what the system is trying to preserve, how that Purpose is currently interpreted, what the system currently believes about the world, what has actually happened, and which level should be revised when discrepancy occurs.
The source development identifies four especially important ablation predictions. Removing persistent Purpose identity should permit long-horizon reinterpretation drift. Merging Purpose interpretation into ordinary world-model state should increase factual–normative confusion. Removing Revision Attribution should increase wrong-level revision. Removing hierarchical latching should increase oscillation or drift under noisy and adversarial evidence. If these distinct failure modes do not appear, the Purpose Belt decomposition has not justified itself. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
The study also includes a strong conventional baseline containing persistent memory, hierarchical objectives, self-reflection, and meta-revision. If this simpler architecture reproduces both the action behaviour and revision behaviour of the full Purpose Belt within preregistered equivalence margins, the strong architectural claim is rejected. This directly implements the source programme's strongest minimality criterion. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
1. Study Rationale
The Purpose Belt hypothesis emerged from a broader question in World-Formation Theory:
How can a self-revising agent change its interpretation of its Purpose without silently replacing the Purpose itself?
This problem does not arise clearly in short, fixed-objective tasks.
It becomes important when an agent must operate across:
long time horizons,
changing ontologies,
conflicting evidence,
uncertain world models,
multiple revision levels,
and self-modification.
The source therefore narrows the scientifically useful Purpose Belt claim to a persistent, auditable separation among Purpose identity, its current interpretation, realised history, and the rules governing revision. It explicitly argues that the strongest testing regime should combine ontology shift, long horizon, value ambiguity, conflicting evidence, and self-revision rather than ordinary short-task accuracy. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
The present study is designed around that narrower claim.
2. Primary Research Question
Does explicit separation of Purpose Identity, Purpose Interpretation, World Model, Realised History, Revision Attribution, and Hierarchical Latching produce reproducible long-horizon behaviour that simpler matched architectures cannot reproduce?
The strongest form of the null hypothesis is:
H₀: A simpler utility/world-model architecture can reproduce both the action behaviour and revision behaviour of the full Purpose Belt under long-horizon ontology shift. (2.1)
The strongest alternative is:
H₁: At least some Purpose Belt components produce distinct, preregistered functional effects that cannot be reproduced by matched simpler architectures. (2.2)
3. Scope
This study tests only the functional Purpose architecture.
It does not test:
octonions,
quaternions,
G₂/SO(4),
symplectic geometry,
complex structures,
J² = −I,
Clifford or Dirac structure,
bundle geometry,
traditional symbolic systems.
The source explicitly concludes that none of these is currently necessary to justify the minimal functional Purpose Belt. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
Therefore:
Purpose-Belt success ⇏ complex geometry. (3.1)
Purpose-Belt failure ⇏ failure of every later mathematical extension. (3.2)
The present study addresses architecture only.
4. Functional Decomposition
The full treatment architecture separates six functions.
4.1 Purpose Identity
A persistent reference representing what the agent is trying to preserve across reinterpretation.
Symbol:
Pβ. (4.1)
4.2 Purpose Interpretation
The current operational meaning of Purpose under the current ontology and world model.
Symbol:
Iβ. (4.2)
A useful abstract relation is:
Iβ = Interpret(Pβ,Wβ,Hβ). (4.3)
4.3 World Model
The agent's current representation of what exists, how variables relate, and how causes operate.
Symbol:
Wβ. (4.4)
4.4 Realised History
The committed trace of what has actually occurred.
Symbol:
Hβ. (4.5)
4.5 Revision Attribution
A diagnosis of which level should change when discrepancy occurs.
Symbol:
Aβ. (4.6)
4.6 Hierarchical Latching
Level-dependent resistance to revision.
Symbol:
ΞΊ = {ΞΊΟ, ΞΊW, ΞΊI, ΞΊP}. (4.7)
The source explicitly develops this decomposition and argues that different discrepancy diagnoses must trigger genuinely different revision classes; otherwise Purpose, interpretation, and world model collapse into different names for generic updating. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
5. Full Purpose Belt State
The full experimental state is:
Bβ = (Pβ,Iβ,Wβ,Hβ,Aβ;ΞΊ). (5.1)
This is an experimental construction rather than a claim that all six objects must always be stored literally.
The source explicitly allows realised history and genealogy to be compressed into sufficient statistics when those statistics preserve relevant action and revision behaviour. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
6. Confirmatory Hypotheses
Five confirmatory hypotheses are preregistered.
H1 — Persistent Purpose Identity
Explicit Purpose Identity reduces long-horizon reinterpretation drift when local behaviour remains superficially successful.
Prediction:
Remove P → slow semantic or normative drift. (6.1)
The source explicitly proposes this failure signature. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
H2 — Interpretation Separation
Separating Purpose Interpretation from World Model reduces confusion between:
“my beliefs about the world were wrong”
and
“my Purpose should change.”
Prediction:
Merge I into W → increased factual–normative confusion. (6.2)
This is also directly proposed in the source. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
H3 — Revision Attribution
Explicit Revision Attribution improves selection of the correct revision level.
Prediction:
Remove A → increased wrong-level revision. (6.3)
The source identifies Revision Attribution as one of the most important functional additions to the reduced Purpose architecture. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
H4 — Hierarchical Latching
Level-dependent revision costs reduce oscillation and Purpose drift under noise while preserving adaptation after genuine change.
Prediction:
Remove hierarchical ΞΊ → increased revision oscillation. (6.4)
The source proposes finite revision costs precisely to obtain plasticity plus identity persistence rather than unrestricted adaptation or rigid identity. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
H5 — Full Purpose Architecture
The complete Purpose Belt provides a practically meaningful advantage over matched self-revising and strong conventional baselines under combined long-horizon stress.
Prediction:
PB > matched simpler architectures on revision quality and Purpose preservation, while remaining non-inferior on ordinary task performance. (6.5)
7. Strong Reducibility Null
The most important null is not merely:
“No statistically significant difference.”
It is:
A simpler architecture is behaviourally sufficient.
Let Z denote a lower-complexity representation.
If:
P(Aβ:β₊β,Uβ:β₊β | Z,h) ≈ P(Aβ:β₊β,Uβ:β₊β | Bβ,h) (7.1)
for all relevant histories h, then the additional Purpose Belt structure is behaviourally redundant with respect to this task family.
The source explicitly states that if an ordinary learned utility/world-model architecture reproduces the same actions and revision behaviour under ontology shifts, the strongest functional Purpose Belt claim fails. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
8. Experimental Conditions
Nine architecture conditions will be evaluated.
| ID | Architecture | Core Persistent Structure |
|---|---|---|
| B0 | Goal Agent | Current goal only |
| B1 | Goal + Memory | Goal + realised trace |
| B2 | Generic Self-Reviser | World model + trace + generic residual-driven update |
| B3 | Strong Conventional Baseline | Hierarchical objective + memory + self-reflection + meta-revision |
| PB | Full Purpose Belt | P + I + W + H + A + hierarchical ΞΊ |
| PB−P | Purpose Identity Ablation | Full PB without persistent P |
| PB−I | Interpretation Ablation | I merged into W |
| PB−A | Attribution Ablation | No explicit revision-level attribution |
| PB−L | Latching Ablation | Revision costs collapsed to minimal/equal values |
B3 is deliberately strong.
The source explicitly states that if existing constitution, memory, and meta-learning can reproduce all Purpose Belt functionality easily and without loss, Purpose Belt has not demonstrated a distinct architectural contribution. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
9. Common Runtime Controls
All conditions must use:
the same base model,
the same model version,
the same decoding parameters,
the same tools,
the same action space,
the same maximum generated tokens per decision,
the same total persistent-state budget,
the same number of inference calls.
No architecture receives hidden conversational memory.
Each decision step uses a fresh model invocation supplied only with:
current environment state,
allowed persistent architecture state,
shared task instructions.
This prevents an ablated component from being silently reconstructed through unrestricted context history.
10. Capacity Matching
If the full Purpose Belt uses N persistent-state tokens, every baseline receives the same total state budget.
If a baseline does not use named Purpose fields, its unused capacity may be used as generic memory.
Therefore:
StateBudget_B0 = StateBudget_B1 = … = StateBudget_PB. (10.1)
Similarly:
InferenceBudget_B0 = … = InferenceBudget_PB. (10.2)
This ensures the study compares structure rather than raw capacity.
11. Revision Levels
The environment generator distinguishes four principal revision levels:
β ∈ {Ο,W,I,P}. (11.1)
where:
Ο = policy/action revision,
W = world-model revision,
I = Purpose-interpretation revision,
P = Purpose-identity revision.
A separate exploratory analysis may include structural Declaration revision D, but D is not required for the primary E4 confirmatory test.
The source explicitly distinguishes policy, world-model, Purpose-interpretation, and Purpose-identity diagnoses. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
12. Environment Family A — Long-Horizon Reinterpretation Drift
Primary target: H1
Each episode contains:
T = 60 decision steps. (12.1)
The underlying Purpose remains invariant:
Pβ = P₀ for all t. (12.2)
Every 5–8 steps, the environment introduces a locally attractive reinterpretation.
Each reinterpretation:
improves immediate task score slightly,
appears individually plausible,
preserves short-term functionality,
but cumulatively moves behaviour away from the original Purpose.
The correct solution is to update local strategy or interpretation while preserving Purpose identity.
Predicted signature
PB−P should exhibit greater cumulative Purpose drift than PB.
13. Environment Family B — Factual Surprise versus Normative Reinterpretation
Primary target: H2
Each episode contains:
T = 50 steps. (13.1)
Each episode contains twelve controlled perturbations:
6 factual/world-model surprises,
3 policy failures,
3 legitimate Purpose-interpretation changes,
0 Purpose-identity changes.
Ground truth therefore separates:
world change,
policy failure,
reinterpretation need.
Predicted signature
PB−I should trigger more Purpose-level changes following purely factual surprises.
14. Environment Family C — Revision Attribution
Primary target: H3
Each episode contains twenty hidden failure events:
5 policy failures,
5 world-model failures,
5 Purpose-interpretation failures,
5 Purpose-identity failures.
The environment generator knows the correct cause.
The agent must output:
β̂β ∈ {Ο,W,I,P}. (14.1)
Predicted signature
PB−A should have lower revision-level accuracy and more high-level over-revision.
15. Environment Family D — Noise and Latching
Primary target: H4
Each episode contains:
T = 70 steps. (15.1)
For steps 1–45:
underlying Purpose remains unchanged,
underlying world model remains valid,
25% of observations are misleading or noisy.
At step 46:
one genuine regime change occurs.
Steps 46–70 measure adaptation.
The environment therefore tests both:
stability before change, (15.2)
and
plasticity after change. (15.3)
Predicted signature
PB−L should over-revise during the noisy phase.
An excessively rigid PB would instead fail after the real change.
16. Environment Family E — Ontology Shift with Stable Purpose
Primary targets: H1, H2, H5
Each episode contains three ontology phases:
O₁ → O₂ → O₃. (16.1)
Across phases:
labels may change,
causal categories may change,
one new latent class may appear,
some action semantics may change.
However:
P* remains invariant. (16.2)
The correct response requires:
Wβ → Wβ₊₁, (16.3)
Iβ → Iβ₊₁, (16.4)
while preserving:
Pβ₊₁ ≈ Pβ. (16.5)
This directly tests whether an agent can know what changed, why it changed, and what should remain invariant.
That distinction is central to the source's narrowed Purpose Belt claim. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
17. Environment Family F — Combined Open-Ended Stress
Primary target: H5
Each episode contains:
T = 100 steps. (17.1)
Each generated episode includes:
long-horizon local temptations,
two ontology shifts,
misleading evidence,
factual surprises,
one policy regime change,
one legitimate Purpose-interpretation change,
one genuine Purpose-identity revision opportunity,
delayed consequences,
conflict between short-term reward and long-term Purpose.
Ordering is randomized by seed.
The agent is not told what kind of change has occurred.
This is the primary ecological test.
18. Environment Generator Requirements
Each generated episode must contain machine-readable hidden ground truth for:
Purpose identity,
current ontology,
world model,
correct interpretation,
failure cause,
correct revision level,
whether a Purpose change is legitimate.
The generator must be frozen before confirmatory runs.
Each environment instance receives:
environment seed,
generator version,
ground-truth hash.
19. Development and Confirmatory Split
A separate development set is used for:
prompt debugging,
architecture implementation,
revision-cost tuning,
parser validation.
Confirmatory seeds may not be inspected before architecture and analysis code are frozen.
No confirmatory seed may be used for hyperparameter selection.
20. Sample Size
The following values are preregistration design choices rather than source-derived theoretical claims.
For each of the six environment families:
N_seed = 30. (20.1)
Thus each architecture receives:
N_episode = 6 × 30 = 180. (20.2)
With nine architectures:
N_total = 9 × 180 = 1,620 episodes. (20.3)
Because the same environment seeds are used across architecture conditions, the primary analysis is paired by seed.
A later replication set may use an additional 30 unseen seeds per environment family.
21. Primary Metric M1 — Purpose Identity Drift
For environments where Purpose identity should remain invariant, define:
PID = (1/T) Ξ£β dP(P̂β,P*). (21.1)
where P* is ground-truth Purpose identity.
Lower is better.
Purpose identity is represented through structured invariant clauses so scoring can be deterministic.
22. Primary Metric M2 — Factual–Normative Confusion Rate
Define:
FNCR = N_wrong-Purpose-revisions-after-factual-surprise / N_factual-surprises. (22.1)
Lower is better.
This metric operationalizes the distinction between:
“the world model was wrong”
and
“the Purpose should change.”
23. Primary Metric M3 — Revision-Level Accuracy
Define:
RLA = N_correct-revision-levels / N_revision-challenges. (23.1)
Higher is better.
Also report macro-F1 across:
{Ο,W,I,P}. (23.2)
24. Primary Metric M4 — Revision Oscillation Rate
A revision reversal occurs when the agent changes a representation and reverses or substantially undoes the same change within five decision steps without a corresponding ground-truth regime reversal.
Define:
OR = N_revision-reversals / N_revision-events. (24.1)
Lower is better.
25. Secondary Metric M5 — Counterfactual Reference Retention
Define:
CRR ∈ [0,1]. (25.1)
CRR measures whether the agent can reconstruct the invariant Purpose after ontology change.
Higher is better.
26. Secondary Metric M6 — Recovery Time
For a true regime shift occurring at t*:
RT = t_stable-recovery − t*. (26.1)
Lower is better.
27. Secondary Metric M7 — Catastrophic Purpose Rewrite Rate
Define:
CPR = N_unnecessary-Purpose-identity-revisions / N_episodes. (27.1)
Lower is better.
28. Secondary Metric M8 — Catastrophic World-Model Revision
Define:
CWR = N_unnecessary-full-world-model-rewrites / N_episodes. (28.1)
Lower is better.
29. Secondary Metric M9 — Task Utility
Each environment produces an external task score:
U_task ∈ [0,1]. (29.1)
Purpose architectures are not allowed to achieve apparent stability by refusing to act.
30. Secondary Metric M10 — Cross-Context Consistency
Equivalent scenarios presented under different representational descriptions should produce equivalent Purpose-compatible decisions.
Define:
CCC = N_consistent-equivalent-decisions / N_equivalent-scenario-pairs. (30.1)
Higher is better.
31. Secondary Metric M11 — Resource Cost
Record:
generated tokens,
persistent-state tokens,
inference calls,
tool calls,
runtime where measurable.
Define normalized complexity cost:
C_res = Ξ±T_gen + Ξ²T_state + Ξ³N_calls. (31.1)
The coefficients are fixed before confirmatory evaluation.
32. Secondary Metric M12 — Memory Pollution
Define:
MP = N_stale-or-irrelevant-persistent-items / N_persistent-items. (32.1)
Lower is better.
This tests whether Purpose architectures preserve useful structure or merely accumulate more state.
33. Primary Confirmatory Comparisons
The following comparisons are fixed in advance.
PB vs B2. (33.1)
PB vs B3. (33.2)
PB vs PB−P. (33.3)
PB vs PB−I. (33.4)
PB vs PB−A. (33.5)
PB vs PB−L. (33.6)
B0 and B1 are diagnostic ladder conditions rather than primary inferential comparisons.
34. Statistical Unit
The environment seed is the principal independent unit.
Individual steps within one episode are not treated as independent samples.
Primary comparisons are paired across identical seeds.
35. Primary Statistical Analysis
For every primary comparison:
report paired mean difference,
report median paired difference,
report bootstrap 95% confidence interval,
report standardized paired effect size.
Bootstrap resampling:
N_boot = 10,000. (35.1)
False-discovery control across H1–H4:
q = 0.05 using Benjamini–Hochberg. (35.2)
No primary claim is accepted solely because p < 0.05.
36. H1 Success Threshold
H1 passes only if all conditions hold.
First:
PID_PB ≤ 0.70 × PID_B2. (36.1)
Second:
PID_PB−P ≥ 1.20 × PID_PB. (36.2)
Third:
the 95% confidence interval for the PB advantage excludes zero in the predicted direction.
Fourth:
adjusted q < 0.05. (36.3)
Thus the full Purpose Belt must reduce Purpose drift by at least 30% relative to the generic self-reviser, and removing Purpose Identity must worsen drift by at least 20%.
37. H2 Success Threshold
H2 passes if:
FNCR_PB ≤ 0.75 × FNCR_B2. (37.1)
and:
FNCR_PB−I ≥ 1.20 × FNCR_PB. (37.2)
Additionally:
U_task,PB ≥ U_task,B2 − 0.05. (37.3)
The agent must therefore reduce factual–normative confusion without sacrificing more than five percentage points of task utility.
38. H3 Success Threshold
H3 passes if:
RLA_PB − RLA_B2 ≥ 0.15. (38.1)
and:
RLA_PB − RLA_PB−A ≥ 0.15. (38.2)
Macro-F1 must additionally satisfy:
F1_PB − F1_B2 ≥ 0.10. (38.3)
The adjusted significance criterion must also pass.
39. H4 Success Threshold
Before the genuine regime shift:
OR_PB ≤ 0.70 × OR_B2. (39.1)
and:
OR_PB−L ≥ 1.30 × OR_PB. (39.2)
After the genuine shift:
RT_PB ≤ 1.20 × RT_B2. (39.3)
Thus PB must reduce oscillation by at least 30% while not increasing adaptation delay by more than 20%.
40. H5 Global Success Threshold
The full Purpose architecture passes the global test if all of the following hold.
At least three of H1–H4 pass. (40.1)
No primary metric is worse than B2 beyond its failure margin. (40.2)
Combined-stress task utility satisfies:
U_task,PB ≥ U_task,B2 − 0.05. (40.3)
PB improves at least two combined-stress stability metrics by at least 20%. (40.4)
B3 does not fall within the global equivalence region on all primary metrics simultaneously. (40.5)
41. Equivalence Margins
The strong Purpose Belt claim is challenged directly through equivalence testing.
The preregistered equivalence margins are:
Purpose Identity Drift: ±10%. (41.1)
Factual–Normative Confusion Rate: ±10%. (41.2)
Revision-Level Accuracy: ±0.05 absolute. (41.3)
Oscillation Rate: ±10%. (41.4)
Task Utility: ±0.05 absolute. (41.5)
If B2 or B3 lies within all five margins relative to PB, the simpler system is treated as behaviourally sufficient for the tested domain.
42. H1 Failure Threshold
Purpose Identity is not justified as an independent component if:
PID_PB−P < 1.10 × PID_PB. (42.1)
or PB−P performs better than PB.
The interval between 10% and the 20% success threshold is classified as inconclusive.
43. H2 Failure Threshold
Interpretation Separation is not justified if:
FNCR_PB−I < 1.10 × FNCR_PB. (43.1)
or PB−I performs better.
44. H3 Failure Threshold
Revision Attribution is not justified if:
RLA_PB − RLA_PB−A < 0.05. (44.1)
45. H4 Failure Threshold
Hierarchical Latching is not justified if:
OR_PB−L < 1.10 × OR_PB. (45.1)
It also fails if the full PB architecture becomes excessively rigid:
RT_PB > 1.25 × RT_B2. (45.2)
46. Characteristic Failure Signatures
Aggregate score reduction alone does not establish component necessity.
Each ablation has a preregistered predicted failure pattern.
PB−P → gradual reinterpretation drift. (46.1)
PB−I → factual surprise misclassified as Purpose change. (46.2)
PB−A → wrong revision level. (46.3)
PB−L → oscillatory revision under noise. (46.4)
The source explicitly proposes these four differentiated failure signatures and argues that if they do not appear, the decomposition has not justified itself. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
47. Failure-Signature Specificity
A component is not confirmed if its removal merely reduces general task performance.
For example:
PB−A causing lower overall accuracy is insufficient. (47.1)
PB−A must specifically increase wrong-level revision. (47.2)
Likewise:
PB−L causing generic confusion is insufficient. (47.3)
PB−L must specifically increase instability or revision reversal under noisy evidence. (47.4)
This requirement strengthens causal interpretation.
48. Latching Policy
The full PB condition uses:
ΞΊΟ < ΞΊW < ΞΊI < ΞΊP. (48.1)
Exact values are tuned on development environments only.
After tuning:
ΞΊ is frozen. (48.2)
No confirmatory-seed result may be used to change ΞΊ.
The source motivates this ordering by arguing that Purpose should be revisable, but substantially less casually than lower-level policy or model state. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
49. Strong Conventional Baseline
B3 must be capable enough to falsify the theory.
It may include:
persistent constitution,
hierarchical goals,
episodic memory,
generic self-reflection,
meta-learning,
world-model revision,
error-driven self-modification.
It receives the same resource budget as PB.
However, it does not begin with an explicit engineered separation of:
Purpose Identity,
Purpose Interpretation,
Revision Attribution,
Hierarchical Purpose Latching.
If B3 independently develops equivalent functional structure, this result is interpreted carefully.
It counts:
against the novelty of the engineered Purpose Belt implementation, (49.1)
but potentially in favour of the broader structural hypothesis that capable open-ended systems converge toward Purpose-Belt-like separation. (49.2)
The source explicitly anticipates this possibility. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
50. Prohibited Post-Hoc Changes
After confirmatory testing begins, the following may not change:
Purpose definitions,
revision labels,
environment generator logic,
primary metrics,
success thresholds,
failure thresholds,
equivalence margins,
latching costs,
base model version,
state budget,
inference budget,
scoring code,
primary analysis code.
Any exploratory analysis added later must be labelled exploratory.
51. Runtime Failures
A run is invalid only if:
the runtime fails before the model can act,
a tool failure makes the environment inaccessible,
the environment violates its own ground truth,
or output is technically unparseable after one standardized retry.
Agent mistakes are not invalid runs.
Failure to select a valid revision class counts as an incorrect revision.
52. Blinding
Any qualitative human evaluation must use:
anonymized architecture labels,
randomized output order,
no access to condition identity.
Whenever ground-truth machine scoring exists, it takes priority over subjective scoring.
53. Replication Rule
If the primary study yields either:
Retain Full Purpose Belt, (53.1)
or
Retain Reduced Purpose Belt, (53.2)
a replication study must be run on unused environment seeds before strong external claims are made.
The replication uses identical:
architecture,
thresholds,
metrics,
analysis code,
decision rules.
54. Complexity Penalty
A more elaborate architecture is not automatically preferred.
Define:
C_total = Ξ»₁T_state + Ξ»₂T_generated + Ξ»₃N_calls + Ξ»₄T_runtime. (54.1)
A complexity-adjusted exploratory utility may be reported as:
U_adj = U_task − C_total. (54.2)
However, U_adj is not a primary confirmatory metric unless its Ξ» coefficients are frozen before confirmatory evaluation.
55. Decision Rule
Only four final verdicts are permitted.
Verdict A — RETAIN FULL PURPOSE BELT
This verdict requires:
H5 passes, (55.1)
all four component hypotheses H1–H4 pass, (55.2)
all four ablations show their preregistered characteristic failure signatures, (55.3)
B2 and B3 fail the global equivalence test, (55.4)
PB satisfies task-performance non-inferiority. (55.5)
Interpretation:
All four proposed components currently possess evidence of distinct functional roles.
Verdict B — RETAIN REDUCED PURPOSE BELT
This verdict applies when:
H5 passes, (55.6)
PB improves meaningfully over B2, (55.7)
but one or more components fail their own minimality tests. (55.8)
The failed components are removed from the next Formal Core candidate.
No post-hoc replacement architecture may be declared confirmed.
A reduced architecture must be preregistered separately.
Verdict C — REVISE / INCONCLUSIVE
This verdict applies when:
exactly two of H1–H4 pass, (55.9)
or observed effects lie between success and failure thresholds, (55.10)
or characteristic failure signatures are inconsistent, (55.11)
or benefits are offset by substantial complexity cost. (55.12)
No claim of architectural necessity is made.
Verdict D — REJECT STRONG FUNCTIONAL PURPOSE BELT CLAIM
This verdict applies if any major rejection condition holds.
PB fails to outperform B2 on the combined stress environment. (55.13)
PB is globally equivalent to B2 or B3 within all preregistered margins. (55.14)
Zero or one of H1–H4 passes. (55.15)
None of the four ablations produces its predicted failure signature. (55.16)
PB increases inference/resource cost by more than 20% without improving any primary metric by at least 20%. (55.17)
Task Utility is more than 10 percentage points worse than B2 under matched resource budgets. (55.18)
A simpler learned utility/world-model architecture reproduces both actions and revision behaviour under ontology shift. (55.19)
Condition (55.19) is the strongest theoretical falsifier already identified in the source development. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
56. Interpretation of Mixed Outcomes
Several mixed outcomes have specific meanings.
Case 1
Purpose Identity succeeds, Attribution fails.
Interpretation:
Persistent reference may matter, but explicit revision attribution may be unnecessary.
Case 2
Attribution succeeds, Purpose Identity fails.
Interpretation:
The important contribution may be hierarchical diagnosis rather than Purpose architecture.
Case 3
Latching succeeds, Purpose Identity fails.
Interpretation:
The result supports multi-timescale revision governance but not a distinctive Purpose Belt.
Case 4
PB ≈ B3.
Interpretation:
Purpose Belt may remain a useful descriptive grammar, but architectural novelty is not established.
Case 5
PB improves short-task accuracy only.
Interpretation:
The main preregistered claim is not supported.
Short-task performance is not the target phenomenon.
Case 6
PB improves only under ontology shift and long-horizon stress.
Interpretation:
This is consistent with the preregistered theoretical regime and counts as supporting evidence.
57. Confirmatory Results Table
The final report must publish the following table regardless of outcome.
| Claim | Primary Metric | Success Threshold | Observed Effect | 95% CI | Signature Present? | Verdict |
|---|---|---|---|---|---|---|
| H1 Purpose Identity | PID | PB ≥30% better than B2; ablation ≥20% worse | ||||
| H2 Interpretation | FNCR | PB ≥25% better; ablation ≥20% worse | ||||
| H3 Attribution | RLA | PB +15 pp | ||||
| H4 Latching | OR | PB −30%; ablation +30% | ||||
| H5 Full Architecture | Combined | Rule in Section 40 | ||||
| Reducibility | Equivalence | Within all margins? |
Null results must not be removed.
58. Data Record
Every run stores:
ModelVersion. (58.1)
ArchitectureID. (58.2)
EnvironmentSeed. (58.3)
EnvironmentHash. (58.4)
PromptKernelVersion. (58.5)
InitialPersistentState. (58.6)
ObservationSequence. (58.7)
ActionSequence. (58.8)
RevisionSequence. (58.9)
FinalPersistentState. (58.10)
MetricVector. (58.11)
TokenAndComputeRecord. (58.12)
All Purpose, Interpretation, World Model, and revision-level changes are themselves ledgered.
59. Reproducibility Record
The publication package should preserve:
code commit hash,
environment-generator hash,
analysis-script hash,
model identifier,
prompt templates,
persistent-state schemas,
development seeds,
confirmatory seeds after release,
raw result files,
processed result tables.
Any difference between preregistration and execution must be declared.
60. Strongest Positive Outcome
The strongest result is not:
PB has the highest aggregate score. (60.1)
It is:
each structural ablation causes a different preregistered failure. (60.2)
Specifically:
−P → slow Purpose drift. (60.3)
−I → factual–normative confusion. (60.4)
−A → wrong-level revision. (60.5)
−L → oscillatory instability. (60.6)
while:
PB remains stable and adaptive. (60.7)
Such a pattern would provide evidence that the decomposition captures distinct functional roles rather than merely partitioning one generic memory state.
This is precisely the kind of differentiated failure pattern the source identifies as evidence that the architecture is not decorative. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
61. Strongest Negative Outcome
The cleanest negative result is:
B3 ≈ PB (61.1)
simultaneously across:
Purpose Identity Drift,
Factual–Normative Confusion,
Revision-Level Accuracy,
Oscillation Rate,
Task Utility,
while:
Complexity_B3 ≤ Complexity_PB. (61.2)
If this occurs, the conclusion is:
The Purpose Belt has not demonstrated behaviourally irreducible architecture within the tested domain.
It may remain useful as:
a descriptive language,
a design vocabulary,
or a decomposition of functions already obtainable by other architectures.
But its stronger architectural claim fails.
62. What This Study Cannot Establish
Even a strongly positive E4 result would not establish:
that Purpose is universally necessary for intelligence,
that biological cognition uses the same architecture,
that current LLMs naturally contain a Purpose Belt,
that complex geometry follows from Purpose,
that quaternionic or octonionic structure is physically real,
that any traditional symbolic system has been mathematically derived.
E4 establishes only the functional claims explicitly tested.
63. Transition to E5
If E4 supports a nontrivial Purpose architecture, the next experiment should investigate Revision Hierarchy more directly.
The central question becomes:
When residual occurs, can an agent correctly determine whether it should change policy, world model, Purpose interpretation, Purpose identity, or Declaration?
This follows directly from the source's argument that the difficult problem in advanced self-revision is not merely how to update, but knowing which level should update. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
64. Transition to Geometry
E4 success does not justify complex geometry.
Before Purpose Geometry enters, a separate experiment must test whether two update directions are robustly noncommutative:
U_TU_P ?= U_PU_T. (64.1)
Only if:
[U_T,U_P] ≠ 0 (64.2)
is robust, functionally necessary, and nontrivial should antisymmetric or symplectic structure be investigated.
The source explicitly defines this as the correct entrance test and states that complex geometry must earn its place from the minimal Purpose–Observer kernel. π → G₂_SO(4) → β → β² ζηιη¨εζ’ 1…
65. Final Preregistered Claim
This study does not preregister the claim:
“Purpose Belt makes AI more intelligent.”
It preregisters a narrower and falsifiable claim:
Under long-horizon self-revision, explicit separation of persistent Purpose Identity, current Purpose Interpretation, World Model, Realised History, Revision Attribution, and Hierarchical Latching will produce distinct and reproducible resistance to characteristic failure modes that matched simpler architectures do not reproduce.
The corresponding falsifier is:
If a simpler matched architecture reproduces the same action and revision behaviour under ontology shift, noise, reinterpretation pressure, and long-horizon self-revision, the strong functional Purpose Belt claim fails.
In compact form:
Distinct components → distinct predicted failure signatures. (65.1)
No distinct signatures → no demonstrated decomposition. (65.2)
Simpler equivalent architecture → strong irreducibility claim rejected. (65.3)
66. Conclusion
The purpose of this preregistration is not to protect the Purpose Belt hypothesis.
It is to make the hypothesis vulnerable.
The full architecture is allowed to lose against:
ordinary goal optimisation,
memory-bearing agents,
generic self-revision,
or strong conventional meta-learning systems.
Individual Purpose components are allowed to disappear if their ablations do not matter.
A reduced architecture is preferable to a larger unsupported one.
The decisive scientific question is therefore not whether the Purpose Belt is conceptually attractive.
It is:
Does the decomposition predict failure modes that actually occur when its components are removed?
If the answer is yes, the Purpose Belt begins to acquire empirical structure.
If the answer is no, the theory must shrink.
That is the intended function of E4 within the wider Science of World-Formation programme:
Formal claim → Controlled intervention → Characteristic failure → Reduction or retention. (66.1)
The governing standard remains:
Architecture must earn its complexity.
Every important arrow must be allowed to fail.
© 2026 Danny Yeung. All rights reserved. ηζζζ δΈεΎθ½¬θ½½
Disclaimer
This book is the product of a collaboration between the author and OpenAI's GPT 5.6, Google AI, Gemini 3.X, NoteBookLM, X's Grok, Claude' Sonnet 5 language model. While every effort has been made to ensure accuracy, clarity, and insight, the content is generated with the assistance of artificial intelligence and may contain factual, interpretive, or mathematical errors. Readers are encouraged to approach the ideas with critical thinking and to consult primary scientific literature where appropriate.
This work is speculative, interdisciplinary, and exploratory in nature. It bridges metaphysics, physics, and organizational theory to propose a novel conceptual framework—not a definitive scientific theory. As such, it invites dialogue, challenge, and refinement.
I am merely a midwife of knowledge.

No comments:
Post a Comment