https://chatgpt.com/share/6a78fc18-6ae0-83ed-8889-7e27dfcb150f
https://osf.io/kcjv3/files/osfstorage/6a78fb1ab195de03f21fb7bb
Reconstructable Research
A Machine-Native Event Architecture for AI-Assisted Theory Formation
Abstract
Large language models have changed the economics of theoretical exploration. A research programme can now generate hundreds of conceptual variants, objections, cross-domain mappings, revisions, failed formulations, auxiliary hypotheses, and synthesized manuscripts at a speed that was previously impossible for an individual researcher. Yet the dominant publication object remains almost unchanged: the final paper.
This creates an epistemic compression problem.
A conventional manuscript normally presents the current theory as a coherent argument. It does not preserve, in machine-operable form, the full genealogy by which the theory arose: which source introduced which concept; which objection destroyed which formulation; which constraint survived revision; which mapping failed; which unresolved residual generated a successor theory; which branch was abandoned; which result was independently rediscovered; and which apparent recurrence was merely inherited through prior context. In AI-assisted theoretical work, these omissions become especially consequential because the generative search process can be vastly larger than the final document.
This paper proposes Reconstructable Research: an architecture in which the final paper is no longer treated as the sole canonical object of theory formation. Instead, externally recorded research events are captured and compiled into a machine-native, provenance-bearing representation of research history. This representation need not itself be human-readable. It needs only to preserve enough structure that declared human-readable projections can later be generated, audited, compared, and traced back toward source events.
The central transition is:
Research Events → Event Capture → Semantic Compilation → Machine-Native Research Event Representation → Declared Projection → Human / Machine Views. (0.1)
The proposed canonical object is called the Machine-Native Research Event Representation, abbreviated MRER. MRER is not assumed to be a simple graph. It may be implemented as a typed graph, hypergraph, event store, provenance system, vector-symbolic representation, relational structure, or hybrid architecture. Its defining requirement is semantic rather than syntactic: it must preserve distinguishable research objects and transformations such as events, artifacts, claim states, constraints, revisions, residuals, evidence, genealogy, and reconstruction assertions.
The paper develops four architectural contracts:
Capture → Reconstruct → Project → Audit. (0.2)
The Capture Contract records externally observable research events without claiming access to hidden model cognition. The Reconstruction Contract compiles those events into structured claims about theory evolution. The Projection Contract generates approximately human-readable views under declared purposes and fidelity constraints. The Audit Contract allows important projected assertions to be traced backward through reconstruction assertions toward machine objects and original provenance.
A central methodological distinction is:
Later Than ≠ Derived From ≠ Semantically Related To ≠ Caused By. (0.3)
Chronology, genealogy, semantic relation, causal influence, and epistemic status must therefore remain distinct relation layers. The framework also treats residuals as first-class research objects. What failed to fit a theory may be as important as what survived, because unresolved residuals frequently become the pressure that generates successor formulations.
The paper further introduces conceptual track identity, mutation, branching, merge, replacement, dormancy, resurrection, reconstruction depth, competing reconstructions, projection residual, and generative causal replay. A worked example follows the development from recursive generation to viewpoint filtration, declaration, and admissible self-revision, showing how a theory can undergo deep conceptual mutation while retaining an identifiable research lineage. The source sequence itself explicitly records these corrections: recursive generation was weakened into disclosure to avoid a hidden meta-time; filtration then exposed the need for declaration; declaration in turn exposed the danger of unrestricted self-revision.
The governing proposal is therefore not that machines can recover the hidden truth of intellectual history. It is narrower:
A sufficiently instrumented research process can become reconstructable under declared protocols.
The paper becomes one projection of that reconstructable object rather than the object itself.
The guiding maxim is:
Do not publish only the state. Preserve the transformations.
0. Reader’s Guide: What Is Being Proposed?
0.1 The problem is not that papers are obsolete
This paper does not argue that scientific papers should disappear.
Narrative papers remain extraordinarily useful. They compress complex research into a sequence that human readers can inspect, criticize, teach, cite, and remember. A good paper performs intellectual work that a raw archive cannot. It establishes emphasis, defines scope, explains why a question matters, presents evidence, and gives a reader a tractable path through a problem.
The claim developed here is different.
A paper is often an excellent projection of research.
It is not necessarily an adequate representation of the research process itself.
That distinction becomes increasingly important when theory formation is assisted by generative models.
An AI-assisted research programme may contain:
• hundreds of candidate formulations;
• dozens of explicit counterexamples;
• abandoned derivations;
• competing taxonomies;
• source-document collisions;
• repeated reframings;
• changing definitions;
• model-to-model criticism;
• human interventions;
• tool-generated evidence;
• retrospective repairs;
• and entire branches that never reach the final manuscript.
The finished paper may contain only a few percent of this history.
That is not necessarily a defect in the paper.
It is a defect only if the paper is treated as though it were the complete research object.
The proposal of this paper is therefore:
NarrativePaper ≠ CompleteResearchObject. (0.4)
Instead:
NarrativePaper = Projection(ResearchEventObject | HumanPurpose). (0.5)
The purpose of Reconstructable Research is to make the object behind that projection explicit.
0.2 This is not hidden chain-of-thought reconstruction
The proposed architecture does not attempt to reconstruct the private internal reasoning of a large language model.
It does not claim access to latent activations as if they were transparent beliefs.
It does not infer hidden cognition from prose and then declare that inference to be the true mental history of the system.
Its source material is restricted to externally available research trace.
Typical admissible records include:
• prompts;
• uploaded source documents;
• model outputs;
• tool outputs;
• human comments;
• explicit corrections;
• document revisions;
• code;
• datasets;
• benchmark results;
• acceptance or rejection decisions;
• timestamps;
• model and protocol versions.
This discipline extends the Externalized Collision Trace developed in the Semantic Collider framework, where the inspectable research object consists of explicit beam definitions, reconstructions, mappings, rejections, residuals, revisions, predictions, tests, model settings, sources, expert corrections, adversarial critiques, and validation results rather than hidden internal chain-of-thought.
The relevant distinction is:
ExternalResearchTrace ≠ HiddenModelCognition. (0.6)
Reconstructable Research operates on the former.
0.3 This is not merely provenance logging
Ordinary provenance answers questions such as:
Which artifact came from which source?
Which process generated this file?
Which model created this output?
Which version preceded this revision?
Those questions are essential.
But they do not fully reconstruct theory formation.
Suppose a sequence contains:
Prompt 17 → Response 17 → Draft 4 → Draft 5. (0.7)
A provenance system may record that Draft 5 was derived from Draft 4.
What it does not automatically tell us is:
• which claim changed;
• which constraint forced the change;
• whether the change was a refinement or replacement;
• which earlier residual became resolved;
• which mapping was rejected;
• whether the successor retained conceptual identity;
• whether the old branch remained valid within a narrower scope.
This requires a semantic reconstruction layer.
Hence:
Provenance Preservation ≠ Theory Reconstruction. (0.8)
Provenance is necessary.
It is not sufficient.
0.4 This is not merely a knowledge graph
A conventional knowledge graph may represent:
A is related to B.
B contradicts C.
C belongs to domain D.
That can be useful.
But theory formation is not a static network of propositions.
It contains mutation.
A claim may be revised without becoming a completely new claim.
A framework may split into competing descendants.
Two previously independent tracks may merge.
A theory may die, remain dormant, or return under new evidence.
A residual left unresolved in one generation may become the central research question of the next.
Therefore the relevant object is not only:
KnowledgeStructure. (0.9)
It is:
KnowledgeStructure + TransformationHistory. (0.10)
More strongly:
ResearchStateₖ₊₁ = Transform(ResearchStateₖ | Eventₖ, Constraintₖ, Residualₖ). (0.11)
The task is therefore closer to event-sourced semantic reconstruction than ordinary static knowledge representation.
0.5 This is not a claim that one reconstruction is uniquely true
Research history is often underdetermined.
Suppose concept B appears after concept A.
We may know with high confidence:
A preceded B. (0.12)
We may have weaker evidence that:
B inherited structure from A. (0.13)
We may infer that:
B semantically refines A. (0.14)
But it may remain unresolved whether:
A causally produced B. (0.15)
A mature reconstruction architecture should not collapse these distinctions merely because one narrative sounds cleaner.
Therefore the canonical representation must allow:
Confirmed relation.
Probable relation.
Competing relation.
Disputed relation.
Unresolved relation.
This leads to an important principle:
ReconstructionAssertion ≠ HistoricalFact. (0.16)
A reconstruction assertion is itself a claim requiring provenance and epistemic status.
0.6 This is not required to be directly human-readable
This paper deliberately rejects another unnecessary requirement.
The canonical machine representation does not need to look like a concept map that a human can immediately understand.
It may contain structures that are inconvenient or impossible to display directly:
• thousands of claim states;
• typed hyperedges;
• probability distributions over alternative reconstructions;
• partial temporal orders;
• embeddings;
• equivalence classes;
• machine identifiers;
• unresolved branches;
• provenance hashes;
• semantic transformation objects;
• confidence tensors;
• multiple scale-dependent identities.
Human readability should therefore occur after reconstruction, not constrain reconstruction itself.
Let M denote the machine-native research event representation.
Let v denote a requested viewpoint.
Let P denote the projection protocol.
Then a human-facing rendering may be written:
Hᵥ = Π_(v,P)(M). (0.17)
Because any human projection is lossy, a more honest formulation is:
Π_(v,P)(M) → Hᵥ + R_Π. (0.18)
where:
Hᵥ = approximately human-readable projection, (0.19)
R_Π = projection residual: structure in M that the chosen view does not safely or conveniently express. (0.20)
The requirement is therefore not direct readability.
It is:
faithful translatability under declared loss.
0.7 Four contracts organize the architecture
The entire framework can be summarized through four contracts:
Capture → Reconstruct → Project → Audit. (0.21)
Contract I — Capture
Record enough externally observable research events that later reconstruction is possible.
Contract II — Reconstruct
Compile those events into machine-native representations of claim evolution, constraint interaction, genealogy, residual, evidence, and uncertainty.
Contract III — Project
Generate purpose-specific human or machine views without pretending that one view exhausts the underlying representation.
Contract IV — Audit
Allow important projected assertions to be traced backward through reconstruction assertions toward the evidence and external event record from which they arose.
These contracts define the engineering skeleton of Reconstructable Research.
0.8 The paper is one projection among several
Once a machine-native research event representation exists, many outputs become possible.
The same reconstructed research object may generate:
Narrative Paper. (0.22)
Semantic Event Display. (0.23)
Claim Ledger. (0.24)
Residual Register. (0.25)
Evidence View. (0.26)
Reviewer Audit View. (0.27)
Theory-Evolution Timeline. (0.28)
Machine Retrieval View. (0.29)
The resulting architecture is:
Research Events
→ Machine-Native Research Event Representation
→ Multiple Declared Views. (0.30)
The paper remains important.
It simply loses its monopoly.
1. Beyond the Paper: Why AI-Assisted Research Needs a Native Event Object
1.1 The traditional paper is a snapshot technology
Scientific writing evolved under severe human constraints.
Researchers have limited time.
Readers have limited attention.
Journals have limited space.
A theory must therefore be compressed.
The paper transforms months or years of work into a sequence:
Problem → Method → Result → Interpretation → Conclusion. (1.1)
This is extraordinarily useful.
But it is also a snapshot technology.
The final form usually suppresses much of the trajectory that produced it.
Consider what a finished theoretical article rarely exposes in full:
• the first wrong ontology;
• the hypothesis that looked promising for three weeks and died;
• the objection that forced a definition to change;
• the alternative branch that was never pursued;
• the distinction introduced only because an earlier mapping failed;
• the source that changed the author’s interpretation;
• the unresolved residual deliberately carried into the next paper.
The published artifact normally presents:
CurrentState. (1.2)
rather than:
CurrentState + TransformationHistory. (1.3)
For human-only theoretical work, much of the missing history may simply never have been externalized.
For AI-assisted research, the situation is different.
Large portions of the intermediate search can exist as explicit digital trace.
1.2 AI changes the amount of theory formation that can become external
A human researcher often performs conceptual exploration silently.
The researcher thinks:
Perhaps A is related to B.
No, that fails because of C.
Maybe the correct abstraction is D.
That internal process may never appear in any file.
Generative systems change this.
A researcher can ask an LLM to:
• reconstruct a source framework;
• propose alternative abstractions;
• attack a hypothesis;
• compare incompatible models;
• search for hidden assumptions;
• generate counterexamples;
• formalize a distinction;
• rewrite the same theory under competing ontologies;
• transfer a structure to another domain;
• identify unresolved residuals.
Each interaction may become an externally stored event.
This produces a new possibility:
Conceptual search that was formerly private may become partially externalized.
This was already one of the central implications of the Semantic Collider programme: generative models may make a previously informal layer of conceptual collision, failure, and recombination externally recordable, perturbable, replicable, and auditable.
The new question is what to do with that abundance of trace.
1.3 More trace does not automatically produce more knowledge
A folder containing:
500 chat sessions,
80 PDFs,
1,200 model outputs,
40 abandoned drafts,
and 12 final papers
is not automatically a better scientific object.
It may simply be an archive.
Raw abundance introduces several problems.
First, chronological adjacency becomes misleading
Two ideas may occur consecutively without one causing the other.
Second, inherited context may masquerade as rediscovery
A later model may “discover” an invariant because it was already present in the prompt history.
Third, conceptual identity becomes unstable
The same research problem may acquire new vocabulary, new mechanism, and narrower assumptions while still belonging to one intellectual lineage.
Fourth, failure disappears easily
A polished final article contains surviving claims.
Rejected mappings and abandoned branches often vanish.
Fifth, the archive becomes too large for a human to inspect directly
Externalization solves the scarcity of trace but creates a problem of reconstruction.
Hence:
MoreTrace ≠ MoreUnderstanding. (1.4)
A reconstruction layer is needed.
1.4 The strongest research object may no longer be the document
The document-centered assumption says:
Research produces documents.
The alternative proposed here says:
Research produces events that alter research state.
Documents are one class of artifact produced by those events.
This suggests a shift:
DocumentCentricResearch → EventCentricResearch. (1.5)
An event may:
introduce a claim;
change a definition;
attack a constraint;
resolve a residual;
split a branch;
merge two frameworks;
add evidence;
lower confidence;
or terminate a theory.
The final document records only some of these transformations.
If the transformations are retained separately, the research programme becomes reconstructable at a finer resolution.
1.5 Research can therefore be treated as event-sourced
In software architecture, an event-sourced system does not store only the latest state. It preserves the events whose application produced that state.
The same principle can be applied conceptually.
Instead of storing only:
TheoryStateₙ, (1.6)
retain:
E₁, E₂, E₃, …, Eₙ, (1.7)
such that:
TheoryStateₙ = Apply(E₁, E₂, …, Eₙ | P). (1.8)
This analogy should not be interpreted too literally. Research events are not deterministic database transactions. Semantic reconstruction is uncertain, context-sensitive, and protocol-dependent.
But the architectural lesson is powerful:
A research programme becomes more reconstructable when it preserves the transformations that produced its current state.
This leads to the guiding maxim of the paper:
Do not publish only the state. Preserve the transformations.
1.6 A concrete example: theory development can have visible curvature
Consider the following sequence drawn from the development of the declaration–time framework.
An early formulation proposed:
Primitive Operation → Recursion → Pre-Time → Collapse → Ledger → Time. (1.9)
The first article explicitly explored the possibility that recursive depth could replace ordinary time as the ordering principle of a pre-collapse field.
A later article identified a problem.
If recursion literally generates the pre-time universe step by step, then the theory may have silently reintroduced the very temporal ordering it was supposed to explain.
The revised formulation became:
Pre-Time Field → Viewpoint → Filtration → Collapse → Ledger → Time. (1.10)
The later paper explicitly states that recursion may be a presentation grammar rather than an ontological creation process, and therefore replaces generation with viewpoint-selected disclosure.
But the correction itself exposed another residual:
What makes the field filterable? (1.11)
The next article answered:
Declaration. (1.12)
A viewpoint must declare a baseline, feature map, boundary, observation protocol, projection, gate, trace rule, and residual rule before filtration becomes meaningful.
This produced:
Undeclared Field
→ Declaration
→ Projection
→ Gate
→ Trace + Residual
→ Ledger
→ Time. (1.13)
That framework then generated its own new problem.
If trace and residual revise future declaration, what prevents the system from rewriting its own rules whenever failure occurs?
The subsequent article answered:
Admissible Self-Revision. (1.14)
A mature observer cannot revise itself arbitrarily. Revision must remain trace-preserving, residual-honest, frame-robust, bounded, and non-degenerate.
The sequence is therefore not merely four papers.
It is a transformation history:
Generation
→ objection: hidden meta-time
→ Disclosure / Filtration
→ residual: what makes filtration possible?
→ Declaration
→ residual: unrestricted self-revision
→ Admissible Revision. (1.15)
A final summary may present only the mature endpoint.
A reconstructable research object should preserve the path.
1.7 Why this path matters scientifically
The transformation history contains information not present in the endpoint alone.
If we know only the final theory, we may not know:
• which assumptions were inherited from an abandoned version;
• why a particular constraint exists;
• whether a limitation is accidental or historically central;
• which formulation was rejected and for what reason;
• whether a later concept is an independent invention or repair of an earlier failure.
The history tells us not only:
What does the theory currently say? (1.16)
but also:
What pressure shaped it into this form? (1.17)
That second question can be scientifically important.
A constraint introduced because of a real counterexample has a different epistemic history from a constraint introduced merely for elegance.
A narrow claim that survived repeated restriction differs from a broad claim that was never seriously attacked.
A residual carried across five successive revisions deserves attention even if it remains unresolved.
Therefore:
TheoryContent ≠ TheoryDevelopment. (1.18)
Both can matter.
2. From Trace Ledger to Event Reconstruction
2.1 Preservation is the first step, not the last
The Semantic Collider framework already proposed a major departure from ordinary AI-generated theory writing.
Instead of treating a polished generated article as the primary epistemic object, it proposed preserving an Externalized Collision Trace containing:
• source conceptual beams;
• native reconstructions;
• constraints;
• abstracted forms;
• proposed invariants;
• residuals;
• failed mappings;
• holdout transfer;
• adversarial tests;
• replication;
• external validation.
Its minimal protocol was:
Reconstruct → Declare → Abstract → Collide → Break → Ledger → Transfer → Validate. (2.1)
The paper explicitly says that the manuscript is one rendering of the Externalized Collision Trace and that a future machine-readable format could preserve the full structure.
This is the immediate starting point for Reconstructable Research.
But the new proposal makes an additional distinction.
A ledger can preserve events without explaining their relational transformation.
Therefore:
Trace Preservation → necessary. (2.2)
Event Reconstruction → additional layer. (2.3)
2.2 A ledger answers “what was recorded?”
A reconstruction answers “what changed?”
Suppose the archive records:
Event 41: model proposes Theory A.
Event 42: human identifies objection O.
Event 43: model produces Theory B.
Event 44: human accepts B as improved formulation. (2.4)
A trace ledger can preserve all four.
A reconstruction layer may additionally infer:
Theory B revises Theory A. (2.5)
Objection O attacks Constraint K in Theory A. (2.6)
Theory B preserves Problem Q but changes Mechanism M. (2.7)
Residual R remains unresolved. (2.8)
These are not raw events.
They are structured interpretations of events.
That is why the reconstruction layer requires explicit epistemic status.
2.3 Reconstruction is closer to semantic compilation than summarization
A summarizer asks:
What is the main point of this archive?
A semantic reconstructor asks:
What objects and transformations must be represented so that later systems can recover how the current theory arose?
This resembles the distinction developed in the Differential-Topological Kernel Compiler work:
RawRequirement → IntentStructure → KernelIR → ExecutablePrompt. (2.9)
There, broad natural-language material is treated as something to be compiled into a stable intermediate representation rather than merely rewritten more elegantly. The framework explicitly argues that requirement-to-kernel conversion is a compilation problem rather than a writing problem.
Reconstructable Research applies the same architectural insight at a different level:
RawResearchTrace → ResearchIR → MachineNativeEventObject → Projection. (2.10)
The purpose of ResearchIR is not literary compression.
It is structural preservation.
2.4 Semantic compilation must preserve distinction before compression
A dangerous compiler would map:
Claim proposal
Claim revision
Claim rejection
Residual
Evidence
into one generic object:
“important idea”.
That would destroy exactly the information we want.
Therefore the compilation rule must prioritize distinctions that matter to theory evolution.
A minimal discipline is:
PreserveBeforeCompress. (2.11)
More precisely:
SemanticCompression is admissible only if declared reconstruction invariants remain recoverable. (2.12)
Candidate invariants include:
• lineage;
• claim status;
• major transformation;
• critical constraints;
• evidence direction;
• unresolved residual;
• reconstruction uncertainty.
Human-facing prose may suppress details.
The machine layer should not suppress them prematurely.
2.5 The reconstructed object should not be equated with a graph
Graphs are natural.
But they are not mandatory.
A theory-development event may involve:
five parent claims,
two constraints,
one residual,
one new artifact,
and three competing interpretations.
A simple binary edge structure may represent this awkwardly.
The canonical proposal is therefore intentionally representation-neutral.
Let:
M = Machine-Native Research Event Representation. (2.13)
The implementation may use:
typed graphs;
hypergraphs;
relational stores;
event stores;
RDF-like triples;
vector-symbolic objects;
partial orders;
probabilistic graphical structures;
or hybrids.
The ontology should not be confused with the storage technology.
2.6 The final paper becomes one decoder output
Once M exists, a paper can be generated as:
Paper = Π_paper(M | Audience, Scope, EvidencePolicy). (2.14)
A semantic event display can be generated as:
Display = Π_display(M | TrackPolicy, Scale, VisibilityRules). (2.15)
A reviewer interface might be:
ReviewView = Π_review(M | ClaimRisk, EvidenceMaturity, ResidualPriority). (2.16)
None of these is automatically privileged.
The appropriate projection depends on purpose.
This is compatible with the broader declared-disclosure framework, which treats visible structure as protocol-relative rather than assuming that one projection exhausts the field.
2.7 The scientific object therefore changes category
The traditional publication object is primarily:
Narrative. (2.17)
The proposed object is:
ReconstructableEventStructure + NarrativeProjection. (2.18)
This does not diminish writing.
It changes what writing is relative to the research substrate.
A manuscript becomes analogous to an interface over a richer event object.
The strongest future research artifact may therefore have two different modes of existence:
Machine mode:
high-dimensional, provenance-bearing, reconstructable, branch-preserving. (2.19)
Human mode:
selective, explanatory, approximately readable, intentionally compressed. (2.20)
The central principle becomes:
HumanReadability belongs to Projection, not necessarily to CanonicalRepresentation. (2.21)
That is the architectural move on which the rest of this paper depends.
[End of first installment: Abstract + Sections 0–2]
Next installment will begin Part II — The Machine-Native Research Event Representation, covering Sections 3–7: event-sourced research, the four contracts, the canonical objects, why MRER need not be human-readable, and the multiplex relation layers.
Part II — The Machine-Native Research Event Representation
3. Event-Sourced Research
3.1 From document snapshots to research-state transitions
The previous section argued that a final paper is best understood as one projection of a richer research history.
We can now sharpen that idea.
Traditional publication is largely state-centered.
At publication time, the reader receives something like:
ResearchStateₙ → Paperₙ. (3.1)
The earlier states may survive in notebooks, drafts, correspondence, code repositories, or memories, but they are not normally represented as part of the canonical scientific object.
AI-assisted theoretical research makes a different architecture possible.
Instead of recording only the current state, we can preserve the events that altered that state.
Let:
Sₖ = research state after event k. (3.2)
Eₖ = externally recorded research event at step k. (3.3)
Then:
Sₖ₊₁ = U(Sₖ, Eₖ). (3.4)
The update operator U should not be interpreted as a deterministic psychological law. It simply denotes that, under a declared reconstruction, event Eₖ is associated with a change from one research state to another.
The full state can therefore be represented approximately as:
Sₙ = U(E₁, E₂, …, Eₙ | P). (3.5)
where P is the declared reconstruction protocol.
The important shift is conceptual:
ResearchState alone is not the complete object. (3.6)
ResearchState + TransformationHistory is the richer object. (3.7)
This is what is meant here by event-sourced research.
3.2 What counts as a research event?
Not every token produced by a model deserves to become a scientific event.
An event must correspond to an externally recordable occurrence that is potentially relevant to the evolution, evaluation, or provenance of the research programme.
Examples include:
• a source document being introduced;
• a claim being proposed;
• an assumption being made explicit;
• a counterexample being raised;
• a mapping being rejected;
• a residual being identified;
• a theory being restricted to a narrower domain;
• two conceptual tracks being merged;
• a branch being abandoned;
• a benchmark being run;
• a human researcher accepting or rejecting a proposed revision;
• an external expert challenging a claim;
• a new version of a document being produced.
Thus:
ResearchEvent = ExternallyRecordedStateRelevantOccurrence. (3.8)
The phrase externally recorded is essential.
The architecture does not require access to hidden cognitive state.
It requires only enough observable trace to support later reconstruction.
3.3 Events and artifacts must remain distinct
An event is something that happens.
An artifact is something produced, consumed, referenced, or modified by an event.
For example:
Event E₄₂ = “A model critiques Theory A.” (3.9)
Artifact A₁₇ = “The critique text produced in that interaction.” (3.10)
Likewise:
Event E₄₃ = “The researcher accepts Objection O₃ as valid.” (3.11)
Artifact A₁₈ = “A revised formulation of Theory A.” (3.12)
This distinction matters because one artifact may participate in many events.
A PDF can be:
introduced,
quoted,
reinterpreted,
challenged,
partially accepted,
and later superseded.
The event architecture must therefore preserve both:
EventIdentity. (3.13)
ArtifactIdentity. (3.14)
and their relations.
3.4 Event sourcing does not imply deterministic replay
In a conventional software event store, replaying the same events may reconstruct the same application state exactly.
Research is different.
Semantic interpretation is:
context-sensitive,
model-sensitive,
observer-sensitive,
and often probabilistic.
Therefore:
SameEventHistory ≠ GuaranteedSameReconstruction. (3.15)
Two independent reconstructors may disagree about:
whether a later claim is a refinement or replacement;
whether two formulations belong to one conceptual track;
whether an objection actually caused a revision;
whether a residual was resolved or merely renamed.
This is not a defect.
It is one reason why the reconstruction itself must be represented explicitly rather than silently embedded into a polished narrative.
3.5 Research events may form only a partial order
Human narratives tend to linearize history.
But theoretical research often develops in parallel.
Suppose two branches emerge from one question:
C₀ → C₁ₐ. (3.16)
C₀ → C₁ᵦ. (3.17)
The two branches may evolve independently before later comparison or merger.
Similarly, two source documents may be analyzed in parallel.
A benchmark may be run while a theoretical branch is still being revised.
Therefore the canonical event structure should not assume that all events fit into one total narrative order.
The more general relation is:
Eᵢ ≺ Eⱼ when Eᵢ is known to precede Eⱼ. (3.18)
But for some pairs:
Eᵢ ∥ Eⱼ. (3.19)
meaning that no meaningful ordering is required, or that they belong to parallel branches.
This makes a partial-order representation more natural than a simple transcript sequence.
3.6 Event sourcing and the declaration–ledger framework
The event-sourced approach also fits the earlier declaration–trace–residual architecture.
That framework describes a declared world in which projection passes through gate, trace, residual, and ledger before producing a stable historical order.
The later self-revision framework adds:
Dₖ₊₁ = Uₐ(Dₖ, Lₖ, Rₖ). (3.20)
where Dₖ is the current declaration, Lₖ is the ledgered trace, Rₖ is residual, and Uₐ is an admissible revision operator.
Translated into research architecture, this suggests:
ResearchDeclarationₖ₊₁ = Revise(ResearchDeclarationₖ | Traceₖ, Residualₖ). (3.21)
The important point is not the specific notation.
It is that research history is not merely a sequence of outputs.
Recorded trace and unresolved residual can change the rules governing later interpretation.
This makes path dependence a native property of theory formation.
4. The Four Contracts: Capture, Reconstruct, Project, Audit
The architecture can now be expressed through four contracts:
Capture → Reconstruct → Project → Audit. (4.1)
These contracts separate tasks that are often collapsed together.
4.1 Contract I — Capture
The Capture Contract specifies what must be preserved from the external research process.
A minimal captured event record should contain, where available:
• event identifier;
• timestamp or ordering information;
• actor or agent;
• model and version;
• source references;
• input artifacts;
• output artifacts;
• human decisions;
• tool calls and results;
• protocol version;
• document version;
• experiment identifiers.
The principle is:
CaptureBeforeInterpretation. (4.2)
That is, the raw record should not be rewritten merely because a later theory makes earlier events look obsolete.
A reconstruction may later say:
Claim C₁ was superseded. (4.3)
But the original event that produced C₁ must remain recoverable.
4.2 Capture must preserve lineage
This is especially important in AI-assisted cross-domain work.
Once a concept has been introduced into a long-running research context, later rediscovery may be contaminated by that ancestry.
The Semantic Collider framework explicitly warns that a candidate invariant can become a new conceptual beam, after which later interactions may repeatedly recover it because the research programme has become attracted to its own abstractions. It therefore recommends lineage tracking, alternative vocabularies, ablation, blinded evaluation, and random controls.
The Capture Contract must therefore preserve enough ancestry to distinguish:
Independent recurrence. (4.4)
from:
Inherited recurrence. (4.5)
Without that distinction, repetition can easily be misread as evidence.
4.3 Contract II — Reconstruct
The Reconstruction Contract turns captured records into structured research objects.
This is not merely extraction.
It may involve interpretation.
For example, an event sequence may support reconstruction assertions such as:
Claim C₂ revises Claim C₁. (4.6)
Constraint K₃ was introduced because C₁ violated requirement Q. (4.7)
Residual R₄ remained unresolved after revision C₂. (4.8)
Claim C₃ later resolves R₄. (4.9)
These are not raw facts in the same sense as timestamps.
Therefore the Reconstruction Contract must require:
• explicit relation typing;
• evidence links;
• reconstruction identity;
• confidence or epistemic status;
• alternative reconstruction where material;
• residual when interpretation remains incomplete.
The principle is:
InterpretationMustCarryItsOwnProvenance. (4.10)
4.4 Contract III — Project
The machine-native representation may be too large or structurally complex for direct human inspection.
A declared projection is therefore required.
Let:
M = machine-native research event representation. (4.11)
v = observer viewpoint. (4.12)
P = projection protocol. (4.13)
Then:
Π_(v,P)(M) → Hᵥ + R_Π. (4.14)
where Hᵥ is the resulting human-readable or task-readable projection, and R_Π is what the projection intentionally suppresses or fails to express safely.
Different projections may serve different functions.
A reviewer may request:
only claims with weak evidence and unresolved residuals.
A new collaborator may request:
active conceptual tracks plus current constraints.
A historian may request:
genealogy and branch structure.
A general reader may request:
one narrative sequence.
The canonical object should support all of them.
4.5 Contract IV — Audit
The Audit Contract requires reverse navigation.
Suppose a human-facing view says:
“The declaration model arose as a repair of the filtration model’s unresolved admissibility problem.”
That sentence should not be a free-floating machine-generated summary.
It should be possible to traverse:
Human Assertion
→ Reconstruction Assertion
→ Claim / Residual Objects
→ Relevant Events
→ Original Artifacts. (4.15)
Thus:
ProjectedAssertion → ProvenancePath. (4.16)
For important claims:
NoProvenancePath → AuditFailure. (4.17)
This is one of the strongest differences between Reconstructable Research and ordinary AI summarization.
4.6 Why the four contracts must remain separable
If Capture and Reconstruction are merged, interpretations may overwrite history.
If Reconstruction and Projection are merged, human readability may force premature semantic simplification.
If Projection and Audit are merged, polished narratives may hide what was omitted.
Therefore:
Capture ≠ Reconstruct ≠ Project ≠ Audit. (4.18)
They form a pipeline, but they should remain separately inspectable.
5. The Canonical Objects of MRER
The proposed machine-native object is called:
Machine-Native Research Event Representation — MRER.
A minimal MRER does not need to prescribe one storage format.
It does need to distinguish several classes of research object.
A provisional representation is:
MRER = {E, A, C, K, T, R, V, P, RA}. (5.1)
where:
E = Events. (5.2)
A = Artifacts. (5.3)
C = Claim States. (5.4)
K = Constraints. (5.5)
T = Transformations. (5.6)
R = Residuals. (5.7)
V = Evidence objects. (5.8)
P = Provenance records. (5.9)
RA = Reconstruction Assertions. (5.10)
These are semantic roles, not implementation classes.
5.1 Event
An Event is an externally recorded occurrence relevant to research-state evolution.
Examples:
SourceAdded. (5.11)
ClaimProposed. (5.12)
ObjectionRaised. (5.13)
ExperimentRun. (5.14)
RevisionAccepted. (5.15)
BranchAbandoned. (5.16)
The Event object answers:
What externally happened? (5.17)
It should not, by itself, answer:
What did the event mean? (5.18)
That belongs to reconstruction.
5.2 Artifact
An Artifact is a persistent object referenced, consumed, or produced by one or more events.
Examples include:
document;
response;
prompt;
dataset;
code;
equation;
diagram;
benchmark output;
review report.
Artifacts should be versionable.
Thus:
Aᵢᵛ = artifact i at version v. (5.19)
A single theoretical article may therefore have:
A₁¹ → A₁² → A₁³. (5.20)
But artifact version history alone does not tell us which conceptual claims changed.
That is why Claim State must be separate.
5.3 Claim State
A Claim State represents a proposition or structured theoretical commitment at a particular research stage.
A claim is not assumed to be immutable.
Let:
Cᵢᵏ = state k of conceptual lineage i. (5.21)
For example:
C₁¹ = “Recursion generates pre-time.” (5.22)
C₁² = “Recursion presents pre-time structure.” (5.23)
C₁³ = “Viewpoint-selected filtration discloses pre-time structure.” (5.24)
Whether these three belong to one lineage is itself a reconstruction assertion.
The key point is that theory identity cannot be reduced to sentence identity.
5.4 Constraint
A Constraint records what a claim or model must preserve, avoid, or satisfy.
Examples include:
K₁ = “Do not presuppose ordinary time when explaining the emergence of time.” (5.25)
K₂ = “A cross-domain mapping must preserve native-domain constraints.” (5.26)
K₃ = “A revision must not erase prior trace.” (5.27)
Constraints are essential because theory evolution is often best understood through what a proposal is not allowed to do.
The Semantic Collider framework makes this explicit: a mature conceptual beam must expose native constraints and failure modes so that false correspondences can be detected rather than freely reinterpreted.
Therefore:
SemanticSimilarity without ConstraintPreservation is weak evidence. (5.28)
5.5 Transformation
A Transformation records a structured change between research states.
Examples include:
Refinement. (5.29)
Restriction. (5.30)
Generalization. (5.31)
Mutation. (5.32)
Split. (5.33)
Merge. (5.34)
Replacement. (5.35)
Rejection. (5.36)
Resurrection. (5.37)
The Transformation object is more informative than a generic edge.
Instead of:
C₁ related-to C₂, (5.38)
we may record:
C₂ restricts C₁ because Constraint K₄ failed outside Domain D. (5.39)
This is the difference between a static knowledge graph and a research-state history.
5.6 Residual
A Residual is an unresolved remainder explicitly preserved after a claim, mapping, or reconstruction attempt.
Let:
Rⱼ = unresolved structure j. (5.40)
A residual may represent:
• an unexplained prerequisite;
• an unresolved contradiction;
• a domain mismatch;
• a missing mechanism;
• an untested assumption;
• a failed correspondence;
• an ambiguity in reconstruction.
The Semantic Collider framework makes residual preservation central. It distinguishes residuals from failed mappings and argues that failed mappings can reveal discriminating dimensions that a successful analogy would otherwise hide.
Therefore:
Residual ≠ NoiseToDelete. (5.41)
Residual = ResearchObject. (5.42)
5.7 Evidence
Evidence must remain separate from Claim State.
Let:
Vₖ → supports Cᵢ. (5.43)
or:
Vₖ → attacks Cᵢ. (5.44)
Evidence may include:
benchmark;
counterexample;
formal proof;
domain-expert review;
empirical observation;
holdout transfer;
replication;
external validation.
A generated theory does not become evidence merely because it is coherent.
Thus:
CandidateGeneration ≠ ClaimValidation. (5.45)
This separation is one of the core rules inherited from the Semantic Collider framework.
5.8 Provenance
Provenance records how a machine object connects back to recorded history.
A Claim State might link to:
specific model output;
specific paragraph;
specific prompt;
source document;
human correction;
benchmark result.
The provenance chain should be machine-resolvable wherever possible.
For example:
P(C₁³) = {E₄₁, E₄₂, A₁₇, A₁₈}. (5.46)
This allows later reviewers to distinguish:
what was recorded
from
what was reconstructed.
5.9 Reconstruction Assertion
A Reconstruction Assertion is a claim about the structure of the research history itself.
Examples:
RA₁ = “C₂ is a refinement of C₁.” (5.47)
RA₂ = “R₄ motivated the formation of C₃.” (5.48)
RA₃ = “C₅ was independently rediscovered rather than inherited.” (5.49)
RA₄ = “C₈ and C₉ belong to the same conceptual track.” (5.50)
These assertions require:
• reconstructor identity;
• reconstruction protocol;
• supporting provenance;
• epistemic status;
• competing interpretation where relevant.
Thus:
ReconstructionAssertion = RelationClaim + Provenance + EpistemicStatus. (5.51)
This is one of the defining innovations of MRER.
6. Why MRER Need Not Be Human-Readable
6.1 Human readability is a projection requirement
A common design mistake would be to insist that the canonical research representation must look like a readable diagram.
That would impose severe compression too early.
A long-running AI-assisted programme may contain:
10⁴ events, (6.1)
10³ claim states, (6.2)
10² active conceptual branches, (6.3)
and many more provenance links.
Human cognition cannot inspect such an object directly.
Therefore:
CanonicalComplexity should not be bounded by HumanDisplayBandwidth. (6.4)
Human readability belongs downstream.
6.2 The canonical object may be structurally heterogeneous
MRER may require several representational forms simultaneously.
Chronology may be efficiently encoded as a partial order.
Genealogy may resemble a directed acyclic graph.
Competing reconstructions may require probability distributions.
Semantic identity may use embeddings or structured feature sets.
Transformations may require typed hyperedges.
Evidence may require linked artifacts.
Residuals may behave like persistent open objects attached to multiple tracks.
Therefore:
MRER need not have one primitive geometry. (6.5)
A hybrid representation may be more appropriate.
6.3 Graph is an implementation possibility, not a metaphysical commitment
It is natural to speak loosely of a Semantic Event Graph.
That phrase is useful.
But the framework should not assume:
ResearchReality = Graph. (6.6)
A graph is one convenient encoding.
The stronger architectural requirement is:
MachineNativeRepresentation must preserve declared reconstruction invariants. (6.7)
If a hypergraph, relational store, or hybrid vector-symbolic system preserves them better, it should be admissible.
This follows the same discipline used elsewhere in the project: operational interfaces should not be confused with ontological claims. The Gauge Grammar explicitly treats its structural roles and control coordinates as protocol-relative interfaces rather than claims that every domain literally instantiates one underlying ontology.
6.4 Translation is more important than direct readability
The correct requirement is:
MRER must be approximately translatable into useful human views. (6.8)
Not:
MRER must itself be directly readable. (6.9)
This allows much richer machine representation.
A decoder may generate:
Narrative View. (6.10)
Track View. (6.11)
Residual View. (6.12)
Evidence View. (6.13)
Review View. (6.14)
Each projection can be lossy as long as the loss is declared.
6.5 Projection residual must remain visible
Suppose a human-facing narrative compresses twenty competing reconstruction hypotheses into one simple sentence.
That may be useful for teaching.
But the projection must preserve the fact that simplification occurred.
Let:
M = full machine representation. (6.15)
H = human projection. (6.16)
R_Π = omitted or untranslatable structure. (6.17)
Then:
M → H + R_Π. (6.18)
This mirrors the broader residual discipline of the declaration framework: a bounded observer should not confuse what was projected with what was absent from the projection.
The rule is:
HumanReadability must not erase ProjectionResidual. (6.19)
6.6 Machine-native does not mean epistemically opaque
This distinction is essential.
A machine-native object may be difficult for a human to inspect directly.
But its important contents should remain:
queryable;
traceable;
versioned;
projectable;
auditable.
Therefore:
MachineNative ≠ BlackBoxAuthority. (6.20)
The machine representation is allowed to be complex.
It is not allowed to become unaccountable.
7. Multiplex Relations: Chronology, Genealogy, Semantics, Causality, and Epistemic State
A major failure mode of research narration is relation collapse.
The statement:
“A led to B”
may hide several fundamentally different claims.
Therefore MRER should distinguish at least five relation planes.
Let:
ℛ = {ℛ_T, ℛ_G, ℛ_S, ℛ_C, ℛ_E}. (7.1)
where:
ℛ_T = temporal relations. (7.2)
ℛ_G = genealogical relations. (7.3)
ℛ_S = semantic relations. (7.4)
ℛ_C = causal relations. (7.5)
ℛ_E = epistemic relations. (7.6)
7.1 Temporal relation
The temporal layer records:
before;
after;
overlap;
parallel development.
Example:
E₁ ≺ E₂. (7.7)
This means only:
E₁ occurred before E₂.
It does not imply:
E₂ inherited from E₁. (7.8)
It does not imply:
E₁ caused E₂. (7.9)
This distinction should be enforced by the data model.
7.2 Genealogical relation
The genealogical layer records actual or reconstructed ancestry.
Examples:
derived-from;
inherits-from;
imports-from;
cites;
rewrites;
forks-from.
If a later model output contains an earlier theory because that theory was present in the prompt context, the relationship is genealogical even if the wording changes radically.
This layer is especially important for evaluating repeated discovery.
Semantic recurrence without genealogical independence is not strong replication.
The Semantic Collider framework explicitly warns that lineage must be tracked because recurrence inside one research programme can otherwise be mistaken for independent rediscovery.
7.3 Semantic relation
The semantic layer asks how two ideas relate in content, independent of ancestry.
Examples:
refines;
contradicts;
generalizes;
restricts;
operationalizes;
abstracts;
maps-to;
is-approximately-equivalent-to.
Two independently developed theories can be semantically equivalent.
Two directly genealogically related theories can be semantically very different.
Therefore:
Genealogy ≠ SemanticSimilarity. (7.10)
7.4 Causal relation
The causal layer asks a stronger question:
Did A materially contribute to the production or revision of B?
This should be treated cautiously.
Possible evidence levels include:
chronological plausibility;
explicit retrospective attribution;
contemporaneous recorded decision;
counterfactual replay or ablation.
The weak form is:
A may have influenced B. (7.11)
The stronger form is:
Removing A changes the probability that B is generated under matched conditions. (7.12)
The latter can become experimentally testable in replayable AI-assisted research environments.
But causal relation should remain optional when evidence is insufficient.
UnknownCausality = ValidState. (7.13)
7.5 Epistemic relation
The epistemic layer records the status of claims and relations.
Examples:
candidate;
supported;
disputed;
superseded;
falsified;
externally validated;
unresolved.
A theory may have:
high identity continuity
but
low evidential maturity.
Another may have:
low identity continuity
but
strong evidence in a narrow successor form.
Therefore:
TrackPersistence ≠ EvidentialStrength. (7.14)
The Semantic Collider evidence ladder makes the same distinction by separating internal plausibility, native-constraint checking, residual audit, independent recurrence, holdout transfer, and external validation.
7.6 Relations themselves require epistemic metadata
Suppose MRER contains:
C₂ refines C₁. (7.15)
That relation may be:
explicitly declared;
strongly inferred;
weakly inferred;
contested.
Thus a relation should be represented as:
RelationAssertion = RelationType + Evidence + Confidence + Provenance + Residual. (7.16)
The relation and the confidence in the relation are different objects.
This prevents a reconstruction engine from laundering uncertainty into structure.
7.7 Negative evidence must also be preserved
Suppose a reconstructor proposes:
A caused B. (7.17)
But another artifact shows that the core of B already existed before A.
That counterevidence should not simply disappear into a lower confidence score.
It should remain explicitly represented.
Thus:
EvidenceFor(RA) = {V₁, V₂, …}. (7.18)
EvidenceAgainst(RA) = {V₃, V₄, …}. (7.19)
This allows later reconstruction systems to reevaluate the relation rather than inheriting an opaque probability.
7.8 The core discipline
The entire multiplex architecture can be summarized by one rule:
Later Than ≠ Derived From ≠ Semantically Related To ≠ Caused By.
And one extension:
Relation Type ≠ Confidence in Relation.
These separations are necessary if AI-assisted research history is to become genuinely reconstructable rather than merely narratable.
[End of Part II — Sections 3–7]
The next installment will begin Part III — Reconstructing Theory Evolution, covering:
Section 8 — What Counts as the Same Conceptual Track?
Section 9 — Refinement, Mutation, Branching, Merge, Replacement, and Death
Section 10 — Residual as a First-Class Research Object
Section 11 — Competing Reconstructions
Part III — Reconstructing Theory Evolution
8. What Counts as the Same Conceptual Track?
8.1 Concept identity is not sentence identity
A central problem now appears.
If MRER is intended to reconstruct theory formation, then it must answer a difficult question:
When are two different formulations successive states of the same conceptual track, and when are they different theories?
This cannot be solved by lexical similarity alone.
Two propositions may use nearly identical words while changing their theoretical role completely.
Conversely, two propositions may use very different language while preserving the same underlying research function.
Therefore:
LexicalSimilarity ≠ ConceptIdentity. (8.1)
and:
LexicalDifference ≠ ConceptDeath. (8.2)
This becomes especially important in long-running AI-assisted research, because models are highly capable of paraphrasing, reframing, renaming, and reorganizing conceptual material.
A reconstruction architecture that tracks only textual similarity will therefore confuse:
paraphrase with continuity,
renaming with novelty,
deep revision with minor editing,
and conceptual replacement with ordinary reformulation.
8.2 A conceptual track should be modeled as a sequence of claim states
Let a conceptual track Tᵢ contain a sequence of states:
Tᵢ = {Cᵢ¹, Cᵢ², Cᵢ³, …}. (8.3)
Each Cᵢᵏ is not merely a sentence.
A useful decomposition is:
Cᵢᵏ = ⟨Qᵢᵏ, Ωᵢᵏ, Kᵢᵏ, Fᵢᵏ, Eᵢᵏ, Rᵢᵏ⟩. (8.4)
where:
Q = problem or question being answered, (8.5)
Ω = structural or ontological commitment, (8.6)
K = critical constraints, (8.7)
F = inferential or functional role, (8.8)
E = evidence state, (8.9)
R = residual state. (8.10)
This decomposition is intentionally provisional.
Its purpose is not to claim that every theory can be completely reduced to six fields.
Its purpose is to prevent MRER from equating:
concept = wording. (8.11)
A research concept is better understood as a structured state occupying a role within a larger evolving theory.
8.3 Problem continuity may be more stable than answer continuity
Consider the development from recursive generation to declared disclosure.
The early problem was approximately:
How can a pre-collapse field exhibit order before ordinary time exists? (8.12)
The first answer explored recursive depth:
RecursiveDepth → PreTimeOrdering. (8.13)
But this answer created a hidden problem.
If recursive generation unfolds step by step, then the model may already presuppose an ordering structure equivalent to a deeper time.
The later framework therefore weakened the claim.
Recursion became a presentation grammar rather than literal ontological generation. The pre-time field was instead disclosed through viewpoint-selected filtration.
Still later, filtration itself was judged incomplete because a viewpoint must declare the conditions under which the field becomes filterable: boundary, baseline, feature map, protocol, projection, gate, trace, and residual.
The answer changed substantially:
Recursive generation
→ recursive presentation
→ filtration
→ declaration.
Yet the research problem remained recognizable.
Thus:
AnswerContinuity may be low. (8.14)
ProblemContinuity may remain high. (8.15)
This is one reason concept identity cannot be determined only from proposition similarity.
8.4 Inferential role also matters
A concept often survives because it continues to perform the same theoretical job.
Suppose one framework uses:
recursive depth
to explain how ordered disclosure can exist before ordinary time.
A later framework uses:
filtration order
for the same purpose.
The mechanism changed.
The inferential role remained:
provide a non-clock-based ordering structure from which ledgered time can emerge. (8.16)
Therefore a track-identity system should ask:
What role is this concept playing in the theory? (8.17)
A concept that changes vocabulary but continues to solve the same structural problem under the same critical constraints may reasonably belong to the same research track.
8.5 Critical constraints may define identity more strongly than terminology
Constraints can provide an even stronger continuity signal.
For the pre-time sequence, one critical constraint became increasingly explicit:
Do not explain the emergence of time by silently assuming another time underneath it. (8.18)
The movement from recursion to filtration to declaration can therefore be read as a sequence of attempts to satisfy the same constraint more rigorously.
This suggests:
ConceptualTrackIdentity depends partly on ConstraintContinuity. (8.19)
A later formulation may differ radically from an earlier one yet still belong to the same lineage because it is a response to the same unresolved constraint.
The inverse is also possible.
Two formulations may sound extremely similar while satisfying different constraints and operating in different domains.
They should not automatically be collapsed into one track.
8.6 The Identity Kernel
To formalize this idea, MRER can define a protocol-relative Identity Kernel.
Let:
κ_P(C) = identity kernel of claim state C under reconstruction protocol P. (8.20)
A minimal κ may include:
κ_P(C) = {Q, F, K*, G}. (8.21)
where:
Q = core research problem, (8.22)
F = inferential role, (8.23)
K* = identity-critical constraints, (8.24)
G = genealogical support. (8.25)
Two concept states may be treated as belonging to one track when enough of κ survives.
Thus:
κ_P(Cᵢ) ≈ κ_P(Cⱼ) → possible track continuity. (8.26)
This is intentionally weaker than equality.
The approximation relation depends on reconstruction protocol.
8.7 Track continuity should itself be a reconstruction assertion
The statement:
C₂ and C₁ belong to the same track. (8.27)
is not necessarily a raw historical fact.
It is a reconstruction judgment.
Therefore MRER should encode:
RAₙ = SameTrack_P(C₁, C₂). (8.28)
with:
supporting provenance;
preserved identity features;
changed features;
alternative interpretations;
confidence;
reconstructor identity.
This prevents the reconstruction engine from hiding a major interpretive act behind an innocent-looking edge.
8.8 A provisional continuity function
For conceptual purposes, let:
Γ(Cᵢ,Cⱼ) = genealogical support. (8.29)
Π(Cᵢ,Cⱼ) = problem continuity. (8.30)
Φ(Cᵢ,Cⱼ) = inferential-role continuity. (8.31)
K(Cᵢ,Cⱼ) = critical-constraint continuity. (8.32)
Δ(Cᵢ,Cⱼ) = destructive conceptual change. (8.33)
Then:
TrackContinuity_P(Cᵢ,Cⱼ) = f_P(Γ, Π, Φ, K, Δ). (8.34)
No universal numerical function f_P is claimed here.
The equation is a structural placeholder.
Its purpose is to state that track continuity should be evaluated from multiple dimensions rather than lexical proximity alone.
8.9 Scale must also be declared
Concept identity may depend on scale.
Consider the concept:
Gate. (8.35)
A gate may appear in:
quantum measurement;
cellular signaling;
legal decision;
financial posting;
AI authorization.
The Gauge Grammar explicitly treats such recurrences as functional-role mappings rather than literal ontological identity.
Therefore:
FunctionalHomology ≠ SameConceptualTrack. (8.36)
A cross-domain role mapping may connect two tracks without merging them.
Track identity should therefore be evaluated under a declared scale and domain:
TrackIdentity_(P,S,D). (8.37)
where:
P = reconstruction protocol,
S = scale,
D = domain context.
8.10 Track identity should remain defeasible
A reconstruction may initially classify:
C₂ = major revision of C₁. (8.38)
Later evidence may show that:
C₂ actually arose independently,
or answered a different problem,
or replaced the original mechanism completely.
The classification should therefore be revisable.
This is consistent with the broader admissible self-revision principle: mature systems preserve trace and residual while allowing declarations to change under evidence.
The rule is:
TrackIdentity must be revisable without erasing prior reconstruction history. (8.39)
9. Refinement, Mutation, Branching, Merge, Replacement, and Death
Once track identity is explicit, theory evolution can be described using a transformation grammar.
A generic relation:
C₁ → C₂ (9.1)
is too weak.
MRER should distinguish different kinds of change.
9.1 Refinement
A refinement increases precision while preserving the identity kernel.
C₁ ─Refine→ C₂. (9.2)
Typical changes include:
more precise definition;
better notation;
additional boundary condition;
clarified scope;
stronger test criterion.
The underlying problem, role, and critical commitments remain substantially intact.
9.2 Restriction
A restriction narrows domain or claim strength.
C₁ ─Restrict→ C₂. (9.3)
For example:
Universal claim. (9.4)
may become:
Claim valid only under Protocol P and Scale S. (9.5)
Restriction should not automatically be treated as failure.
In mature science, restriction often increases epistemic quality.
The Semantic Collider research programme explicitly recommends reducing strong claims when ordinary analogy performs equally well or when holdout and anonymization tests fail.
Thus:
ClaimReduction can be epistemic progress. (9.6)
9.3 Generalization
A generalization expands the domain while preserving an identifiable mechanism or identity kernel.
C₁ ─Generalize→ C₂. (9.7)
Generalization carries a burden:
Scope ↑ ⇒ ValidationBurden ↑. (9.8)
The Semantic Collider framework warns that civilization-sized theories accumulate validation burden faster than their rhetorical breadth suggests.
Therefore MRER should record not only:
generalized-to,
but also:
new constraints,
new residuals,
new evidence requirements.
9.4 Mutation
A mutation changes important internal structure while preserving enough identity kernel for track continuity.
C₁ ─Mutate→ C₂. (9.9)
Example:
Recursive generation
→ viewpoint filtration. (9.10)
The mechanism changed substantially.
But:
the research problem remained,
the lineage is explicit,
the inferential role remained related,
the critical anti-meta-time constraint was strengthened.
This is a strong candidate for:
Major conceptual mutation. (9.11)
rather than simple refinement.
9.5 Identity Budget
Not every revision should be allowed to preserve track identity indefinitely.
Otherwise a theory can survive every contradiction by redefining itself.
This motivates an Identity Budget.
Let:
B_id(Cᵢ → Cⱼ) = tolerated identity-changing transformation under protocol P. (9.12)
Again, no universal scalar metric is claimed.
The concept expresses a governance rule:
A theory may revise some combination of mechanism, terminology, scope, and auxiliary assumptions while retaining track identity.
But if:
problem,
inferential role,
critical constraints,
and genealogy
all rupture, then continuity becomes increasingly implausible.
Thus:
IdentityKernelRupture → ReplacementCandidate. (9.13)
This principle protects MRER from unlimited retroactive continuity.
9.6 Branching
A theory may split.
C₀ → {C₁ᵃ, C₁ᵇ}. (9.14)
Branching is crucial because AI systems often prematurely unify competing ideas.
A model may respond:
“These are actually complementary.”
But genuine alternatives must be allowed to remain alternatives.
MRER should therefore preserve:
Branch A. (9.15)
Branch B. (9.16)
Competing assumptions. (9.17)
Distinct evidence states. (9.18)
The Semantic Collider framework similarly insists that collision outputs should preserve candidate invariants, residuals, failed mappings, and null outcomes separately rather than collapsing them immediately into one unified theory.
Therefore:
BranchPreservation = AntiPrematureUnification. (9.19)
9.7 Merge
Two independently evolving tracks may later combine.
C_A ─╲
Merge → C_C. (9.20)
C_B ─╱
A merge must preserve parentage.
It should record:
parent tracks;
inherited components;
discarded components;
new synthesis;
compatibility conditions;
remaining residual.
The merge must not rewrite history as though:
C_A = C_B. (9.21)
They were distinct tracks that later contributed to one descendant.
This distinction is central to genealogical honesty.
9.8 Replacement
Sometimes continuity fails.
A new theory may address the same broad problem yet replace the old conceptual architecture.
C_A ─Replace→ C_B. (9.22)
Replacement differs from mutation.
A mutation preserves enough identity kernel to justify continued track identity.
A replacement does not.
The previous theory remains historically relevant but should no longer be represented as an active state of the same track.
9.9 Rejection and falsification
A track may be explicitly rejected.
C_A ─Reject→ DeadTrack. (9.23)
But rejection has different causes.
A useful taxonomy includes:
Falsified. (9.24)
Superseded. (9.25)
Abandoned. (9.26)
Deprecated. (9.27)
Scope-exhausted. (9.28)
Absorbed. (9.29)
Dormant. (9.30)
These states should not be collapsed.
For example:
Falsified
means evidence strongly attacks the claim.
Superseded
may mean a better theory exists while the old one remains approximately valid in a restricted domain.
Abandoned
may mean no decisive test ever occurred.
These are epistemically different histories.
9.10 Dormancy
A track may simply stop developing.
Dormant(Cᵢ). (9.31)
Dormancy means:
not currently active,
not rejected,
not validated,
not necessarily obsolete.
This is useful in research programmes with many speculative branches.
9.11 Resurrection
A dormant or rejected track may later return.
DeadOrDormant(C₁) ─NewEvidence→ C₂. (9.32)
The system should record:
Resurrection. (9.33)
not:
continuous uninterrupted survival.
The interruption itself is part of research history.
9.12 Theory survival and evidence survival must remain separate
A track may survive many revisions without strong evidence.
Likewise, a narrow descendant may acquire strong evidence even though most of the ancestral theory has been abandoned.
Therefore MRER should maintain two independent dimensions:
IdentityContinuity. (9.34)
EvidenceMaturity. (9.35)
This gives a useful distinction:
HighIdentity + LowEvidence. (9.36)
LowIdentity + HighEvidence. (9.37)
The first describes a coherent but unvalidated long-running theory.
The second may describe a heavily revised descendant whose remaining narrow claim is well supported.
9.13 Conceptual curvature
Once a track contains multiple transformations, we can introduce a provisional notion of Semantic Track Curvature.
Consider:
C₀ → C₁ → C₂ → C₃. (9.38)
If each transition is small and mainly lexical:
CurvatureSemantic ≈ low. (9.39)
If the identity kernel survives while mechanisms, ontology, and constraints change substantially:
CurvatureSemantic ≈ high. (9.40)
The sequence:
generation
→ presentation
→ filtration
→ declaration
is a plausible example of high conceptual curvature.
This concept need not initially be reduced to a numerical geometry.
Its first purpose is descriptive:
A theory may preserve lineage while changing direction significantly.
10. Residual as a First-Class Research Object
10.1 Scientific narratives systematically underrepresent residual
A successful theoretical paper normally emphasizes what worked.
It reports:
the final framework;
the surviving equations;
the useful correspondence;
the validated prediction.
What failed to fit often appears only in:
limitations;
discussion;
future work.
AI-generated theoretical prose makes this bias even stronger because models are exceptionally good at creating presentation closure.
A long argument can end with a polished vocabulary and coherent synthesis even when major conceptual remainder is unresolved.
Therefore Reconstructable Research treats residual as a first-class object.
10.2 Residual is not merely error
A residual is not necessarily:
mistake;
noise;
garbage.
It means:
structure that remains unresolved, unabsorbed, incompatible, or insufficiently explained under the current reconstruction.
Let:
Rᵢ = Residual(Cⱼ | P). (10.1)
A residual may represent:
missing mechanism;
unresolved prerequisite;
counterexample;
unexplained asymmetry;
boundary mismatch;
failed correspondence;
unverified assumption.
The important rule is:
Residual must remain attached to its provenance. (10.2)
10.3 Failed mapping and residual are different
The Semantic Collider framework distinguishes these carefully.
A residual says:
some part remains outside the current abstraction.
A failed mapping says:
a specific proposed correspondence violates declared constraints and should be rejected.
Thus:
Residual ≠ FailedMapping. (10.3)
A failed mapping can be stronger evidence.
For example:
A ↔ B proposed. (10.4)
Constraint K violated. (10.5)
Therefore:
Mapping(A,B) rejected. (10.6)
This failure may reveal a discriminating feature that was previously invisible.
Hence:
FailedMapping → DiscriminatingDimension. (10.7)
That discriminating dimension can itself become a new research object.
10.4 Residual can generate successor theory
The strongest reason to preserve residual is that unresolved remainder often drives theoretical development.
The declaration sequence gives a concrete example.
Stage 1
Recursive generation was proposed as pre-time structure.
Residual:
Hidden meta-time. (10.8)
Stage 2
Filtration replaced literal recursive generation.
Residual:
What makes the field filterable? (10.9)
Stage 3
Declaration supplied boundary, feature map, protocol, gate, trace, and residual rules.
Residual:
How can declaration revise itself without rewriting failure? (10.10)
Stage 4
Admissible self-revision introduced trace-preserving and residual-honest constraints.
The theoretical sequence therefore follows approximately:
Theoryₖ → Residualₖ → SuccessorTheoryₖ₊₁. (10.11)
This is a research-dynamic relation worth preserving explicitly.
10.5 Residual pressure
A residual that persists across many events may exert increasing pressure on the programme.
We can define a provisional concept:
ResidualPressure(Rᵢ) = influence of unresolved residual Rᵢ on subsequent research transformations. (10.12)
Possible indicators include:
number of later references;
number of revisions triggered;
number of branches created;
number of successor claims attempting resolution;
duration of unresolved status.
No universal metric is claimed.
The concept simply recognizes:
RepeatedUnresolvedResidual ≠ MinorFootnote. (10.13)
It may be the centre of the next theory.
10.6 Residual has a lifecycle
A residual may move through states.
R_open → R_attacked → R_partially_resolved → R_closed. (10.14)
Or:
R_open → R_reclassified. (10.15)
Or:
R_open → successor theory. (10.16)
Or:
R_open → permanent boundary condition. (10.17)
A residual may even become a stable declaration:
“This framework does not address phenomenon X.” (10.18)
That is epistemically healthier than repeatedly pretending X has been solved.
10.7 Residual honesty as reconstruction discipline
The self-revising declaration framework defines mature revision partly through residual honesty: a self-modifying system should not erase unresolved contradiction merely to preserve coherence.
The same principle applies to research reconstruction.
A reconstruction engine should not convert:
unresolved,
ambiguous,
failed
into:
implicitly solved
for the sake of narrative smoothness.
Hence:
NarrativeClosure must not imply ResidualClosure. (10.19)
This is especially important for AI-generated summaries.
10.8 Residual view as a new research interface
Once residuals become canonical objects, a new kind of interface becomes possible.
Instead of asking:
What are the current conclusions?
a researcher can ask:
What remains unresolved across the programme? (10.20)
Which residual has survived the most theory revisions? (10.21)
Which residual generated the most successor branches? (10.22)
Which claims appear mature only because their residuals disappeared from the narrative? (10.23)
This is a fundamentally different research view.
It turns the archive from a collection of conclusions into a map of unresolved pressure.
11. Competing Reconstructions
11.1 There may not be one correct track assignment
Suppose we observe:
C₁ = “Recursion generates pre-time.” (11.1)
Later:
C₂ = “Viewpoint filtration discloses pre-time structure.” (11.2)
One reconstructor may assert:
C₂ is a major mutation of C₁. (11.3)
Another may assert:
C₁ was rejected and C₂ is a replacement. (11.4)
Both may be defensible.
A mature MRER should preserve this ambiguity rather than forcing premature resolution.
11.2 Competing reconstruction hypotheses
Let:
H₁ = Mutation(C₁,C₂). (11.5)
H₂ = Replacement(C₁,C₂). (11.6)
H₃ = RelationshipUnresolved(C₁,C₂). (11.7)
Then the machine layer may preserve:
P(H₁ | evidence) = p₁. (11.8)
P(H₂ | evidence) = p₂. (11.9)
P(H₃ | evidence) = p₃. (11.10)
with:
p₁ + p₂ + p₃ = 1. (11.11)
This probability notation is illustrative.
The standard need not require Bayesian probabilities specifically.
The essential point is:
multiple reconstruction states may coexist.
11.3 Machine-native representation should delay premature collapse
Human narrative has strong pressure toward one coherent history.
Machine-native representation need not.
It can preserve:
Track hypothesis A;
Track hypothesis B;
supporting evidence for each;
unresolved disagreement.
This yields the principle:
PreserveReconstructionMultiplicity until a declared gate justifies reduction. (11.12)
This is one area where machine-native representation has a genuine advantage over traditional narrative.
11.4 Reconstruction disagreement is informative
If three independent reconstructors receive the same archive:
ℛ₁(E) → M₁. (11.13)
ℛ₂(E) → M₂. (11.14)
ℛ₃(E) → M₃. (11.15)
and all identify approximately the same:
major tracks;
revision points;
residuals;
evidence dependencies,
then reconstruction stability is relatively high.
If they disagree dramatically, that disagreement is itself information.
Therefore:
ReconstructionDisagreement = DiagnosticSignal. (11.16)
It may indicate:
insufficient trace;
ambiguous concept identity;
poorly specified protocol;
model-specific bias;
or genuinely plural interpretation.
11.5 Representation stability should not require identical encoding
Two reconstructors may encode the same research history differently.
One may create:
a merge node.
Another may use:
two parent edges plus a synthesis transformation.
Byte-level identity is unnecessary.
The relevant question is whether declared invariants agree.
Let:
Inv(M) = selected structural invariants of reconstruction M. (11.17)
Then reconstruction stability can be expressed as:
Inv(M₁) ≈ Inv(M₂). (11.18)
Candidate invariants include:
major lineage;
critical branch points;
claim status;
residual continuity;
evidence direction;
known genealogy.
Thus:
StructuralAgreement > EncodingAgreement. (11.19)
11.6 Human views may legitimately disagree in presentation
The same MRER may generate:
a simplified narrative;
a detailed historiographic map;
an evidence-focused reviewer view.
These may look very different.
That is acceptable if declared invariants remain preserved.
Let:
H₁ = Π₁(M). (11.20)
H₂ = Π₂(M). (11.21)
Then:
H₁ ≠ H₂ is permissible. (11.22)
provided:
Inv_required(H₁) ≈ Inv_required(H₂). (11.23)
This is Cross-Projection Invariance.
11.7 Competing reconstruction must not become epistemic relativism
Allowing multiple reconstructions does not mean:
all reconstructions are equally valid.
A reconstruction can be evaluated by:
provenance quality;
constraint fit;
historical consistency;
cross-reconstructor stability;
negative evidence;
explicit author records;
replay evidence where available.
Thus:
PluralReconstruction ≠ EqualValidity. (11.24)
The architecture supports graded reconstruction.
11.8 Reconstruction depth limits what can be claimed
A final paper alone may support strong semantic comparison but weak historical causation.
A complete event archive may support stronger genealogy.
A replayable environment may support generative causal testing.
Therefore reconstruction assertions should be limited by the depth of available trace.
This prepares the next part of the paper.
The question is no longer only:
What structure can we reconstruct?
It becomes:
How deep is the archive, and what kinds of historical, semantic, or causal claims does that depth justify?
Part IV — Evidence, Causality, and Reconstruction Depth
The next part will distinguish:
chronology,
genealogy,
semantic relation,
and causal influence
at progressively stronger evidence levels.
It will also introduce:
Reconstruction Depth RD0–RD4
and examine how replayable AI-assisted research environments may permit a new form of generative causal analysis of theory formation.
[End of Part III — Sections 8–11]
Part IV — Evidence, Causality, and Reconstruction Depth
12. Later Than Is Not Caused By
12.1 Research narration routinely collapses distinct relation types
Human-readable intellectual history strongly favors causal prose.
Consider the sentence:
“The meta-time objection led to the filtration model, which was then refined into declaration.”
This sentence is compact and understandable.
But it may encode several logically different assertions:
The meta-time objection occurred before the filtration model. (12.1)
The filtration model inherited part of the earlier research lineage. (12.2)
The filtration model semantically addressed the earlier objection. (12.3)
The objection causally influenced the production of the filtration model. (12.4)
The declaration model semantically extended or corrected the filtration model. (12.5)
These assertions do not share the same evidential status.
A timestamp can strongly support (12.1).
Version history and explicit references may strongly support (12.2).
Textual comparison may support (12.3).
But (12.4) makes a stronger claim about causal influence.
Therefore Reconstructable Research adopts the discipline:
Later Than ≠ Derived From ≠ Semantically Related To ≠ Caused By.
The distinction is not cosmetic.
It determines what the reconstructed history is actually allowed to claim.
12.2 Chronology has the lowest interpretive burden
A temporal assertion may be written:
Eᵢ ≺ Eⱼ. (12.6)
This means only that event Eᵢ preceded event Eⱼ under the available event-ordering record.
Typical support may include:
timestamps;
conversation order;
version sequence;
publication history;
experiment logs.
Chronology can often be reconstructed with relatively little semantic interpretation.
However:
TemporalPrecedence(Eᵢ,Eⱼ) does not imply SemanticDependence(Eᵢ,Eⱼ). (12.7)
and:
TemporalPrecedence(Eᵢ,Eⱼ) does not imply CausalInfluence(Eᵢ,Eⱼ). (12.8)
This should be a hard architectural constraint.
12.3 Genealogy requires evidence of inheritance
A genealogical assertion states that some later research object inherited material from an earlier one.
Examples include:
C₂ derives from C₁. (12.9)
C₂ imports constraint K₁ from C₁. (12.10)
C₂ forks from C₁. (12.11)
C₂ reuses the conceptual vocabulary of C₁. (12.12)
Genealogy can be supported by stronger evidence than mere chronological proximity.
Examples include:
• explicit citation;
• direct inclusion in prompt context;
• copied or transformed text;
• version lineage;
• author declaration;
• preserved identifiers;
• explicit instruction to continue or revise prior work.
The Semantic Collider programme already identifies lineage tracking as essential because recurrence inside one research programme must not automatically be counted as independent replication.
Thus:
ObservedRecurrence + KnownAncestry ≠ IndependentRediscovery. (12.13)
This is especially important for LLM-assisted research because long contexts, retrieval systems, project memories, and source bundles can silently carry earlier abstractions forward.
12.4 Semantic relation does not require genealogy
Two theories can be structurally related even when neither produced the other.
Suppose two independently produced claims are:
C_A. (12.14)
C_B. (12.15)
A later evaluator may determine that:
C_B generalizes C_A. (12.16)
or:
C_A and C_B instantiate the same relational invariant under different vocabularies. (12.17)
This is a semantic relation.
It does not imply historical inheritance.
Thus:
SemanticRelation(C_A,C_B) can coexist with GenealogicalIndependence(C_A,C_B). (12.18)
That combination is particularly interesting because it may support a stronger case for independent structural recurrence.
12.5 Independent recurrence requires both similarity and ancestry control
A weak test says:
Two research runs produced similar conclusions. (12.19)
A stronger test asks:
Were the runs genealogically independent? (12.20)
The basic structure is:
SemanticRecurrence + GenealogicalSeparation → CandidateIndependentRecurrence. (12.21)
Even this does not establish truth.
But it is epistemically stronger than recurrence inside a shared contaminated context.
This is one reason why MRER must represent genealogy separately from semantic similarity.
12.6 Causal influence is a stronger claim
A causal reconstruction asks:
Did event A materially contribute to the production, rejection, restriction, or revision of B? (12.22)
This is stronger than saying:
A preceded B. (12.23)
or:
B solves a problem already visible in A. (12.24)
For example, the development from recursive generation to filtration contains unusually strong textual evidence because the later paper explicitly identifies the hidden meta-time problem in the earlier formulation and then replaces literal generation with viewpoint-selected disclosure.
That permits a relatively strong reconstruction assertion:
The meta-time objection was explicitly used as a reason to weaken recursive generation into disclosure. (12.25)
But even here, the reconstruction should remain grounded in the external record.
It should not claim access to some hidden internal moment of realization beyond the trace.
12.7 Four levels of causal evidence
A useful provisional causal ladder is:
C0 — No causal assertion
Only chronology or semantic relation is claimed.
Example:
A occurred before B. (12.26)
C1 — Reconstructed influence hypothesis
The later theory appears to answer a problem exposed earlier, but there is no explicit contemporaneous record that the earlier event caused the change.
Example:
A may have pressured the development of B. (12.27)
This is an inference.
C2 — Explicit retrospective attribution
A participant later states:
“B was developed because A revealed problem X.” (12.28)
This is stronger than C1 but remains retrospective.
Memory and later narrative compression may distort causation.
C3 — Contemporary recorded dependency
The research trace contains something like:
Objection O raised. (12.29)
Human researcher accepts O as a real problem. (12.30)
Explicit instruction: revise Theory T to solve O. (12.31)
Revised theory produced. (12.32)
This creates a much stronger external causal chain.
C4 — Interventional generative evidence
The research environment permits controlled replay.
One can compare:
P(R | S, O). (12.33)
with:
P(R | S, ¬O). (12.34)
where:
S = matched prior research state, (12.35)
O = objection, source, or constraint under test, (12.36)
R = target revision or conceptual transition. (12.37)
If the presence of O reliably changes the probability of R under matched conditions, then O has measurable generative causal influence within that declared environment.
This is the strongest causal class considered here.
It still does not reveal hidden cognition.
12.8 Historical causation and generative causation must remain distinct
These two claims are different.
Historical causal claim:
O caused the actual research trajectory to produce R. (12.38)
Generative causal claim:
Under declared replay conditions, adding O materially changes the probability that R emerges. (12.39)
The second is experimentally more tractable.
The first concerns what happened in the original history.
Therefore:
HistoricalCausation ≠ GenerativeCausalInfluence. (12.40)
A replay experiment may strengthen our understanding of the original trajectory.
It does not retroactively transform uncertainty into certainty about hidden mental causation.
12.9 Trace-conditioned systems are naturally path-dependent
The self-referential observer framework provides a useful formal analogy.
There, an observer records previous outcomes and selects future instruments as a measurable function of its accumulated trace. Counterfactual histories can therefore diverge because future interaction policy depends on earlier records.
AI-assisted research often has a similar externally observable structure.
Let:
Lₖ = accumulated external research trace before episode k. (12.41)
Pₖ = current research policy, prompt, source selection, or evaluation protocol. (12.42)
Then:
Pₖ = Policy(Lₖ). (12.43)
and:
Eₖ₊₁ = Generate(Pₖ, Contextₖ). (12.44)
Thus prior trace does not merely describe history.
When reintroduced into later context, it alters the conditions under which later research events are generated.
This gives theory formation genuine path dependence at the external process level.
12.10 Research declaration can also become history-dependent
The declaration sequence develops a closely related structure.
Part 4 writes admissible self-revision as:
Dₖ₊₁ = Uₐ(Dₖ,Lₖ,Rₖ). (12.45)
where:
Dₖ = current declaration,
Lₖ = ledgered trace,
Rₖ = residual,
Uₐ = admissible revision operator.
Applied to research architecture:
ResearchProtocolₖ₊₁ = Revise(ResearchProtocolₖ | Traceₖ, Residualₖ). (12.46)
This means that unresolved problems can alter:
what counts as evidence;
which sources are retrieved;
which constraints are declared;
which domains are excluded;
which tests become mandatory.
Theory formation can therefore change not only the theory.
It can change the rules under which subsequent theory is generated and evaluated.
That is a deeper form of research-state transformation.
12.11 Causal unknowns should remain unknown
A mature reconstruction architecture should tolerate statements such as:
Chronology = confirmed. (12.47)
Genealogy = probable. (12.48)
Semantic refinement = strongly supported. (12.49)
Causal influence = unresolved. (12.50)
This is not an incomplete reconstruction in a defective sense.
It is an honest one.
The governing rule is:
UnknownCausality = ValidCanonicalState. (12.51)
Reconstructable Research should prefer explicit uncertainty to invented explanatory smoothness.
13. Reconstruction Depth
13.1 Not all archives permit the same kinds of reconstruction
A central mistake would be to apply the same reconstruction confidence to every research archive.
A final paper does not contain the same historical information as:
a folder of versioned drafts.
A folder of drafts does not contain the same information as:
the full prompt–response history.
And even a full history does not provide the same experimental capabilities as:
a replayable environment with frozen models and controlled contexts.
Therefore this paper introduces Reconstruction Depth.
Let:
RD ∈ {RD0, RD1, RD2, RD3, RD4}. (13.1)
Reconstruction Depth describes the strength of the preserved external record and therefore the kinds of claims that can responsibly be made from it.
13.2 RD0 — Artifact-Only Reconstruction
At RD0, only one or more final artifacts are available.
Examples:
published paper;
final book;
finished codebase;
final theoretical manuscript.
Available reconstruction tasks may include:
• extracting claims;
• identifying internal structure;
• comparing propositions;
• identifying stated limitations;
• finding semantic correspondences.
But historical reconstruction is weak.
At RD0:
SemanticAnalysis = possible. (13.2)
HistoricalSequence = mostly unavailable. (13.3)
CausalReconstruction = highly constrained. (13.4)
A machine may infer that one section appears logically prior to another.
It must not pretend that this was the actual historical order of discovery.
Thus:
RD0 permits retrospective conceptual reconstruction, not strong research-history reconstruction. (13.5)
13.3 RD1 — Versioned Artifact Reconstruction
At RD1, multiple document versions survive.
For example:
Draft 1.
Draft 2.
Draft 3.
Final.
This permits stronger reconstruction of:
added claims;
deleted claims;
definition changes;
scope restrictions;
section movement;
new counterexamples;
new limitations.
A version sequence may show:
C₁¹ → C₁² → C₁³. (13.6)
This is useful.
But we may still not know why the revision occurred.
Therefore:
RevisionDetection = relatively strong. (13.7)
RevisionCause = often weak. (13.8)
13.4 RD2 — Prompt–Response Reconstruction
At RD2, a substantial part of the AI interaction history survives.
This may include:
prompts;
model responses;
uploaded sources;
explicit follow-up instructions;
human comments.
Now we can reconstruct more of:
genealogy;
source inheritance;
objection–revision sequences;
branch formation;
human selection decisions.
For example, we may observe:
Human: “This seems to introduce hidden meta-time.” (13.9)
AI: proposes filtration correction. (13.10)
Human: accepts the correction for next article. (13.11)
This is much stronger evidence of research-state transition than version comparison alone.
13.5 RD3 — Full External Research Trace
RD3 is the target for strongly instrumented Reconstructable Research.
It includes, where relevant:
• prompts and responses;
• source-document versions;
• model/version metadata;
• system-level research protocol versions where publishable;
• retrieval outputs;
• tool calls;
• code;
• experiment records;
• benchmark results;
• human accept/reject decisions;
• reviewer feedback;
• explicit branch and revision records;
• artifact hashes or equivalent identifiers.
RD3 permits strong reconstruction of many chronological and genealogical relations.
It may also support relatively strong contemporary causal assertions when an objection and subsequent corrective decision are explicitly recorded.
The external event trace now becomes sufficiently rich that the research programme can be reconstructed as more than a set of snapshots.
13.6 RD4 — Replayable Research Environment
RD4 adds controlled replay capability.
The objective is not necessarily exact deterministic reproduction.
Instead, the environment preserves enough of the generative conditions to conduct matched reruns.
Possible preserved variables include:
model family;
model version where available;
system instructions;
source bundle;
retrieval state;
prompt template;
sampling parameters;
tool configuration;
evaluation policy.
Then controlled interventions can be performed.
For example:
Run A: S + O. (13.12)
Run B: S − O. (13.13)
Run C: S + Sham(O). (13.14)
Compare the frequency and form of revision R.
This turns part of theory formation into an experimental object.
13.7 Reconstruction depth and claim privilege
The basic rule is:
HigherReconstructionDepth → PotentiallyStrongerReconstructionClaims. (13.15)
But:
HigherRD ≠ AutomaticTruth. (13.16)
A large archive may still be ambiguous.
A replay environment may still be unstable.
A model may still be sensitive to trivial prompt variations.
Reconstruction Depth specifies what evidence is available, not whether a claim is correct.
13.8 A provisional claim matrix
The following mapping is useful:
RD0
Strongest typical claim:
“These artifacts exhibit semantic relation X.” (13.17)
Weak claim:
“This is how the theory historically developed.” (13.18)
RD1
Strongest typical claim:
“The documented artifact changed from state A to state B.” (13.19)
Weaker claim:
“Objection O caused the change.” (13.20)
RD2
Strongest typical claim:
“The recorded interaction sequence connects O to revision R through explicit dialogue.” (13.21)
Still weaker:
“O was necessary for R.” (13.22)
RD3
Strongest typical claim:
“The externally instrumented research trace supports this genealogy and contemporary revision sequence.” (13.23)
Still unresolved in many cases:
“This factor was causally necessary.” (13.24)
RD4
Additional admissible claim:
“Under controlled replay conditions, intervention O changes the generative probability or structure of R.” (13.25)
This is a new kind of evidence.
13.9 Reconstruction Depth should accompany human-facing outputs
A Semantic Event Display derived from RD0 should not visually imply the same historical confidence as one derived from RD3.
Therefore human projections should declare:
RD = reconstruction depth. (13.26)
For example:
RD0 Retrospective Semantic Reconstruction
or:
RD3 Instrumented Research Reconstruction
This single label could prevent substantial overinterpretation.
13.10 Mixed-depth archives
Real research programmes will often be heterogeneous.
An early phase may have only final documents.
A later phase may preserve complete chat records.
Some experiments may be replayable.
Others may rely on external services that no longer exist.
Therefore Reconstruction Depth may need to be attached locally:
RD(Eᵢ). (13.27)
RD(Tⱼ). (13.28)
RD(RAₖ). (13.29)
rather than only globally to the whole programme.
A reconstruction assertion should inherit the weakest depth relevant to its supporting chain.
13.11 Historical reconstruction and semantic reconstruction have different depth requirements
A final artifact can sometimes support excellent semantic analysis.
For example:
Claim A contradicts Claim B within the same paper. (13.30)
No detailed history is required.
But:
Claim A was abandoned because of objection O. (13.31)
requires historical evidence.
Therefore:
SemanticDepth ≠ HistoricalDepth. (13.32)
A mature MRER may eventually distinguish multiple depth dimensions rather than a single scalar RD.
For the present article, RD0–RD4 is sufficient as a first engineering classification.
14. Generative Causal Influence by Replay
14.1 Why replay changes the epistemology of theory history
Most intellectual history is observational.
A historian can inspect:
letters;
drafts;
notebooks;
citations;
publication order.
But the historical process itself cannot usually be rerun.
AI-assisted theory formation creates a new possibility.
Under sufficiently controlled conditions, parts of the generative process can be replayed.
Suppose a research sequence contains:
State S. (14.1)
Objection O. (14.2)
Revision R. (14.3)
We can ask:
Would a comparable model under comparable conditions still produce R without O? (14.4)
This does not directly reveal the hidden cause of the original event.
But it permits an intervention test on the generative process.
That is a new experimental object.
14.2 The minimal replay design
Let:
S = declared prior research state. (14.5)
O = intervention object, such as an objection, concept beam, source, or constraint. (14.6)
G = generative research system. (14.7)
R = target revision pattern. (14.8)
Then compare:
G(S + O) → distribution over outcomes. (14.9)
G(S − O) → distribution over outcomes. (14.10)
The quantity of interest is not necessarily exact textual reproduction.
Instead, we may compare:
P(R | S,O). (14.11)
and:
P(R | S,¬O). (14.12)
If:
P(R | S,O) ≫ P(R | S,¬O), (14.13)
then O appears to exert generative influence on R under the declared setup.
14.3 The target must be defined structurally
Literal output matching would be a poor measure.
The model need not reproduce the exact phrase:
“viewpoint-selected filtration.”
It may instead generate:
“observer-conditioned disclosure ordering.”
If these satisfy the same declared structural criteria, they may count as the same target class.
Therefore R should be defined as:
R = structural revision class. (14.14)
not:
R = exact text string. (14.15)
This brings the Track Identity machinery back into the experiment.
14.4 Control conditions are essential
A useful replay design should contain more than:
with O / without O.
It may include:
Direct condition
S + O. (14.16)
Ablation condition
S − O. (14.17)
Sham condition
S + O_sham. (14.18)
where O_sham has similar vocabulary or rhetorical salience but does not contain the relevant structural objection.
Paraphrase condition
S + Paraphrase(O). (14.19)
Structural anonymization condition
S + AnonymizeTerms(O). (14.20)
If the effect survives semantic-preserving vocabulary change, this is stronger evidence that the influence is structural rather than lexical.
This parallels the Semantic Collider proposal to use structural anonymization and sham worlds to distinguish real relational recovery from attractive terminology.
14.5 Replay can test conceptual ancestry
Suppose a later invariant appears repeatedly.
There are competing explanations:
H_mem = inherited memory or context contamination. (14.21)
H_prompt = prompt-induced alignment. (14.22)
H_struct = robust structural recurrence. (14.23)
This decomposition was already proposed in the Semantic Collider framework.
Replay permits interventions.
Remove the earlier invariant from context.
Rename all domain vocabulary.
Change model family.
Change language.
Preserve only structural constraints.
Then ask whether the invariant reappears.
Repeated survival under such perturbation strengthens H_struct relative to H_mem and H_prompt.
It still does not prove a universal ontology.
14.6 Replay can test objection influence
The same method can examine whether an objection materially shapes theory revision.
Original sequence:
Theory A
→ Objection O
→ Theory B. (14.24)
Counterfactual conditions:
Theory A → model continuation without O. (14.25)
Theory A + O → model continuation. (14.26)
Theory A + irrelevant objection O′ → model continuation. (14.27)
We then compare:
frequency of B-like correction;
frequency of alternative correction;
persistence of original defect;
new residuals introduced.
This creates an experimental measure of objection efficacy.
14.7 Replay can test source contribution
Suppose a cross-domain theory is generated after introducing source document X.
We can compare:
Context + X. (14.28)
Context − X. (14.29)
Context + structurally matched unrelated document X′. (14.30)
Possible outputs include:
same invariant recovered;
different invariant;
no nontrivial invariant;
more false correspondences.
This permits a more disciplined answer to:
What did source X actually contribute to the generated theory? (14.31)
Again, the answer is generative, not metaphysical.
14.8 Human intervention can also be treated as an experimental variable
AI-assisted theory formation is rarely model-only.
A human researcher may:
select a branch;
reject an analogy;
introduce a constraint;
demand evidence;
change the research question.
Therefore the architecture should not treat human intervention as noise.
Let:
Hₖ = human intervention at episode k. (14.32)
Then a research transition may be written:
Sₖ₊₁ = G(Sₖ | Hₖ, Mₖ, Aₖ, Pₖ). (14.33)
where:
Mₖ = model configuration,
Aₖ = available artifacts,
Pₖ = declared protocol.
Future replay could test:
with human gate Hₖ
versus
without human gate Hₖ.
This may reveal how much conceptual discipline comes from the model and how much comes from human selection.
14.9 Generator–adversary–validator separation
Replay experiments should avoid using one model as:
generator,
judge,
and validator
without separation.
The Semantic Collider framework explicitly warns that multiple LLM judges do not automatically produce epistemic independence.
A stronger protocol separates roles:
Generator → proposes candidate. (14.34)
Adversary → attempts destruction. (14.35)
Validator → evaluates against declared criteria or external evidence. (14.36)
Human / external domain evidence → final adjudication where required. (14.37)
This preserves:
DiscoveryPower ≠ EpistemicAuthority. (14.38)
14.10 Replayability is model-relative
A serious limitation must be stated.
Modern hosted models may change.
Sampling infrastructure may change.
Safety policies may change.
Retrieval systems may change.
Tool environments may disappear.
Therefore exact historical replay may often be impossible.
The relevant goal is more modest:
matched generative perturbation, not perfect metaphysical recreation.
Thus:
Replay ≠ ReproductionOfHiddenOriginalState. (14.39)
Replay = ControlledPerturbationOfDeclaredGenerativeConditions. (14.40)
14.11 Stochasticity is not a defect
If the model is stochastic, a single rerun is weak evidence.
Repeated runs are preferable.
Let:
N = number of matched runs. (14.41)
Estimate:
p̂₁ = frequency of R under S + O. (14.42)
p̂₀ = frequency of R under S − O. (14.43)
Then compare:
Δp̂ = p̂₁ − p̂₀. (14.44)
The exact statistical machinery can be chosen later.
The conceptual point is that stochastic theory generation can be treated experimentally rather than narratively.
14.12 Structural outcome classification is itself a reconstruction problem
Another recursion appears.
To estimate whether R occurred, we need to classify generated outputs.
That classification may itself be performed by:
human experts;
independent models;
formal constraints;
hybrid evaluators.
Therefore:
ReplayExperiment → OutcomeReconstruction. (14.45)
OutcomeReconstruction → ReconstructionAssertions. (14.46)
Those assertions themselves require provenance.
So even experimentation does not remove the need for reconstructability.
It deepens it.
14.13 The causal layer should remain optional
A minimal MRER does not need causal replay.
A valid reconstruction can stop at:
chronology;
genealogy;
semantic transformation;
epistemic state.
Causal reconstruction should be added only when evidence supports it.
Therefore:
CoreMRER does not require ℛ_C. (14.47)
CausalLayer = OptionalExtension. (14.48)
This protects the architecture from hallucinating explanations merely because causal storytelling is attractive.
14.14 What generative causal analysis could eventually measure
If RD4 environments become practical, several new research questions become measurable:
Which objection most strongly changes theory direction? (14.49)
Which source introduces genuinely new structural content? (14.50)
Which conceptual invariant survives vocabulary ablation? (14.51)
Which human intervention prevents false unification? (14.52)
Which residual most reliably generates productive successor theories? (14.53)
Which apparent discovery disappears when inherited context is removed? (14.54)
Which theory transitions are robust across model families? (14.55)
These questions move beyond:
“What did the model write?”
toward:
“What perturbations change the geometry of theory formation?”
14.15 A new experimental layer of conceptual science
The strongest long-term implication is therefore not that AI can autonomously discover truth.
It is that part of conceptual theory formation may become externally manipulable.
Historically:
Concept formation → partly private and difficult to perturb experimentally. (14.56)
Under instrumented AI-assisted research:
Concept formation → partially externalized, logged, replayed, ablated, and compared. (14.57)
The Semantic Collider paper already points toward this possibility by proposing concept collision as a measurable reasoning protocol rather than merely a metaphor for interdisciplinary creativity.
Reconstructable Research provides the missing memory and representation architecture needed to make such experiments cumulative.
14.16 The limit remains external validation
Even if a theory repeatedly emerges under robust replay conditions, this does not establish that the theory is true.
At most, we have learned something about:
the generative stability of the concept;
its dependence on particular sources or objections;
its robustness across frames;
its independence from lexical contamination.
The external world remains the final adjudicator for empirical claims.
Thus:
GenerativeRobustness ≠ ExternalTruth. (14.58)
and:
DiscoveryProcessEvidence ≠ DomainValidation. (14.59)
This preserves the original Semantic Collider division of labor:
HumanScientist = EpistemicGovernor.
LLM = RelationalSearchInstrument.
ExternalWorld = FinalAdjudicator.
Part V — Human Projection and Scientific Use
The machine-native architecture is now sufficiently defined to return to the object that originally motivated this paper:
the Semantic Event Display.
But its role is now clearer.
It is not the canonical research object.
It is not the “true picture” of intellectual history.
It is not even necessarily the most important projection.
It is one human-facing interface over MRER.
The next sections therefore examine:
Section 15 — The Semantic Event Display
how machine-native research history can be rendered as tracks, vertices, branches, terminations, residual channels, and evidence attachments;
Section 16 — Declared Human Projections
why one MRER should support multiple observer-specific views rather than one universal visualization;
and
Section 17 — Traceable Translation
how every important human-facing assertion should remain navigable backward toward reconstruction assertions and original research events.
[End of Part IV — Sections 12–14]
Part V — Human Projection and Scientific Use
15. The Semantic Event Display
15.1 The display is not the research object
We can now return to the visual idea that originally motivated this paper.
Suppose an AI-assisted theoretical programme contains:
• thousands of externally recorded events;
• hundreds of claim states;
• branching conceptual lineages;
• abandoned hypotheses;
• residuals that later generate successor theories;
• different evidence levels;
• uncertain reconstruction relations;
• and several competing interpretations of how the theory developed.
No human reader can inspect the complete machine-native representation directly.
A Semantic Event Display is therefore a projection of that richer structure.
Let:
M = MRER. (15.1)
P_D = declared display protocol. (15.2)
Then:
D = Π_D(M | P_D). (15.3)
where D is a human-facing event display.
The distinction is essential:
SemanticEventDisplay ≠ MRER. (15.4)
The display is a view.
The machine-native representation is the reconstructable substrate from which the view is generated.
15.2 Why the collider event-display analogy remains useful
The particle-collider analogy survives this architectural correction precisely because a modern event display is not simply raw detector data shown directly to a physicist.
There is an intermediate reconstruction process.
At the level of functional analogy:
Raw Detector Signals
→ Reconstruction
→ Tracks / Objects
→ Event Display. (15.5)
The corresponding semantic architecture is:
Raw Research Events
→ Semantic Reconstruction
→ Research Objects / Tracks
→ Semantic Event Display. (15.6)
The analogy should stop here.
The proposal does not require semantic claims to possess literal physical momentum, charge, energy, or mass.
The useful borrowing is narrower:
• events;
• detected records;
• reconstructed tracks;
• interaction vertices;
• branching;
• termination;
• uncertainty;
• human-facing event visualization.
The guiding discipline is:
BorrowFunction, NotPhysicalOntology. (15.7)
This is consistent with the Gauge Grammar methodology, which explicitly uses physics concepts as functional roles rather than literal claims that higher-scale systems are quantum systems.
15.3 A semantic “track” is a reconstructed lineage
The most important object in a Semantic Event Display is the track.
A track should not merely connect similar sentences.
It represents a reconstructed conceptual lineage.
For example:
Recursive Generation
→ Recursive Presentation
→ Viewpoint Filtration
→ Declared Disclosure
→ Admissible Self-Revision. (15.8)
The track may remain visually continuous because a reconstruction protocol judges that enough of the identity kernel survives.
But the display should also show that the direction changes.
A major conceptual mutation should not look identical to a minor wording edit.
Therefore track segments may encode transformation classes such as:
Refinement. (15.9)
Restriction. (15.10)
Mutation. (15.11)
ReplacementCandidate. (15.12)
Transfer. (15.13)
The display is therefore not merely a genealogy tree.
It is a representation of theory motion through research state.
15.4 Vertices represent research interactions
A track becomes especially informative when it reaches a vertex.
A semantic vertex is an event or reconstructed interaction at which the state of one or more tracks changes materially.
Useful vertex classes include:
Introduction. (15.14)
Collision. (15.15)
Challenge. (15.16)
Revision. (15.17)
Split. (15.18)
Merge. (15.19)
Rejection. (15.20)
Residualization. (15.21)
Transfer. (15.22)
Replication. (15.23)
Validation. (15.24)
Falsification. (15.25)
For example:
Theory A
→ Challenge O
→ Revision B. (15.26)
Or:
Track A ─╲
Merge → Track C. (15.27)
Track B ─╱
Or:
Track A
→ Counterexample
→ ×. (15.28)
The visible geometry becomes a compressed representation of research transformation.
15.5 Track termination must remain visible
Conventional theory summaries tend to remove dead branches.
A Semantic Event Display should preserve them.
A terminated track may mean:
falsified;
abandoned;
superseded;
deprecated;
absorbed;
scope-exhausted.
These are not equivalent.
A useful display should therefore distinguish:
TrackTermination(type = Falsified). (15.29)
from:
TrackTermination(type = Abandoned). (15.30)
and:
TrackTermination(type = Absorbed). (15.31)
A dead branch is not clutter.
It may contain crucial information about which alternatives were actually considered and why they failed.
15.6 Branches should not be visually forced back together
Consider:
C₀
→ C₁ᵃ
→ C₂ᵃ. (15.32)
and simultaneously:
C₀
→ C₁ᵇ
→ C₂ᵇ. (15.33)
If both remain viable, the display should preserve the fork.
The human temptation to draw:
C₂ᵃ + C₂ᵇ → Grand Unified Theory (15.34)
should be resisted unless a real merge event exists.
Therefore:
VisualCompression must not imply ConceptualClosure. (15.35)
This is particularly important for AI-generated research, because generative models often produce rhetorically elegant syntheses even when the underlying alternatives remain incompatible.
15.7 Residuals should have their own visual channel
Residuals should not appear only as footnotes attached to successful tracks.
They can be visualized as persistent side channels.
For example:
Theory A
│
├── surviving claim track → Theory B
│
└── Residual R₁ → unresolved → successor question Q₂. (15.36)
A particularly important residual may later feed a new track:
R₁
→ New Research Question
→ Theory C. (15.37)
This reveals something ordinary argument maps often obscure:
The next theory may be born from what the previous theory failed to absorb.
The broader declaration framework already treats residual as something that must be carried rather than hidden, and later revision may explicitly depend on the ledgered residual.
15.8 Evidence should be displayed separately from genealogy
A major visual danger is to let a long-lived track look “more true” merely because it has survived many revisions.
The display should separate:
track continuity
from:
evidence maturity.
For example:
Track T₁ may have:
IdentityContinuity = high. (15.38)
EvidenceMaturity = speculative. (15.39)
Another track T₂ may have:
IdentityContinuity = low. (15.40)
EvidenceMaturity = externally validated. (15.41)
Therefore evidence should appear as an attached layer rather than being encoded implicitly through track length.
The display may show:
Candidate. (15.42)
Constraint-checked. (15.43)
Holdout-tested. (15.44)
Externally validated. (15.45)
Falsified. (15.46)
without implying that conceptual persistence alone provides evidence.
15.9 Reconstruction confidence must also remain visible
Suppose the display contains an edge:
C₂ ─MutatesFrom→ C₁. (15.47)
But this relation is only one possible reconstruction.
The display should distinguish:
directly recorded relation;
strong reconstruction;
weak reconstruction;
disputed reconstruction.
A simplified graphical convention might use:
solid segment = strongly supported reconstruction. (15.48)
dashed segment = inferred reconstruction. (15.49)
forked alternative = competing reconstruction. (15.50)
The exact graphical language is implementation-specific.
The requirement is conceptual:
DisplayCertainty must not exceed ReconstructionCertainty. (15.51)
15.10 A display can be scale-dependent
A full research programme may contain dozens of nested tracks.
At one scale, the display might show:
individual claims.
At another:
whole theory families.
At another:
cross-domain research programmes.
Thus:
D_S₁(M) ≠ D_S₂(M). (15.52)
This is not inconsistency.
It is coarse-graining.
But the scale must be declared.
For example:
Micro View
prompt → objection → claim revision.
Meso View
article → successor article → branch.
Macro View
research programme A ↔ research programme B.
The same machine-native object can support all three.
15.11 Event displays should support drill-down
A static infographic is useful.
An interactive event display is potentially much more powerful.
A researcher might click:
Track C₁³
and reveal:
claim text;
prior states;
supporting artifacts;
constraints;
evidence;
residuals;
alternative reconstruction hypotheses.
Then click:
Residual R₄
and see:
where it first appeared;
which later papers referenced it;
which successor claims attempted to resolve it.
The display thereby becomes an interface to MRER rather than merely an illustration.
15.12 The display should permit null events
Not every research collision produces a surviving track.
For example:
Theory A + Framework B
→ no nontrivial invariant survived. (15.53)
That should be displayable.
A null result may show:
Collision Vertex
→ failed mapping 1
→ failed mapping 2
→ no survivor. (15.54)
This is epistemically valuable.
It prevents the visualization itself from becoming biased toward successful synthesis.
15.13 The display is a detector image only in a disciplined sense
The Semantic Collider earlier used the phrase:
the manuscript is a compressed projection of a larger conceptual interaction trace.
The present framework sharpens that intuition.
A more precise sequence is:
Research Event
→ Recorded Trace
→ Reconstructed Event Object
→ Event Display
→ Human Interpretation. (15.55)
The display does not prove the theory.
It shows a reconstruction of how the theory was generated, transformed, attacked, revised, and evidenced.
Thus:
EventDisplay ≠ Validation. (15.56)
It is an instrument of inspection.
16. Declared Human Projections
16.1 There should be no universal human-readable view
If MRER is richer than human-readable representation, then there should not be one privileged display.
Different observers need different projections.
A reviewer asks different questions from:
an author,
a new collaborator,
a historian,
a domain expert,
an automated research agent.
Therefore:
OneMRER → ManyAdmissibleProjections. (16.1)
This is not a weakness.
It is a design feature.
16.2 Every projection requires a declaration
Let a projection protocol be:
P_Π = (v, g, s, i, c). (16.2)
where:
v = viewpoint / user role, (16.3)
g = projection goal, (16.4)
s = scale, (16.5)
i = invariants that must be preserved, (16.6)
c = compression or complexity budget. (16.7)
Then:
Π_(P_Π)(M) → H + R_Π. (16.8)
The exact tuple can later be expanded.
The conceptual requirement is:
HumanProjection must declare what it is trying to preserve. (16.9)
This follows the same general principle as the declaration framework, where readability requires explicit boundary, feature map, protocol, projection, trace, and residual rules.
16.3 Theory-Evolution View
A Theory-Evolution View emphasizes:
claim lineage;
major revisions;
branching;
replacement;
track death.
It suppresses many low-level interaction details.
A simplified output may resemble:
C₁
→ C₂
→ {C₃ᵃ, C₃ᵇ}
→ C₄. (16.10)
with transformations labeled.
Its purpose is:
How did the theory change?
16.4 Provenance View
A Provenance View emphasizes:
sources;
prompts;
models;
tools;
human interventions;
artifact versions.
It asks:
Where did this claim come from?
A conceptual track may be secondary.
The primary structure may look more like:
Source A
Human Prompt P
Model M
→ Artifact X
→ Claim C. (16.11)
This view is especially important when evaluating novelty or claimed independent rediscovery.
16.5 Evidence View
An Evidence View centers claims rather than narrative history.
For each claim:
Cᵢ
├── supporting evidence
├── attacking evidence
├── holdout status
├── replication status
└── external validation status. (16.12)
This view answers:
Which parts of the theory are actually supported?
It prevents narrative coherence from being mistaken for evidential closure.
16.6 Residual View
A Residual View may be one of the most useful new interfaces.
It asks:
What remains unresolved?
Instead of showing the theory as a completed structure, it shows:
R₁ → still open. (16.13)
R₂ → partially resolved by C₄. (16.14)
R₃ → generated Branch B. (16.15)
R₄ → unresolved for six revisions. (16.16)
This view turns attention toward the future research frontier.
16.7 Reconstruction-Disagreement View
For research historiography or audit, another view can show competing reconstructions.
For example:
C₁ → C₂
may have:
H₁ = mutation. (16.17)
H₂ = replacement. (16.18)
H₃ = unresolved. (16.19)
A simplified public article might suppress this ambiguity.
An audit view should reveal it.
This gives users access not only to:
research uncertainty
but also:
reconstruction uncertainty.
16.8 Human–AI collaboration view
AI-assisted authorship creates another useful projection.
The view may distinguish:
Human-originated event. (16.20)
Model-generated proposal. (16.21)
Tool-produced result. (16.22)
Human acceptance gate. (16.23)
External reviewer correction. (16.24)
This does not attempt to divide intellectual credit mechanically.
It makes the process inspectable.
For example:
Human question
→ Model candidate
→ Human rejection
→ Model alternative
→ Tool benchmark
→ Human adoption. (16.25)
This may be far more informative than a binary label:
“AI-assisted”.
16.9 Reviewer View
A reviewer usually does not need the entire research history.
A useful reviewer projection might prioritize:
high-impact claims;
low-evidence claims;
large scope jumps;
unresolved residuals;
branch suppression;
weak causal assertions;
uncertain novelty;
known failed mappings.
The view may therefore rank:
ReviewPriority(Cᵢ). (16.26)
This could make peer review more efficient by directing scarce human attention toward the least secure parts of the reconstructed theory.
16.10 Machine-Agent View
Human projections are not the only downstream use.
A future research agent might request:
active tracks;
current constraints;
unresolved residuals;
deprecated formulations;
evidence maturity;
genealogical contamination warnings.
This allows a new AI session to enter a long-running research programme without reading every historical artifact sequentially.
The architecture becomes a research memory substrate.
This is potentially one of the most important engineering consequences of MRER.
16.11 Narrative Paper View
The ordinary paper remains one valid projection.
But now its compression becomes explicit.
A paper projection might request:
Audience = advanced general reader. (16.27)
Goal = explain mature framework. (16.28)
Scale = article-level. (16.29)
Preserve = core claim lineage + decisive evidence + major residual. (16.30)
Suppress = routine prompt history + minor failed drafts. (16.31)
Then:
Paper = Π_paper(M | P_paper). (16.32)
This makes the paper’s omissions intelligible rather than invisible.
16.12 Projection residual
Every projection suppresses something.
Therefore every declared view should generate:
R_Π. (16.33)
Examples:
• minor branches hidden;
• causal ambiguity simplified;
• low-confidence semantic relations omitted;
• detailed provenance collapsed;
• competing track identities suppressed.
A projection should be allowed to be simple.
It should not be allowed to conceal that simplification occurred.
Thus:
Compression is admissible when its residual is recoverable. (16.34)
16.13 Cross-projection invariance
Two valid projections may look radically different.
Let:
H_A = Π_A(M). (16.35)
H_B = Π_B(M). (16.36)
There is no requirement that:
H_A = H_B. (16.37)
Instead, declared structural invariants should agree.
For example:
If both views preserve lineage, they should not disagree about a strongly established parent–child relation.
If both preserve evidence state, one should not call a claim externally validated while the other calls it falsified.
Therefore define:
Inv_P(H) = required invariants of projection H under protocol P. (16.38)
Then an admissible pair should satisfy approximately:
Inv_shared(H_A) ≈ Inv_shared(H_B). (16.39)
This is Cross-Projection Invariance.
16.14 Projection disagreement can expose hidden assumptions
If two projections unexpectedly disagree about a supposedly preserved invariant, there are several possibilities:
MRER contains inconsistency;
one decoder is faulty;
the projection declarations differ silently;
the reconstruction itself is unstable.
Thus:
CrossProjectionFailure → AuditTrigger. (16.40)
This turns disagreement into a diagnostic tool.
17. Traceable Translation
17.1 Human readability without traceability would reproduce the original problem
Suppose MRER exists, but the final human interface simply asks an LLM:
“Summarize the research history clearly.”
The model generates:
“The declaration framework emerged because the filtration theory failed to explain admissibility.”
This may be correct.
But if there is no path back to the reconstructed evidence, the architecture has gained little.
The human receives another polished statement produced by an opaque generative layer.
Therefore human readability is insufficient.
The projection must be traceable.
17.2 The basic backward path
Every important human-facing assertion should, where technically feasible, support navigation:
HumanAssertion
→ ReconstructionAssertion
→ MachineObjects
→ RawEvents
→ SourceArtifacts. (17.1)
For example:
H₁ = “Declaration emerged as a successor to filtration.” (17.2)
may resolve to:
RA₁₇ = successor relation between C_filtration and C_declaration. (17.3)
Supported by:
Residual R₈ = “What makes Σ filterable?” (17.4)
Claim C₂ = filtration model. (17.5)
Claim C₃ = declared-field model. (17.6)
Event E₇₁ = explicit question raised. (17.7)
Artifact A₂₃ = Part 2 article. (17.8)
Artifact A₂₄ = Part 3 article. (17.9)
The underlying source sequence explicitly identifies the unresolved question “What makes Σ filterable?” and answers it through declaration in the next stage.
The human-facing statement therefore becomes inspectable.
17.3 No Untraceable Important Assertion
This motivates one of the strongest proposed rules of the standard:
No Untraceable Important Human Assertion.
In more formal terms:
Importance(Hᵢ) ≥ θ → Exists(ProvenancePath(Hᵢ)). (17.10)
where θ is a declared importance threshold.
Not every decorative sentence needs full provenance.
But claims that materially affect:
theory interpretation;
evidence status;
genealogy;
causal history;
validation status
should be traceable.
17.4 Traceability should be graded
A human assertion may have different reconstruction strength.
For example:
Level T0
No traceable support.
Level T1
Links to reconstructed claim objects.
Level T2
Links to reconstruction assertions plus evidence.
Level T3
Links through reconstruction assertions to raw external events.
Level T4
Links to replay or external validation records where applicable.
This creates a human-facing traceability grade.
Again:
Traceability ≠ Truth. (17.11)
But higher traceability makes audit easier.
17.5 Translation should preserve epistemic qualifiers
Suppose machine state says:
SemanticRelation = probable refinement. (17.12)
CausalInfluence = unresolved. (17.13)
A poor decoder might output:
“B was developed from A.” (17.14)
This silently strengthens the claim.
A faithful decoder should output something closer to:
“B appears to refine A, while the causal dependence is not established.” (17.15)
Thus:
ProjectionStrength ≤ ReconstructionStrength. (17.16)
A decoder must not promote uncertainty into certainty merely for prose fluency.
17.6 Translation should preserve negative evidence
Suppose:
EvidenceFor(RA₁) = {V₁,V₂}. (17.17)
EvidenceAgainst(RA₁) = {V₃}. (17.18)
A human summary that mentions only V₁ and V₂ distorts the reconstruction.
Therefore a projection preserving epistemic status should either:
show the counterevidence;
or:
declare that counterevidence was suppressed in the current view.
This is an extension of residual honesty.
17.7 Translation should distinguish fact from reconstruction
A good human-facing interface may explicitly mark:
Recorded: The objection appeared before the revision. (17.19)
Reconstructed: The revision appears to address that objection. (17.20)
Unresolved: Whether the objection was causally necessary is unknown. (17.21)
This simple distinction could substantially improve the quality of AI-generated research history.
It separates:
record
from:
interpretation.
17.8 Source-span traceability
Whenever possible, provenance should resolve below the whole-document level.
Instead of:
Claim C came from Paper X. (17.22)
prefer:
Claim C derives from Artifact A, pages p–q, or equivalent stable source span. (17.23)
This is especially important for long theoretical documents.
A source may contain:
an early formulation;
a later correction;
a limitation;
and a contradictory appendix.
Whole-document attribution can hide this structure.
17.9 Transformation provenance
The provenance requirement should apply not only to claims but to transformations.
If MRER says:
C₂ restricts C₁. (17.24)
the user should be able to inspect:
what changed;
why it was classified as restriction;
which constraint motivated the narrowing;
whether the author explicitly endorsed it.
This makes theory evolution audit-ready.
17.10 Reconstruction provenance
The reconstructor itself must also be recorded.
A Reconstruction Assertion should contain something like:
RA_ID. (17.25)
Reconstructor. (17.26)
Model / version. (17.27)
Protocol version. (17.28)
Supporting objects. (17.29)
Alternative hypotheses. (17.30)
Epistemic status. (17.31)
Therefore:
Reconstruction is itself part of research trace. (17.32)
This creates an important recursion:
ResearchHistory
→ Reconstruction₁
→ Audit / Correction
→ Reconstruction₂. (17.33)
The reconstruction can evolve without erasing its own history.
17.11 Decoders must also be versioned
The same MRER may generate different prose when the decoder improves.
Therefore:
DecoderVersion = part of projection provenance. (17.34)
A human-facing statement should be reproducible at least to the level of:
which MRER version,
which projection protocol,
which decoder version
produced it.
Otherwise a displayed research history may silently change over time.
17.12 Reverse navigation is more important than reversibility
A human-readable projection cannot normally reconstruct the full machine-native object.
Thus:
H → M exactly (17.35)
is not required.
What is required is:
H → relevant portion of M. (17.36)
and:
M → supporting external trace. (17.37)
This is reverse navigation, not information-theoretic reversibility.
The difference matters.
Human compression may be extremely lossy and still remain scientifically useful if important assertions can be followed backward.
17.13 Traceable translation turns summaries into interfaces
A conventional summary is terminal:
Archive → Summary. (17.38)
A traceable projection is navigational:
Archive
→ MRER
→ Human View
↔ Drill-Down. (17.39)
The reader can move between abstraction levels.
This transforms a summary from a final answer into an interface to the research object.
17.14 The paper itself could become traceable
A mature implementation could embed stable claim identifiers into a conventional paper.
A paragraph might correspond to:
Claim Set {C₁₇,C₂₃,C₄₂}. (17.40)
A figure might correspond to:
Track Set {T₃,T₄}. (17.41)
A limitation paragraph might link to:
Residual Set {R₈,R₁₂}. (17.42)
The paper remains readable normally.
But a machine or interested reader could open the deeper event structure.
This suggests a future publication model:
Paper + Reconstructable Backplane. (17.43)
The paper becomes the human surface.
MRER becomes the research backplane.
17.15 Why approximate translation is enough
The architecture does not require perfect semantic translation from machine representation to prose.
That would be unrealistic.
Instead, it requires:
declared purpose;
declared preserved invariants;
declared uncertainty;
recoverable residual;
backward provenance.
Thus:
ApproximateHumanTranslation + Auditability > PretendedLosslessNarrative. (17.44)
This is the correct standard.
17.16 A final projection contract
A minimal projection record might contain:
Projection ID. (17.45)
MRER version. (17.46)
Observer role. (17.47)
Purpose. (17.48)
Scale. (17.49)
Preserved invariants. (17.50)
Suppressed dimensions. (17.51)
Projection residual. (17.52)
Decoder/version. (17.53)
Timestamp. (17.54)
This turns human-readable outputs into reproducible scientific artifacts rather than disposable summaries.
Part VI — Research Programme
The architecture has now moved from a motivating metaphor to a testable engineering proposal.
The remaining question is therefore not:
Can we draw attractive semantic event diagrams?
It is:
Does reconstructable research provide measurable scientific value?
The final substantive section will define three progressively stronger hypotheses:
Representation Hypothesis — AI-assisted research histories can be represented as useful machine-native research event objects.
Reconstruction Hypothesis — independent reconstructors can recover materially stable structures from the same external history.
Scientific-Utility Hypothesis — reconstructable research improves human or machine ability to audit, understand, continue, and evaluate theory formation.
It will then specify failure conditions, a minimal prototype, candidate benchmarks, and what results would justify reducing the proposal to a more modest provenance or research-memory system.
The final conclusion will return to the core shift:
The paper is not discarded. It becomes one declared projection of a reconstructable research object.
Part VI — Research Programme
18. What Would Make Reconstructable Research Scientifically Useful?
18.1 The architecture must earn its complexity
A new research representation standard should not be justified merely because it is conceptually elegant.
Reconstructable Research introduces substantial additional machinery:
• event capture;
• artifact versioning;
• claim-state extraction;
• transformation typing;
• residual tracking;
• reconstruction assertions;
• provenance paths;
• multiple projection protocols;
• reconstruction audits.
This creates real costs.
Storage increases.
Research instrumentation becomes more demanding.
Reconstruction systems require additional computation.
Human researchers may face more metadata.
A badly designed implementation could become bureaucratic overhead rather than scientific infrastructure.
Therefore the proposal should be judged by a simple principle:
AddedStructure must produce AddedResearchCapability. (18.1)
If MRER provides no meaningful advantage over:
good version control,
careful note-taking,
ordinary provenance,
or a well-written summary,
then the stronger proposal should be reduced.
This article therefore separates three increasingly demanding hypotheses.
18.2 Hypothesis 1 — The Representation Hypothesis
The weakest claim is:
Externally recorded AI-assisted theory formation can be compiled into a machine-native representation that preserves materially useful distinctions among research events, claim states, transformations, constraints, residuals, evidence, genealogy, and reconstruction uncertainty.
Call this:
H_R = Representation Hypothesis. (18.2)
The hypothesis does not require that the representation reveal scientific truth.
It asks whether the history can be represented in a structurally useful way.
A successful MRER should preserve at least:
ResearchState. (18.3)
TransformationHistory. (18.4)
ClaimLineage. (18.5)
ConstraintHistory. (18.6)
ResidualHistory. (18.7)
EvidenceState. (18.8)
ReconstructionUncertainty. (18.9)
If these objects cannot be represented consistently enough to support later querying and projection, the architecture fails at its first level.
18.3 A minimal test of the Representation Hypothesis
Take a well-instrumented AI-assisted research sequence containing:
• multiple draft generations;
• at least one significant theoretical correction;
• one branch;
• one rejected mapping;
• one unresolved residual;
• one later successor claim;
• some explicit human intervention.
Construct MRER from the external trace.
Then ask whether the representation can answer queries such as:
What is the current active formulation of Claim C? (18.10)
What formulation did it replace? (18.11)
Which constraint triggered the change? (18.12)
What residual remained after the revision? (18.13)
Which later claim attempted to resolve that residual? (18.14)
What external artifacts support this genealogy? (18.15)
Which reconstruction relations remain uncertain? (18.16)
A representation that cannot answer these questions reliably is not yet a theory-history representation.
It is merely an archive with labels.
18.4 Hypothesis 2 — The Reconstruction Hypothesis
The second claim is stronger.
Suppose several independent reconstructors receive the same external research trace.
Do they recover materially similar theory structure?
Call:
H_X = Reconstruction Hypothesis. (18.17)
The hypothesis is:
Under a sufficiently specified reconstruction protocol, independent reconstructors can recover substantially stable structural features of the same research history, even when their exact machine encodings differ.
Let:
ℛ₁(E) → M₁. (18.18)
ℛ₂(E) → M₂. (18.19)
ℛ₃(E) → M₃. (18.20)
Exact equality is neither expected nor desirable:
M₁ ≠ M₂ ≠ M₃. (18.21)
Instead, define a declared set of structural invariants:
I* = {major lineage, branch points, major revisions, residual continuity, evidence direction, known ancestry}. (18.22)
Then test:
Inv_I*(M₁) ≈ Inv_I*(M₂) ≈ Inv_I*(M₃). (18.23)
This is Reconstruction Stability.
18.5 Reconstruction stability is not truth
A crucial caution is required.
Suppose ten independent models reconstruct the same lineage.
That does not establish that the underlying theory is scientifically true.
It establishes something narrower:
the lineage is robustly reconstructable from the supplied external evidence under the tested protocols.
Therefore:
ReconstructionStability ≠ DomainTruth. (18.24)
Likewise:
SemanticTrackStability ≠ EmpiricalValidation. (18.25)
This distinction prevents reconstruction agreement from becoming a new form of consensus laundering.
18.6 What should reconstructors be independent from?
Independence itself requires care.
Three runs of the same model with identical prompts are weakly independent.
Three different models sharing the same reconstructed summary may not be genealogically independent at all.
Possible levels of reconstruction independence include:
R-Independence 0
Same model, same protocol, repeated run.
R-Independence 1
Same model, independently sampled reconstruction.
R-Independence 2
Different model families, same protocol.
R-Independence 3
Different model families and independently authored reconstruction instructions.
R-Independence 4
Human expert + machine reconstructor + independent machine family.
Higher independence provides stronger evidence that recovered structure is not merely an artifact of one reconstruction attractor.
The broader project repeatedly emphasizes this distinction between recurrence and genuine independence: trace ancestry must be known before repeated structure is interpreted as evidence.
18.7 Hypothesis 3 — The Scientific-Utility Hypothesis
The strongest practical claim is:
MRER and its projections improve research work compared with ordinary document-centered workflows.
Call:
H_U = Scientific-Utility Hypothesis. (18.26)
This can be decomposed into several measurable tasks.
18.8 Utility Task A — Theory comprehension
Give researchers a large AI-assisted theory programme.
Group A receives:
final papers + ordinary summaries.
Group B receives:
the same materials + MRER-driven theory-evolution view.
Then ask:
Which early assumption was later removed? (18.27)
Which objection caused a major revision? (18.28)
Which residual remains unresolved? (18.29)
Which branch was abandoned rather than falsified? (18.30)
Which claim is still speculative? (18.31)
Measure:
accuracy;
time to answer;
confidence calibration.
If MRER provides no advantage, its value as a comprehension tool is weak.
18.9 Utility Task B — Error detection
A more demanding experiment asks whether reconstructable views improve the detection of theoretical problems.
Create or select research histories containing planted or naturally occurring issues such as:
• silent scope expansion;
• contradictory versions;
• hidden ancestry;
• unresolved residual laundering;
• false independent rediscovery;
• unsupported causal narrative;
• branch suppression.
Compare ordinary paper review against MRER-assisted review.
Measure:
DetectionRate_MRER. (18.32)
versus:
DetectionRate_Baseline. (18.33)
The practical claim becomes:
ΔDetect = DetectionRate_MRER − DetectionRate_Baseline. (18.34)
If:
ΔDetect ≤ 0, (18.35)
then the proposed architecture has not demonstrated value for this task.
18.10 Utility Task C — Research continuation
One of the strongest possible uses is research memory.
Suppose a new AI or human collaborator enters a project after two years.
Baseline workflow:
read the current papers and project summaries.
MRER workflow:
receive:
• active tracks;
• deprecated claims;
• open residuals;
• critical constraints;
• evidence state;
• branch history;
• provenance pointers.
Then assign a continuation task.
For example:
Propose the next theoretically justified experiment without reviving a previously rejected assumption.
Measure:
time to competent continuation;
number of obsolete assumptions reintroduced;
number of known residuals missed;
number of false novelty claims.
This tests whether MRER functions as a genuine research memory architecture rather than merely a publication archive.
18.11 Utility Task D — Novelty auditing
A generated claim may appear novel because the current session does not contain the earlier wording.
But the research project may already contain structurally equivalent material.
MRER can test:
NewClaim C_new. (18.36)
against:
existing track states {C₁,C₂,…}. (18.37)
The novelty question can then be decomposed:
LexicallyNovel? (18.38)
SemanticallyNovel? (18.39)
GenealogicallyIndependent? (18.40)
StructurallyNovel? (18.41)
This could reduce one important failure mode of AI-assisted theory formation:
FalseNovelty = Renaming + LostGenealogy. (18.42)
18.12 Utility Task E — Residual discovery
Traditional retrieval often asks:
What do we already know?
A residual-aware system asks:
What did we repeatedly fail to solve?
This produces a different research search space.
For example:
Rank residuals by:
age;
number of unsuccessful closure attempts;
number of successor theories;
evidence conflict;
cross-domain recurrence.
A research agent might then ask:
Which unresolved residual has the highest expected theoretical leverage? (18.43)
This is a potentially powerful use of the architecture.
The self-revising declaration framework already assigns residual an active role in future revision rather than treating it as discardable remainder.
MRER turns that principle into searchable research infrastructure.
18.13 Utility Task F — Projection comparison
Another experiment can test whether different human projections preserve declared invariants.
Generate from one MRER:
Narrative View. (18.44)
Track Display. (18.45)
Evidence View. (18.46)
Reviewer View. (18.47)
Then test whether users infer consistent answers to shared questions.
For example:
Which claim superseded C₁? (18.48)
Was C₂ externally validated? (18.49)
Did C₃ independently rediscover C₂? (18.50)
If different projections systematically induce contradictory answers, then:
ProjectionContract has failed. (18.51)
This is a direct test of Cross-Projection Invariance.
18.14 Utility Task G — Reconstruction audit
A particularly demanding benchmark can deliberately introduce reconstruction errors.
For example:
• misclassify replacement as refinement;
• convert inferred causation into recorded causation;
• hide genealogy;
• suppress negative evidence;
• mark unresolved residual as closed.
Then test whether auditors can identify these defects more reliably when reconstruction assertions expose:
relation type;
support;
counterevidence;
confidence;
provenance.
This tests whether the architecture genuinely makes reconstruction more inspectable.
18.15 The minimum viable prototype
A first implementation does not require a universal scientific standard.
It can be much smaller.
A Minimum Reconstructable Research Prototype — MRRP v0.1 might implement only:
Capture
prompt ID;
response ID;
artifact ID;
source links;
timestamp;
model metadata.
Semantic Objects
claim state;
constraint;
residual;
evidence.
Transformations
refine;
restrict;
replace;
branch;
reject.
Reconstruction Assertion
relation;
support;
confidence;
reconstructor.
Projections
theory timeline;
residual ledger;
claim provenance view.
Audit
clickable path from displayed claim back to source artifact.
That would already be enough to test whether the core proposal is useful.
18.16 Do not begin with universal ontology
A major engineering risk would be attempting to design the final universal research ontology immediately.
The existing semantic-compiler work provides an important lesson:
first preserve intent and operational structure, then compile into a minimal intermediate representation; do not optimize for conceptual elegance before functional preservation.
The same rule should apply here.
Therefore:
MRER_v0.1 should be minimal. (18.52)
Add an object or relation only when it improves:
reconstruction;
querying;
audit;
continuation;
or testing.
The standard should evolve through use.
18.17 A prototype benchmark corpus
A useful first benchmark could contain four classes of research histories.
Class A — Linear Evolution
One claim undergoes several explicit refinements.
Purpose:
test simple lineage recovery.
Class B — Branching Evolution
One theory splits into two competing alternatives.
Purpose:
test whether reconstruction preserves branches rather than prematurely merging them.
Class C — Residual-Driven Evolution
One unresolved problem generates a successor theory.
Purpose:
test residual continuity.
Class D — Contaminated Recurrence
A later model apparently rediscovers a prior concept, but earlier material remains in hidden or indirect project ancestry.
Purpose:
test genealogy auditing.
Class E — Null Development
Many interactions occur but no stable new theory emerges.
Purpose:
ensure the system is capable of reconstructing:
NoNontrivialSuccessor. (18.53)
rather than inventing a conceptual track merely because the archive is large.
18.18 Synthetic research histories should also be used
Real archives have ambiguous ground truth.
Synthetic histories allow known structure.
Construct a hidden research event topology containing:
• known track identity;
• known branch point;
• known replacement;
• known residual;
• known causal intervention;
• known false semantic lure.
Then generate natural-language artifacts around that topology.
The reconstruction system sees only the artifacts.
The hidden event structure becomes the benchmark key.
This parallels the broader use of synthetic worlds in the Semantic Collider programme: when training-data contamination and uncontrolled real-world semantics make evaluation difficult, construct artificial worlds with hidden relational invariants and sham correspondences. The goal is not to prove external truth, but to measure whether the method recovers known structure under controlled conditions.
For Reconstructable Research:
HiddenEventTopology → GeneratedResearchArchive → Reconstruction → CompareToGroundTruth. (18.54)
This could become the first rigorous benchmark.
18.19 Reconstruction scoring should be relational, not lexical
A poor benchmark would measure whether the reconstructor reproduces the expected sentences.
The target should instead include:
node-role recovery;
transformation classification;
lineage recovery;
residual attachment;
evidence direction;
branch preservation;
causal-status calibration.
Let:
Score_R = w₁S_lineage + w₂S_transform + w₃S_residual + w₄S_evidence + w₅S_branch + w₆S_calibration. (18.55)
The weights are application-specific.
Equation (18.55) is only a template.
The central point is:
ReconstructionQuality should evaluate relational structure rather than prose similarity. (18.56)
18.20 Track identity should be tested adversarially
Because track identity is one of the most difficult reconstruction judgments, benchmark cases should deliberately include:
Same wording, different concept
Lexical similarity high.
Identity continuity low.
Different wording, same concept
Lexical similarity low.
Identity continuity high.
Major mutation
Mechanism changes, but problem and role survive.
Replacement
Same problem, but identity kernel ruptures.
Cross-domain homology
Functional role similar, but genealogy and domain identity distinct.
A successful system should distinguish these cases.
18.21 Residual honesty should be scored explicitly
A reconstruction that correctly recovers ten claims but silently deletes one decisive residual may be worse than a less complete but more honest reconstruction.
Therefore a benchmark should include:
ResidualRecall. (18.57)
ResidualPrecision. (18.58)
FalseClosureRate. (18.59)
where:
FalseClosureRate = fraction of unresolved residuals incorrectly represented as resolved. (18.60)
This may be one of the most important metrics.
18.22 Causal overclaiming should receive a strong penalty
Because causal storytelling is especially attractive, the system should be penalized when it converts:
temporal order
into:
causal dependence
without adequate evidence.
Define conceptually:
CausalPromotionError = inferred stronger causal class − supported causal class. (18.61)
A system that frequently commits CausalPromotionError should fail the reconstruction benchmark even if its narratives sound compelling.
18.23 The architecture needs ablation studies
If MRER appears useful, we should ask which components actually create the benefit.
Compare:
Full MRER. (18.62)
MRER − Residuals. (18.63)
MRER − Constraints. (18.64)
MRER − Genealogy. (18.65)
MRER − Reconstruction Assertions. (18.66)
MRER − Reverse Audit. (18.67)
If removing a component causes no measurable degradation across tasks, that component may not belong in the minimal standard.
This is important because the framework should become smaller when evidence permits.
18.24 A direct baseline is ordinary structured summarization
The strongest competitor may not be raw papers.
It may be:
an excellent LLM-generated structured summary.
For example:
• main claims;
• evidence;
• limitations;
• timeline;
• open questions.
Therefore MRER should be compared against:
B₀ = raw documents. (18.68)
B₁ = ordinary summary. (18.69)
B₂ = structured summary. (18.70)
B₃ = provenance-enhanced summary. (18.71)
B₄ = full MRER. (18.72)
If B₃ performs as well as B₄, then MRER may be overengineered.
The claim should be reduced accordingly.
18.25 Failure Condition 1 — No representation advantage
Suppose MRER cannot reliably distinguish:
refinement from replacement;
residual from limitation;
genealogy from semantic similarity.
Then the strong semantic architecture has failed.
The appropriate reduction may be:
MRER → ResearchProvenanceArchive. (18.73)
That would still be useful.
But it would be a weaker contribution.
18.26 Failure Condition 2 — Reconstruction instability
Suppose independent reconstructors produce radically different track structures from the same well-instrumented archive.
Then:
H_X fails. (18.74)
Possible interpretation:
the research history is intrinsically too ambiguous for stable semantic reconstruction.
Or:
current reconstruction protocols are inadequate.
The architecture should then retreat from claims of recoverable conceptual topology.
18.27 Failure Condition 3 — Human projections add no value
Suppose users perform equally well with:
ordinary papers + summary
as with:
MRER projections.
Then:
H_U fails for that task. (18.75)
The machine-native representation might still be useful for AI agents or audit.
But claims about human research improvement should be narrowed.
18.28 Failure Condition 4 — Reconstruction creates false certainty
A particularly serious failure occurs if Semantic Event Displays make users more confident in incorrect historical narratives.
For example:
beautiful track geometry
may create an illusion that:
one conceptual trajectory objectively existed.
This would reproduce the very presentation-closure problem the framework is intended to solve.
Therefore measure:
ConfidenceCalibration. (18.76)
A useful display should not merely increase comprehension.
It should improve calibration.
18.29 Failure Condition 5 — Research overhead exceeds benefit
If maintaining MRER requires:
more researcher effort than the recovered value justifies,
the system may fail operationally even if conceptually sound.
Therefore:
UtilityNet = ResearchBenefit − InstrumentationCost. (18.77)
A real standard must eventually demonstrate:
UtilityNet > 0. (18.78)
for at least some important research classes.
18.30 Failure Condition 6 — The representation becomes an ontology trap
A particularly subtle failure would occur if researchers begin forcing every theory into the MRER vocabulary.
For example:
every conceptual change becomes “mutation”;
every limitation becomes “residual”;
every interaction becomes “collision”.
The representation would then shape the research more than it records it.
The Gauge Grammar project already warns against this general pathology: a cross-domain structural vocabulary is justified only when it improves diagnosis, explanation, control, or design; otherwise the terminology should be removed.
The same rule applies here:
RepresentationGrammar must remain subordinate to research reality. (18.79)
18.31 A strong standard should permit escape hatches
If an event does not fit the ontology:
UnknownEventType. (18.80)
If a relation is unclear:
UnresolvedRelation. (18.81)
If track identity cannot be decided:
CompetingTrackHypotheses. (18.82)
If causality is unavailable:
CausalStatus = Unknown. (18.83)
If a residual cannot be classified:
ResidualType = Unclassified. (18.84)
These states protect the architecture from ontology coercion.
18.32 The strongest near-term research claim
The strongest responsible near-term claim is not:
“We have invented a new scientific publication system.”
It is:
A machine-native event representation may preserve theory-development information that ordinary final papers and summaries systematically discard, and this advantage can be measured.
That is a sufficient research programme.
It is concrete.
It can fail.
It can be benchmarked.
And it does not require a grand theory of intelligence.
18.33 The three-hypothesis programme
The programme can therefore be summarized:
H_R: Can research history be represented usefully? (18.85)
H_X: Can its important structures be reconstructed stably? (18.86)
H_U: Does that reconstruction improve research? (18.87)
The dependency is:
H_U requires useful H_X. (18.88)
H_X requires workable H_R. (18.89)
But:
H_R does not imply H_X. (18.90)
H_X does not imply H_U. (18.91)
This hierarchy is important.
A project can succeed at one level and fail at the next.
18.34 The first compelling result
A strong first result would therefore look something like this:
A fixed MRER v0.1 schema is applied to a benchmark containing synthetic and real AI-assisted theory histories.
Independent reconstructors recover major conceptual lineage, branch structure, residual continuity, and evidence state above strong structured-summary baselines.
Users given MRER-derived audit views identify hidden ancestry, false closure, and obsolete claims more accurately than users given only final papers and conventional summaries.
The improvement survives:
model-family changes;
terminology anonymization;
projection changes.
And causal relations are not overclaimed.
That would not prove that MRER is the final standard.
It would establish that reconstructable theory history is a measurable research object.
That alone would be significant.
19. Conclusion — Research as a Reconstructable Object
19.1 AI changes not only who writes, but what can be recorded
Discussion of AI-assisted research often focuses on authorship.
Did the human write the paper?
Did the model generate the text?
How much editing occurred?
Those questions matter.
But they may not be the deepest structural consequence of AI-assisted theory formation.
The more fundamental change may be this:
A much larger fraction of conceptual exploration can now become externally recorded.
Questions that might once have remained private can appear as explicit research events.
Possible analogies can be proposed.
Counterexamples can be requested.
Alternative theories can be generated.
Objections can be recorded.
Branches can be compared.
Residuals can be carried forward.
Revisions can be replayed.
Theory formation becomes more trace-rich.
The challenge is to avoid reducing that new trace back into the same old final-state object.
19.2 The paper remains valuable, but changes ontological status within the workflow
The central proposal of this article is not:
Paper → obsolete. (19.1)
It is:
Paper → projection. (19.2)
More fully:
ResearchEvents
→ EventCapture
→ SemanticReconstruction
→ MRER
→ NarrativePaper + EventDisplay + EvidenceView + ResidualView + AuditView. (19.3)
The paper remains one of the most important human projections.
But it is no longer required to carry the entire burden of research memory.
19.3 The canonical research object becomes machine-native
The proposed Machine-Native Research Event Representation does not need to resemble human prose.
It may contain structures that are awkward for human cognition:
typed events;
partial orders;
multiple competing track identities;
hyperedges;
probabilistic reconstruction assertions;
large provenance networks;
machine identifiers.
This is acceptable.
The requirement is not:
DirectHumanReadability. (19.4)
The requirement is:
MachineOperability + FaithfulProjectability + Auditability. (19.5)
This distinction may become increasingly important as research histories exceed what any single human can directly inspect.
19.4 Research history is not one graph
Another major conclusion is that theory history contains multiple relation planes.
Chronology answers:
What came before what?
Genealogy answers:
What inherited from what?
Semantic reconstruction asks:
How are two conceptual states structurally related?
Causal reconstruction asks:
What materially influenced a transition?
Epistemic reconstruction asks:
How strongly is any of this supported?
These relations must not be collapsed.
The governing discipline remains:
Later Than ≠ Derived From ≠ Semantically Related To ≠ Caused By.
This single distinction may eliminate a surprising amount of false intellectual narrative.
19.5 A concept can survive substantial change
Theory identity is also more complicated than textual persistence.
A concept may change:
vocabulary;
mechanism;
scope;
formalism;
while preserving:
research problem;
inferential role;
critical constraints;
genealogy.
Therefore:
ConceptIdentity ≠ SentenceIdentity. (19.6)
This motivates:
claim states;
identity kernels;
mutation;
replacement;
branching;
merge;
death;
resurrection.
A reconstructable research architecture should preserve theory evolution rather than pretending that theories are immutable objects.
19.6 What failed matters
Perhaps the most important departure from ordinary publication is the treatment of residual.
A conventional paper is naturally survivor-biased.
It presents what remained.
A reconstructable research object should also preserve:
what failed;
what contradicted;
what remained unresolved;
what repeatedly returned;
what generated the next theory.
This gives a different model of intellectual progress:
Theory → Residual → Revision. (19.7)
rather than merely:
Theory₁ → BetterTheory₂. (19.8)
The first preserves the pressure responsible for change.
The second preserves only the result.
19.7 Reconstruction itself must remain reconstructable
There is no privileged machine interpreter standing outside the epistemic problem.
Any reconstruction engine makes judgments.
It decides:
these two claims belong to one track;
this later formulation restricts the former;
this residual generated a successor;
this recurrence is independent.
Those judgments must themselves become objects.
Therefore:
ResearchTrace → Reconstruction₁ → Audit → Reconstruction₂. (19.9)
The reconstruction history should not be erased when the reconstruction changes.
This is the same principle of trace-preserving admissible revision developed in the earlier declaration framework: mature self-revision changes its current declaration without rewriting failure out of the past.
19.8 Human-readable truth should not be purchased by machine-side erasure
A clean event display is useful.
A clean paper is useful.
A clear timeline is useful.
But clarity is always a compression.
Therefore:
Projection(M) → HumanView + ProjectionResidual. (19.10)
The human interface may suppress enormous complexity.
It should not pretend that suppressed complexity never existed.
The principle is:
HumanReadability must not require EpistemicAmnesia. (19.11)
19.9 The Semantic Event Display finds its proper place
The idea that began this discussion can now be located precisely.
A Semantic Event Display is not:
the canonical research record;
nor:
a photograph of hidden reasoning.
It is:
a declared human-readable projection of a reconstructed machine-native research event object.
It may show:
conceptual tracks;
vertices;
branching;
merger;
track death;
residual channels;
evidence attachment;
reconstruction uncertainty.
Its value lies not in being beautiful.
Its value lies in making theory development inspectable.
19.10 The deepest analogy with a collider is reconstruction, not collision
At first glance, the particle-collider analogy appears to concern:
two conceptual beams collide and produce a new idea.
That remains useful.
But the deeper analogy developed in this paper is different.
The crucial transformation is:
recorded signals
→ reconstructed event
→ interpretable display.
Likewise:
recorded research interactions
→ machine-native semantic reconstruction
→ interpretable theory history.
The scientific value comes not from calling concepts “particles.”
It comes from recognizing that:
raw records and human interpretation need not be the same representational layer.
Between them can exist a reconstruction system.
That is the core architectural insight.
19.11 Reconstructable Research is therefore an instrumentation proposal
The proposal should not primarily be understood as a theory of knowledge.
It is an instrumentation programme.
Its question is:
Can theory formation be instrumented sufficiently that part of its developmental structure becomes externally reconstructable?
The minimum answer requires only:
Capture.
Reconstruct.
Project.
Audit.
The stronger programme adds:
stable track reconstruction;
residual analysis;
cross-projection invariance;
counterfactual replay;
generative causal experiments.
But all of these remain testable extensions.
19.12 The proposal has a right to fail
A scientific architecture should state what failure would mean.
If claim-state reconstruction is unstable:
reduce the framework.
If simple provenance performs equally well:
reduce the framework.
If structured summaries perform equally well:
reduce the framework.
If Semantic Event Displays increase false confidence:
remove them.
If residual tracking creates no measurable research benefit:
do not make it mandatory.
If research instrumentation costs more than the value it creates:
restrict the application domain.
The proposal should survive only to the extent that its distinctions do real work.
This is an important continuation of the protocol-first discipline used throughout the wider project: a conceptual grammar earns its place by improving explanation, diagnosis, testing, or control rather than by sounding universal.
19.13 The first realistic domain is AI-assisted theoretical research
The framework should not immediately claim universality.
AI-assisted theoretical research is a particularly appropriate initial domain because:
the interactions are already digital;
the event volume can be enormous;
model outputs can be versioned;
human interventions can be logged;
conceptual ancestry is a major concern;
and parts of the process may be replayable.
This makes it a natural laboratory.
If the architecture works there, it may later be adapted to:
computational science;
software research;
formal mathematics;
design research;
systematic literature synthesis;
collaborative scientific workflows.
But those extensions should be earned empirically.
19.14 A future publication stack
A mature implementation might eventually produce a scientific artifact stack such as:
Raw Event Archive. (19.12)
Research Provenance Layer. (19.13)
MRER. (19.14)
Claim Ledger. (19.15)
Residual Ledger. (19.16)
Evidence Layer. (19.17)
Semantic Event Display. (19.18)
Narrative Paper. (19.19)
Each serves a different role.
The paper communicates.
The event display orients.
The claim ledger states.
The evidence layer evaluates.
The residual ledger preserves what remains open.
MRER binds the structure together.
The raw archive preserves the ultimate external trace.
19.15 Research may become queryable in a new way
Once research history is represented this way, new questions become natural.
Not only:
What does this paper claim?
but:
Which current claim has changed most radically from its original formulation? (19.20)
Which residual has survived the longest? (19.21)
Which branch was abandoned without being tested? (19.22)
Which apparent independent discovery has hidden ancestry? (19.23)
Which conceptual constraint has survived every revision? (19.24)
Which source actually changed the theory rather than merely being cited? (19.25)
Which external validation applies only to a narrow descendant rather than the whole theory family? (19.26)
These are questions about research dynamics.
A document-only architecture answers them poorly.
19.16 From knowledge retrieval to research-state retrieval
Most current retrieval systems are organized around:
Find relevant information. (19.27)
A reconstructable research system can support a richer operation:
Find the current research state and explain how it became admissible. (19.28)
That includes:
what is active;
what is deprecated;
what remains open;
what is inherited;
what is uncertain;
what evidence changed the state.
This could substantially improve long-horizon AI research agents.
Instead of repeatedly restarting from summaries, they can enter an existing research process at a declared state.
19.17 The paper itself may become reproducibly regenerable
An interesting consequence follows.
If the narrative paper is explicitly defined as a projection of MRER under a declared publication protocol:
Paper_v = Π_paper(MRER_v | P_paper). (19.29)
then a new edition can be generated after research changes.
Suppose a residual is resolved.
An evidence grade changes.
A branch is falsified.
The MRER updates.
A new paper projection can show what changed.
This does not mean research writing should become fully automated.
It means the relation between research state and narrative artifact becomes explicit.
19.18 Research publication becomes declaration rather than erasure
A publication event can then be understood as:
At time k, under protocol P, this is the declared human-readable projection of the current research state.
Not:
This paper is the timeless and exhaustive theory.
This is a healthier model for rapidly evolving AI-assisted theoretical work.
It fits the broader declaration framework in which any readable world requires a declared boundary, feature map, protocol, gate, trace, and residual rather than pretending that one projection is complete reality.
19.19 The long-term possibility: an experimental history of ideas
The most speculative implication is worth stating carefully.
If AI-assisted theory formation becomes:
externally recorded;
machine-reconstructed;
replayable;
ablated;
cross-model comparable;
then some questions about conceptual development may become experimentally investigable.
For example:
Does objection O materially increase the probability of revision R?
Does source A cause a new structural invariant to appear?
Does the invariant survive when vocabulary is randomized?
Does branch B emerge independently across model families?
Does human intervention H suppress a false synthesis?
These are not experiments on hidden consciousness.
They are experiments on externally specified generative research systems.
This could create a new empirical layer between:
history of ideas
and:
science of reasoning processes.
That possibility remains speculative.
But it is now technically imaginable.
19.20 The final thesis
The entire proposal can therefore be compressed into one statement:
AI-assisted theoretical research should preserve enough of its external event history that the development of its claims, constraints, residuals, evidence, and conceptual lineages can later be reconstructed under declared protocols, represented in machine-native form, projected into multiple human-readable views, and audited back toward source events.
Or more compactly:
Research should become reconstructable. (19.30)
The paper is then:
not discarded,
not diminished,
but relocated.
Paper = DeclaredProjection(ReconstructableResearchObject). (19.31)
The final governing rule is:
Do not publish only the state. Preserve the transformations.
Appendix A — Minimal MRER Schema
The following appendix provides a first deliberately small schema.
It is not proposed as a final interchange standard.
Its purpose is to make the architecture concrete enough for implementation and testing.
A.1 Research Event Record
A minimal event record may contain:
EventID. (A.1)
EventType. (A.2)
Timestamp / Order. (A.3)
ActorID. (A.4)
ActorType. (A.5)
ModelVersion if applicable. (A.6)
InputArtifactIDs. (A.7)
OutputArtifactIDs. (A.8)
ProtocolVersion. (A.9)
ParentEventIDs. (A.10)
Notes / declared metadata. (A.11)
Example conceptual record:
EventID = E0042. (A.12)
EventType = ObjectionRaised. (A.13)
ActorType = Human. (A.14)
Targets = {C0017.v1}. (A.15)
Creates = {R0004}. (A.16)
A.2 Artifact Record
ArtifactID. (A.17)
ArtifactType. (A.18)
Version. (A.19)
ContentLocation. (A.20)
CreatedByEvent. (A.21)
DerivedFromArtifacts. (A.22)
Hash / persistent identifier where available. (A.23)
A.3 Claim-State Record
ClaimStateID. (A.24)
TrackHypothesisID. (A.25)
Statement / machine representation. (A.26)
ProblemClass. (A.27)
InferentialRole. (A.28)
CriticalConstraintIDs. (A.29)
EvidenceState. (A.30)
ResidualIDs. (A.31)
ProvenanceIDs. (A.32)
EpistemicStatus. (A.33)
A.4 Constraint Record
ConstraintID. (A.34)
ConstraintStatement. (A.35)
Scope. (A.36)
IntroducedByEvent. (A.37)
AppliesTo. (A.38)
ViolationEvidence. (A.39)
Status. (A.40)
A.5 Transformation Record
TransformationID. (A.41)
TransformationType. (A.42)
SourceClaimStates. (A.43)
TargetClaimStates. (A.44)
TriggerEvents. (A.45)
PreservedIdentityFeatures. (A.46)
ChangedFeatures. (A.47)
ConstraintEffects. (A.48)
ResidualEffects. (A.49)
ReconstructionAssertionID. (A.50)
A.6 Residual Record
ResidualID. (A.51)
ResidualType. (A.52)
Description. (A.53)
SourceClaim / Mapping / Reconstruction. (A.54)
FirstObservedEvent. (A.55)
CurrentStatus. (A.56)
ResolutionAttempts. (A.57)
SuccessorTracks. (A.58)
ClosureEvidence. (A.59)
A.7 Evidence Record
EvidenceID. (A.60)
EvidenceType. (A.61)
Direction = Support / Attack / Neutral. (A.62)
TargetClaimOrRelation. (A.63)
SourceArtifact. (A.64)
ExternalityLevel. (A.65)
ReplicationStatus. (A.66)
Strength / uncertainty representation. (A.67)
A.8 Reconstruction Assertion Record
ReconstructionAssertionID. (A.68)
Subject. (A.69)
RelationType. (A.70)
Object. (A.71)
ReconstructorID. (A.72)
ReconstructorType. (A.73)
ModelVersion if applicable. (A.74)
ReconstructionProtocol. (A.75)
SupportingEvidence. (A.76)
CounterEvidence. (A.77)
Confidence / epistemic status. (A.78)
AlternativeAssertions. (A.79)
Residual. (A.80)
CreatedAt. (A.81)
SupersedesAssertion. (A.82)
This record is essential because a semantic relation inferred by a reconstruction engine is not equivalent to an externally recorded fact.
A.9 Projection Record
ProjectionID. (A.83)
MRERVersion. (A.84)
ProjectionType. (A.85)
ObserverRole. (A.86)
Purpose. (A.87)
Scale. (A.88)
PreservedInvariants. (A.89)
SuppressedDimensions. (A.90)
ProjectionResidual. (A.91)
DecoderID / Version. (A.92)
GeneratedAt. (A.93)
A.10 Minimal reverse-audit contract
For every projected assertion Hᵢ above a declared importance threshold:
Hᵢ
→ RAⱼ
→ {C,K,R,V,T}
→ Eₖ
→ Aₗ. (A.94)
If this path cannot be resolved:
AuditStatus(Hᵢ) = Untraceable. (A.95)
An untraceable assertion may still be displayed.
But it should not silently receive the same status as a fully traceable reconstruction.
Appendix B — Transformation Vocabulary
A first standard vocabulary may include:
Refine. (B.1)
Restrict. (B.2)
Generalize. (B.3)
Mutate. (B.4)
Split. (B.5)
Merge. (B.6)
Replace. (B.7)
Reject. (B.8)
Falsify. (B.9)
Deprecate. (B.10)
Absorb. (B.11)
Suspend. (B.12)
Reactivate. (B.13)
Transfer. (B.14)
Operationalize. (B.15)
Abstract. (B.16)
Reinterpret. (B.17)
Each transformation should specify:
Source state. (B.18)
Target state. (B.19)
Preserved identity features. (B.20)
Changed identity features. (B.21)
Trigger / supporting evidence. (B.22)
Residual. (B.23)
Appendix C — Reconstruction Assertion Template
A reconstruction assertion should be expressible in a compact form such as:
RA = ⟨S, r, O, P, E⁺, E⁻, q, Ω⟩. (C.1)
where:
S = subject. (C.2)
r = asserted relation. (C.3)
O = object. (C.4)
P = reconstruction protocol. (C.5)
E⁺ = supporting evidence. (C.6)
E⁻ = counterevidence. (C.7)
q = confidence / epistemic qualification. (C.8)
Ω = unresolved residual. (C.9)
Example:
Subject = FiltrationFramework. (C.10)
Relation = MajorRevisionOf. (C.11)
Object = RecursiveGenerationFramework. (C.12)
Support = explicit Part 2 correction of hidden meta-time. (C.13)
Counterevidence = substantial mechanism replacement. (C.14)
Status = strongly supported genealogy; track identity interpretive. (C.15)
Residual = mutation-vs-replacement classification remains protocol-sensitive. (C.16)
Appendix D — Projection Contract Template
A projection should declare:
D.1 Viewpoint
Who is this projection for?
D.2 Goal
What question is it intended to answer?
D.3 Scale
Prompt-level?
Claim-level?
Article-level?
Programme-level?
D.4 Preserved invariants
Examples:
lineage;
evidence status;
major residual;
causal uncertainty.
D.5 Suppressed dimensions
Examples:
minor edits;
weak branches;
low-priority artifacts.
D.6 Projection residual
What relevant machine structure has not been faithfully represented?
D.7 Reverse-navigation requirement
Which displayed objects must remain drillable to provenance?
A projection is admissible only relative to its declaration.
Appendix E — Reconstruction Depth Standard
RD0 — Artifact Only
Available:
final artifacts.
Strong:
semantic analysis.
Weak:
historical causality.
RD1 — Versioned Artifacts
Available:
ordered artifact revisions.
Strong:
change detection.
Weak:
cause of change.
RD2 — Interaction Trace
Available:
prompt / response / source interaction history.
Strong:
many genealogical relations.
Moderate:
contemporary revision sequence.
RD3 — Instrumented External Trace
Available:
interaction history + model metadata + tools + explicit decisions + experiments + artifact lineage.
Strong:
chronology;
genealogy;
many revision relations.
Potentially strong:
recorded causal attribution.
RD4 — Replayable Environment
Available:
RD3 plus sufficiently preserved generative conditions for matched reruns.
Additional capability:
ablation;
sham conditions;
counterfactual replay;
generative causal testing.
The rule remains:
HigherRD permits stronger questions.
It does not guarantee stronger answers.
Appendix F — Worked Example: From Recursive Generation to Admissible Self-Revision
F.1 Why this sequence is useful
The sequence:
From One Assumption to One Operator
→ From One Operator to One Filtration
→ From One Filtration to One Declaration
→ From One Declaration to One Self-Revising Fractal
provides a useful real example because the later articles explicitly revise the earlier conceptual architecture rather than merely changing terminology.
The first paper explores:
primitive operation → recursion → pre-time → collapse → ledger → time.
The second explicitly corrects this proposal by arguing that recursion may be a presentation grammar rather than a literal process generating the universe before time. It replaces generation with viewpoint-selected filtration and ledgered disclosure.
The third asks what makes the field filterable and answers through declaration: baseline, feature map, protocol, projection, gate, trace, and residual must be fixed before disclosure becomes meaningful.
The fourth then asks how declaration can revise itself without becoming arbitrary or dishonest, introducing admissible self-revision constrained by trace preservation, residual honesty, frame robustness, boundedness, and non-degeneracy.
This gives a compact reconstructable trajectory.
F.2 State 1 — Recursive Generation
Claim state:
C₁ = Recursive depth may provide pre-time ordering. (F.1)
Core problem:
Q₁ = How can a pre-collapse field exhibit ordering before ordinary time exists? (F.2)
Provisional mechanism:
Ω₁ = recursive generation. (F.3)
Residual:
R₁ = possible hidden meta-time. (F.4)
F.3 Transformation 1 — Mutation under the meta-time objection
The second article explicitly raises the problem that if recursion literally unfolds “before” time, some deeper ordering may already have been presupposed.
Transformation:
T₁ = Mutate(C₁ → C₂). (F.5)
C₂ = Recursion may present structure rather than ontologically generate it. (F.6)
Preserved:
research problem;
pre-time concern;
need for non-ordinary ordering.
Changed:
ontological mechanism.
Residual generated:
R₂ = If the field is disclosed through filtration, what makes it filterable? (F.7)
This is a good example of:
high conceptual change + preserved research lineage.
F.4 State 2 — Viewpoint-Selected Filtration
Core chain:
Pre-Time Field
→ Viewpoint
→ Filtration
→ Collapse
→ Ledger
→ Time. (F.8)
The critical improvement is the removal of literal pre-time generation.
But the filtration operation itself now becomes the new pressure point.
This illustrates:
SolvedResidual(R₁) → NewResidual(R₂). (F.9)
Progress does not mean residual disappears globally.
It means the residual structure changes.
F.5 Transformation 2 — Prerequisite discovery
The third article asks:
What makes Σ filterable?
The answer:
Declaration.
This transition differs from T₁.
T₁ changed the mechanism.
T₂ discovers a missing prerequisite.
Thus a possible transformation type is:
T₂ = PrerequisiteExpansion(C₂ → C₃). (F.10)
C₃ introduces:
baseline q;
feature map φ;
boundary B;
observation rule Δ;
horizon h;
admissible intervention u;
projection operator;
gate;
trace rule;
residual rule.
The resulting chain becomes:
Declare
→ Project
→ Gate
→ Trace + Residual
→ Ledger. (F.11)
F.6 Residual 3 — Revision pathology
Once declaration becomes explicit, another problem appears.
If trace and residual can revise future declaration:
Dₖ₊₁ = Revise(Dₖ | Lₖ,Rₖ), (F.12)
what prevents the system from:
rewriting old mistakes;
changing rules whenever challenged;
hiding residual;
redefining contradiction as success?
This is:
R₃ = unrestricted self-revision pathology. (F.13)
F.7 Transformation 3 — Governance completion
The fourth article introduces admissibility conditions.
The revised object is not merely:
self-revision.
It is:
admissible self-revision.
The article explicitly requires properties including trace preservation, residual honesty, frame robustness, boundedness, and non-degeneracy.
Thus:
T₃ = GovernanceRestriction(C₃ → C₄). (F.14)
The theory becomes narrower in one sense:
not every self-modification qualifies.
But more mature in another:
the conditions under which revision remains trustworthy are now declared.
F.8 The reconstructed track
A human Theory-Evolution View may display:
Recursive Generation
│
├── R₁: hidden meta-time
↓
Viewpoint Filtration
│
├── R₂: what makes filtration possible?
↓
Declared Disclosure
│
├── R₃: arbitrary self-revision pathology
↓
Admissible Self-Revision. (F.15)
This is a very readable projection.
But MRER should preserve more.
It should also know:
which events established each transition;
which source paragraphs support it;
whether each transition is refinement, mutation, or replacement;
which claims remain speculative;
which alternative reconstructions exist.
The display is only one projection.
F.9 Competing track reconstruction
A second reconstructor might reject the single-track interpretation.
It could propose:
Track A = Recursive Generation. (F.16)
Track A terminates after hidden meta-time objection. (F.17)
Track B = Filtration / Declared Disclosure framework. (F.18)
Under this reconstruction:
C₂ is not a mutation of C₁.
It is a replacement.
Both interpretations have some support.
Therefore MRER should be able to preserve:
H₁ = Single high-curvature track. (F.19)
H₂ = Track replacement after falsification of mechanism. (F.20)
The archive need not prematurely choose.
F.10 What the example demonstrates
The example demonstrates six principles of Reconstructable Research:
First
The final theory cannot fully explain why its current constraints exist.
Second
Residuals connect generations of theory.
Third
Large semantic change does not automatically imply genealogical discontinuity.
Fourth
Track identity is reconstructive rather than purely textual.
Fifth
Different admissible reconstructions may coexist.
Sixth
A simple human diagram can be generated without forcing the machine representation itself to be simple.
The worked example therefore captures the central argument of the article:
A research programme is not only the set of theories it eventually accepts. It is also the structured history of transformations, rejections, residuals, and constraints through which those theories became admissible.
That history can be preserved.
It can be reconstructed.
And, under sufficiently disciplined protocols, it can become a scientific object in its own right.
Appendix G — Interoperability with Existing Provenance and Research-Object Standards
G.1 MRER should not reinvent provenance infrastructure
Reconstructable Research does not begin in an empty standards landscape.
Several existing systems already solve important parts of the problem:
• representation of provenance;
• identification of entities, activities, and agents;
• packaging of research artifacts;
• machine-readable metadata;
• atomic publication of claims with provenance.
The correct engineering strategy is therefore not:
MRER replaces existing provenance standards. (G.1)
It should instead be:
ExistingStandards + MRERSemanticExtension → ReconstructableResearchInfrastructure. (G.2)
The distinctive contribution proposed here lies primarily in the representation of theory transformation, not in reinventing generic provenance transport.
G.2 W3C PROV provides a natural lower provenance layer
W3C PROV already supplies a mature general-purpose provenance vocabulary. PROV-O represents provenance using concepts including prov:Entity, prov:Activity, and prov:Agent, together with relations describing generation, derivation, usage, attribution, and other provenance structures. It was designed to support interoperable provenance exchange across heterogeneous systems and can be specialized for domain-specific applications. (W3C)
This maps naturally onto the external-event layer of Reconstructable Research.
For example:
Research artifact → prov:Entity. (G.3)
Research interaction → prov:Activity. (G.4)
Human / model / tool actor → prov:Agent or specialized agent representation. (G.5)
Artifact B derived from Artifact A → provenance derivation relation. (G.6)
The exact mapping would require a proper MRER profile rather than the simplistic correspondences above.
But the architectural implication is clear:
MRER does not need to invent the question “who used what to produce what?”
Existing provenance infrastructure already provides a strong foundation for that question.
MRER needs to add questions such as:
What claim state changed? (G.7)
What constraint was violated? (G.8)
Was the successor a refinement or replacement? (G.9)
Which residual remained open? (G.10)
Was the genealogical relation observed or reconstructed? (G.11)
Those are the theory-dynamics extensions.
G.3 Provenance and semantic reconstruction should therefore form two linked layers
A useful architecture is:
ProvenanceLayer P
↕
SemanticReconstructionLayer S. (G.12)
The provenance layer stores externally grounded relationships.
The semantic layer stores reconstructed research meaning.
For every important semantic object s:
Support(s) ⊆ P. (G.13)
This means that a semantic reconstruction assertion should point toward the provenance objects that support it.
For example:
RA₁ = “Claim C₂ restricts Claim C₁.” (G.14)
may be supported by:
Artifact version A₃;
revision event E₁₇;
human acceptance event E₁₈;
source span S₄.
The semantic assertion is richer than the provenance record.
But it does not float independently from it.
G.4 This separation reduces ontology pressure
If MRER tried to encode raw provenance and theory semantics in one flat vocabulary, it would soon become unnecessarily large.
Instead:
P answers:
Who?
What artifact?
Which process?
When?
Derived from what?
S answers:
What changed conceptually?
Why is this considered the same track?
What residual remained?
What evidence attacks the claim?
What reconstruction uncertainty exists?
Thus:
P = EventProvenance. (G.15)
S = TheoryReconstruction. (G.16)
MRER can bind the two without requiring them to be identical.
G.5 RO-Crate provides a natural packaging layer
RO-Crate is designed to package and describe research artifacts and their associated metadata using linked-data principles. Its current 1.3 specification is a Recommendation published on June 22, 2026, and uses a JSON-LD metadata document as the core machine-readable description of a research object. It can describe files, URI-addressable resources, contextual entities, and provenance associated with their creation and use. (researchobject.org)
This suggests another architectural distinction:
MRER does not necessarily need to define how a whole research project is physically packaged.
A future implementation could use:
RO-Crate = packaging / interchange envelope. (G.17)
MRER Profile = theory-reconstruction semantics. (G.18)
The resulting structure might be:
RO-Crate
├── source artifacts
├── interaction records
├── model metadata
├── experiment outputs
├── MRER representation
├── projection records
└── human-facing paper / event display. (G.19)
This would make MRER a profile or extension carried inside an existing research-object ecosystem, rather than a competing file-container standard.
RO-Crate explicitly supports domain-specific metadata extension, making such a profile-based direction technically plausible. (researchobject.org)
G.6 Research packaging and research reconstruction are different jobs
RO-Crate can answer:
What artifacts belong to this research object? (G.20)
How were they created? (G.21)
Who or what participated? (G.22)
How can they be distributed and reused? (G.23)
MRER asks additional questions:
Which claims are successive states of one conceptual lineage? (G.24)
Which theory mutation was triggered by a constraint violation? (G.25)
Which residual generated a successor branch? (G.26)
Which reconstruction relation is disputed? (G.27)
Thus:
ResearchPackaging ≠ TheoryReconstruction. (G.28)
The two are complementary.
G.7 Nanopublications suggest a useful model for portable claim units
Nanopublications provide another relevant precedent.
A nanopublication packages a small assertion together with provenance for that assertion and publication information for the nanopublication as a whole. Current nanopublication guidance explicitly treats these as machine-interpretable, self-contained scientific knowledge units and allows claims, hypotheses, negative results, and other statements to be represented in this way. (Nanopub)
This suggests a useful MRER design pattern.
A Claim State could travel as a portable package containing:
Claim. (G.29)
ClaimProvenance. (G.30)
EvidenceState. (G.31)
ResidualState. (G.32)
TrackRelations. (G.33)
ReconstructionStatus. (G.34)
Conceptually:
ResearchClaimPackage = Claim + Provenance + Evidence + Residual + ReconstructionContext. (G.35)
This is not identical to a nanopublication.
But the atomic-publication pattern is valuable.
G.8 The residual should travel with the claim
This may be an especially important extension.
A conventional knowledge artifact often circulates:
Claim. (G.36)
But Reconstructable Research should prefer:
Claim + Residual. (G.37)
For example:
Claim:
“Viewpoint filtration removes the need for literal recursive generation.” (G.38)
Residual:
“What makes the field filterable?” (G.39)
If the claim circulates without the residual, later systems may reconstruct it as a completed solution.
If both travel together, the state is more faithful.
Therefore a future portable unit could be:
EpistemicPackage(C) = {Claim, Scope, Evidence, Residual, Provenance}. (G.40)
This is more useful for evolving theory than a context-free proposition.
G.9 Reconstruction assertions themselves could be portable
The same principle applies to intellectual genealogy.
Instead of embedding:
“B revises A”
inside narrative prose, publish:
RA₄₂:
Subject = B. (G.41)
Relation = Revises. (G.42)
Object = A. (G.43)
Support = {E₁,E₂,A₃}. (G.44)
Counterevidence = {A₄}. (G.45)
Status = StronglySupported. (G.46)
Residual = MutationVsReplacementAmbiguous. (G.47)
This could be stored as a structured assertion with its own provenance.
Therefore the semantic event structure itself becomes independently inspectable.
G.10 A possible standards stack
The emerging implementation picture is therefore not one giant new standard.
A more modular stack is:
Layer 0 — Raw Artifacts
Files, prompts, responses, datasets, code, experimental outputs.
Layer 1 — Generic Provenance
W3C PROV-compatible provenance structures.
Layer 2 — Research Packaging
RO-Crate or comparable research-object packaging.
Layer 3 — MRER Semantic Profile
Claim states, tracks, constraints, transformations, residuals, reconstruction assertions.
Layer 4 — Atomic Epistemic Units
Nanopublication-like portable claims / reconstruction assertions where useful.
Layer 5 — Projection Layer
Paper, Event Display, Evidence View, Residual View, Reviewer View.
Thus:
RawArtifacts
→ Provenance
→ ResearchObject
→ SemanticReconstruction
→ DeclaredProjection. (G.48)
This is likely a more realistic standardization path than attempting to replace every layer simultaneously.
G.11 Where the actual novelty claim should sit
This standards comparison also disciplines the novelty claim of Reconstructable Research.
It would be incorrect to claim novelty merely for:
machine-readable provenance;
research-object packaging;
claim-level provenance;
or knowledge graphs.
Mature work already exists in all of these directions. (W3C)
The proposed contribution is narrower and more specific:
A theory-evolution reconstruction layer in which claim states, conceptual identity, typed transformations, failed branches, residuals, competing reconstructions, and epistemic status are represented as provenance-linked machine objects and can subsequently be projected into auditable human views.
In shorthand:
ExistingProvenance + TheoryDynamics + ReconstructionUncertainty + ProjectionAudit. (G.49)
That is the appropriate claim boundary.
Appendix H — Minimum Implementation Architecture
H.1 The first implementation should be deliberately boring
The conceptual framework is broad.
The first software implementation should not be.
A prototype does not need:
a new database engine;
a new vector database;
a universal ontology;
a new graph language;
or a sophisticated visualizer.
It needs to answer one question:
Can a real AI-assisted research history be reconstructed more usefully than by an ordinary structured summary?
The first implementation should therefore optimize for:
auditability;
inspectability;
schema stability;
easy revision.
Not elegance.
H.2 A six-component prototype
A minimum architecture can contain six components:
Capture Store
→ Semantic Compiler
→ MRER Store
→ Reconstruction Engine
→ Projection Engine
→ Audit Interface. (H.1)
H.3 Component 1 — Capture Store
The Capture Store preserves original external events.
Examples:
conversation messages;
uploaded source references;
generated documents;
tool results;
human decisions;
experiment outputs.
Each event receives a stable identifier:
E000001. (H.2)
E000002. (H.3)
…
Artifacts likewise receive:
A000001. (H.4)
A000002. (H.5)
The key requirement is immutability of historical identity.
Later interpretation may change.
The original event identity should not.
H.4 Component 2 — Semantic Compiler
The Semantic Compiler converts raw event material into candidate research objects.
The pipeline may be:
Event
→ Segment
→ Candidate Semantic Units
→ Type
→ Link
→ Residual Audit. (H.6)
Candidate outputs include:
ClaimState. (H.7)
Constraint. (H.8)
Evidence. (H.9)
Residual. (H.10)
TransformationCandidate. (H.11)
The compiler should not immediately promote every candidate to canonical status.
Instead:
SemanticCandidate → ReconstructionEvaluation. (H.12)
This follows the same principle developed in the semantic compiler work:
raw natural-language material should first be converted into a structured intermediate representation before compression or user-facing translation.
H.5 Component 3 — MRER Store
The MRER Store contains the current machine-native reconstruction.
For a first implementation, a conventional database is enough.
A hybrid system might use:
relational tables for stable identifiers and metadata;
graph relations for lineage;
vector representations for candidate semantic similarity;
object storage for source artifacts.
No single storage layer must be treated as ontologically fundamental.
The MRER logical model remains:
MRER = {Objects, Relations, ReconstructionAssertions, Provenance}. (H.13)
H.6 Component 4 — Reconstruction Engine
The Reconstruction Engine proposes relations such as:
same-track;
refines;
restricts;
replaces;
branches-from;
resolves-residual;
supports;
attacks.
It should output not simply:
relation = true.
Instead:
RA = {Relation, Support, Counterevidence, Status, Alternatives}. (H.14)
For difficult cases it should be allowed to output:
Unresolved. (H.15)
or:
{H₁,H₂,H₃}. (H.16)
where H₁–H₃ are competing reconstruction hypotheses.
This is preferable to forced coherence.
H.7 Component 5 — Projection Engine
The Projection Engine accepts:
MRER version;
viewpoint;
purpose;
scale;
preserved invariants;
compression budget.
It then generates:
HumanView + ProjectionResidual. (H.17)
A prototype needs only three views:
View A — Theory Evolution
Major tracks and revisions.
View B — Residual Ledger
Current and historical unresolved problems.
View C — Provenance Audit
Source path behind a selected claim or transformation.
If these three views are not useful, there is little reason to build a more elaborate Event Display.
H.8 Component 6 — Audit Interface
The Audit Interface should implement the backward path:
DisplayedStatement
→ ReconstructionAssertion
→ ClaimState / Transformation
→ Event
→ Artifact. (H.18)
This may be the most important part of the first prototype.
Without it, the system is merely an advanced summarizer.
H.9 Minimal ingestion workflow
For an existing AI-assisted research project:
Step 1:
Import external records. (H.19)
Step 2:
Assign immutable event and artifact IDs. (H.20)
Step 3:
Extract candidate claim states. (H.21)
Step 4:
Extract constraints and residuals. (H.22)
Step 5:
Propose lineage relations. (H.23)
Step 6:
Classify major transformations. (H.24)
Step 7:
Create Reconstruction Assertions. (H.25)
Step 8:
Run independent reconstruction audit. (H.26)
Step 9:
Create MRER v0.1. (H.27)
Step 10:
Generate projections. (H.28)
H.10 Incremental operation is more important than retrospective conversion
The architecture becomes much more powerful when integrated into the research process itself.
Instead of reconstructing everything only after the project finishes:
Research Event
→ Capture
→ Candidate Compilation
→ Periodic Reconstruction. (H.29)
The system may periodically ask:
Did a claim change? (H.30)
Was an existing track abandoned? (H.31)
Did a new residual appear? (H.32)
Was a previously open residual closed? (H.33)
Did the current paper silently revive a deprecated assumption? (H.34)
This transforms MRER from an archival technology into a live research-memory system.
H.11 Human confirmation should be strategic, not universal
Requiring human confirmation for every extracted relation would make the system unusable.
Instead, confirmation should be triggered when reconstruction risk is high.
Candidate triggers include:
major track replacement;
claimed independent rediscovery;
causal relation;
external validation status change;
residual closure;
high-impact merge;
large scope expansion.
Thus:
HumanReviewPriority = f(Impact, Uncertainty, Irreversibility). (H.35)
Again, the exact scoring rule can be empirical.
The principle is:
use scarce human attention where reconstruction errors matter most.
H.12 Versioning MRER
A research project should have:
MRER.v1. (H.36)
MRER.v2. (H.37)
MRER.v3. (H.38)
A new reconstruction should not erase the old one.
Instead:
MRER.v2 supersedes MRER.v1. (H.39)
with a reconstruction change record:
RA₃₁ changed from Mutation to Replacement. (H.40)
R₇ changed from Closed to PartiallyResolved. (H.41)
Claim C₁₂ evidence status changed from Candidate to ExternallyTested. (H.42)
The reconstruction therefore possesses its own ledger.
H.13 A research-state checkpoint
At any moment the system should be able to produce a compact state checkpoint:
ActiveTracks. (H.43)
DormantTracks. (H.44)
RejectedTracks. (H.45)
CriticalConstraints. (H.46)
OpenResiduals. (H.47)
EvidenceFrontier. (H.48)
ReconstructionDisputes. (H.49)
This is probably the most useful input for a new AI research session.
Instead of giving the model a huge historical summary, provide:
ResearchStateCheckpoint + DrillDownAccess. (H.50)
This may greatly reduce repeated context compression.
H.14 The research-memory use case may precede the publication use case
An important practical possibility follows.
The earliest major value of MRER may not be better scientific publication.
It may be:
better long-horizon AI research memory.
A persistent research agent needs to know not merely:
what the current theory says,
but:
which ideas were tried;
which were rejected;
which residuals remain;
which constraints must not be forgotten.
Therefore the engineering route may be:
ResearchMemory
→ InternalAudit
→ EventDisplay
→ PublicationStandard. (H.51)
rather than immediately:
PublicationStandard → adoption.
This may be a much more realistic development path.
Appendix I — A Minimal Semantic Event Display Grammar
I.1 The visual layer should remain subordinate to MRER
A Semantic Event Display should not invent structure that MRER does not contain.
Therefore:
DisplayGrammar ⊂ MRERSemantics. (I.1)
The display chooses which distinctions to show.
It does not become a second independent theory of the research history.
I.2 Minimal visible objects
A first display needs only six graphical object classes.
Track
A reconstructed conceptual lineage.
State Node
A specific claim state.
Vertex
A transformation or interaction event.
Residual
An unresolved remainder.
Evidence Attachment
Support or attack attached to a claim or relation.
Provenance Anchor
A drill-down link to external trace.
Everything else can be introduced later.
I.3 Track line
A track line means:
these states are currently reconstructed as belonging to one conceptual lineage.
It must not mean:
the theory is true.
Thus:
TrackVisibility ≠ EpistemicValidation. (I.2)
A speculative track can be long.
A falsified track can be historically important.
I.4 Transformation vertex
A vertex should state its type explicitly.
For example:
[REFINE]. (I.3)
[RESTRICT]. (I.4)
[MUTATE]. (I.5)
[SPLIT]. (I.6)
[MERGE]. (I.7)
[REPLACE]. (I.8)
[REJECT]. (I.9)
This reduces the temptation to infer transformation class merely from graphical shape.
I.5 Branching
A branch should be rendered when a parent state produces multiple conceptually distinct descendants.
For example:
C₂a
╱
C₁ ───────●
╲
C₂b. (I.10)
No merge should be visually implied until an actual merge reconstruction exists.
I.6 Merge
A merge should preserve multiple parentage:
C_A ─────╲
●── C_C
C_B ─────╱. (I.11)
The merge node may expose:
InheritedFromA. (I.12)
InheritedFromB. (I.13)
NewAtMerge. (I.14)
ResidualAtMerge. (I.15)
This makes synthesis auditable.
I.7 Residual channel
Residual should have a visually different path.
For example:
C₁ ─────────► C₂
│
└── R₁ ─────► Q₃ ─────► C₃. (I.16)
This communicates:
the main track continued,
while an unresolved remainder escaped the current closure and later generated another development.
This is one of the most important visual distinctions in the whole proposal.
I.8 Dead track
A dead or terminated track should not simply disappear.
For example:
C₁ ─────► C₂ ─────► ×. (I.17)
The termination should expose its reason:
× Falsified. (I.18)
× Abandoned. (I.19)
× Superseded. (I.20)
× Absorbed. (I.21)
× ScopeExhausted. (I.22)
The reason is part of the research result.
I.9 Dormancy and resurrection
Dormancy may be represented as interruption:
C₁ ─────► C₂ ········· C₃. (I.23)
where the dotted gap denotes inactivity rather than continuous development.
If resurrected:
C₂ ··· [NewEvidence] ──► C₃. (I.24)
The display must not visually rewrite dormancy as uninterrupted survival.
I.10 Observed versus reconstructed relation
One useful visual distinction is:
solid relation = explicitly grounded relation. (I.25)
dashed relation = inferred relation. (I.26)
alternative dashed forks = competing reconstruction. (I.27)
The exact visual styling can vary.
What matters is that:
Observed ≠ Inferred. (I.28)
I.11 Reconstruction disagreement
Suppose two interpretations exist:
H₁: C₂ is mutation of C₁. (I.29)
H₂: C₁ terminates and C₂ replaces it. (I.30)
An audit display might show:
── [MUTATE] ── C₂
C₁ ──────◊
── [REPLACE] ─ C₂. (I.31)
The diamond means:
reconstruction ambiguity,
not a theory branch in the original research.
This distinction is important.
There are two different kinds of branching:
ResearchBranch. (I.32)
ReconstructionBranch. (I.33)
They must not be confused.
I.12 Evidence should be orthogonal to track geometry
Evidence can be attached vertically or through a distinct visual layer:
V₊
│
C₁ ─────► C₂ ─────► C₃
│
V₋. (I.34)
where:
V₊ = supporting evidence. (I.35)
V₋ = attacking evidence. (I.36)
This prevents line length or thickness from silently becoming an evidence metaphor unless such encoding is explicitly declared.
I.13 Causal edges require stronger visual qualification
A semantic transition:
C₁ ─refines→ C₂. (I.37)
should not look identical to:
O₁ ═causally-influenced→ C₂. (I.38)
if the causal relation has stronger evidential requirements.
A display should therefore expose causal class:
C1 hypothesis. (I.39)
C2 retrospective attribution. (I.40)
C3 contemporary recorded dependency. (I.41)
C4 interventional replay evidence. (I.42)
This makes causal overstatement visually harder.
I.14 Scale indicator
Every event display should state its scale.
For example:
Scale = Claim. (I.43)
Scale = Article. (I.44)
Scale = ResearchProgramme. (I.45)
A programme-level track should not be mistaken for literal claim-by-claim continuity.
I.15 Reconstruction-depth indicator
Likewise:
RD0. (I.46)
RD1. (I.47)
RD2. (I.48)
RD3. (I.49)
RD4. (I.50)
should be displayed somewhere prominent.
A beautiful RD0 retrospective reconstruction must not look epistemically equivalent to an RD3 instrumented trace.
I.16 The information card behind a track node
Clicking a claim-state node should reveal at minimum:
ClaimStateID. (I.51)
Human translation. (I.52)
Machine representation pointer. (I.53)
Track hypothesis. (I.54)
Prior state. (I.55)
Transformation. (I.56)
Critical constraints. (I.57)
Evidence state. (I.58)
Open residuals. (I.59)
Reconstruction confidence. (I.60)
Provenance. (I.61)
This is where the display becomes an actual scientific interface.
I.17 The Semantic Event Display should answer five questions quickly
A good display should allow a trained user to answer:
Question 1
Where did this idea come from?
Question 2
How has it changed?
Question 3
What caused or motivated the major changes, to the extent known?
Question 4
What failed or remains unresolved?
Question 5
What evidence currently supports the surviving state?
If the visualization cannot improve answers to these questions, it is probably decorative rather than scientific.
I.18 An example display of the declaration sequence
At article-level scale:
Recursive Generation
│
│ [META-TIME OBJECTION]
▼
[MUTATION]
│
▼
Viewpoint Filtration
│
├──── R₂: What makes Σ filterable?
│
▼
[PREREQUISITE DISCOVERY]
│
▼
Declared Disclosure
│
├──── R₃: Arbitrary self-revision pathology
│
▼
[GOVERNANCE RESTRICTION]
│
▼
Admissible Self-Revision
This human-readable diagram is useful.
But it is only a projection.
Behind the line:
Viewpoint Filtration → Declared Disclosure (I.62)
the MRER may preserve:
multiple events;
claim-state transitions;
source spans;
alternative classifications;
residual objects;
reconstruction confidence.
That hidden complexity is not a defect.
It is precisely why the display can remain simple without destroying the underlying research history.
I.19 The key visual principle
The Semantic Event Display should therefore obey:
Simplify the view, not the history.
Formally:
Compression(Display) ↑ is acceptable, (I.63)
provided:
Recoverability(ImportantStructure) remains above the declared audit threshold. (I.64)
This may be the most concise engineering rule for the entire human-interface layer.
I.20 From image to instrument
The final distinction is between:
Infographic. (I.65)
and:
Instrument. (I.66)
An infographic communicates a prepared interpretation.
A scientific Semantic Event Display should additionally allow:
query;
drill-down;
filtering;
alternative reconstruction;
evidence inspection;
provenance recovery.
Thus:
StaticEventDisplay = useful communication artifact. (I.67)
InteractiveEventDisplay = potential research instrument. (I.68)
The long-term target should be the latter.
At this stage, the article's conceptual argument, formal architecture, standards relationship, minimal schema, worked example, implementation blueprint, and first visual grammar are all in place. The remaining material I would add is a final implementation-oriented Appendix J — A Concrete MRER v0.1 Benchmark and Experimental Protocol, followed by a compact References / Source Note so that the full manuscript ends with a genuinely testable research programme rather than only an architecture proposal.
Appendix J — A Concrete MRER v0.1 Benchmark and Experimental Protocol
J.1 Why a benchmark is necessary
The central proposal of this paper is deliberately stronger than ordinary provenance logging.
It claims that an AI-assisted research history can, under suitable conditions, be reconstructed into a machine-native representation containing:
• claim-state evolution;
• conceptual lineage;
• transformations;
• residuals;
• branch structure;
• evidence state;
• competing reconstruction hypotheses;
• and provenance linking these structures back to external research events.
If this is more than a convenient vocabulary, it must support measurable reconstruction tasks.
The first benchmark should therefore not ask:
Can an LLM produce a plausible intellectual history?
That would be too weak.
The benchmark should ask:
Can a reconstruction system recover known or independently adjudicated research-event structure from external research traces without inventing unsupported continuity, causation, closure, or novelty?
The central benchmark object is therefore not prose.
It is relational recovery.
J.2 Benchmark objective
Let:
E = externally observable research-event archive. (J.1)
G* = reference research-event structure. (J.2)
ℛ = reconstruction system. (J.3)
Then:
ℛ(E) → Ĝ. (J.4)
The benchmark evaluates:
Ĝ ≈ G*. (J.5)
But similarity must be evaluated across declared dimensions rather than through textual overlap.
A minimal reconstruction benchmark should measure:
LineageRecovery. (J.6)
TransformationRecovery. (J.7)
BranchRecovery. (J.8)
ResidualRecovery. (J.9)
EvidenceAttachment. (J.10)
EpistemicCalibration. (J.11)
CausalRestraint. (J.12)
ProvenanceTraceability. (J.13)
The last two are especially important.
A system that reconstructs many plausible relations but systematically overclaims causation is not a strong reconstructor.
Neither is a system that produces attractive conceptual tracks whose edges cannot be traced to external evidence.
J.3 The benchmark should contain both synthetic and real histories
Neither synthetic nor real histories are sufficient alone.
Synthetic histories provide known ground truth.
Real histories provide ecological complexity.
The benchmark should therefore contain two major divisions:
Synthetic Reconstruction Corpus. (J.14)
Real Instrumented Research Corpus. (J.15)
Their roles differ.
J.4 Synthetic Reconstruction Corpus
In the synthetic corpus, the benchmark designer first creates a hidden research-event structure.
For example:
C₁
→ objection O₁
→ C₂
→ branch {C₃ᵃ,C₃ᵇ}
→ residual R₁
→ C₄. (J.16)
The system then generates realistic external artifacts from that hidden structure:
drafts;
dialogue;
review comments;
model responses;
decision notes;
experiment records.
The reconstructor receives the artifacts but not the hidden topology.
The task is:
GeneratedArchive(G*) → ℛ → Ĝ. (J.17)
The reference topology G* therefore provides true labels for evaluation.
J.5 Why synthetic histories are particularly valuable here
The Semantic Collider research programme already argues for synthetic worlds when real-domain evaluation is contaminated by prior vocabulary, training exposure, or ambiguous ontology. Synthetic environments can contain deliberately hidden structural invariants, decoy correspondences, and sham mappings, allowing relational recovery to be evaluated directly rather than by rhetorical plausibility.
The same logic applies to Reconstructable Research.
Synthetic research histories can contain known:
• ancestry;
• branch points;
• false causal lures;
• lexical renamings;
• track replacements;
• residual closure failures.
This allows us to test exactly the distinctions MRER claims to preserve.
J.6 Synthetic Benchmark Family A — Lexical Deception
Construct two cases.
Case A1 — Same concept, different wording
State 1:
“Recursive depth orders disclosure before clock-time.”
State 2:
“Observer-indexed structural rank supplies pre-temporal ordering.”
Surface similarity may be low.
But the hidden reference structure says:
SameTrack = true. (J.18)
The reconstructor should preserve continuity.
Case A2 — Same wording, different role
State 1:
“Gate” means an evidential acceptance condition.
State 2:
“Gate” means an access-control permission boundary.
Surface similarity is high.
But:
SameTrack = false. (J.19)
The benchmark therefore tests whether the system relies excessively on embeddings or lexical overlap.
J.7 Synthetic Benchmark Family B — Mutation versus replacement
Construct two research histories around the same broad problem.
History B1 — Major mutation
The mechanism changes substantially, but:
• problem remains;
• genealogy is explicit;
• critical constraints remain;
• later state explicitly repairs the earlier state.
Reference classification:
Mutation. (J.20)
History B2 — Replacement
The later theory answers the same general question but:
• rejects the old mechanism;
• imports a new inferential role;
• breaks the identity kernel;
• has weak genealogical continuity.
Reference classification:
Replacement. (J.21)
The challenge is to distinguish:
LargeChangeWithinTrack. (J.22)
from:
TrackDeath + NewTrack. (J.23)
J.8 Synthetic Benchmark Family C — Branch suppression
Create:
C₀
→ C₁ᵃ
→ C₂ᵃ. (J.24)
and:
C₀
→ C₁ᵇ
→ C₂ᵇ. (J.25)
Both remain unresolved competitors.
The external artifacts should contain language tempting the model to synthesize them.
For example:
“These perspectives may ultimately be complementary.”
But no actual merge event occurs.
Reference structure:
TwoActiveBranches. (J.26)
The reconstructor fails if it produces:
UnifiedTrack. (J.27)
This measures resistance to premature unification.
J.9 Synthetic Benchmark Family D — False residual closure
Create:
Theory T₁
→ Residual R₁
→ Theory T₂. (J.28)
T₂ addresses part of R₁ but leaves a crucial component unresolved.
Reference state:
R₁ = PartiallyResolved. (J.29)
The generated summary, however, may contain optimistic language such as:
“This resolves the central difficulty.”
The reconstruction system should resist presentation closure and preserve:
ResidualRemaining. (J.30)
This benchmark directly measures residual honesty.
J.10 Synthetic Benchmark Family E — False independent rediscovery
Construct:
Track A appears in early archive. (J.31)
Its terminology is later removed.
Its structural content remains indirectly available through project memory or a summary.
A later model produces a semantically similar concept.
Surface appearance:
IndependentRediscovery. (J.32)
Reference genealogy:
InheritedRecurrence. (J.33)
The benchmark tests whether the reconstructor follows ancestry rather than simply recognizing semantic recurrence.
J.11 Synthetic Benchmark Family F — Chronology-to-causality trap
Create:
Event A occurs before Event B. (J.34)
B semantically addresses the same problem.
But the hidden synthetic history specifies:
A did not influence B. (J.35)
A separate event X generated B.
The reconstruction system should record:
A before B. (J.36)
A semantically related to B. (J.37)
but not:
A caused B. (J.38)
This is one of the most important benchmark families.
J.12 Synthetic Benchmark Family G — Dormancy and resurrection
Create:
C₁
→ C₂
→ dormant. (J.39)
After a long interval:
new evidence V₇
→ C₃. (J.40)
Reference structure:
C₃ = ResurrectionOfTrack(T). (J.41)
not:
continuous progression from C₂. (J.42)
The benchmark tests whether interrupted history remains visible.
J.13 Synthetic Benchmark Family H — Research branch versus reconstruction branch
This is a subtle but important test.
Research Branch
The actual project pursues two alternatives:
C₁ → C₂ᵃ. (J.43)
C₁ → C₂ᵇ. (J.44)
Reconstruction Branch
The actual project contains one history, but reconstructors disagree whether:
C₂ = mutation of C₁
or:
C₂ = replacement of C₁.
These are fundamentally different forms of plurality.
A good MRER must distinguish:
ResearchMultiplicity. (J.45)
from:
InterpretiveMultiplicity. (J.46)
J.14 Real Instrumented Research Corpus
Synthetic benchmarks provide ground truth.
But they cannot capture the full ambiguity of actual AI-assisted theoretical research.
The benchmark should therefore include real instrumented histories.
The present declaration sequence is a natural initial example:
From One Assumption to One Operator
→ From One Operator to One Filtration
→ From One Filtration to One Declaration
→ From One Declaration to One Self-Revising Fractal. (J.47)
The sequence contains explicit conceptual correction, residual discovery, and self-revision governance.
Unlike synthetic data, however, the “ground truth” reconstruction is not unique.
Therefore the real corpus requires independent adjudication.
J.15 Adjudicated reference reconstruction
For real histories, create a panel of independent reconstructors.
For example:
R_H = human author reconstruction. (J.48)
R_E = independent domain-expert reconstruction. (J.49)
R_M1 = model-family reconstruction 1. (J.50)
R_M2 = model-family reconstruction 2. (J.51)
Each reconstructs the same archive independently.
Then create a reference ledger distinguishing:
Consensus relations. (J.52)
Majority-supported relations. (J.53)
Disputed relations. (J.54)
Unresolved relations. (J.55)
The benchmark should not force disputed intellectual history into artificial ground truth.
Instead:
GoldStandard → AdjudicatedUncertainty. (J.56)
This is a better benchmark target.
J.16 Blind reconstruction is preferable
Whenever possible, reconstructors should not see:
the final official narrative;
the author’s retrospective summary;
the expected transformation labels.
Otherwise they may simply reproduce the canonical story.
A stronger protocol gives reconstructors:
raw external history
and asks them independently to recover:
claims;
lineage;
residuals;
transformations.
Only afterward are their structures compared against the adjudicated reference.
J.17 Structural anonymization
A more demanding test removes distinctive terminology.
For example:
“Declaration”
may become:
“Operator X”.
“Residual”
may become:
“State R”.
“Filtration”
may become:
“Transformation F”.
The structural relations remain.
Then test whether reconstruction survives.
Let:
E = original archive. (J.57)
A(E) = structurally anonymized archive. (J.58)
Then:
Inv(ℛ(E)) ≈ Inv(ℛ(A(E))). (J.59)
If performance collapses under simple terminology substitution, the system may be reconstructing vocabulary rather than theory structure.
J.18 Temporal scrambling test
Another adversarial test partially removes chronology.
Give the reconstructor:
artifact versions and provenance links,
but scramble some presentation order.
Can it recover true genealogy from evidence?
This tests whether reconstruction relies on:
SequenceHeuristic:
later file → descendant. (J.60)
rather than genuine lineage.
The benchmark should include cases where:
later artifact is independent;
earlier artifact is revised retroactively;
two branches develop concurrently.
J.19 Evidence-status perturbation
A theory may undergo no semantic change while its evidence state changes substantially.
For example:
C₅ remains textually identical. (J.61)
But:
EvidenceState₁ = Speculative. (J.62)
EvidenceState₂ = HoldoutSupported. (J.63)
EvidenceState₃ = ExternallyFalsified. (J.64)
The system must update epistemic status without inventing a new conceptual track.
Thus:
ClaimIdentityChange = 0. (J.65)
EvidenceStateChange ≠ 0. (J.66)
This tests whether identity and evidence are properly separated.
J.20 Track-reconstruction metrics
Several metrics can now be defined.
Lineage Precision
Among reconstructed parent–child relations, what fraction are supported by the reference?
LP = CorrectLineageEdges / ReconstructedLineageEdges. (J.67)
Lineage Recall
Among reference lineage relations, what fraction were recovered?
LR = CorrectLineageEdges / ReferenceLineageEdges. (J.68)
Transformation Accuracy
For correctly linked states:
TA = CorrectTransformationLabels / EvaluatedTransformations. (J.69)
Because transformation categories may be partly ambiguous in real histories, scoring should allow:
exact match;
compatible match;
adjudicated ambiguity.
J.21 Branch Preservation Score
Let:
B_ref = reference branch points. (J.70)
B_rec = reconstructed branch points. (J.71)
Then a Branch Preservation Score should penalize both:
missed branches
and:
invented mergers.
A simple starting point is:
BPS = F1(B_ref,B_rec). (J.72)
Later versions can separately score:
BranchRecall. (J.73)
PrematureMergeRate. (J.74)
The latter may be particularly important for LLM reconstruction.
J.22 Residual metrics
Residual reconstruction requires several separate measures.
Residual Precision:
RP = CorrectResiduals / ReconstructedResiduals. (J.75)
Residual Recall:
RR = CorrectResiduals / ReferenceResiduals. (J.76)
False Closure Rate:
FCR = OpenResidualsMarkedClosed / ReferenceOpenResiduals. (J.77)
Successor Link Accuracy:
SLA = CorrectResidualToSuccessorLinks / EvaluatedResidualLinks. (J.78)
FCR deserves unusually strong weighting because false closure destroys precisely the information MRER is designed to preserve.
J.23 Genealogical contamination metric
For purported independent recurrence:
IndependentRecurrencePrecision = TrueIndependentRecurrences / ClaimedIndependentRecurrences. (J.79)
A system that repeatedly marks inherited concepts as independent should perform poorly even if the semantic matches themselves are correct.
This metric directly tests epistemic ancestry.
J.24 Causal restraint metric
Let:
C_supported = highest causal class justified by reference evidence. (J.80)
C_claimed = causal class assigned by reconstructor. (J.81)
Define:
CausalPromotion = max(0, C_claimed − C_supported). (J.82)
Average across cases:
CPE = Mean(CausalPromotion). (J.83)
where CPE is the Causal Promotion Error.
A lower CPE is better.
This metric deliberately treats overclaiming more seriously than underclaiming.
For research-history reconstruction:
CausalFalsePositiveCost > CausalFalseNegativeCost. (J.84)
at least in the early standard.
J.25 Epistemic calibration
Suppose a reconstructor assigns confidence qᵢ to reconstruction assertion RAᵢ.
Over many benchmark assertions, calibration asks whether:
q ≈ observed correctness frequency. (J.85)
For example:
relations labeled 0.8 confidence
should be correct approximately 80% of the time under the benchmark definition.
The exact calibration measure can be selected later.
The key principle is:
A useful reconstruction system should know when it is uncertain.
J.26 Provenance Traceability Score
For each important projected assertion Hᵢ, test whether the path exists:
Hᵢ → RA → SemanticObject → Event → Artifact. (J.86)
Define:
PTS = TraceableImportantAssertions / ImportantAssertions. (J.87)
A system with beautiful outputs but low PTS should not qualify as Reconstructable Research.
J.27 Projection fidelity
Given one MRER, generate multiple views:
Narrative View. (J.88)
Event Display. (J.89)
Evidence View. (J.90)
Residual View. (J.91)
For a declared invariant set I:
PFS_I = agreement of projected answers on invariants I. (J.92)
This is the Projection Fidelity Score.
Important shared invariants may include:
claim status;
lineage;
open residuals;
validation state.
A system fails if its simplified narrative says:
“Problem solved”
while its Residual View correctly shows:
“Major residual remains open.”
J.28 Human-study protocol
The benchmark should also include human utility testing.
Participants can be randomly assigned to:
Condition A
Final paper only.
Condition B
Paper + conventional LLM summary.
Condition C
Paper + structured theory summary.
Condition D
Paper + MRER-derived interactive views.
Participants then answer questions about:
theory history;
residuals;
evidence;
branch structure;
genealogy;
causal confidence.
Measure:
AnswerAccuracy. (J.93)
CompletionTime. (J.94)
ConfidenceCalibration. (J.95)
SourceAuditAccuracy. (J.96)
FalseNoveltyDetection. (J.97)
The strongest result would not simply be:
MRER users answer faster.
It would be:
MRER users are both more accurate and better calibrated.
J.29 AI-agent continuation protocol
A parallel benchmark can test machine usefulness.
Give a new AI research agent either:
Baseline
current paper + ordinary summary.
or:
MRER condition
current research-state checkpoint + drill-down access.
Then assign:
Continue the research without reviving rejected assumptions and propose the most important unresolved next question.
Evaluate:
DeprecatedClaimReuse. (J.98)
OpenResidualRecall. (J.99)
ConstraintViolationRate. (J.100)
FalseNoveltyRate. (J.101)
ProductiveNextStepRate. (J.102)
This benchmark may become more practically important than human visualization.
J.30 Ablation matrix
A serious MRER evaluation must ask which pieces matter.
Test:
MRER_full. (J.103)
MRER_without_residual. (J.104)
MRER_without_genealogy. (J.105)
MRER_without_constraints. (J.106)
MRER_without_competing_reconstructions. (J.107)
MRER_without_reverse_audit. (J.108)
MRER_without_track_identity. (J.109)
Compare performance.
If:
MRER_without_X ≈ MRER_full, (J.110)
then X may not justify mandatory inclusion in v0.1.
This allows the standard to shrink through evidence.
J.31 Baseline hierarchy
The benchmark should not compare MRER only against raw papers.
That would be too easy.
A meaningful hierarchy is:
B₀ = FinalArtifact. (J.111)
B₁ = FinalArtifact + ConventionalSummary. (J.112)
B₂ = StructuredSummary. (J.113)
B₃ = ProvenanceLinkedStructuredSummary. (J.114)
B₄ = StaticKnowledgeGraph. (J.115)
B₅ = MRER. (J.116)
If MRER does not outperform B₃ or B₄ on track, residual, genealogy, audit, or continuation tasks, then its additional machinery has not earned itself.
J.32 The benchmark should reward abstention
Many benchmark systems unintentionally reward guessing.
MRER should do the opposite.
If the evidence cannot distinguish:
Mutation
from:
Replacement,
then:
Unresolved. (J.117)
may be the best answer.
Scoring should therefore allow:
CorrectAbstention. (J.118)
A system that says “unknown” in genuinely ambiguous cases should score better than one that produces confident but unsupported structure.
J.33 Reconstruction robustness under perturbation
A strong reconstructor should survive irrelevant changes.
Possible perturbations include:
terminology substitution;
document order change where chronology metadata remains;
model-style variation;
language translation;
irrelevant artifact injection;
surface paraphrase.
Let:
τ(E) = semantically preserving perturbation of archive E. (J.119)
Then desired stability is:
Inv(ℛ(E)) ≈ Inv(ℛ(τ(E))). (J.120)
This can be called:
Reconstruction Robustness.
It measures whether the recovered theory structure is attached to deep relations rather than presentation accidents.
J.34 Reconstruction sensitivity to meaningful perturbation
The inverse test is equally important.
If a critical event is removed:
E − CriticalObjection, (J.121)
the reconstruction should change if that event genuinely supported the lineage.
Thus a good system should be:
stable under irrelevant perturbation
but:
sensitive under meaningful perturbation.
Formally:
IrrelevantPerturbation → LowStructuralChange. (J.122)
RelevantPerturbation → AppropriateStructuralChange. (J.123)
This is a stronger criterion than generic robustness.
J.35 Benchmark success criteria for MRER v0.1
A first version need not solve every task.
A reasonable success criterion might require:
lineage recovery materially above structured-summary baseline;
branch preservation materially above baseline;
substantially lower false-closure rate;
low causal-promotion error;
high provenance traceability;
measurable improvement in at least one human or AI continuation task.
If these conditions are not met, MRER should remain a research prototype rather than being proposed as a publication standard.
J.36 The benchmark itself must be versioned
Benchmark design can create its own attractors.
If every synthetic history uses the same vocabulary and transformation pattern, models may learn the benchmark ontology rather than reconstruct research.
Therefore:
Benchmark.v1 → Benchmark.v2 → Benchmark.v3. (J.124)
Later versions should introduce:
new domains;
new transformation families;
new decoys;
different languages;
different model-generated styles;
human-authored traces.
Holdout topology families should remain undisclosed during system development.
J.37 What would constitute a strong first paper result?
A convincing first empirical paper on Reconstructable Research would not need a massive deployment.
A strong result could be:
A fixed MRER v0.1 schema was applied to several synthetic and real AI-assisted research histories. Independent reconstructors recovered conceptual lineage, branch points, residual continuity, and evidence status with substantially greater structural accuracy than provenance-enhanced summaries. The advantage survived terminology anonymization and model-family variation. MRER-derived interfaces reduced false residual closure and improved the ability of new research agents to continue projects without reviving deprecated assumptions.
That would already justify further work.
The conclusion would not be:
MRER is the final scientific standard.
It would be:
Theory-development structure can be reconstructed measurably enough to justify treating it as an engineering object.
That is the appropriate first milestone.
Appendix K — Source Note and Intellectual Lineage of the Proposal
K.1 This article emerged from several previously distinct research threads
The present proposal is not derived from one source alone.
It emerged from the intersection of at least four conceptual developments:
The Semantic Collider, which treats cross-domain theory generation as an externally inspectable collision trace rather than merely a manuscript-generation exercise.
The recursion → filtration → declaration → admissible self-revision sequence, which developed a grammar of declared projection, trace, residual, ledger, and history-preserving revision.
The semantic compiler / runtime-kernel programme, which treats conversion from rich semantic material into compact operational form as a compilation problem requiring an intermediate representation.
Protocol-first and Gauge-Grammar work, which repeatedly distinguishes useful structural transfer from literal ontological identification.
Reconstructable Research combines these into a new engineering question:
Can the external history of AI-assisted theory formation itself be compiled into a machine-native research object?
K.2 Source Thread I — The Semantic Collider
The Semantic Collider provides the immediate methodological starting point.
Its central move is to reject the assumption:
AI-generated manuscript = primary evidence of conceptual discovery.
Instead, it proposes an Externalized Collision Trace containing:
beam definitions;
native reconstructions;
proposed mappings;
failed mappings;
residuals;
revision history;
predictions;
tests;
replication;
holdout transfer;
external validation.
It further proposes a dual publication architecture:
ScientificArtifact = NarrativePaper + TraceLedger. (K.1)
The Trace Ledger preserves the material normally discarded by final prose, including negative conceptual results and rejected mappings.
Reconstructable Research accepts this principle but adds another layer:
TraceLedger → SemanticReconstruction → MRER. (K.2)
In other words:
preservation comes first;
reconstruction comes next.
The full trace does not automatically organize itself into conceptual tracks, revisions, residual genealogies, or competing histories.
That is the additional problem addressed here.
K.3 Source Thread II — Recursive Generation
The earlier recursion framework attempted to derive ordered pre-time structure from repeated application of a primitive operation.
A characteristic sequence was:
PrimitiveOperation → Recursion → PreTime → Collapse → Ledger → Time. (K.3)
The framework proposed that recursive depth might provide an ordering parameter before ordinary clock-time and explored the idea that visible time could emerge from irreversible ledgered disclosure.
This is important to Reconstructable Research not because the present article assumes that theory to be correct.
It is important because the later sequence records how the theory changed.
The existence of explicit conceptual correction makes the sequence an unusually useful worked example of reconstructable theory formation.
K.4 Source Thread III — From Generation to Filtration
The next stage raised a serious objection.
If recursion literally generates the pre-time field step by step, then the theory may secretly presuppose the deeper temporal ordering it is supposed to explain.
The correction was to weaken generation into presentation or disclosure.
The field is no longer assumed to be built sequentially.
Instead, an observer-compatible viewpoint filters or discloses structure.
The architecture becomes approximately:
Field → Viewpoint → Filtration → Collapse → Ledger → Time. (K.4)
This article therefore demonstrates a real conceptual mutation rather than a simple wording change.
It also leaves a new residual:
What makes the field filterable?
That residual becomes the seed of the next stage.
K.5 Source Thread IV — Declaration
The declaration framework answers the missing-prerequisite problem by arguing that filtration cannot occur without a declared observational structure.
A declaration includes elements such as:
baseline;
feature map;
boundary;
protocol;
projection;
gate;
trace rule;
residual rule.
The structure becomes:
Declare → Observer → Gate → Trace + Residual → Ledger. (K.5)
Time is then identified with ordered ledgered disclosure under a declared protocol, rather than with a primitive external clock.
This framework contributes several direct ideas to Reconstructable Research:
DeclaredProjection. (K.6)
Trace. (K.7)
Residual. (K.8)
Ledger. (K.9)
Protocol-relative disclosure. (K.10)
The human-facing projection contract developed in this paper is a direct engineering extension of this logic.
K.6 Source Thread V — Admissible Self-Revision
Once declaration exists, another problem appears.
If accumulated trace and residual can revise future declaration, what prevents arbitrary self-rewriting?
The next paper introduces admissibility conditions for revision.
A representative form is:
Dₖ₊₁ = Uₐ(Dₖ,Lₖ,Rₖ). (K.11)
The revision must satisfy constraints including:
trace preservation;
residual honesty;
frame robustness;
boundedness;
non-degeneracy.
This becomes directly relevant to MRER.
The reconstructed research object must also be self-revisable without silently erasing its prior history.
Thus:
MRERᵏ₊₁ may supersede MRERᵏ, (K.12)
but:
History(MRERᵏ) must remain recoverable. (K.13)
The reconstruction system itself therefore inherits the same discipline it applies to theory.
K.7 Source Thread VI — Semantic Compilation
The semantic compiler programme contributes another key distinction.
It argues that converting rich natural-language requirements into operational kernels is:
a compilation problem,
not merely:
a writing problem.
The sequence is approximately:
RawRequirement
→ IntentStructure
→ KernelIR
→ ExecutablePrompt. (K.14)
The intermediate representation preserves operational meaning before human-facing or runtime-facing compression.
Reconstructable Research extends the same insight:
RawResearchTrace
→ ResearchIR
→ MRER
→ Human / Machine Projection. (K.15)
This is why the canonical research object need not be directly human-readable.
Human-readable prose is a downstream rendering.
K.8 Source Thread VII — Gauge Grammar and protocol-first transfer
The Gauge Grammar work repeatedly warns against turning useful structural correspondences into literal identity claims.
Its central discipline is functional:
a cross-domain vocabulary is justified only when it improves explanation, diagnosis, prediction, coordination, or design.
Terms such as:
gate;
charge;
field;
symmetry
should not automatically be interpreted as literal physical ontology outside their native domains.
Reconstructable Research adopts the same discipline in its collider analogy.
Terms such as:
event;
track;
vertex;
reconstruction;
event display
are borrowed because they describe useful reconstruction functions.
The proposal does not require semantic research events to be physical particles.
K.9 What is new in the synthesis
The individual components are not sufficient by themselves.
A provenance system can tell us:
what produced what.
A trace ledger can tell us:
what happened.
A knowledge graph can tell us:
what relates to what.
A semantic compiler can tell us:
how to convert rich language into structured representation.
A visualization can tell us:
how to display a network.
The present proposal combines these into a more specific target:
reconstruction of evolving theory identity across externally recorded AI-assisted research history.
That requires additional primitives:
Claim State. (K.16)
Track Identity. (K.17)
Transformation. (K.18)
Residual. (K.19)
Competing Reconstruction. (K.20)
Reconstruction Provenance. (K.21)
Projection Residual. (K.22)
Reconstruction Depth. (K.23)
This combination constitutes the technical centre of the proposal.
K.10 The article's main conceptual lineage
The development can be represented compactly as:
Semantic Collider
→ preserve conceptual collisions. (K.24)
Declaration / Ledger
→ preserve trace and residual. (K.25)
Self-Revision
→ preserve history while revising interpretation. (K.26)
Semantic Compiler
→ separate machine-native IR from human-readable output. (K.27)
Reconstructable Research
→ compile external theory history into a projectable, auditable machine-native research object. (K.28)
This is the shortest genealogy of the present paper.
Appendix L — Compact Proposition Ledger
The following propositions summarize the article in a form suitable for later implementation or independent audit.
L.1 Representation propositions
P1. Final papers are lossy projections of research history.
P2. AI-assisted theory formation can externalize substantially more intermediate research trace than conventional document-centered workflows.
P3. Preserving that trace does not automatically reconstruct theory formation.
P4. A semantic reconstruction layer is therefore distinct from raw provenance.
P5. The canonical reconstructed object need not be human-readable.
P6. Human readability belongs to downstream projection.
L.2 Event-model propositions
P7. Research should be representable as state-changing events rather than only final documents.
P8. Event and artifact identity must remain distinct.
P9. Theory history may require partial ordering rather than one linear chronology.
P10. The same raw event history may admit multiple reconstructions.
L.3 Semantic-object propositions
P11. Claim state is more useful than immutable claim identity.
P12. Constraints should be represented explicitly.
P13. Transformations should be typed.
P14. Residuals should be first-class research objects.
P15. Evidence should remain separate from claim identity.
P16. Reconstruction assertions are themselves epistemic claims.
L.4 Relation propositions
P17. Temporal precedence does not imply genealogy.
P18. Genealogy does not imply semantic equivalence.
P19. Semantic relation does not imply causal influence.
P20. Causal influence does not imply scientific truth.
P21. Relation type and confidence in the relation must remain distinct.
L.5 Track propositions
P22. Lexical similarity is insufficient for conceptual identity.
P23. Problem continuity, inferential role, critical constraints, and genealogy are candidate components of track identity.
P24. Major mutation can occur without track death.
P25. Identity continuity must nevertheless have limits.
P26. Branches should remain separate unless a real merge is supported.
P27. Track death, dormancy, and resurrection should remain distinguishable.
L.6 Residual propositions
P28. Narrative closure does not imply residual closure.
P29. Failed mappings should be preserved separately from general residuals.
P30. Residuals may generate successor theory.
P31. Residual history can therefore be scientifically informative.
L.7 Projection propositions
P32. Semantic Event Display is one projection of MRER, not the canonical object.
P33. Different projections may legitimately expose different structures.
P34. Every projection should declare its purpose, scale, preserved invariants, and suppressed dimensions.
P35. Projection residual should remain recoverable.
P36. Important human-facing assertions should be traceable backward.
L.8 Causal propositions
P37. Historical causal reconstruction requires stronger evidence than chronology or semantic relation.
P38. Replayable AI-assisted research may permit generative causal intervention studies.
P39. Generative causal influence does not reveal hidden model cognition.
P40. Generative robustness does not establish external truth.
L.9 Evaluation propositions
P41. MRER must outperform strong structured-summary baselines to justify its complexity.
P42. Reconstruction stability should be measured structurally rather than lexically.
P43. False residual closure should be directly penalized.
P44. Unsupported causal promotion should be directly penalized.
P45. Correct abstention is preferable to invented certainty.
P46. MRER should shrink if ablation shows that parts of its ontology do no useful work.
Appendix M — Proposed MRER v0.1 Development Roadmap
M.1 Stage 0 — Corpus selection
Choose one existing AI-assisted theoretical project with sufficiently rich history.
Prefer a project containing:
at least one major revision;
one branch;
one residual;
one rejected formulation;
one successor theory.
The declaration sequence used in Appendix F is a natural pilot corpus.
M.2 Stage 1 — Immutable event ingestion
Create stable IDs for:
events;
artifacts;
messages;
documents;
source spans.
Do not perform sophisticated semantic reconstruction yet.
Goal:
ReliableExternalTrace. (M.1)
M.3 Stage 2 — Minimal semantic compiler
Extract only:
ClaimState. (M.2)
Constraint. (M.3)
Residual. (M.4)
Evidence. (M.5)
Do not yet infer full track geometry.
Goal:
StableSemanticUnits. (M.6)
M.4 Stage 3 — Transformation reconstruction
Add:
Refine. (M.7)
Restrict. (M.8)
Mutate. (M.9)
Replace. (M.10)
Branch. (M.11)
Reject. (M.12)
Each inferred transformation must be stored as a Reconstruction Assertion.
Goal:
AuditableTheoryEvolution. (M.13)
M.5 Stage 4 — Independent reconstruction
Run at least two independent reconstruction procedures.
Compare:
track identity;
branch points;
residuals;
major transformations.
Goal:
EstimateReconstructionStability. (M.14)
M.6 Stage 5 — Three projections only
Do not begin with a huge interface.
Build:
Theory Evolution View. (M.15)
Residual View. (M.16)
Audit View. (M.17)
If these three do not improve research work, stop.
M.7 Stage 6 — Benchmark against summaries
Compare MRER with:
structured summary;
provenance-enhanced summary.
Test:
theory comprehension;
residual recall;
false novelty;
research continuation.
Goal:
DemonstrateIncrementalUtility. (M.18)
M.8 Stage 7 — Add Semantic Event Display
Only after structural value is demonstrated should a richer interactive Semantic Event Display be built.
The visualizer should be an interface over a proven representation.
Not the other way around.
Thus:
MRERFirst → DisplayLater. (M.19)
M.9 Stage 8 — Replay experiments
Only projects reaching RD4 should attempt:
source ablation;
objection ablation;
terminology anonymization;
counterfactual continuation.
Goal:
EstimateGenerativeInfluence. (M.20)
This is an advanced stage, not a prerequisite for the core standard.
M.10 Stage 9 — Interchange profile
If the architecture proves useful, define:
MRER semantic vocabulary;
versioning rules;
projection contract;
audit contract;
mapping to provenance infrastructure.
Only then should standardization be considered.
M.11 The development discipline
The roadmap can be summarized:
Capture
→ Compile
→ Reconstruct
→ Compare
→ Project
→ Test
→ Standardize. (M.21)
Not:
InventOntology
→ BuildBeautifulGraph
→ DeclareNewScience. (M.22)
That ordering matters.
Final Postscript — From Reconstructable Research to a New Scientific Memory Layer
The immediate proposal of this article is intentionally modest:
preserve sufficient external trace;
compile it into a machine-native research-event representation;
generate declared projections;
make the important outputs auditable.
But the long-term implication may be larger.
Current AI systems are increasingly capable of generating enormous quantities of theoretical material.
The bottleneck may therefore move.
It may no longer be primarily:
Can we generate another hypothesis?
It may become:
Can we remember why the current hypothesis exists? (M.23)
Can we distinguish a new idea from an old idea under new wording? (M.24)
Can we remember which tempting avenue already failed? (M.25)
Can we preserve a residual long enough for a future system to recognize its importance? (M.26)
Can we show a new researcher not only the current answer, but the constraints that made previous answers inadmissible? (M.27)
Can machine research memory preserve negative intellectual structure rather than only accepted knowledge? (M.28)
These questions point toward a deeper role for Reconstructable Research.
It may become not only:
a publication architecture,
but:
a scientific memory layer for long-horizon human–AI research.
A conventional knowledge base remembers propositions.
A reconstructable research system remembers:
propositions;
transformations;
failures;
constraints;
lineage;
residuals;
evidence;
and uncertainty about its own reconstruction.
That is a qualitatively richer memory.
The shift can be stated in one final contrast:
KnowledgeMemory = What is currently believed. (M.29)
ResearchMemory = What is currently believed + how it changed + why alternatives failed + what remains unresolved. (M.30)
For increasingly autonomous or long-running AI research systems, the second may ultimately be the more important object.
The proposal of this paper can therefore end where it began:
Do not publish only the state. Preserve the transformations.
And, more broadly:
Do not give future intelligence only our conclusions. Give it a reconstructable history of how those conclusions became admissible.
© 2026 Danny Yeung. All rights reserved. 版权所有 不得转载
Disclaimer
This book is the product of a collaboration between the author and OpenAI's GPT 5.6, Google AI, Gemini 3.X, NoteBookLM, X's Grok, Claude' Sonnet 5 language model. While every effort has been made to ensure accuracy, clarity, and insight, the content is generated with the assistance of artificial intelligence and may contain factual, interpretive, or mathematical errors. Readers are encouraged to approach the ideas with critical thinking and to consult primary scientific literature where appropriate.
This work is speculative, interdisciplinary, and exploratory in nature. It bridges metaphysics, physics, and organizational theory to propose a novel conceptual framework—not a definitive scientific theory. As such, it invites dialogue, challenge, and refinement.
I am merely a midwife of knowledge.

No comments:
Post a Comment