https://chatgpt.com/share/6a78fc18-6ae0-83ed-8889-7e27dfcb150f
https://osf.io/kcjv3/files/osfstorage/6a78fb1ab195de03f21fb7bb
Reconstructable Research
A Machine-Native Event Architecture for AI-Assisted Theory Formation
Abstract
Large language models have changed the economics of theoretical exploration. A research programme can now generate hundreds of conceptual variants, objections, cross-domain mappings, revisions, failed formulations, auxiliary hypotheses, and synthesized manuscripts at a speed that was previously impossible for an individual researcher. Yet the dominant publication object remains almost unchanged: the final paper.
This creates an epistemic compression problem.
A conventional manuscript normally presents the current theory as a coherent argument. It does not preserve, in machine-operable form, the full genealogy by which the theory arose: which source introduced which concept; which objection destroyed which formulation; which constraint survived revision; which mapping failed; which unresolved residual generated a successor theory; which branch was abandoned; which result was independently rediscovered; and which apparent recurrence was merely inherited through prior context. In AI-assisted theoretical work, these omissions become especially consequential because the generative search process can be vastly larger than the final document.
This paper proposes Reconstructable Research: an architecture in which the final paper is no longer treated as the sole canonical object of theory formation. Instead, externally recorded research events are captured and compiled into a machine-native, provenance-bearing representation of research history. This representation need not itself be human-readable. It needs only to preserve enough structure that declared human-readable projections can later be generated, audited, compared, and traced back toward source events.
The central transition is:
Research Events → Event Capture → Semantic Compilation → Machine-Native Research Event Representation → Declared Projection → Human / Machine Views. (0.1)
The proposed canonical object is called the Machine-Native Research Event Representation, abbreviated MRER. MRER is not assumed to be a simple graph. It may be implemented as a typed graph, hypergraph, event store, provenance system, vector-symbolic representation, relational structure, or hybrid architecture. Its defining requirement is semantic rather than syntactic: it must preserve distinguishable research objects and transformations such as events, artifacts, claim states, constraints, revisions, residuals, evidence, genealogy, and reconstruction assertions.
The paper develops four architectural contracts:
Capture → Reconstruct → Project → Audit. (0.2)
The Capture Contract records externally observable research events without claiming access to hidden model cognition. The Reconstruction Contract compiles those events into structured claims about theory evolution. The Projection Contract generates approximately human-readable views under declared purposes and fidelity constraints. The Audit Contract allows important projected assertions to be traced backward through reconstruction assertions toward machine objects and original provenance.
A central methodological distinction is:
Later Than ≠ Derived From ≠ Semantically Related To ≠ Caused By. (0.3)
Chronology, genealogy, semantic relation, causal influence, and epistemic status must therefore remain distinct relation layers. The framework also treats residuals as first-class research objects. What failed to fit a theory may be as important as what survived, because unresolved residuals frequently become the pressure that generates successor formulations.
The paper further introduces conceptual track identity, mutation, branching, merge, replacement, dormancy, resurrection, reconstruction depth, competing reconstructions, projection residual, and generative causal replay. A worked example follows the development from recursive generation to viewpoint filtration, declaration, and admissible self-revision, showing how a theory can undergo deep conceptual mutation while retaining an identifiable research lineage. The source sequence itself explicitly records these corrections: recursive generation was weakened into disclosure to avoid a hidden meta-time; filtration then exposed the need for declaration; declaration in turn exposed the danger of unrestricted self-revision.
The governing proposal is therefore not that machines can recover the hidden truth of intellectual history. It is narrower:
A sufficiently instrumented research process can become reconstructable under declared protocols.
The paper becomes one projection of that reconstructable object rather than the object itself.
The guiding maxim is:
Do not publish only the state. Preserve the transformations.







