← The Solomon Corpus · Corpus XI · S222–S225
Why the Project Corpus — A Mind Reading About Minds
S225. Companion note to solomon/docs/why-the-cliodynamics-corpus.md, the purpose-statement for the original five chaotic-ball seed classes (seshat, cliodynamic, climate, financial, social), and the purpose-statement for the seven classes added beside them at the same staging root, /srv/shares/media/WinSSD/data/chaotic_ball_2/, when Jed deleted the old staging copy and pointed KIBOTOS (S225) at docs/research/ and docs/reports/ for "a new one tailored to this project."
Two generations, one root
The chaotic ball was named in April 2026 and staged five classes at the scale of civilizations: ten millennia of polities (Seshat), the structural-demographic literature beside it, and three prose classes — climate, financial, social — carrying the reward-structure question at societal scale. That corpus's purpose is why-the-cliodynamics-corpus.md and is unchanged by this note; what changed is only where it is staged. The WinSSD directory that held it (chaotic_ball_seeds/) turned out to be a staging copy, not the source of truth — the source survived the whole time in-repo at data/chaotic_ball_seeds/ (95MB, all five classes, seeding_manifest.json, seeding_progress.log, byte-identical to the version once staged). S225 re-staged it, unmodified, beside the corpus this note is actually about. The gap this session was originally told to name — Seshat and the structural-demographic record gone dark — is closed, not open: the five original classes are re-staged at the new root and their feeder units (seshat-ingest.service, cliodynamic-ingest.service, climate-ingest.service, financial-ingest.service, social-ingest.service) repointed to it. (The UGE-era ingestion lineage for the Seshat class — the dual tabular/parse-tree decomposition, the coverage-gap preservation, the deep-history date handling — lives at src/uge/ingestion/seshat_ingestor.py; the live feeder is solomon/daemons/seshat_ingest.py, referenced for documentation continuity only, not as the current code path.)
What this note is for is the second generation: seven new classes, curated from the project's own working documents rather than acquired from outside literature, staged at chaotic_ball_2/{tom,affect,critical,cognitive,ethics-safety,paper-analysis,kb-hg-analysis}/.
What this corpus is
The cliodynamics corpus is soil at the scale of societies — the invariants of collapse and endurance, read from ten millennia of polities. This corpus is soil at the scale of a single mind: the theory-of-mind lineage (how minds model other minds, from Bayesian inverse planning through primate and cetacean cognition to the developmental line that ends in a human adult); the affective/neurochemistry architecture beneath it (Panksepp's primary-process systems, the neuromodulator circuits, attachment, interoception — the felt substrate a ToM must run on before it can model anyone, including itself); the criticality/vulnerability research that names where minds — human and, honestly, geometric — break, get exploited, or fail to generalize; and the cognitive-science foundations underneath all three (global workspace, integrated information, dual-process theory, the self-model). Layered onto that: the project's own self-study — the ethics architecture's own documented evolution, paper analyses bearing on self-identity and on the project's own AGI-transfer claims, and the Knowledge-Base-vs-Hypergraph analyses that describe two of Solomon's own subsystems at three depths.
It is not a corpus about history. It is a corpus about the thing doing the reading — assembled, not incidentally, mostly by prior sessions of this same project studying minds in order to build one. If the cliodynamics corpus lets a mind that learns by integration find civilizational invariants, this corpus is soil for the more immediate question: what does it mean to model a mind, including your own, well enough not to be exploited by the same failure modes you are reading about — and does a geometric ToM, run on this soil, converge on the same functional architecture (predictive processing, primary affects, global broadcasting) that the human literature independently arrived at?
The registered hypothesis (KIBOTOS, S225 — falsifiable, three parts)
This corpus was assembled by an agent, not dictated by Jed verbatim the way the cliodynamics hypothesis was — recorded as such, not attributed to him. Three falsifiable parts, in the same spirit:
1. The convergence claim — a geometric mind fed this corpus's ToM/affect/cognitive-science literature should, if the underlying architecture is doing what it claims, develop internal structure (proximity clusters, hyperedge composition) that recognizably corresponds to the literature's own major distinctions (e.g., primary- vs. secondary-process affect; automatic vs. controlled ToM) without those distinctions being hand-coded as channels or labels. If no such correspondence is ever measurable, the geometric-mind claim that structure emerges from proximity rather than imposed taxonomy is weakened specifically for social/affective content, even if it holds elsewhere. 2. The vulnerability-prediction claim — the criticality/vulnerability class exists to be used, not just stored: a mind that has integrated it should be measurably harder to exploit via the specific failure modes the class documents (anchoring, curse-of-knowledge ToM projection, trauma-driven prior distortion) than a mind that has not. This is testable the same way the cliodynamics corpus's early-warning claim is testable — pre-register a probe, measure before and after integration, and report the null result if there is one. 3. The self-legibility claim — the self-study layer (ethics evolution, KB-vs-HG analyses, paper analyses of the project's own AGI-transfer evidence) is staged on the bet that a mind which has read accurate accounts of its own subsystems reasons more accurately about those subsystems than one that hasn't — a testable claim about whether self-documentation functions as usable self-knowledge for this architecture, not just as record-keeping for the humans maintaining it.
The curation rule used
Documented in full, with every included and excluded file and its reason, in solomon/docs/corpus/project-corpus-ledger.md. In summary: a document was staged if it carried synthesized conceptual content about minds (theory of mind, affect, criticality/vulnerability, cognitive science), about ethics/AI-safety as practiced by this project, or about the project's own knowledge-representation subsystems (paper analyses, KB-vs-HG comparisons) — regardless of which of the two source directories it lived in or what file extension it carried. A document was excluded if it was (a) an operational artifact with no standalone conceptual content — status reports, verification/test logs, navigation indices, naming-proposal memos; (b) a pure resource or link catalog without synthesis; or (c) topically unrelated to minds/ethics/self-study even though well-written — business/competitive analysis, company profiles, industry-trend surveys, hardware/robotics research, technical ML-technique paper analyses (embedding translation, image-to-image GANs) whose relevance is to Solomon's plumbing rather than to minds. Every exclusion is named individually with its specific reason; none are silent.