AbstractPhil commited on
Commit
65e4c84
·
verified ·
1 Parent(s): 4d81e3a

exp018_r12 + exp019_cr + exp020_gen packages (protocol ship: sub-article READMEs, standalone scrubbed code, self-asserting builders, ledgers, all specimens). exp018: keystone init avenues priced (neutral) / tied readout scoped (reconstruction-regime). exp019: capacity cliff, zero-to-negative distillation tax, anchors carry no bytes, universal forgetting. exp020: rule-induction plateau structure-independent; aleph bottleneck best generalizer; inverse law (memorization ease substitutes for induction)

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. README.md +39 -0
  2. exp018_r12/README.md +64 -0
  3. exp018_r12/ar_differentiation_bed.py +487 -0
  4. exp018_r12/build_results.py +36 -0
  5. exp018_r12/exp014_genetic_distillation.py +515 -0
  6. exp018_r12/exp018_reran12.py +176 -0
  7. exp018_r12/geolip_vitals.py +219 -0
  8. exp018_r12/read_codebook.py +159 -0
  9. exp018_r12/repro.py +27 -0
  10. exp018_r12/results/ledger.jsonl +8 -0
  11. exp018_r12/results/results.json +34 -0
  12. exp018_r12/specimens/r12_farmed_init_s0.pt +3 -0
  13. exp018_r12/specimens/r12_farmed_init_s1.pt +3 -0
  14. exp018_r12/specimens/r12_keystone_s0.pt +3 -0
  15. exp018_r12/specimens/r12_keystone_s1.pt +3 -0
  16. exp018_r12/specimens/r12_penta_init_s0.pt +3 -0
  17. exp018_r12/specimens/r12_penta_init_s1.pt +3 -0
  18. exp018_r12/specimens/r12_tied_s0.pt +3 -0
  19. exp018_r12/specimens/r12_tied_s1.pt +3 -0
  20. exp019_cr/README.md +82 -0
  21. exp019_cr/ar_differentiation_bed.py +487 -0
  22. exp019_cr/build_results.py +81 -0
  23. exp019_cr/exp014_genetic_distillation.py +515 -0
  24. exp019_cr/exp019_content_retention.py +402 -0
  25. exp019_cr/geolip_vitals.py +219 -0
  26. exp019_cr/read_codebook.py +159 -0
  27. exp019_cr/repro.py +29 -0
  28. exp019_cr/results/ledger.jsonl +36 -0
  29. exp019_cr/results/results.json +21 -0
  30. exp019_cr/specimens/rule_direct_s0.pt +3 -0
  31. exp019_cr/specimens/rule_direct_s1.pt +3 -0
  32. exp019_cr/specimens/rule_kd_facts_s0.pt +3 -0
  33. exp019_cr/specimens/rule_kd_facts_s1.pt +3 -0
  34. exp019_cr/specimens/rule_kd_general_s0.pt +3 -0
  35. exp019_cr/specimens/rule_kd_general_s1.pt +3 -0
  36. exp019_cr/specimens/rule_teacher_s0.pt +3 -0
  37. exp019_cr/specimens/rule_teacher_s1.pt +3 -0
  38. exp019_cr/specimens/teacher_N1024_s0.pt +3 -0
  39. exp019_cr/specimens/teacher_N1024_s1.pt +3 -0
  40. exp019_cr/specimens/teacher_N256_s0.pt +3 -0
  41. exp019_cr/specimens/teacher_N256_s1.pt +3 -0
  42. exp019_cr/specimens/teacher_N64_s0.pt +3 -0
  43. exp019_cr/specimens/teacher_N64_s1.pt +3 -0
  44. exp020_gen/README.md +73 -0
  45. exp020_gen/ar_differentiation_bed.py +487 -0
  46. exp020_gen/build_results.py +52 -0
  47. exp020_gen/exp014_genetic_distillation.py +515 -0
  48. exp020_gen/exp017_aleph_constellation.py +266 -0
  49. exp020_gen/exp019_content_retention.py +402 -0
  50. exp020_gen/exp020_generalization.py +231 -0
README.md CHANGED
@@ -198,6 +198,45 @@ trainable bank. Value case: the structured 768 address (conditioning/routing/
198
  lookup), not raw bpb. 8-row ledger + 6 checkpoints;
199
  write-up in [exp017_ac/README.md](./exp017_ac/README.md).
200
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
201
  ## Reproducibility
202
 
203
  Every experiment package (`exp012_ar/`, `exp013_aug/`, `exp014_gd/`,
 
198
  lookup), not raw bpb. 8-row ledger + 6 checkpoints;
199
  write-up in [exp017_ac/README.md](./exp017_ac/README.md).
200
 
201
+ ## exp018 — experiment 18: exp012 rerun with keystone parameters (July 11, 2026)
202
+
203
+ [`exp018_r12/`](./exp018_r12): a labeled rerun (exp012's record untouched)
204
+ pricing the two keystone avenues the original bed never exercised. Verdicts:
205
+ the init avenue is task-neutral (geovocab2 regular-pentachoron vertices and a
206
+ farmed recon-real codebook both land in the certified band; farmed anchors
207
+ drift less, accelerate nothing); the tied M̂ readout fails in autoregression
208
+ (+1.0 bpb, codebook starved) — a reconstruction-regime device. Net: exp012's
209
+ original parameters stand. Write-up: [exp018_r12/README.md](./exp018_r12/README.md).
210
+
211
+ ## exp019 — the capacity for distilled content retention (July 11, 2026)
212
+
213
+ [`exp019_cr/`](./exp019_cr): fact corpora in the byte stream; exact-match
214
+ recall as the retention gauge; distillation channels vs direct learning across
215
+ N ∈ {64, 256, 1024}, plus a rule-content generalization block. Verdicts: a
216
+ capacity cliff between 256 and 1024 facts; **the distillation tax is
217
+ zero-to-negative** (teacher logits alone transfer rote content at parity+ and
218
+ rule content better than ground truth, 2/2 seeds — with a steep clean-bpb
219
+ cost); no logit leakage without exposure; **anchors carry no bytes** (a
220
+ trained codebook from a 256-fact teacher transfers none of them); **universal
221
+ catastrophic forgetting** (every channel → 0.000 exact after 1k clean steps) —
222
+ persistence, not transfer, is the unsolved axis. Rules are learned only
223
+ fragmentarily and content is format-locked (KD students less so). Write-up:
224
+ [exp019_cr/README.md](./exp019_cr/README.md).
225
+
226
+ ## exp020 — generalization: structure vs capacity on a hidden rule (July 11, 2026)
227
+
228
+ [`exp020_gen/`](./exp020_gen): the exp019 rule task raced across six
229
+ structural arms with an exact param-matched control. Verdicts: no structure
230
+ lifts rule induction off the ~0.26 plateau (exact 0.000 in all 12 cells); the
231
+ **aleph bottleneck is the best generalizer** (+25% over its matched free head,
232
+ both seeds); **the inverse law** — generalization ordering reverses modeling
233
+ strength (the constellation head models best and generalizes worst;
234
+ memorization ease substitutes for rule induction); format diversity covers
235
+ seen formats at full rule level but not novel ones. The measured
236
+ memorization↔generalization axis doubles as a component registry for
237
+ slider/registry-style composites. Write-up:
238
+ [exp020_gen/README.md](./exp020_gen/README.md).
239
+
240
  ## Reproducibility
241
 
242
  Every experiment package (`exp012_ar/`, `exp013_aug/`, `exp014_gd/`,
exp018_r12/README.md ADDED
@@ -0,0 +1,64 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # exp018_r12 — experiment 18: exp012 rerun with the keystone parameters
2
+
3
+ A labeled rerun, not a replacement: [exp012](../exp012_ar/)'s certified record
4
+ stands untouched. A process review of the exp012 bed against the aleph
5
+ keystone ([aleph-void article](https://huggingface.co/blog/AbstractPhil/geometric-vocabulary-patchwork-aleph-void))
6
+ found the `AlephAddress` primitive faithful (K=64, D=4, τ=0.1, M̂ reads,
7
+ pure Adam) but two keystone avenues never exercised. exp018 prices both.
8
+
9
+ ## Arms (2 seeds, 2000 steps, otherwise the certified bed unchanged)
10
+
11
+ | arm | s0 | s1 | drift s0/s1 |
12
+ |---|---|---|---|
13
+ | penta_init — geovocab2 regular pentachoron vertices as codebook init | 2.4957 | 2.4389 | .242 / .211 |
14
+ | farmed_init — an exp012 specimen's trained (recon-real) codebook as init | 2.4934 | 2.4614 | .179 / .199 |
15
+ | tied — zero-parameter M̂ readout (decode tied through the projection + byte embedding) | 3.5131 | 3.4997 | .019 / .020 |
16
+ | keystone — penta_init + tied | 3.5169 | 3.5006 | .020 / .021 |
17
+
18
+ exp012 certified reference: addr_msl64 3-seed mean 2.469; s0 bed 2.4990.
19
+ `build_results.py` re-asserts every claim from `results/ledger.jsonl`.
20
+
21
+ ## Verdicts
22
+
23
+ 1. **The init avenue is task-neutral.** Both the exact regular-pentachoron
24
+ init (the factory's 5-dim construction, centered and SVD-projected to R⁴ —
25
+ pairwise cos = −¼ asserted in the smoke) and the farmed recon-real
26
+ codebook land inside the certified band. Consistent with the
27
+ basin-set-at-init and interchangeable-scaffold results: codebook identity
28
+ does not price into bits-per-byte on this bed at this budget.
29
+ 2. **Farmed codebooks drift less** (0.18–0.20 vs 0.21–0.24) — mature anchors
30
+ are more stationary — **and accelerate nothing**, echoing the exp014
31
+ implant studies.
32
+ 3. **The tied M̂ readout fails in autoregression** (+1.0 bpb, both seeds, with
33
+ or without the pentachoron init) **and starves the codebook** (drift ~0.02,
34
+ binding fraction 0): it is a reconstruction-regime device; as a
35
+ zero-parameter AR head it cannot shape a 256-way distribution and passes
36
+ almost no cultivating gradient into the address.
37
+ 4. **Net: exp012's original parameters stand.** Random init + free linear
38
+ head was not a wrong configuration — the two unexercised keystone avenues
39
+ are now priced (one neutral, one negative in this regime).
40
+
41
+ ## Files
42
+ - `exp018_reran12.py` — pentachoron/farmed inits, the tied readout, runner,
43
+ smoke (asserts exact simplex regularity).
44
+ - `geolip_vitals.py` / `ar_differentiation_bed.py` /
45
+ `exp014_genetic_distillation.py` / `read_codebook.py` — this package's own
46
+ copies of the shared harness. Standalone.
47
+ - `repro.py` — loads the code files from this folder and runs them.
48
+ - `build_results.py` → `results/results.json` — re-asserts every claim above.
49
+ - `results/ledger.jsonl` — 8 rows. `specimens/` — all 8 checkpoints.
50
+
51
+ ## Reproduce (from inside this folder)
52
+ ```bash
53
+ pip install torch --index-url https://download.pytorch.org/whl/cu128
54
+ pip install pyarrow huggingface_hub
55
+ pip install "git+https://github.com/AbstractEyes/geolip-svae" # geovocab2 (penta init)
56
+ python repro.py # CPU smoke
57
+ python repro.py --run # 4 arms x 2 seeds (GPU, ~40 min)
58
+ python build_results.py # re-assert every claim from the ledger
59
+ ```
60
+ Data lands in `./data` (override `GEOLIP_DATA`); the farmed donor is
61
+ `../exp012_ar/specimens/addr_msl64_s0_t2000.pt` (in this repo) or set
62
+ `GEOLIP_FARMED_SPECIMEN`.
63
+
64
+ License: MIT · AbstractPhil + Claude Fable 5 · July 11, 2026
exp018_r12/ar_differentiation_bed.py ADDED
@@ -0,0 +1,487 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """ar_differentiation_bed.py — exp012: autoregressive differentiation of the aleph.
2
+
3
+ Differentiation is cultivated by PREDICTIVE pressure along the sequence — the
4
+ address parameterizing the next-byte distribution (Law 2: chain-rule advantage pays
5
+ ONLY where the composed address directly parameterizes the predictive distribution).
6
+ This bed puts the aleph in the autoregressive gradient path and measures what
7
+ differentiates. The head arms enforce the employment law at its maximum: the
8
+ ENTIRE next-byte distribution is parameterized by the address.
9
+
10
+ Byte-level causal LM on wikitext-2-raw (HF parquet, CDN-fast), block 256. ARMS:
11
+ sdpa — standard causal transformer control (matched trunk).
12
+ hub — attention replaced by CAUSAL HUB: linear attention whose feature map
13
+ is the 2K-oriented aleph address, prefix-sum memories (no selection
14
+ event; O(n*K*d)). Differentiation cultivated INSIDE attention.
15
+ addr_head — sdpa trunk, but the OUTPUT HEAD reads ONLY the signed aleph
16
+ coefficient vector w_k = sinh(u_k)/sum_j cosh(u_j) of the final
17
+ hidden state (K -> 256 logits). The address MUST carry every bit of
18
+ next-byte information — the hardest Law-2 bottleneck.
19
+
20
+ JUDGED BY: val bits-per-byte per arm (task) + CULTIVATION VITALS on every aleph
21
+ codebook (readouts, never losses): axis aliveness/hppl, drift-from-init +
22
+ binding fraction @0.29154, winner-|cos| saturation (sign-code emergence), shadow
23
+ path diversity (fixed high-bits hash). Never by recon.
24
+
25
+ Riders: pure Adam wd=0; no BN/Dropout/GAP on geometric paths; orthogonal init;
26
+ Colab-cell-safe (paste-ahead imports, no bare argparse, no __file__ reliance);
27
+ GPU-only for verdict runs.
28
+
29
+ Terminal: python ar_differentiation_bed.py # shapes/parse smoke
30
+ python ar_differentiation_bed.py --train # verdict run
31
+ Colab: paste geolip_vitals.py cell, then this file (smoke auto-runs),
32
+ then train(steps=2000, data_root="/content/data") in the next cell.
33
+ """
34
+ from __future__ import annotations
35
+ import math
36
+ import torch
37
+ import torch.nn as nn
38
+ import torch.nn.functional as F
39
+
40
+ if "anchor_drift" not in globals():
41
+ try:
42
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
43
+ except ImportError:
44
+ _here = globals().get("__file__")
45
+ if _here is not None:
46
+ import sys, pathlib
47
+ sys.path.insert(0, str(pathlib.Path(_here).parent))
48
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
49
+ else:
50
+ raise ImportError(
51
+ "geolip_vitals not found — paste/run its cell first, or "
52
+ "hf_hub_download exp012_ar/geolip_vitals.py from "
53
+ "AbstractPhil/geolip-aleph-differentiation.")
54
+
55
+ VOCAB = 256 # bytes
56
+
57
+
58
+ # ------------------------------------------------------------------ aleph address
59
+ def _super_fibonacci_s3(n: int) -> torch.Tensor:
60
+ """Near-uniform unit quaternions (Alexa CVPR'22) —
61
+ starts the codebook INSIDE the RP^3 attractor basin. D=4 only."""
62
+ PHI, PSI = math.sqrt(2.0), 1.533751168755204288118041
63
+ i = torch.arange(n, dtype=torch.float64)
64
+ s = (i + 0.5) / n
65
+ r, R = torch.sqrt(s), torch.sqrt(1.0 - s)
66
+ a, b = 2 * math.pi * i / PHI, 2 * math.pi * i / PSI
67
+ q = torch.stack([r * torch.sin(a), r * torch.cos(a),
68
+ R * torch.sin(b), R * torch.cos(b)], dim=-1)
69
+ return F.normalize(q, dim=-1).float()
70
+
71
+
72
+ class AlephAddress(nn.Module):
73
+ """Closed-form aleph over 2K oriented half-axes (aleph-void article).
74
+ signed(x): (..., K) w_k = sinh(u_k)/sum_j cosh(u_j) — the Law-2 head feature.
75
+ oriented(x): ((..., K), (..., K)) positive halves of the 2K softmax — HUB map."""
76
+
77
+ def __init__(self, K: int, D: int, tau: float = 0.1, init: str = "random"):
78
+ super().__init__()
79
+ self.K, self.D, self.tau = K, D, tau
80
+ if init == "fibonacci":
81
+ assert D == 4, "fibonacci init lives on S^3 (D=4)"
82
+ A = _super_fibonacci_s3(K)
83
+ else:
84
+ A = F.normalize(torch.randn(K, D), dim=-1)
85
+ self.codebook = nn.Parameter(A)
86
+ self.register_buffer("home", self.codebook.detach().clone())
87
+
88
+ def _u(self, x):
89
+ A = F.normalize(self.codebook, dim=-1)
90
+ return (F.normalize(x, dim=-1) @ A.transpose(-1, -2)) / self.tau
91
+
92
+ def oriented(self, x):
93
+ u = self._u(x)
94
+ m = u.abs().amax(dim=-1, keepdim=True)
95
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
96
+ Z = (ep + en).sum(dim=-1, keepdim=True)
97
+ return ep / Z, en / Z
98
+
99
+ def signed(self, x):
100
+ u = self._u(x)
101
+ m = u.abs().amax(dim=-1, keepdim=True)
102
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
103
+ return (ep - en) / (ep + en).sum(dim=-1, keepdim=True)
104
+
105
+ def signed_at(self, x, taus):
106
+ """Multi-tau stroboscope (rule of 3): signed coefficients at several
107
+ temperatures, concatenated — softer taus keep the vector dense while a
108
+ hard tau supplies the sign-code sharpness. v2 refinement (b)."""
109
+ A = F.normalize(self.codebook, dim=-1)
110
+ cos = F.normalize(x, dim=-1) @ A.transpose(-1, -2)
111
+ outs = []
112
+ for t in taus:
113
+ u = cos / t
114
+ m = u.abs().amax(dim=-1, keepdim=True)
115
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
116
+ outs.append((ep - en) / (ep + en).sum(dim=-1, keepdim=True))
117
+ return torch.cat(outs, dim=-1)
118
+
119
+ def m_hat(self, x):
120
+ """Closed-form soft read (decoders read M_hat, never M). v2 control (c)."""
121
+ u = self._u(x)
122
+ m = u.abs().amax(dim=-1, keepdim=True)
123
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
124
+ A = F.normalize(self.codebook, dim=-1)
125
+ return ((ep - en) @ A) / (ep + en).sum(dim=-1, keepdim=True)
126
+
127
+ def m_hard_ste(self, x):
128
+ """Hard mode (aleph-void article): M_hard = sign(cos_win) * A[win], straight-through to
129
+ the soft read — forward fully discrete SIGN CODE, backward soft gradient.
130
+ Legal per theme A (reconstructive sign code, not a one-hot roster pick)."""
131
+ u = self._u(x)
132
+ soft = self.m_hat(x)
133
+ win = u.abs().argmax(dim=-1)
134
+ A = F.normalize(self.codebook, dim=-1)
135
+ sign = torch.sign(torch.gather(u, -1, win.unsqueeze(-1))).squeeze(-1)
136
+ hard = sign.unsqueeze(-1) * A[win]
137
+ return hard + soft - soft.detach()
138
+
139
+ @torch.no_grad()
140
+ def vitals(self, x_sample) -> dict:
141
+ u = self._u(x_sample.reshape(-1, x_sample.shape[-1]))
142
+ p, n = self.oriented(x_sample.reshape(-1, x_sample.shape[-1]))
143
+ two_k = torch.cat([p, n], dim=-1)
144
+ win = two_k.argmax(dim=-1)
145
+ cos_win = (u.abs().amax(dim=-1) * self.tau) # winner |cos| — sign-code sat.
146
+ d = anchor_drift(self.codebook, self.home)
147
+ return {"drift": round(d["mean"], 4),
148
+ "binding_frac": round(d["binding_fraction"], 4),
149
+ "aliveness": axis_aliveness(two_k),
150
+ "win_cos_mean": round(cos_win.mean().item(), 4),
151
+ "paths": path_diversity(win)}
152
+
153
+
154
+ # ------------------------------------------------------------------------- blocks
155
+ class CausalSDPA(nn.Module):
156
+ def __init__(self, d: int, heads: int = 4):
157
+ super().__init__()
158
+ self.h = heads
159
+ self.qkv = nn.Linear(d, 3 * d, bias=False)
160
+ self.o = nn.Linear(d, d, bias=False)
161
+ nn.init.orthogonal_(self.qkv.weight); nn.init.orthogonal_(self.o.weight)
162
+
163
+ def forward(self, x):
164
+ B, n, d = x.shape
165
+ q, k, v = self.qkv(x).chunk(3, dim=-1)
166
+ q, k, v = (t.view(B, n, self.h, d // self.h).transpose(1, 2) for t in (q, k, v))
167
+ y = F.scaled_dot_product_attention(q, k, v, is_causal=True)
168
+ return self.o(y.transpose(1, 2).reshape(B, n, d))
169
+
170
+
171
+ class CausalHUB(nn.Module):
172
+ """Causal aleph linear attention: prefix-sum memories over the two K-wide
173
+ halves of the oriented address; 2K never materialized; no selection event."""
174
+
175
+ def __init__(self, d: int, K: int = 32, D: int = 4, tau: float = 0.1):
176
+ super().__init__()
177
+ self.addr = AlephAddress(K, D, tau)
178
+ self.q = nn.Linear(d, D, bias=False)
179
+ self.k = nn.Linear(d, D, bias=False)
180
+ self.v = nn.Linear(d, d, bias=False)
181
+ self.o = nn.Linear(d, d, bias=False)
182
+ for m in (self.q, self.k, self.v, self.o):
183
+ nn.init.orthogonal_(m.weight)
184
+
185
+ def forward(self, x):
186
+ qp, qn = self.addr.oriented(self.q(x)) # (B, n, K)
187
+ kp, kn = self.addr.oriented(self.k(x))
188
+ v = self.v(x) # (B, n, d)
189
+ Sp = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kp, v), dim=1)
190
+ Sn = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kn, v), dim=1)
191
+ zp = torch.cumsum(kp, dim=1)
192
+ zn = torch.cumsum(kn, dim=1)
193
+ num = torch.einsum("bnk,bnkd->bnd", qp, Sp) + torch.einsum("bnk,bnkd->bnd", qn, Sn)
194
+ den = (qp * zp).sum(-1, keepdim=True) + (qn * zn).sum(-1, keepdim=True)
195
+ return self.o(num / den.clamp_min(1e-12))
196
+
197
+
198
+ class MslRelay(nn.Module):
199
+ """Depth-composition unit (chain-rule probe): multi-slot M_hat read entering
200
+ the trunk as a NEAR-ZERO gated residual (gate init -3.0, sigma~0.047 — theme D:
201
+ geometry enters as a nudge and grows only if it earns gradient)."""
202
+
203
+ def __init__(self, d: int, n_slots: int = 16, K: int = 64):
204
+ super().__init__()
205
+ self.n_slots = n_slots
206
+ self.proj = nn.Linear(d, n_slots * 4, bias=False)
207
+ self.out = nn.Linear(n_slots * 4, d, bias=False)
208
+ nn.init.orthogonal_(self.proj.weight)
209
+ nn.init.orthogonal_(self.out.weight)
210
+ self.addr = AlephAddress(K, 4)
211
+ self.gate = nn.Parameter(torch.tensor(-3.0))
212
+
213
+ def forward(self, x):
214
+ B, n, _ = x.shape
215
+ slots = self.proj(x).view(B, n, self.n_slots, 4)
216
+ m = self.addr.m_hat(slots).reshape(B, n, -1)
217
+ return x + self.gate.sigmoid() * self.out(m)
218
+
219
+
220
+ class Block(nn.Module):
221
+ def __init__(self, d: int, attn: nn.Module):
222
+ super().__init__()
223
+ self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)
224
+ self.attn = attn
225
+ self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
226
+
227
+ def forward(self, x):
228
+ x = x + self.attn(self.n1(x))
229
+ return x + self.mlp(self.n2(x))
230
+
231
+
232
+ class ByteLM(nn.Module):
233
+ def __init__(self, arm: str, d: int = 192, layers: int = 4, block: int = 256,
234
+ K: int = 32, D: int = 4):
235
+ super().__init__()
236
+ # "<arm>_tri" suffix = trigram byte embedding (AlephLM byte_emb x3 lineage):
237
+ # token embedding is the sum of embeddings of bytes t, t-1, t-2.
238
+ self.trigram = arm.endswith("_tri")
239
+ if self.trigram:
240
+ arm = arm[:-4]
241
+ # "_fib" = super-Fibonacci S^3 codebook init (basin test: starts INSIDE
242
+ # the RP^3 attractor; primary observable is init->final geodesic drift).
243
+ self.fib = arm.endswith("_fib")
244
+ if self.fib:
245
+ arm = arm[:-4]
246
+ # "relay*" = stacked addresses in depth: MslRelay after every block.
247
+ # relay -> sdpa trunk + standard head; relay_msl64 -> + addressed head.
248
+ self.use_relay = arm.startswith("relay")
249
+ if arm == "relay":
250
+ arm = "sdpa"
251
+ elif arm == "relay_msl64":
252
+ arm = "addr_msl64"
253
+ self.arm, self.block = arm, block
254
+ self.emb = nn.Embedding(VOCAB, d)
255
+ if self.trigram:
256
+ self.emb1 = nn.Embedding(VOCAB, d)
257
+ self.emb2 = nn.Embedding(VOCAB, d)
258
+ self.pos = nn.Parameter(torch.zeros(1, block, d) + 0.01 * torch.randn(1, block, d))
259
+ mk_attn = (lambda: CausalHUB(d, K, D)) if arm == "hub" else (lambda: CausalSDPA(d))
260
+ self.blocks = nn.ModuleList([Block(d, mk_attn()) for _ in range(layers)])
261
+ if self.use_relay:
262
+ self.relays = nn.ModuleList([MslRelay(d) for _ in range(layers)])
263
+ self.nf = nn.LayerNorm(d)
264
+ if arm == "addr_head":
265
+ self.head_addr = AlephAddress(K, d) # v1: codebook in model dim — COLLAPSED
266
+ self.head = nn.Linear(K, VOCAB, bias=True)
267
+ elif arm in ("addr_d4", "addr_3tau", "addr_mhat"):
268
+ # v2 refinements: LOW-D HOME — learned projection to the native D=4 home
269
+ # before addressing (mirrors the healthy HUB arms), K=64.
270
+ self.head_proj = nn.Linear(d, 4, bias=False)
271
+ nn.init.orthogonal_(self.head_proj.weight)
272
+ self.head_addr = AlephAddress(64, 4)
273
+ if arm == "addr_d4":
274
+ self.head = nn.Linear(64, VOCAB, bias=True) # w alone, D=4 home
275
+ elif arm == "addr_3tau":
276
+ self.taus = (0.05, 0.1, 0.3) # rule-of-3 strobe
277
+ self.head = nn.Linear(64 * 3, VOCAB, bias=True)
278
+ else: # addr_mhat
279
+ self.head = nn.Linear(4, VOCAB, bias=True) # tightest: M_hat
280
+ elif arm.startswith("addr_msl"):
281
+ # v3: MULTI-SLOT heads — the 16s funnel widening: P parallel D=4 slots
282
+ # over a SHARED codebook. addr_msl consumes the reconstructive M_hat per
283
+ # slot (Px4 dims); addr_msl_w consumes signed w per slot (Px64) — tests
284
+ # whether slot-parallel consumption alone rescues the coefficient path.
285
+ # addr_msl<P> = slot-count dose-response. addr_mslh<P> = HARD sign-code
286
+ # consumption (straight-through M_hard per slot).
287
+ self.hard = arm.startswith("addr_mslh")
288
+ if arm in ("addr_msl", "addr_msl_w"):
289
+ self.n_slots = 16
290
+ else:
291
+ self.n_slots = int(arm[len("addr_mslh" if self.hard else "addr_msl"):])
292
+ self.head_proj = nn.Linear(d, self.n_slots * 4, bias=False)
293
+ nn.init.orthogonal_(self.head_proj.weight)
294
+ self.head_addr = AlephAddress(
295
+ 64, 4, init="fibonacci" if self.fib else "random")
296
+ width = self.n_slots * (64 if arm == "addr_msl_w" else 4)
297
+ self.head = nn.Linear(width, VOCAB, bias=True)
298
+ elif arm == "addr_3tau_mhat":
299
+ # v3: combine the two v2 winners — 3-tau stroboscope + reconstructive read.
300
+ self.head_proj = nn.Linear(d, 4, bias=False)
301
+ nn.init.orthogonal_(self.head_proj.weight)
302
+ self.head_addr = AlephAddress(64, 4)
303
+ self.taus = (0.05, 0.1, 0.3)
304
+ self.head = nn.Linear(64 * 3 + 4, VOCAB, bias=True)
305
+ else:
306
+ self.head = nn.Linear(d, VOCAB, bias=True)
307
+ self._last_h = None
308
+
309
+ def forward(self, idx):
310
+ x = self.emb(idx)
311
+ if self.trigram: # past-only shifts — causality preserved
312
+ x = x + self.emb1(F.pad(idx, (1, 0), value=0)[:, :-1]) \
313
+ + self.emb2(F.pad(idx, (2, 0), value=0)[:, :-2])
314
+ x = x + self.pos[:, : idx.shape[1]]
315
+ if self.use_relay:
316
+ for b, r in zip(self.blocks, self.relays):
317
+ x = r(b(x))
318
+ else:
319
+ for b in self.blocks:
320
+ x = b(x)
321
+ h = self.nf(x)
322
+ self._last_h = h.detach()
323
+ if self.arm == "addr_head":
324
+ return self.head(self.head_addr.signed(h))
325
+ if self.arm == "addr_d4":
326
+ return self.head(self.head_addr.signed(self.head_proj(h)))
327
+ if self.arm == "addr_3tau":
328
+ return self.head(self.head_addr.signed_at(self.head_proj(h), self.taus))
329
+ if self.arm == "addr_mhat":
330
+ return self.head(self.head_addr.m_hat(self.head_proj(h)))
331
+ if self.arm.startswith("addr_msl"):
332
+ B, n, _ = h.shape
333
+ slots = self.head_proj(h).view(B, n, self.n_slots, 4)
334
+ if self.arm == "addr_msl_w":
335
+ feats = self.head_addr.signed(slots).reshape(B, n, -1)
336
+ elif getattr(self, "hard", False):
337
+ feats = self.head_addr.m_hard_ste(slots).reshape(B, n, -1)
338
+ else:
339
+ feats = self.head_addr.m_hat(slots).reshape(B, n, -1)
340
+ return self.head(feats)
341
+ if self.arm == "addr_3tau_mhat":
342
+ p = self.head_proj(h)
343
+ feats = torch.cat([self.head_addr.signed_at(p, self.taus),
344
+ self.head_addr.m_hat(p)], dim=-1)
345
+ return self.head(feats)
346
+ return self.head(h)
347
+
348
+ @torch.no_grad()
349
+ def vitals(self) -> dict:
350
+ out = {}
351
+ if self.arm == "hub":
352
+ for i, b in enumerate(self.blocks):
353
+ if self._last_h is not None:
354
+ out[f"L{i}"] = b.attn.addr.vitals(b.attn.q(self._last_h[:2]))
355
+ elif self.arm == "addr_head" and self._last_h is not None:
356
+ out["head"] = self.head_addr.vitals(self._last_h[:2])
357
+ elif self.arm in ("addr_d4", "addr_3tau", "addr_mhat",
358
+ "addr_3tau_mhat") and self._last_h is not None:
359
+ out["head"] = self.head_addr.vitals(self.head_proj(self._last_h[:2]))
360
+ elif self.arm.startswith("addr_msl") and self._last_h is not None:
361
+ slots = self.head_proj(self._last_h[:2])
362
+ out["head"] = self.head_addr.vitals(
363
+ slots.reshape(*slots.shape[:-1], self.n_slots, 4))
364
+ if self.use_relay and self._last_h is not None:
365
+ for i, r in enumerate(self.relays):
366
+ s = r.proj(self._last_h[:2])
367
+ v = r.addr.vitals(s.reshape(*s.shape[:-1], r.n_slots, 4))
368
+ out[f"relay{i}"] = {"gate": round(r.gate.sigmoid().item(), 4),
369
+ "drift": v["drift"],
370
+ "binding_frac": v["binding_frac"],
371
+ "ppl": round(v["aliveness"]["usage_ppl"], 1)}
372
+ return out
373
+
374
+
375
+ # --------------------------------------------------------------------------- data
376
+ def _wikitext_bytes(data_root: str):
377
+ """wikitext-2-raw as flat uint8 tensors via the HF parquet CDN."""
378
+ from huggingface_hub import hf_hub_download
379
+ import pyarrow.parquet as pq
380
+
381
+ def load(split):
382
+ p = hf_hub_download("Salesforce/wikitext",
383
+ f"wikitext-2-raw-v1/{split}-00000-of-00001.parquet",
384
+ repo_type="dataset", local_dir=data_root)
385
+ text = "".join(pq.read_table(p).column("text").to_pylist())
386
+ return torch.frombuffer(bytearray(text.encode("utf-8")), dtype=torch.uint8).clone()
387
+
388
+ return load("train"), load("validation")
389
+
390
+
391
+ def _batch(data: torch.Tensor, batch: int, block: int, device, g: torch.Generator):
392
+ ix = torch.randint(0, data.numel() - block - 1, (batch,), generator=g)
393
+ x = torch.stack([data[i:i + block] for i in ix]).long().to(device)
394
+ y = torch.stack([data[i + 1:i + block + 1] for i in ix]).long().to(device)
395
+ return x, y
396
+
397
+
398
+ # -------------------------------------------------------------------- train/smoke
399
+ def train(arms=("sdpa", "hub", "addr_head"), steps: int = 2000, batch: int = 32,
400
+ block: int = 256, device: str = "cuda", data_root: str = "./data",
401
+ seed: int = 0, eval_every: int = 500, save: bool = True):
402
+ """Verdict run — GPU only. Pure Adam wd=0. Reports val bits-per-byte + vitals.
403
+ save=True writes {data_root}/ar_ckpts/{arm}_s{seed}_t{steps}.pt per arm —
404
+ the cultivated codebooks are SPECIMENS for the projective reading instruments."""
405
+ import os
406
+ if device == "cuda" and not torch.cuda.is_available():
407
+ raise RuntimeError("Verdict runs are GPU-only (never CPU-train for accuracy).")
408
+ ckpt_dir = os.path.join(data_root, "ar_ckpts")
409
+ os.makedirs(ckpt_dir, exist_ok=True)
410
+ tr, va = _wikitext_bytes(data_root)
411
+ print(f"data ready: train {tr.numel():,} bytes, val {va.numel():,} bytes", flush=True)
412
+ results = {}
413
+ for arm in arms:
414
+ torch.manual_seed(seed)
415
+ g = torch.Generator().manual_seed(seed)
416
+ model = ByteLM(arm, block=block).to(device)
417
+ n_params = sum(p.numel() for p in model.parameters())
418
+ opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
419
+ for step in range(1, steps + 1):
420
+ x, y = _batch(tr, batch, block, device, g)
421
+ logits = model(x)
422
+ loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
423
+ opt.zero_grad(set_to_none=True)
424
+ loss.backward()
425
+ opt.step()
426
+ if step % eval_every == 0 or step == steps:
427
+ model.eval()
428
+ with torch.no_grad():
429
+ losses = []
430
+ for _ in range(20):
431
+ xv, yv = _batch(va, batch, block, device, g)
432
+ lv = F.cross_entropy(model(xv).reshape(-1, VOCAB),
433
+ yv.reshape(-1))
434
+ losses.append(lv.item())
435
+ bpb = sum(losses) / len(losses) / math.log(2)
436
+ print(f"[{arm}] step {step} val_bpb={bpb:.4f} vitals={model.vitals()}",
437
+ flush=True)
438
+ model.train()
439
+ results[arm] = {"val_bpb": bpb, "params": n_params, "vitals": model.vitals()}
440
+ if save:
441
+ path = os.path.join(ckpt_dir, f"{arm}_s{seed}_t{steps}.pt")
442
+ torch.save({"arm": arm, "seed": seed, "steps": steps, "val_bpb": bpb,
443
+ "state_dict": {k: v.cpu() for k, v in
444
+ model.state_dict().items()}}, path)
445
+ print(f"saved specimen: {path}", flush=True)
446
+ print(results, flush=True)
447
+ return results
448
+
449
+
450
+ def smoke():
451
+ """Shapes/parse only — no accuracy claims."""
452
+ x = torch.randint(0, VOCAB, (2, 64))
453
+ for arm in ("sdpa", "hub", "addr_head"):
454
+ m = ByteLM(arm, d=96, layers=2, block=64, K=16)
455
+ logits = m(x)
456
+ assert logits.shape == (2, 64, VOCAB)
457
+ logits.sum().backward()
458
+ # causality check: future byte must not affect past logits
459
+ with torch.no_grad():
460
+ a = m(x)[0, 10]
461
+ x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
462
+ b = m(x2)[0, 10]
463
+ assert torch.allclose(a, b, atol=1e-4), f"{arm} leaks future context"
464
+ print(f"{arm}: OK params={sum(p.numel() for p in m.parameters()):,} "
465
+ f"vitals={m.vitals()}", flush=True)
466
+ print("OK — AR bed smoke passed (verdict run: train() on GPU)", flush=True)
467
+
468
+
469
+ def _in_notebook() -> bool:
470
+ try:
471
+ get_ipython() # type: ignore[name-defined] # noqa: F821
472
+ return True
473
+ except NameError:
474
+ return False
475
+
476
+
477
+ if __name__ == "__main__":
478
+ if _in_notebook():
479
+ smoke()
480
+ print("Notebook mode: call train(steps=2000) in the next cell (GPU).")
481
+ else:
482
+ import argparse
483
+ ap = argparse.ArgumentParser()
484
+ ap.add_argument("--train", action="store_true")
485
+ ap.add_argument("--steps", type=int, default=2000)
486
+ a, _ = ap.parse_known_args()
487
+ train(steps=a.steps) if a.train else smoke()
exp018_r12/build_results.py ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """build_results.py — exp018_r12: read results/ledger.jsonl and RE-ASSERT every
2
+ claim in the README. Run from inside this folder: python build_results.py
3
+ """
4
+ import json
5
+ import os
6
+
7
+ HERE = os.path.dirname(os.path.abspath(__file__))
8
+ rows = [json.loads(l) for l in
9
+ open(os.path.join(HERE, "results", "ledger.jsonl"), encoding="utf-8")]
10
+ assert all(r["exp"] == "18" for r in rows) and len(rows) == 8
11
+ cell = {(r["arm"], r["seed"]): r for r in rows}
12
+ CERT_BAND = (2.43, 2.50) # exp012 certified: 3-seed mean 2.469, s0 bed 2.4990
13
+
14
+ # claim 1: both init avenues land inside the certified band, both seeds
15
+ for arm in ("penta_init", "farmed_init"):
16
+ for s in (0, 1):
17
+ assert CERT_BAND[0] < cell[(arm, s)]["bpb"] < CERT_BAND[1], (arm, s)
18
+
19
+ # claim 2: the farmed codebook drifts LESS than the pentachoron init (maturity
20
+ # = stationarity), both seeds — and neither accelerates
21
+ for s in (0, 1):
22
+ assert cell[("farmed_init", s)]["drift"] < cell[("penta_init", s)]["drift"]
23
+
24
+ # claim 3: the tied M_hat readout fails in AR (+1.0 bpb) and starves the
25
+ # codebook (drift ~0.02, binding 0), both seeds, with or without penta init
26
+ for arm in ("tied", "keystone"):
27
+ for s in (0, 1):
28
+ r = cell[(arm, s)]
29
+ assert r["bpb"] > 3.4 and r["drift"] < 0.03 and r["binding_frac"] == 0.0
30
+
31
+ out = {f"{a}_s{s}": {"bpb": cell[(a, s)]["bpb"], "drift": cell[(a, s)]["drift"]}
32
+ for (a, s) in sorted(cell)}
33
+ json.dump(out, open(os.path.join(HERE, "results", "results.json"), "w",
34
+ encoding="utf-8"), indent=1)
35
+ print(f"{len(rows)} rows -> results/results.json")
36
+ print("all README claims asserted OK")
exp018_r12/exp014_genetic_distillation.py ADDED
@@ -0,0 +1,515 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """exp014_genetic_distillation.py — genetic distillation + memory substrate.
2
+ 14_a: multi-generational tournament (GM3 paradigm) where the aleph codebook is the
3
+ explicit heritable genome. Lineages: ALEPH-FLAT (consensus book + KD) |
4
+ ALEPH-TREE (structured genome: root book + branch books) | MLP-KD
5
+ (traditional: best-parent weights + KD) | NO-INHERIT (evolution floor).
6
+ Both sides intentionally inherit logits (KD); only ours inherits geometry.
7
+ Consensus = Procrustes/GPA alignment of parents' books to mean shape
8
+ (placement by construction — replaces GM3's k-means-on-consensus init).
9
+ 14_b: memory substrate — the D=4 home makes books size-agnostic. Implant books
10
+ cultivated in a small organism into a larger one (frozen / trainable), and
11
+ into GPT-2 relay adapters (cross-architecture frozen distillation).
12
+
13
+ Riders: pure Adam wd=0; KD = KL to detached teacher probs (predictive pressure, no
14
+ contrastive); tree routing is DENSE SOFT (oriented weights; collapse monitor on the
15
+ root); drift-check precedes every freeze claim; GPU-only verdict runs; Colab-safe.
16
+ Founders share a COMMON-ANCESTOR book so GPA row correspondence is inherited.
17
+
18
+ Colab paste order: geolip_vitals.py -> ar_differentiation_bed.py ->
19
+ exp013_augmentation_bed.py (only for run_b2) -> this file.
20
+ """
21
+ from __future__ import annotations
22
+ import copy
23
+ import json
24
+ import math
25
+ import os
26
+ import torch
27
+ import torch.nn as nn
28
+ import torch.nn.functional as F
29
+
30
+ if "anchor_drift" not in globals():
31
+ try:
32
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
33
+ except ImportError:
34
+ _here = globals().get("__file__")
35
+ if _here is None:
36
+ raise ImportError("paste/run geolip_vitals.py first")
37
+ import sys, pathlib
38
+ sys.path.insert(0, str(pathlib.Path(_here).parent))
39
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
40
+ if "ByteLM" not in globals():
41
+ try:
42
+ from ar_differentiation_bed import (ByteLM, AlephAddress, _wikitext_bytes,
43
+ _batch, VOCAB)
44
+ except ImportError:
45
+ raise ImportError("paste/run ar_differentiation_bed.py first")
46
+
47
+ DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
48
+ EXP_DIR = os.path.join(DATA_ROOT, "exp014")
49
+
50
+
51
+ class SquaredReLU(nn.Module):
52
+ def forward(self, x):
53
+ return F.relu(x) ** 2
54
+
55
+
56
+ # ===================================================== consensus (the germline) ===
57
+ @torch.no_grad()
58
+ def procrustes_rotation(A: torch.Tensor, M: torch.Tensor) -> torch.Tensor:
59
+ """Orthogonal R minimizing ||A R - M||_F (rows correspond)."""
60
+ U, _, Vt = torch.linalg.svd(A.T.double() @ M.double())
61
+ return (U @ Vt).float()
62
+
63
+
64
+ @torch.no_grad()
65
+ def align_to(A: torch.Tensor, ref: torch.Tensor, iters: int = 20) -> torch.Tensor:
66
+ """Projective Procrustes (rows correspond, signs free): alternate the
67
+ orthogonal rotation and per-row sign flips (books live on RP^(D-1))."""
68
+ s = torch.ones(A.shape[0], 1)
69
+ for _ in range(iters):
70
+ R = procrustes_rotation(s * A, ref)
71
+ AR = (s * A) @ R
72
+ s_upd = torch.where((AR * ref).sum(-1, keepdim=True) < 0, -s, s)
73
+ if torch.equal(s_upd, s):
74
+ return AR
75
+ s = s_upd
76
+ return (s * A) @ procrustes_rotation(s * A, ref)
77
+
78
+
79
+ @torch.no_grad()
80
+ def consensus_codebook(books: list, iters: int = 50, tol: float = 1e-8):
81
+ """GPA to mean shape (GM3 machinery, applied to aleph books), anchored to the
82
+ FIRST parent's frame. Rows must correspond (common-ancestor convention); signs
83
+ are projective. Returns (consensus, n_iters, delta)."""
84
+ # device-pin to CPU: parent models may live on CUDA after KD teacher moves
85
+ Bs = [F.normalize(b.detach().float().cpu(), dim=-1).clone() for b in books]
86
+ # pairwise projective alignment to parent-0's frame, THEN GPA refinement
87
+ aligned = [Bs[0]] + [align_to(b, Bs[0]) for b in Bs[1:]]
88
+ M = F.normalize(torch.stack(aligned).mean(0), dim=-1)
89
+ delta, it = 0.0, 0
90
+ for it in range(1, iters + 1):
91
+ aligned = [align_to(b, M) for b in Bs]
92
+ M_new = F.normalize(torch.stack(aligned).mean(0), dim=-1)
93
+ delta = (M_new - M).norm().item()
94
+ M = M_new
95
+ if delta < tol:
96
+ break
97
+ # re-anchor to parent-0 (GPA drift of the global frame stays measurable)
98
+ M = align_to(M, Bs[0])
99
+ return M, it, delta
100
+
101
+
102
+ @torch.no_grad()
103
+ def implant_book(addr: "AlephAddress", book: torch.Tensor, trainable: bool = True):
104
+ """Load a book into an AlephAddress: codebook + home (drift measured from the
105
+ implant). Freeze only via trainable=False AFTER a drift-check justifies it."""
106
+ b = F.normalize(book.float(), dim=-1).to(addr.codebook.device)
107
+ assert b.shape == addr.codebook.shape, (b.shape, addr.codebook.shape)
108
+ addr.codebook.data.copy_(b)
109
+ addr.home.copy_(b)
110
+ addr.codebook.requires_grad_(trainable)
111
+
112
+
113
+ # ============================================================= tree head ==========
114
+ class TreeHead(nn.Module):
115
+ """Autoregressive tree (the constellation-anchor analogue, exp011 TREE operator
116
+ in the healthy consumption regime): a ROOT aleph (K=2, D=4) yields the 4
117
+ oriented weights (2K half-axes = the 4 branches, dense soft, sums to 1);
118
+ each BRANCH is a 64-slot... shared slot projection read by a branch-specific
119
+ book (K=64, D=4); output = branch-weighted mixture of branch reads -> vocab.
120
+ Heritable genome: root book (2,4) + 4 branch books (64,4)."""
121
+
122
+ ROOT_SLOTS = 4 # slot-parallel root consumption (the collapse cure)
123
+ ROOT_TAU = 0.3 # softer root temperature (wave-1 fix: single hard-tau
124
+ # root partially collapsed, usage [.85,.12,.01,.02])
125
+
126
+ def __init__(self, d: int, vocab: int = 256, n_slots: int = 16):
127
+ super().__init__()
128
+ self.n_slots = n_slots
129
+ self.root_proj = nn.Linear(d, self.ROOT_SLOTS * 4, bias=False)
130
+ self.slot_proj = nn.Linear(d, n_slots * 4, bias=False)
131
+ nn.init.orthogonal_(self.root_proj.weight)
132
+ nn.init.orthogonal_(self.slot_proj.weight)
133
+ self.root = AlephAddress(2, 4, tau=self.ROOT_TAU)
134
+ self.branches = nn.ModuleList([AlephAddress(64, 4) for _ in range(4)])
135
+ self.out = nn.Linear(n_slots * 4, vocab, bias=True)
136
+ self._last_root = None
137
+
138
+ def forward(self, h):
139
+ B, n, _ = h.shape
140
+ rs = self.root_proj(h).view(B, n, self.ROOT_SLOTS, 4)
141
+ p, m = self.root.oriented(rs) # (B,n,S,2) x2
142
+ w = torch.cat([p, m], dim=-1).mean(dim=-2) # slot-avg -> (B,n,4)
143
+ self._last_root = w.detach()
144
+ slots = self.slot_proj(h).view(B, n, self.n_slots, 4)
145
+ mix = 0
146
+ for b, br in enumerate(self.branches):
147
+ mix = mix + w[..., b:b + 1] * br.m_hat(slots).reshape(B, n, -1)
148
+ return self.out(mix)
149
+
150
+ def genome(self):
151
+ return {"root": self.root.codebook.detach().clone(),
152
+ **{f"branch{i}": br.codebook.detach().clone()
153
+ for i, br in enumerate(self.branches)}}
154
+
155
+ @torch.no_grad()
156
+ def inherit(self, genomes: list):
157
+ c, it, dl = consensus_codebook([g["root"] for g in genomes])
158
+ implant_book(self.root, c)
159
+ for i, br in enumerate(self.branches):
160
+ c, _, _ = consensus_codebook([g[f"branch{i}"] for g in genomes])
161
+ implant_book(br, c)
162
+
163
+ @torch.no_grad()
164
+ def vitals(self):
165
+ out = {"root_drift": round(anchor_drift(self.root.codebook,
166
+ self.root.home)["mean"], 4)}
167
+ if self._last_root is not None:
168
+ w = self._last_root.reshape(-1, 4)
169
+ usage = w.mean(0)
170
+ usage = usage / usage.sum()
171
+ out["root_usage"] = [round(float(u), 3) for u in usage]
172
+ ent = -(usage.clamp_min(1e-9) * usage.clamp_min(1e-9).log()).sum()
173
+ out["root_ppl4"] = round(float(ent.exp()), 3)
174
+ d = [anchor_drift(br.codebook, br.home)["mean"] for br in self.branches]
175
+ out["branch_drift"] = [round(x, 3) for x in d]
176
+ return out
177
+
178
+
179
+ # ============================================================ organisms ===========
180
+ def make_organism(lineage: str, d: int = 192, layers: int = 4, block: int = 256,
181
+ seed: int = 0):
182
+ """lineage in {aleph_flat, aleph_tree, mlp_kd, no_inherit}. no_inherit uses the
183
+ aleph_flat architecture (the control isolates INHERITANCE, not architecture)."""
184
+ torch.manual_seed(seed)
185
+ if lineage in ("aleph_flat", "no_inherit", "aleph_full", "aleph_weights"):
186
+ return ByteLM("addr_msl64", d=d, layers=layers, block=block)
187
+ if lineage == "aleph_tree":
188
+ m = ByteLM("sdpa", d=d, layers=layers, block=block)
189
+ m.head = TreeHead(d)
190
+ return m
191
+ if lineage == "mlp_kd":
192
+ m = ByteLM("sdpa", d=d, layers=layers, block=block)
193
+ m.head = nn.Sequential(nn.Linear(d, 224), SquaredReLU(),
194
+ nn.LayerNorm(224), nn.Linear(224, VOCAB))
195
+ return m
196
+ raise ValueError(lineage)
197
+
198
+
199
+ def genome_of(model):
200
+ """The heritable organ = co-adapted (projection, book) pair(s). Books inherit
201
+ by CONSENSUS (the geometric germline); projections inherit from the BEST
202
+ parent (weight copy) — implanting a book against a random projection puts the
203
+ child below random init (campaign-v2 lesson)."""
204
+ if isinstance(model.head, TreeHead):
205
+ g = model.head.genome()
206
+ g["root_proj"] = model.head.root_proj.weight.detach().cpu().clone()
207
+ g["slot_proj"] = model.head.slot_proj.weight.detach().cpu().clone()
208
+ return g
209
+ if hasattr(model, "head_addr"):
210
+ return {"flat": model.head_addr.codebook.detach().cpu().clone(),
211
+ "proj": model.head_proj.weight.detach().cpu().clone()}
212
+ return None
213
+
214
+
215
+ @torch.no_grad()
216
+ def inherit_genome(model, genomes: list):
217
+ """genomes[0] = the BEST parent (selection order matters)."""
218
+ if isinstance(model.head, TreeHead):
219
+ model.head.inherit(genomes)
220
+ model.head.root_proj.weight.copy_(genomes[0]["root_proj"].to(
221
+ model.head.root_proj.weight.device))
222
+ model.head.slot_proj.weight.copy_(genomes[0]["slot_proj"].to(
223
+ model.head.slot_proj.weight.device))
224
+ elif hasattr(model, "head_addr"):
225
+ c, it, dl = consensus_codebook([g["flat"] for g in genomes])
226
+ implant_book(model.head_addr, c)
227
+ model.head_proj.weight.copy_(genomes[0]["proj"].to(
228
+ model.head_proj.weight.device))
229
+
230
+
231
+ def organism_vitals(model):
232
+ if isinstance(model.head, TreeHead):
233
+ return model.head.vitals()
234
+ return model.vitals() if hasattr(model, "vitals") else {}
235
+
236
+
237
+ # ========================================================= train one member ======
238
+ def train_member(model, tr, va, steps=2000, batch=32, block=256, device="cuda",
239
+ seed=0, teachers=None, kd_alpha=1.0):
240
+ """CE (+ KL to detached mean teacher probs when teachers given). Pure Adam."""
241
+ g = torch.Generator().manual_seed(seed)
242
+ model = model.to(device)
243
+ opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
244
+ if teachers:
245
+ teachers = [t.to(device).eval() for t in teachers]
246
+ for step in range(1, steps + 1):
247
+ x, y = _batch(tr, batch, block, device, g)
248
+ logits = model(x)
249
+ loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
250
+ if teachers:
251
+ with torch.no_grad():
252
+ tp = torch.stack([F.softmax(t(x), -1) for t in teachers]).mean(0)
253
+ loss = loss + kd_alpha * F.kl_div(
254
+ F.log_softmax(logits, -1), tp, reduction="batchmean")
255
+ opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
256
+ model.eval()
257
+ with torch.no_grad():
258
+ ls = []
259
+ for _ in range(20):
260
+ xv, yv = _batch(va, batch, block, device, g)
261
+ ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
262
+ yv.reshape(-1)).item())
263
+ return sum(ls) / len(ls) / math.log(2) # bpb
264
+
265
+
266
+ # ============================================================ the tournament ======
267
+ def run_tournament(lineage: str, gens: int = 4, pop: int = 4, steps: int = 2000,
268
+ seed: int = 0, device: str = "cuda",
269
+ catastrophic_at: int | None = None):
270
+ """One lineage, one tournament seed. Logs per-gen to the ledger; saves the
271
+ champion genome per generation. catastrophic_at=G injects a 0-step random
272
+ parent into the consensus at generation G (the GM3 robustness probe)."""
273
+ if not torch.cuda.is_available():
274
+ raise RuntimeError("verdict runs are GPU-only")
275
+ os.makedirs(EXP_DIR, exist_ok=True)
276
+ tr, va = _wikitext_bytes(DATA_ROOT)
277
+ ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
278
+ # common ancestor: every founder book starts identical within a tournament
279
+ torch.manual_seed(9000 + seed)
280
+ ancestor = make_organism(lineage, seed=9000 + seed)
281
+ anc_genome = genome_of(ancestor)
282
+ parents, parent_models, champion_genomes = [], [], []
283
+ for gen in range(gens):
284
+ members = []
285
+ for i in range(pop):
286
+ mseed = seed * 1000 + gen * 100 + i
287
+ m = make_organism(lineage, seed=mseed)
288
+ if anc_genome and gen == 0:
289
+ inherit_genome(m, [anc_genome]) # common ancestor
290
+ if gen > 0:
291
+ is_fresh = (i == pop - 1) # gene flow founder
292
+ if not is_fresh:
293
+ if lineage in ("aleph_flat", "aleph_tree"):
294
+ gs = [genome_of(pm) for pm in parent_models]
295
+ if catastrophic_at == gen:
296
+ bad = make_organism(lineage, seed=666 + i)
297
+ gs = gs + [genome_of(bad)]
298
+ inherit_genome(m, gs)
299
+ elif lineage == "aleph_full":
300
+ # v4 arm (v3 lesson: continuity is what pays) — inherit the
301
+ # WHOLE best parent, then overwrite the book with the
302
+ # two-parent consensus: germline ON TOP of continuity.
303
+ m.load_state_dict(copy.deepcopy(
304
+ parent_models[0].state_dict()))
305
+ gs = [genome_of(pm) for pm in parent_models]
306
+ if catastrophic_at == gen:
307
+ bad = make_organism(lineage, seed=666 + i)
308
+ gs = gs + [genome_of(bad)]
309
+ c, _, _ = consensus_codebook([g["flat"] for g in gs])
310
+ implant_book(m.head_addr, c)
311
+ elif lineage in ("mlp_kd", "aleph_weights"):
312
+ # pure continuity (no germline op) — aleph_weights is the
313
+ # within-architecture control for aleph_full
314
+ m.load_state_dict(copy.deepcopy(
315
+ parent_models[0].state_dict()))
316
+ # no_inherit: nothing
317
+ # KD: alpha 0.25 (campaign-v2 lesson: alpha=1.0 from near-parity
318
+ # teachers COMPOUNDS DOWNWARD — inverse evolution; fresh-founder
319
+ # control isolated it). Fresh founders get NO KD (clean gene flow).
320
+ is_fresh_now = (gen > 0 and i == pop - 1)
321
+ teachers = parent_models if (gen > 0 and not is_fresh_now
322
+ and lineage != "no_inherit") else None
323
+ bpb = train_member(m, tr, va, steps=steps, device=device,
324
+ seed=mseed, teachers=teachers, kd_alpha=0.25)
325
+ vit = organism_vitals(m)
326
+ members.append((bpb, m))
327
+ rec = {"exp": "14a", "lineage": lineage, "tseed": seed, "gen": gen,
328
+ "member": i, "fresh": gen > 0 and i == pop - 1,
329
+ "catastrophic": catastrophic_at == gen, "bpb": round(bpb, 4),
330
+ "vitals": vit}
331
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
332
+ print(f"[14a {lineage} t{seed} g{gen} m{i}] bpb={bpb:.4f} {vit}",
333
+ flush=True)
334
+ members.sort(key=lambda t: t[0])
335
+ parent_models = [members[0][1].cpu(), members[1][1].cpu()]
336
+ best = members[0][0]
337
+ gene = genome_of(members[0][1])
338
+ if gene:
339
+ torch.save(gene, os.path.join(
340
+ EXP_DIR, f"champion_{lineage}_t{seed}_g{gen}.pt"))
341
+ champion_genomes.append(gene)
342
+ print(f"[14a {lineage} t{seed} g{gen}] BEST={best:.4f} "
343
+ f"mean={sum(b for b, _ in members)/pop:.4f}", flush=True)
344
+ for _, mm in members[2:]:
345
+ del mm
346
+ torch.cuda.empty_cache()
347
+ ledger.close()
348
+ return best
349
+
350
+
351
+ # ============================================================ 14_b implants ======
352
+ def run_b1(steps: int = 2000, seed: int = 0, device: str = "cuda",
353
+ donor_book: torch.Tensor | None = None, tag: str = "small_cultivated"):
354
+ """Cross-size: donor book (default: cultivate in a small organism) implanted
355
+ into a LARGE organism. Arms: fresh | implant-trainable | implant-frozen | mlp."""
356
+ if not torch.cuda.is_available():
357
+ raise RuntimeError("verdict runs are GPU-only")
358
+ os.makedirs(EXP_DIR, exist_ok=True)
359
+ tr, va = _wikitext_bytes(DATA_ROOT)
360
+ ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
361
+ if donor_book is None:
362
+ small = make_organism("aleph_flat", d=128, layers=2, seed=seed)
363
+ bpb_small = train_member(small, tr, va, steps=steps, device=device, seed=seed)
364
+ donor_book = genome_of(small)["flat"]
365
+ print(f"[14b donor small] bpb={bpb_small:.4f}", flush=True)
366
+ results = {}
367
+ for arm in ("fresh", "implant_train", "implant_frozen", "mlp"):
368
+ lineage = "mlp_kd" if arm == "mlp" else "aleph_flat"
369
+ m = make_organism(lineage, d=384, layers=6, seed=seed + 10)
370
+ if arm.startswith("implant"):
371
+ implant_book(m.head_addr, donor_book, trainable=(arm == "implant_train"))
372
+ bpb = train_member(m, tr, va, steps=steps, device=device, seed=seed + 10)
373
+ vit = organism_vitals(m)
374
+ results[arm] = {"bpb": round(bpb, 4), "vitals": vit}
375
+ rec = {"exp": "14b1", "arm": arm, "donor": tag, "seed": seed,
376
+ "bpb": round(bpb, 4), "vitals": vit}
377
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
378
+ print(f"[14b1 {arm} donor={tag}] bpb={bpb:.4f} {vit}", flush=True)
379
+ del m; torch.cuda.empty_cache()
380
+ ledger.close()
381
+ return results, donor_book
382
+
383
+
384
+ def run_b2(donor_book: torch.Tensor, steps: int = 1500, seed: int = 0,
385
+ device: str = "cuda", tag: str = "small_cultivated"):
386
+ """Cross-architecture: implant the donor book into every GPT-2 relay adapter
387
+ (exp013 Track C bed) vs random-init relays. Books are (64,4) — size-agnostic."""
388
+ from exp013_augmentation_bed import _wikitext_lines
389
+ from transformers import GPT2LMHeadModel, GPT2TokenizerFast
390
+ if "MslRelay" not in globals():
391
+ from ar_differentiation_bed import MslRelay
392
+ from exp013_augmentation_bed import _BlockWithAdapter
393
+ tok = GPT2TokenizerFast.from_pretrained("gpt2")
394
+ tr_lines, va_lines = _wikitext_lines(DATA_ROOT)
395
+ stream_tr = tok("\n\n".join(tr_lines[:8000]), return_tensors="pt").input_ids[0]
396
+ stream_va = tok("\n\n".join(va_lines[:1000]), return_tensors="pt").input_ids[0]
397
+ ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
398
+ out = {}
399
+ for arm in ("random_relays", "implanted_relays"):
400
+ torch.manual_seed(seed)
401
+ g = torch.Generator().manual_seed(seed)
402
+ model = GPT2LMHeadModel.from_pretrained("gpt2").to(device)
403
+ for p in model.parameters():
404
+ p.requires_grad_(False)
405
+ adapters = []
406
+ for i, blk in enumerate(model.transformer.h):
407
+ ad = MslRelay(model.config.n_embd).to(device)
408
+ if arm == "implanted_relays":
409
+ implant_book(ad.addr, donor_book, trainable=True)
410
+ model.transformer.h[i] = _BlockWithAdapter(blk, ad)
411
+ adapters.append(ad)
412
+ params = [p for ad in adapters for p in ad.parameters()
413
+ if p.requires_grad]
414
+ opt = torch.optim.Adam(params, lr=1e-3, weight_decay=0.0)
415
+ block = 256
416
+ for step in range(1, steps + 1):
417
+ ix = torch.randint(0, stream_tr.numel() - block - 1, (8,), generator=g)
418
+ x = torch.stack([stream_tr[i:i + block] for i in ix]).to(device)
419
+ loss = model(x, labels=x).loss
420
+ opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
421
+ model.eval()
422
+ with torch.no_grad():
423
+ ls = []
424
+ for j in range(0, stream_va.numel() - block - 1, block * 4):
425
+ x = stream_va[j:j + block].unsqueeze(0).to(device)
426
+ ls.append(model(x, labels=x).loss.item())
427
+ ppl = math.exp(sum(ls) / len(ls))
428
+ gates = [round(ad.gate.sigmoid().item(), 4) for ad in adapters]
429
+ drifts = [round(anchor_drift(ad.addr.codebook, ad.addr.home)["mean"], 3)
430
+ for ad in adapters]
431
+ out[arm] = {"ppl": round(ppl, 3), "gates": gates, "drift": drifts}
432
+ rec = {"exp": "14b2", "arm": arm, "donor": tag, "seed": seed,
433
+ "ppl": round(ppl, 3), "gates": gates, "drift": drifts}
434
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
435
+ print(f"[14b2 {arm} donor={tag}] ppl={ppl:.3f} gates={gates[:3]}.. "
436
+ f"drift={drifts[:3]}..", flush=True)
437
+ del model; torch.cuda.empty_cache()
438
+ ledger.close()
439
+ return out
440
+
441
+
442
+ # ================================================================ smoke ===========
443
+ def smoke():
444
+ """CPU shapes/parse only: GPA ground truth, tree causality, implant, KD path."""
445
+ g = torch.Generator().manual_seed(0)
446
+ # GPA: two rotated (+row-sign-flipped) copies of one book must align back to it
447
+ A = F.normalize(torch.randn(64, 4, generator=g), dim=-1)
448
+ q, _ = torch.linalg.qr(torch.randn(4, 4, generator=g))
449
+ B = A @ q
450
+ B[::3] = -B[::3]
451
+ C, it, dl = consensus_codebook([A, B])
452
+ cos = (F.normalize(C, dim=-1) * A).sum(-1).abs().mean()
453
+ assert cos > 0.999, cos
454
+ print(f"GPA OK (iters={it}, delta={dl:.2e}, |cos to truth|={cos:.5f})")
455
+ x = torch.randint(0, 256, (2, 64))
456
+ for lineage in ("aleph_flat", "aleph_tree", "mlp_kd", "no_inherit"):
457
+ m = make_organism(lineage, d=96, layers=2, block=64, seed=0)
458
+ lg = m(x); assert lg.shape == (2, 64, 256); lg.sum().backward()
459
+ with torch.no_grad():
460
+ a = m(x)[0, 10]; x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
461
+ b = m(x2)[0, 10]
462
+ assert torch.allclose(a, b, atol=1e-4), lineage + " leaks"
463
+ gnm = genome_of(m)
464
+ if gnm:
465
+ inherit_genome(m, [gnm, gnm]) # self-consensus = identity-ish
466
+ print(lineage, "OK params",
467
+ f"{sum(p.numel() for p in m.parameters()):,}",
468
+ organism_vitals(m) if lineage != "mlp_kd" else {})
469
+ # KD path: teacher forward + KL backward
470
+ t = make_organism("mlp_kd", d=96, layers=2, block=64, seed=1)
471
+ s = make_organism("aleph_flat", d=96, layers=2, block=64, seed=2)
472
+ tp = F.softmax(t(x), -1).detach()
473
+ loss = F.kl_div(F.log_softmax(s(x), -1), tp, reduction="batchmean")
474
+ loss.backward()
475
+ print("KD OK — exp014 smoke passed (tournament on GPU: run_tournament(...))")
476
+
477
+
478
+ def _in_notebook():
479
+ try:
480
+ get_ipython() # type: ignore[name-defined] # noqa: F821
481
+ return True
482
+ except NameError:
483
+ return False
484
+
485
+
486
+ if __name__ == "__main__":
487
+ if _in_notebook():
488
+ smoke()
489
+ print("Notebook: run_tournament('aleph_flat'), run_b1(), run_b2(book).")
490
+ else:
491
+ import argparse
492
+ ap = argparse.ArgumentParser()
493
+ ap.add_argument("--mode", default="smoke",
494
+ choices=["smoke", "tournament", "b1", "b2"])
495
+ ap.add_argument("--lineage", default="aleph_full",
496
+ help="tournament lineage: aleph_flat|aleph_full|"
497
+ "aleph_weights|aleph_tree|mlp_kd|no_inherit")
498
+ ap.add_argument("--seed", type=int, default=0)
499
+ ap.add_argument("--steps", type=int, default=2000)
500
+ ap.add_argument("--genome", default="genomes/champion_aleph_full_t0_g3.pt",
501
+ help="donor genome .pt for --mode b1/b2 (uses its 'flat' book)")
502
+ a, _ = ap.parse_known_args()
503
+ if a.mode == "smoke":
504
+ smoke()
505
+ elif a.mode == "tournament":
506
+ run_tournament(a.lineage, steps=a.steps, seed=a.seed)
507
+ elif a.mode == "b1":
508
+ donor = (torch.load(a.genome, map_location="cpu")["flat"]
509
+ if os.path.exists(a.genome) else None)
510
+ run_b1(steps=a.steps, seed=a.seed, donor_book=donor,
511
+ tag=os.path.basename(a.genome) if donor is not None
512
+ else "small_cultivated")
513
+ elif a.mode == "b2":
514
+ donor = torch.load(a.genome, map_location="cpu")["flat"]
515
+ run_b2(donor, seed=a.seed, tag=os.path.basename(a.genome))
exp018_r12/exp018_reran12.py ADDED
@@ -0,0 +1,176 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """exp018_reran12.py — EXPERIMENT 18: exp012 RERUN with the keystone aleph model
2
+ parameters (exp012 and its certified record stay untouched — this is a labeled
3
+ rerun, not a replacement). A process review found the exp012 bed's AlephAddress
4
+ primitive faithful to the keystone (per the aleph-void article) but two keystone avenues
5
+ never exercised: the caller-supplied FARMED/PENTACHORON codebook init and the
6
+ TIED M_hat readout (U=M_hat, S=Omega-token, Vt=I). This rerun exercises them on the
7
+ certified bed, unchanged otherwise (d=192, 4 layers, wikitext-2 bytes, 2000
8
+ steps, pure Adam wd=0, addr_msl64 multi-slot base).
9
+
10
+ Arms (x2 seeds):
11
+ penta_init — codebook initialized from geovocab2 REGULAR PENTACHORON
12
+ vertices (13 randomly-rotated regular 4-simplices -> 64 rows on
13
+ S^3, row-normalized; SimplexFactory is the source per keystone
14
+ "geovocab pentachoron vertices"). Free linear head (isolates
15
+ the init).
16
+ farmed_init — codebook initialized from a FARMED recon-real book: the exp012
17
+ certified specimen addr_msl64_s0_t2000's trained codebook —
18
+ anchors the M_hat gradient itself produced. Free linear head.
19
+ (The direct "farm FROM the aleph logit structure" arm.)
20
+ tied — random init + TIED readout: zero-parameter head; logits =
21
+ M_hat_slots @ head_proj.weight @ emb.weight^T (decode tied back
22
+ through the encoder projection and the byte embedding — the
23
+ AR form of the keystone's tied linear off M_hat).
24
+ keystone — penta_init + tied (the keystone-faithful configuration).
25
+ Baselines (exp012 certified, cited not rerun): addr_msl64 2.469 / sdpa 2.518
26
+ (3-seed means @2k); certified bed reproduction 2.4990 (s0).
27
+ Riders: pure Adam wd=0; GPU-only verdicts; >=2 seeds; drift-check before any
28
+ freeze claim; Colab-safe. Paste order: geolip_vitals ->
29
+ ar_differentiation_bed -> exp014_genetic_distillation -> this file.
30
+ """
31
+ from __future__ import annotations
32
+ import json
33
+ import os
34
+ import torch
35
+ import torch.nn as nn
36
+ import torch.nn.functional as F
37
+
38
+ if "ByteLM" not in globals():
39
+ try:
40
+ from ar_differentiation_bed import ByteLM, _wikitext_bytes
41
+ from exp014_genetic_distillation import train_member, implant_book
42
+ from geolip_vitals import anchor_drift
43
+ except ImportError:
44
+ _here = globals().get("__file__")
45
+ if _here is None:
46
+ raise ImportError("paste geolip_vitals + ar_differentiation_bed + "
47
+ "exp014_genetic_distillation first")
48
+ import sys, pathlib
49
+ sys.path.insert(0, str(pathlib.Path(_here).parent))
50
+ from ar_differentiation_bed import ByteLM, _wikitext_bytes
51
+ from exp014_genetic_distillation import train_member, implant_book
52
+ from geolip_vitals import anchor_drift
53
+
54
+ DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
55
+ RERUN_DIR = os.path.join(DATA_ROOT, "exp018")
56
+ FARMED_SPECIMEN = os.environ.get(
57
+ "GEOLIP_FARMED_SPECIMEN",
58
+ os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "exp012_ar",
59
+ "specimens", "addr_msl64_s0_t2000.pt"))
60
+
61
+
62
+ def pentachoron_book(K: int = 64, seed: int = 0) -> torch.Tensor:
63
+ """K rows on S^3 from geovocab2 regular pentachoron vertices: ceil(K/5)
64
+ regular 4-simplices, each rotated by a seeded random SO(4), stacked and
65
+ row-normalized. The keystone's 'geovocab pentachoron vertices' init."""
66
+ from geovocab2.shapes.factory.simplex_factory import SimplexFactory
67
+ g = torch.Generator().manual_seed(seed)
68
+ # the factory's regular construction lives in k+1 = 5 dims; center it and
69
+ # project isometrically onto its rank-4 subspace (SVD) -> exact regular
70
+ # pentachoron in R^4 (pairwise cos = -1/4 on the sphere)
71
+ fac = SimplexFactory(k=4, embed_dim=5, method="regular")
72
+ base5 = fac.build_torch().double() # (5, 5)
73
+ base5 = base5 - base5.mean(dim=0, keepdim=True)
74
+ _, _, Vt = torch.linalg.svd(base5, full_matrices=False)
75
+ base = (base5 @ Vt[:4].T).float() # (5, 4) regular
76
+ rows = []
77
+ for _ in range((K + 4) // 5):
78
+ q, _ = torch.linalg.qr(torch.randn(4, 4, generator=g, dtype=torch.float64))
79
+ rows.append((base.double() @ q).float())
80
+ book = torch.cat(rows)[:K]
81
+ return F.normalize(book, dim=-1)
82
+
83
+
84
+ def farmed_book() -> torch.Tensor:
85
+ """The exp012 certified specimen's trained codebook — anchors farmed by the
86
+ M_hat recon/CE gradient itself (recon-real by construction)."""
87
+ ck = torch.load(FARMED_SPECIMEN, map_location="cpu", weights_only=True)
88
+ return F.normalize(ck["state_dict"]["head_addr.codebook"].float(), dim=-1)
89
+
90
+
91
+ class TiedReadout(nn.Module):
92
+ """Zero-parameter tied decode: M_hat slots -> back through the encoder
93
+ projection -> byte-embedding transpose. Gradients flow into head_proj and
94
+ emb through both encode and decode paths (the tie)."""
95
+
96
+ def __init__(self, head_proj: nn.Linear, emb: nn.Embedding):
97
+ super().__init__()
98
+ self._proj = [head_proj] # references, not submodules
99
+ self._emb = [emb]
100
+
101
+ def forward(self, feats): # feats (..., P*4) = M_hat per slot
102
+ h = feats @ self._proj[0].weight # (P*4, d) -> (..., d)
103
+ return h @ self._emb[0].weight.T # (..., VOCAB)
104
+
105
+
106
+ def make_rerun_model(arm: str, seed: int = 0, d: int = 192, layers: int = 4,
107
+ block: int = 256):
108
+ torch.manual_seed(seed)
109
+ m = ByteLM("addr_msl64", d=d, layers=layers, block=block)
110
+ if arm in ("penta_init", "keystone"):
111
+ implant_book(m.head_addr, pentachoron_book(64, seed=seed))
112
+ elif arm == "farmed_init":
113
+ implant_book(m.head_addr, farmed_book())
114
+ if arm in ("tied", "keystone"):
115
+ m.head = TiedReadout(m.head_proj, m.emb)
116
+ return m
117
+
118
+
119
+ def run_rerun(arms=("penta_init", "farmed_init", "tied", "keystone"),
120
+ seeds=(0, 1), steps=2000, device="cuda"):
121
+ if not torch.cuda.is_available():
122
+ raise RuntimeError("verdict runs are GPU-only")
123
+ os.makedirs(RERUN_DIR, exist_ok=True)
124
+ tr, va = _wikitext_bytes(DATA_ROOT)
125
+ ledger = open(os.path.join(RERUN_DIR, "ledger.jsonl"), "a", encoding="utf-8")
126
+ for seed in seeds:
127
+ for arm in arms:
128
+ m = make_rerun_model(arm, seed=seed)
129
+ bpb = train_member(m, tr, va, steps=steps, device=device, seed=seed)
130
+ d = anchor_drift(m.head_addr.codebook, m.head_addr.home)
131
+ rec = {"exp": "18", "arm": arm, "seed": seed, "steps": steps,
132
+ "bpb": round(bpb, 4), "drift": round(d["mean"], 4),
133
+ "binding_frac": round(d["binding_fraction"], 4)}
134
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
135
+ print(f"[18r {arm} s{seed}] FINAL bpb={bpb:.4f} "
136
+ f"drift={rec['drift']} bind={rec['binding_frac']}", flush=True)
137
+ torch.save({"arm": arm, "seed": seed, "steps": steps,
138
+ "state_dict": {k: v.cpu() for k, v in
139
+ m.state_dict().items()}},
140
+ os.path.join(RERUN_DIR, f"r12_{arm}_s{seed}.pt"))
141
+ del m
142
+ torch.cuda.empty_cache()
143
+ ledger.close()
144
+
145
+
146
+ def smoke():
147
+ b = pentachoron_book(64, seed=0)
148
+ assert b.shape == (64, 4)
149
+ assert torch.allclose(b.norm(dim=-1), torch.ones(64), atol=1e-5)
150
+ # regular-simplex signature: within one pentachoron, pairwise cos ~ -1/4
151
+ c = (b[:5] @ b[:5].T)
152
+ off = c[~torch.eye(5, dtype=torch.bool)]
153
+ assert (off - (-0.25)).abs().max() < 1e-4, off
154
+ x = torch.randint(0, 256, (2, 64))
155
+ for arm in ("penta_init", "tied", "keystone"):
156
+ m = make_rerun_model(arm, seed=0, d=96, layers=2, block=64)
157
+ lg = m(x); assert lg.shape == (2, 64, 256); lg.sum().backward()
158
+ assert m.head_addr.codebook.grad is not None
159
+ if arm in ("tied", "keystone"):
160
+ assert sum(p.numel() for p in m.head.parameters()) == 0
161
+ assert m.emb.weight.grad is not None
162
+ m.zero_grad()
163
+ print("exp018 (reran 12) smoke passed (farmed_init needs the specimen on disk)")
164
+
165
+
166
+ def _in_notebook():
167
+ try:
168
+ get_ipython() # type: ignore[name-defined] # noqa: F821
169
+ return True
170
+ except NameError:
171
+ return False
172
+
173
+
174
+ if __name__ == "__main__":
175
+ smoke() if not _in_notebook() else (smoke(),
176
+ print("Notebook: run_rerun() on GPU."))
exp018_r12/geolip_vitals.py ADDED
@@ -0,0 +1,219 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """geolip_vitals.py — shared diagnostic harness for the GeoLIP aleph experiments.
2
+ ALL functions are READOUTS: no gradients, no losses. CV is a readout, never a
3
+ force. Addressing is judged by drift->0.29154 and CV->0.20, never by recon cosine
4
+ (judgment criteria per the aleph-void article: https://huggingface.co/blog/AbstractPhil/geometric-vocabulary-patchwork-aleph-void).
5
+
6
+ Vitals provided:
7
+ anchor_drift — geodesic drift of anchors from init; binding fraction @0.29154
8
+ pentachoron_cv — CM 4-volume CV over random 5-row subsets (geovocab2 import)
9
+ axis_aliveness — oriented-address usage: axes alive, hppl, collapse flag
10
+ gate_stats — gate means vs the 0.012-0.03 band
11
+ path_diversity — unique-path counting, FIXED high-bits hash (low-16 bug is the
12
+ retracted artifact — never use the low bits)
13
+ grad_norm_spread — gradient democracy monitor (orders-of-magnitude spread)
14
+ CVScreen — CV@1000-batch early band screen (<0.30 LOW / .35-.50 MID / >.80 HIGH)
15
+
16
+ Smoke on a torch-capable env: python geolip_vitals.py
17
+ """
18
+ from __future__ import annotations
19
+ import math
20
+ import torch
21
+
22
+ BINDING = 0.29154 # radians; the binding/separation constant
23
+ CV_BAND = (0.13, 0.30) # CM CV band (discovery_catalog #4)
24
+ GATE_BAND = (0.012, 0.03) # live invariant candidate (acd_campaign)
25
+ KNUTH32 = 2654435761
26
+
27
+
28
+ # ----------------------------------------------------------------------------- drift
29
+ @torch.no_grad()
30
+ def anchor_drift(current: torch.Tensor, init: torch.Tensor, tol: float = 0.05) -> dict:
31
+ """Geodesic drift (radians) of each row of `current` from its row in `init`,
32
+ both row-normalized. Returns mean/std/per-row drift and the fraction of rows
33
+ within +/-tol of BINDING (the GLFM '46%' readout)."""
34
+ a = torch.nn.functional.normalize(current.float(), dim=-1)
35
+ b = torch.nn.functional.normalize(init.float(), dim=-1)
36
+ cos = (a * b).sum(-1).clamp(-1.0, 1.0)
37
+ drift = torch.arccos(cos)
38
+ frac = ((drift - BINDING).abs() <= tol).float().mean()
39
+ return {"mean": drift.mean().item(), "std": drift.std().item(),
40
+ "per_row": drift, "binding_fraction": frac.item()}
41
+
42
+
43
+ # -------------------------------------------------------------------------------- cv
44
+ @torch.no_grad()
45
+ def _pentachoron_volumes(pts: torch.Tensor) -> torch.Tensor:
46
+ """Batched Cayley-Menger 4-simplex volumes. pts: (B, 5, D) -> (B,) volumes.
47
+ One float64 det over all samples (vol^2 = -det(CM)/9216 for n=4). Built-in
48
+ for speed (the per-sample reference path is ~260x slower in a vitals loop);
49
+ geovocab2 remains the formula's reference implementation, parity-checked
50
+ via cv_reference_check()."""
51
+ B = pts.shape[0]
52
+ d2 = torch.cdist(pts.double(), pts.double()).pow(2) # (B,5,5)
53
+ cm = torch.ones(B, 6, 6, dtype=torch.float64, device=pts.device)
54
+ cm[:, 0, 0] = 0.0
55
+ cm[:, 1:, 1:] = d2
56
+ det = torch.linalg.det(cm)
57
+ return (-det / 9216.0).clamp_min(0.0).sqrt().float()
58
+
59
+
60
+ @torch.no_grad()
61
+ def pentachoron_cv(rows: torch.Tensor, n_samples: int = 200,
62
+ generator: torch.Generator | None = None) -> float:
63
+ """CV (std/mean) of Cayley-Menger 4-simplex volumes over n_samples random
64
+ 5-row subsets. Rows are row-normalized before measurement. Uses the built-in
65
+ batched CM (float64 det); validate against geovocab2 with
66
+ cv_reference_check() after any change to the volume math."""
67
+ x = torch.nn.functional.normalize(rows.float(), dim=-1)
68
+ n = x.shape[0]
69
+ if n < 5:
70
+ raise ValueError(f"pentachoron_cv needs >=5 rows, got {n}")
71
+ g = generator or torch.Generator(device="cpu").manual_seed(0)
72
+ idx = torch.stack([torch.randperm(n, generator=g)[:5]
73
+ for _ in range(n_samples)]) # (B,5)
74
+ v = _pentachoron_volumes(x[idx].cpu())
75
+ return (v.std() / v.mean().clamp_min(1e-12)).item()
76
+
77
+
78
+ @torch.no_grad()
79
+ def cv_reference_check(n_trials: int = 50, tol: float = 1e-5) -> float:
80
+ """Parity check of the built-in batched CM against geovocab2's reference
81
+ implementation (the formula's source of truth). Returns max |rel diff|;
82
+ raises if geovocab2 is absent or parity fails. Run after touching
83
+ _pentachoron_volumes."""
84
+ try:
85
+ from geovocab2.shapes.formula.symbolic.cayley_menger import (
86
+ CayleyMengerFromSimplex)
87
+ except Exception as e: # pragma: no cover
88
+ raise ImportError(
89
+ "cv_reference_check requires geovocab2 (install via the geolip-svae "
90
+ "umbrella: pip install git+https://github.com/AbstractEyes/"
91
+ "geolip-svae).") from e
92
+ ref = CayleyMengerFromSimplex()
93
+ g = torch.Generator().manual_seed(0)
94
+ pts = torch.nn.functional.normalize(
95
+ torch.randn(n_trials, 5, 4, generator=g), dim=-1)
96
+ mine = _pentachoron_volumes(pts)
97
+ # compare at float64: the reference computes in the INPUT dtype, and fp32
98
+ # dets lose up to ~4% on near-degenerate pentachora (measured 2026-07-11)
99
+ theirs = torch.stack([ref.forward(p.double())["volume"].float() for p in pts])
100
+ rel = ((mine - theirs).abs() / theirs.abs().clamp_min(1e-12)).max().item()
101
+ if rel > tol:
102
+ raise AssertionError(f"CM parity vs geovocab2 failed: max rel {rel}")
103
+ return rel
104
+
105
+
106
+ # ------------------------------------------------------------------------- aliveness
107
+ @torch.no_grad()
108
+ def axis_aliveness(oriented_weights: torch.Tensor, alive_thresh: float = 1e-3) -> dict:
109
+ """`oriented_weights`: (..., 2K) nonnegative oriented-softmax address rows
110
+ (sum to 1 on the last dim). Returns axes-alive count, mean-usage perplexity
111
+ (hppl analogue; healthy hosted reference 125-126/128), and a collapse flag.
112
+ Reference behavior: near-uniform aliveness at div_weight=0 (discovery #22)."""
113
+ w = oriented_weights.reshape(-1, oriented_weights.shape[-1]).float()
114
+ usage = w.mean(0)
115
+ usage = usage / usage.sum().clamp_min(1e-12)
116
+ # an axis is alive if its mean usage exceeds alive_thresh x the uniform share
117
+ alive = int((usage > alive_thresh * (1.0 / usage.numel())).sum())
118
+ ent = -(usage.clamp_min(1e-12) * usage.clamp_min(1e-12).log()).sum()
119
+ ppl = float(ent.exp())
120
+ return {"axes_total": usage.numel(), "axes_alive": alive, "usage_ppl": ppl,
121
+ "collapsed": ppl < 0.05 * usage.numel()}
122
+
123
+
124
+ # ------------------------------------------------------------------------------ gates
125
+ @torch.no_grad()
126
+ def gate_stats(gates: torch.Tensor) -> dict:
127
+ """Gate values (post-sigmoid/clamp). Reports mean and whether it sits in the
128
+ 0.012-0.03 band (read-only — the band is a candidate invariant, never a target)."""
129
+ g = gates.float().flatten()
130
+ m = g.mean().item()
131
+ return {"mean": m, "std": g.std().item(),
132
+ "in_band": GATE_BAND[0] <= m <= GATE_BAND[1]}
133
+
134
+
135
+ # ------------------------------------------------------------------------------ paths
136
+ @torch.no_grad()
137
+ def path_diversity(ids: torch.Tensor) -> dict:
138
+ """Unique-path counting with the FIXED multiplicative hash:
139
+ ((ids * 2654435761) % 2^32) >> 16 — Knuth needs the HIGH bits; the low-16
140
+ variant produced a retracted ~1,500 path ceiling in a prior campaign.
141
+ `ids`: integer tensor, one composed path id per row (any shape)."""
142
+ x = ids.reshape(-1).to(torch.int64)
143
+ hashed = ((x * KNUTH32) % (1 << 32)) >> 16
144
+ return {"n": int(x.numel()),
145
+ "unique_raw": int(torch.unique(x).numel()),
146
+ "unique_hashed": int(torch.unique(hashed).numel())}
147
+
148
+
149
+ @torch.no_grad()
150
+ def compose_path_ids(stage_indices: list[torch.Tensor], radix: int) -> torch.Tensor:
151
+ """Compose per-stage discrete indices (each (...,) int in [0, radix)) into a
152
+ single path id, positional base-`radix` — construction, not hashing."""
153
+ out = torch.zeros_like(stage_indices[0], dtype=torch.int64)
154
+ for s in stage_indices:
155
+ out = out * radix + s.to(torch.int64)
156
+ return out
157
+
158
+
159
+ # --------------------------------------------------------------------- grad democracy
160
+ @torch.no_grad()
161
+ def grad_norm_spread(groups: dict[str, list[torch.nn.Parameter]]) -> dict:
162
+ """Gradient-democracy monitor. `groups`: name -> params of one parallel member
163
+ (tower/expert). Reports per-group grad norms and the orders-of-magnitude spread.
164
+ Reference: unequalized heterogeneous towers spread ~20 orders (fibonacci dead at
165
+ 2.25e-21 under helix); equalized ~0.0 (geofractal gradient-democracy result)."""
166
+ norms = {}
167
+ for name, params in groups.items():
168
+ gs = [p.grad for p in params if p.grad is not None]
169
+ norms[name] = float(torch.sqrt(sum((g.float() ** 2).sum() for g in gs)).item()) \
170
+ if gs else 0.0
171
+ vals = [v for v in norms.values() if v > 0]
172
+ spread = (math.log10(max(vals)) - math.log10(min(vals))) if len(vals) >= 2 else 0.0
173
+ return {"norms": norms, "spread_orders": spread, "dead": [k for k, v in norms.items() if v == 0.0]}
174
+
175
+
176
+ # ----------------------------------------------------------------------------- screen
177
+ class CVScreen:
178
+ """CV@N early band screen (tri-band ft1): record pentachoron CV at `step_mark`
179
+ batches; classify <0.30 LOW / 0.35-0.50 MID / >0.80 HIGH. Turns ~2h/config
180
+ into ~7min. Readout only."""
181
+ def __init__(self, step_mark: int = 1000):
182
+ self.step_mark = step_mark
183
+ self.recorded: float | None = None
184
+
185
+ def maybe_record(self, step: int, rows: torch.Tensor) -> float | None:
186
+ if self.recorded is None and step >= self.step_mark:
187
+ self.recorded = pentachoron_cv(rows)
188
+ return self.recorded
189
+
190
+ @property
191
+ def band(self) -> str | None:
192
+ c = self.recorded
193
+ if c is None:
194
+ return None
195
+ if c < 0.30:
196
+ return "LOW"
197
+ if 0.35 <= c <= 0.50:
198
+ return "MID"
199
+ if c > 0.80:
200
+ return "HIGH"
201
+ return "BETWEEN"
202
+
203
+
204
+ # ------------------------------------------------------------------------------ smoke
205
+ if __name__ == "__main__": # shapes/parse smoke ONLY — no training, ever.
206
+ g = torch.Generator().manual_seed(0)
207
+ K, D = 64, 4
208
+ init = torch.nn.functional.normalize(torch.randn(K, D, generator=g), dim=-1)
209
+ cur = torch.nn.functional.normalize(init + 0.29 * torch.randn(K, D, generator=g), dim=-1)
210
+ print("drift:", {k: v for k, v in anchor_drift(cur, init).items() if k != "per_row"})
211
+ w = torch.softmax(torch.randn(32, 2 * K, generator=g), dim=-1)
212
+ print("aliveness:", axis_aliveness(w))
213
+ print("gates:", gate_stats(torch.full((8,), 0.024)))
214
+ ids = compose_path_ids([torch.randint(0, 16, (4096,), generator=g) for _ in range(4)], 16)
215
+ print("paths:", path_diversity(ids))
216
+ lin = torch.nn.Linear(8, 8)
217
+ lin(torch.randn(4, 8)).sum().backward()
218
+ print("democracy:", grad_norm_spread({"a": list(lin.parameters())}))
219
+ print("OK — vitals smoke passed (pentachoron_cv needs geovocab2; run on GPU env)")
exp018_r12/read_codebook.py ADDED
@@ -0,0 +1,159 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """read_codebook.py — projective reading of cultivated aleph codebooks.
2
+ Antipodal-collapse extraction on trained codebooks + projective statistics on
3
+ RP^(D-1) — applied to the exp012 AR-bed specimens.
4
+
5
+ Recipe per the Polygonal Omega article (geometric-tri-band-ft2): collapse = (row_i - row_j)/2 normalized for each MUTUAL-STRONGEST
6
+ pair with cos < -0.9 — "a deterministic tensor operation," not clustering.
7
+ Projective metric ALWAYS arccos|<a,b>| (metric-alignment rule, reading-voids-ft1).
8
+ D=4 scope is the validated regime (D=5 walked back; axis count grows with D).
9
+
10
+ Readouts per specimen:
11
+ pairs / n_axes / unpaired — antipodal structure
12
+ proj_angle mean vs uniform baseline, deviation — near-uniform RP^(D-1)?
13
+ drift from home + binding fraction @0.29154 — cultivation record
14
+ erank of the axis set — spectral occupancy
15
+ verdict: PROJECTIVE-CLEAN (|dev|<0.05, util>0.95, secondary pairs<=3) /
16
+ -MOSTLY / STRUCTURED / DEGENERATE (per Polygonal Omega thresholds)
17
+
18
+ Usage (terminal): python read_codebook.py <ckpt_or_dir> [more paths...]
19
+ Colab: paste geolip_vitals.py cell first (optional), then this file, then
20
+ read_all(r"/content/data/ar_ckpts").
21
+ """
22
+ from __future__ import annotations
23
+ import math
24
+ import sys
25
+ import torch
26
+ import torch.nn.functional as F
27
+
28
+ BINDING = 0.29154
29
+
30
+
31
+ @torch.no_grad()
32
+ def antipodal_collapse(codebook: torch.Tensor, thresh: float = -0.9) -> dict:
33
+ """Mutual-strongest antipodal pairing + collapse to axes on RP^(D-1)."""
34
+ A = F.normalize(codebook.float(), dim=-1)
35
+ K = A.shape[0]
36
+ cos = A @ A.T
37
+ cos.fill_diagonal_(2.0) # exclude self from minima
38
+ nearest_neg = cos.argmin(dim=-1) # most-antipodal partner
39
+ pairs = []
40
+ used = set()
41
+ for i in range(K):
42
+ j = int(nearest_neg[i])
43
+ if i < j and int(nearest_neg[j]) == i and cos[i, j] < thresh:
44
+ pairs.append((i, j))
45
+ used.update((i, j))
46
+ axes = [F.normalize((A[i] - A[j]) / 2.0, dim=-1) for i, j in pairs]
47
+ axes += [A[i] for i in range(K) if i not in used] # unpaired rows as axes
48
+ axes = torch.stack(axes) if axes else A[:0]
49
+ # sign-canon onto RP: first nonzero coordinate positive
50
+ for r in range(axes.shape[0]):
51
+ nz = torch.nonzero(axes[r].abs() > 1e-8)
52
+ if nz.numel() and axes[r, nz[0, 0]] < 0:
53
+ axes[r] = -axes[r]
54
+ return {"pairs": len(pairs), "n_axes": axes.shape[0],
55
+ "unpaired": K - 2 * len(pairs), "axes": axes}
56
+
57
+
58
+ @torch.no_grad()
59
+ def projective_stats(axes: torch.Tensor, n_baseline: int = 20000,
60
+ seed: int = 0) -> dict:
61
+ """Mean projective angle arccos|<a,b>| vs a uniform-RP baseline at same (n, D)."""
62
+ n, D = axes.shape
63
+ if n < 2:
64
+ return {"proj_angle_mean": None, "uniform_baseline": None,
65
+ "deviation": None, "erank": None}
66
+ def mean_angle(rows):
67
+ c = (rows @ rows.T).abs().clamp(max=1.0)
68
+ iu = torch.triu_indices(rows.shape[0], rows.shape[0], offset=1)
69
+ return torch.arccos(c[iu[0], iu[1]]).mean().item()
70
+ obs = mean_angle(axes)
71
+ g = torch.Generator().manual_seed(seed)
72
+ base_angles = []
73
+ m = max(2, n)
74
+ for _ in range(max(1, n_baseline // max(1, m * (m - 1) // 2))):
75
+ r = F.normalize(torch.randn(m, D, generator=g), dim=-1)
76
+ base_angles.append(mean_angle(r))
77
+ base = sum(base_angles) / len(base_angles)
78
+ s = torch.linalg.svdvals(axes)
79
+ p = (s / s.sum().clamp_min(1e-12))
80
+ erank = float(torch.exp(-(p.clamp_min(1e-12) * p.clamp_min(1e-12).log()).sum()))
81
+ return {"proj_angle_mean": round(obs, 4), "uniform_baseline": round(base, 4),
82
+ "deviation": round(obs - base, 4), "erank": round(erank, 3)}
83
+
84
+
85
+ @torch.no_grad()
86
+ def read_specimen(path: str) -> dict:
87
+ ck = torch.load(path, map_location="cpu", weights_only=True)
88
+ out = {"file": path.split("\\")[-1].split("/")[-1],
89
+ "arm": ck.get("arm"), "seed": ck.get("seed"),
90
+ "steps": ck.get("steps"), "val_bpb": round(ck.get("val_bpb", -1), 4)}
91
+ if "state_dict" in ck: # full specimen checkpoint
92
+ sd = ck["state_dict"]
93
+ books = {k[:-len(".codebook")]: sd[k] for k in sd
94
+ if k.endswith("addr.codebook") or k.endswith("head_addr.codebook")}
95
+ homes = {k[:-len(".home")]: sd[k] for k in sd if k.endswith(".home")}
96
+ else: # bare genome dict (exp014+ champion files):
97
+ # books under flat/root/branch* keys; *_proj entries are projections
98
+ books = {k: v for k, v in ck.items()
99
+ if torch.is_tensor(v) and v.ndim == 2
100
+ and (k in ("flat", "root") or k.startswith("branch"))}
101
+ homes = {}
102
+ reads = {}
103
+ for name, cb in books.items():
104
+ col = antipodal_collapse(cb)
105
+ stats = projective_stats(col["axes"])
106
+ home = homes.get(name)
107
+ drift = None
108
+ binding = None
109
+ if home is not None and home.shape == cb.shape:
110
+ a = F.normalize(cb.float(), dim=-1)
111
+ b = F.normalize(home.float(), dim=-1)
112
+ dr = torch.arccos((a * b).sum(-1).clamp(-1, 1))
113
+ drift = round(dr.mean().item(), 4)
114
+ binding = round(((dr - BINDING).abs() <= 0.05).float().mean().item(), 4)
115
+ util = col["n_axes"] / cb.shape[0]
116
+ dev = stats["deviation"]
117
+ if dev is not None and abs(dev) < 0.05 and util > 0.95 and col["pairs"] <= 3:
118
+ verdict = "PROJECTIVE-CLEAN"
119
+ elif dev is not None and abs(dev) < 0.05:
120
+ verdict = "PROJECTIVE-MOSTLY"
121
+ elif dev is not None and dev > 0.05:
122
+ verdict = "STRUCTURED(repulsive)"
123
+ else:
124
+ verdict = "DEGENERATE/CLUMPED" if dev is not None else "TOO-FEW-AXES"
125
+ reads[name] = {
126
+ "pairs": col["pairs"], "n_axes": col["n_axes"], **stats,
127
+ "drift": drift, "binding_frac": binding, "verdict": verdict}
128
+ out["codebooks"] = reads
129
+ return out
130
+
131
+
132
+ def read_all(root: str) -> list:
133
+ import glob, os
134
+ results = []
135
+ for p in sorted(glob.glob(os.path.join(root, "*.pt"))):
136
+ r = read_specimen(p)
137
+ print(r, flush=True)
138
+ results.append(r)
139
+ return results
140
+
141
+
142
+ def _in_notebook() -> bool:
143
+ try:
144
+ get_ipython() # type: ignore[name-defined] # noqa: F821
145
+ return True
146
+ except NameError:
147
+ return False
148
+
149
+
150
+ if __name__ == "__main__":
151
+ if _in_notebook():
152
+ print("Notebook mode: call read_all(r'<data_root>/ar_ckpts') in the next cell.")
153
+ else:
154
+ args = [a for a in sys.argv[1:] if not a.startswith("-")]
155
+ if not args:
156
+ print("usage: python read_codebook.py <ckpt_or_dir> [...]")
157
+ for a in args:
158
+ import os
159
+ read_all(a) if os.path.isdir(a) else print(read_specimen(a))
exp018_r12/repro.py ADDED
@@ -0,0 +1,27 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """repro.py — standalone loader/runner for exp018_r12. Code dependencies live
2
+ in THIS folder (geolip_vitals.py, ar_differentiation_bed.py,
3
+ exp014_genetic_distillation.py, exp018_reran12.py, read_codebook.py).
4
+
5
+ python repro.py # CPU smoke: pentachoron regularity + arm shapes
6
+ python repro.py --run # all 4 arms x 2 seeds (GPU, ~40 min)
7
+
8
+ Data lands in ./data (override with GEOLIP_DATA). The farmed_init arm needs
9
+ the exp012 donor specimen: ../exp012_ar/specimens/addr_msl64_s0_t2000.pt
10
+ (present in this repo) or set GEOLIP_FARMED_SPECIMEN to a path.
11
+ """
12
+ import os
13
+ import sys
14
+
15
+ HERE = os.path.dirname(os.path.abspath(__file__))
16
+ sys.path.insert(0, HERE)
17
+
18
+ if __name__ == "__main__":
19
+ import geolip_vitals # noqa: F401 (paste order)
20
+ import ar_differentiation_bed # noqa: F401
21
+ import exp014_genetic_distillation # noqa: F401
22
+ import exp018_reran12 as r12
23
+ if "--run" in sys.argv[1:]:
24
+ r12.run_rerun()
25
+ else:
26
+ r12.smoke()
27
+ print("repro smoke passed — run with --run for the arms (GPU)")
exp018_r12/results/ledger.jsonl ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {"exp": "18", "arm": "penta_init", "seed": 0, "steps": 2000, "bpb": 2.4957, "drift": 0.2417, "binding_frac": 0.25}
2
+ {"exp": "18", "arm": "farmed_init", "seed": 0, "steps": 2000, "bpb": 2.4934, "drift": 0.1787, "binding_frac": 0.2031}
3
+ {"exp": "18", "arm": "tied", "seed": 0, "steps": 2000, "bpb": 3.5131, "drift": 0.0188, "binding_frac": 0.0}
4
+ {"exp": "18", "arm": "keystone", "seed": 0, "steps": 2000, "bpb": 3.5169, "drift": 0.0203, "binding_frac": 0.0}
5
+ {"exp": "18", "arm": "penta_init", "seed": 1, "steps": 2000, "bpb": 2.4389, "drift": 0.2114, "binding_frac": 0.2969}
6
+ {"exp": "18", "arm": "farmed_init", "seed": 1, "steps": 2000, "bpb": 2.4614, "drift": 0.1986, "binding_frac": 0.2812}
7
+ {"exp": "18", "arm": "tied", "seed": 1, "steps": 2000, "bpb": 3.4997, "drift": 0.0203, "binding_frac": 0.0}
8
+ {"exp": "18", "arm": "keystone", "seed": 1, "steps": 2000, "bpb": 3.5006, "drift": 0.0208, "binding_frac": 0.0}
exp018_r12/results/results.json ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "farmed_init_s0": {
3
+ "bpb": 2.4934,
4
+ "drift": 0.1787
5
+ },
6
+ "farmed_init_s1": {
7
+ "bpb": 2.4614,
8
+ "drift": 0.1986
9
+ },
10
+ "keystone_s0": {
11
+ "bpb": 3.5169,
12
+ "drift": 0.0203
13
+ },
14
+ "keystone_s1": {
15
+ "bpb": 3.5006,
16
+ "drift": 0.0208
17
+ },
18
+ "penta_init_s0": {
19
+ "bpb": 2.4957,
20
+ "drift": 0.2417
21
+ },
22
+ "penta_init_s1": {
23
+ "bpb": 2.4389,
24
+ "drift": 0.2114
25
+ },
26
+ "tied_s0": {
27
+ "bpb": 3.5131,
28
+ "drift": 0.0188
29
+ },
30
+ "tied_s1": {
31
+ "bpb": 3.4997,
32
+ "drift": 0.0203
33
+ }
34
+ }
exp018_r12/specimens/r12_farmed_init_s0.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:863c95c2ada021129ed01776dd0505df338f1c3ba0134f7a2ff036f7bd535890
3
+ size 7977645
exp018_r12/specimens/r12_farmed_init_s1.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0ef0870f9a5ec4067e72dcf1950e183dd65e526d0e5a21316865681853798a16
3
+ size 7977645
exp018_r12/specimens/r12_keystone_s0.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5b4700cd47a40d15928ea8e3f5cc2fbe6a86aa42368e6d9cb7090c8d972a9477
3
+ size 7713726
exp018_r12/specimens/r12_keystone_s1.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1183c1dc78f0c1a2f20b277922103fb7192134d09939265575ecade8b87f524a
3
+ size 7713726
exp018_r12/specimens/r12_penta_init_s0.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6e73f0554426e45fffcabeea895eff4e7d7a47076ca7a57f552c1962d6d11df7
3
+ size 7977590
exp018_r12/specimens/r12_penta_init_s1.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f8b615f8a3ce5358b1286b5184f20c91aa055eb7b0a5b417f0c3b81d78f554f5
3
+ size 7977590
exp018_r12/specimens/r12_tied_s0.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f12ee10ac833b18a3821ca9273ebc575713e334baa2fad06e93d732bace80222
3
+ size 7713514
exp018_r12/specimens/r12_tied_s1.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c1d8da581864f4fb18458b86d51148357c91b5cb7af0f958aae72a190912ae6f
3
+ size 7713514
exp019_cr/README.md ADDED
@@ -0,0 +1,82 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # exp019_cr — the capacity for distilled content retention (+ generalization)
2
+
3
+ When content is **distilled** rather than directly learned, how much transfers,
4
+ how much is retained, and for how long? Sequel to [exp016](../exp016_ct/) and
5
+ the [exp014](../exp014_gd/) memory-substrate studies; the weak-to-strong
6
+ distillation cell made real (the teacher holds content the student lacks —
7
+ true headroom on the content axis).
8
+
9
+ **Instrument**: N fact records `\n@<key6>=<value12>\n` (random alphanumeric —
10
+ uncompletable from language statistics, so exact-match greedy completion IS
11
+ retention), mixed into the wikitext byte stream at rate 0.5. Teacher =
12
+ addr_msl64 bed model, 4000 steps, gated on its own recall. Channels (fresh
13
+ student, 2000 steps): `direct` (ground truth; ceiling) | `kd_facts` (fact rows
14
+ supervised ONLY by teacher logits — pure distilled content, α=1.0) |
15
+ `kd_general` (KD on clean text only — leakage probe) | `book_implant`
16
+ (teacher's trained codebook into a fresh student) | `none` (floor). Axes:
17
+ N ∈ {64, 256, 1024}; retention re-measured after 1000 further clean-stream
18
+ steps (interference). The **19b block** swaps rote facts for **rule-bearing
19
+ content** (value = fixed substitution cipher of the key; 256 train keys, 128
20
+ held out) plus prompt-format variants.
21
+
22
+ `build_results.py` re-asserts every claim below from `results/ledger.jsonl`
23
+ (36 rows).
24
+
25
+ ## Findings — retention sweep
26
+
27
+ 1. **Capacity curve**: teacher recall at 4k steps = 1.00 (N=64), 0.99 (N=256),
28
+ 0.01–0.20 (N=1024) — the ~1.9M model holds ~256 facts near-perfectly and
29
+ hits a cliff before 1024 (cliff edge seed-chaotic).
30
+ 2. **The distillation tax is zero-to-negative.** Where the teacher knows the
31
+ content, teacher logits alone transfer it at parity with ground truth
32
+ (N=64: exact 1.0 = 1.0, both seeds) or better (N=256: 0.953 vs 0.871 s0;
33
+ 0.859 vs 0.856 s1). Soft targets from a teacher with headroom are at least
34
+ as content-efficient as the data itself. The cost is general modeling:
35
+ kd students' clean bpb 4.0–4.3 vs direct 2.8–3.1 vs floor 2.47.
36
+ 3. **No logit leakage**: KD on clean text transfers zero facts (0.000, 6/6
37
+ cells) — content does not cross without exposure at this scale.
38
+ 4. **Anchors carry no bytes**: the trained codebook of a teacher that knows
39
+ 256 facts, implanted into a fresh student, transfers **none** of them
40
+ (0.000, both seeds). The codebook organizes addressing; content lives in
41
+ the trunk.
42
+ 5. **Universal catastrophic forgetting**: after 1000 clean-stream steps, every
43
+ channel's recall is exactly 0.000 — including perfectly-learned direct
44
+ students. At this scale nothing persists without pressure; **persistence,
45
+ not transfer, is the unsolved axis of the memory substrate.**
46
+
47
+ ## Findings — 19b generalization block (rule content)
48
+
49
+ 6. **Rules are learned fragmentarily**: teachers memorize the rule-pairs
50
+ near-perfectly (train exact 0.98–1.00) but reach only ~0.25–0.27 held-out
51
+ byte accuracy (~10× chance) with **zero** exact completions — character
52
+ mappings absorbed, the full cipher never.
53
+ 7. **The logit channel beats direct learning on structured content, 2/2
54
+ seeds** (train exact 0.758/0.844 vs 0.652/0.773) and transfers the partial
55
+ rule at full fidelity — the dark-knowledge advantage, certified on rules.
56
+ 8. **Content is format-locked** — variant-format prompts collapse recall to
57
+ ~0 even on trained keys — but **KD students are consistently less
58
+ format-locked** than direct students (variant byte 0.077–0.091 vs
59
+ 0.000–0.024): soft targets bind content less rigidly to surface form.
60
+ 9. Interference erases the rule too — structure forgets like instances.
61
+
62
+ ## Files
63
+ - `exp019_content_retention.py` — fact/rule corpora, streams, exact-match
64
+ recall (with format variants), channel training, both runners, smoke.
65
+ - `geolip_vitals.py` / `ar_differentiation_bed.py` /
66
+ `exp014_genetic_distillation.py` / `read_codebook.py` — this package's own
67
+ harness copies. Standalone.
68
+ - `repro.py`, `build_results.py`, `results/ledger.jsonl` (36 rows),
69
+ `specimens/` (6 fact teachers + 2 rule teachers + 6 rule students).
70
+
71
+ ## Reproduce (from inside this folder)
72
+ ```bash
73
+ pip install torch --index-url https://download.pytorch.org/whl/cu128
74
+ pip install pyarrow huggingface_hub
75
+ python repro.py # CPU smoke
76
+ python repro.py --run # retention sweep (GPU, ~3h)
77
+ python repro.py --run19b # rule/generalization block (GPU, ~1h)
78
+ python build_results.py # re-assert every claim from the ledger
79
+ ```
80
+ Data lands in `./data` (override with `GEOLIP_DATA`).
81
+
82
+ License: MIT · AbstractPhil + Claude Fable 5 · July 11, 2026
exp019_cr/ar_differentiation_bed.py ADDED
@@ -0,0 +1,487 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """ar_differentiation_bed.py — exp012: autoregressive differentiation of the aleph.
2
+
3
+ Differentiation is cultivated by PREDICTIVE pressure along the sequence — the
4
+ address parameterizing the next-byte distribution (Law 2: chain-rule advantage pays
5
+ ONLY where the composed address directly parameterizes the predictive distribution).
6
+ This bed puts the aleph in the autoregressive gradient path and measures what
7
+ differentiates. The head arms enforce the employment law at its maximum: the
8
+ ENTIRE next-byte distribution is parameterized by the address.
9
+
10
+ Byte-level causal LM on wikitext-2-raw (HF parquet, CDN-fast), block 256. ARMS:
11
+ sdpa — standard causal transformer control (matched trunk).
12
+ hub — attention replaced by CAUSAL HUB: linear attention whose feature map
13
+ is the 2K-oriented aleph address, prefix-sum memories (no selection
14
+ event; O(n*K*d)). Differentiation cultivated INSIDE attention.
15
+ addr_head — sdpa trunk, but the OUTPUT HEAD reads ONLY the signed aleph
16
+ coefficient vector w_k = sinh(u_k)/sum_j cosh(u_j) of the final
17
+ hidden state (K -> 256 logits). The address MUST carry every bit of
18
+ next-byte information — the hardest Law-2 bottleneck.
19
+
20
+ JUDGED BY: val bits-per-byte per arm (task) + CULTIVATION VITALS on every aleph
21
+ codebook (readouts, never losses): axis aliveness/hppl, drift-from-init +
22
+ binding fraction @0.29154, winner-|cos| saturation (sign-code emergence), shadow
23
+ path diversity (fixed high-bits hash). Never by recon.
24
+
25
+ Riders: pure Adam wd=0; no BN/Dropout/GAP on geometric paths; orthogonal init;
26
+ Colab-cell-safe (paste-ahead imports, no bare argparse, no __file__ reliance);
27
+ GPU-only for verdict runs.
28
+
29
+ Terminal: python ar_differentiation_bed.py # shapes/parse smoke
30
+ python ar_differentiation_bed.py --train # verdict run
31
+ Colab: paste geolip_vitals.py cell, then this file (smoke auto-runs),
32
+ then train(steps=2000, data_root="/content/data") in the next cell.
33
+ """
34
+ from __future__ import annotations
35
+ import math
36
+ import torch
37
+ import torch.nn as nn
38
+ import torch.nn.functional as F
39
+
40
+ if "anchor_drift" not in globals():
41
+ try:
42
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
43
+ except ImportError:
44
+ _here = globals().get("__file__")
45
+ if _here is not None:
46
+ import sys, pathlib
47
+ sys.path.insert(0, str(pathlib.Path(_here).parent))
48
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
49
+ else:
50
+ raise ImportError(
51
+ "geolip_vitals not found — paste/run its cell first, or "
52
+ "hf_hub_download exp012_ar/geolip_vitals.py from "
53
+ "AbstractPhil/geolip-aleph-differentiation.")
54
+
55
+ VOCAB = 256 # bytes
56
+
57
+
58
+ # ------------------------------------------------------------------ aleph address
59
+ def _super_fibonacci_s3(n: int) -> torch.Tensor:
60
+ """Near-uniform unit quaternions (Alexa CVPR'22) —
61
+ starts the codebook INSIDE the RP^3 attractor basin. D=4 only."""
62
+ PHI, PSI = math.sqrt(2.0), 1.533751168755204288118041
63
+ i = torch.arange(n, dtype=torch.float64)
64
+ s = (i + 0.5) / n
65
+ r, R = torch.sqrt(s), torch.sqrt(1.0 - s)
66
+ a, b = 2 * math.pi * i / PHI, 2 * math.pi * i / PSI
67
+ q = torch.stack([r * torch.sin(a), r * torch.cos(a),
68
+ R * torch.sin(b), R * torch.cos(b)], dim=-1)
69
+ return F.normalize(q, dim=-1).float()
70
+
71
+
72
+ class AlephAddress(nn.Module):
73
+ """Closed-form aleph over 2K oriented half-axes (aleph-void article).
74
+ signed(x): (..., K) w_k = sinh(u_k)/sum_j cosh(u_j) — the Law-2 head feature.
75
+ oriented(x): ((..., K), (..., K)) positive halves of the 2K softmax — HUB map."""
76
+
77
+ def __init__(self, K: int, D: int, tau: float = 0.1, init: str = "random"):
78
+ super().__init__()
79
+ self.K, self.D, self.tau = K, D, tau
80
+ if init == "fibonacci":
81
+ assert D == 4, "fibonacci init lives on S^3 (D=4)"
82
+ A = _super_fibonacci_s3(K)
83
+ else:
84
+ A = F.normalize(torch.randn(K, D), dim=-1)
85
+ self.codebook = nn.Parameter(A)
86
+ self.register_buffer("home", self.codebook.detach().clone())
87
+
88
+ def _u(self, x):
89
+ A = F.normalize(self.codebook, dim=-1)
90
+ return (F.normalize(x, dim=-1) @ A.transpose(-1, -2)) / self.tau
91
+
92
+ def oriented(self, x):
93
+ u = self._u(x)
94
+ m = u.abs().amax(dim=-1, keepdim=True)
95
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
96
+ Z = (ep + en).sum(dim=-1, keepdim=True)
97
+ return ep / Z, en / Z
98
+
99
+ def signed(self, x):
100
+ u = self._u(x)
101
+ m = u.abs().amax(dim=-1, keepdim=True)
102
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
103
+ return (ep - en) / (ep + en).sum(dim=-1, keepdim=True)
104
+
105
+ def signed_at(self, x, taus):
106
+ """Multi-tau stroboscope (rule of 3): signed coefficients at several
107
+ temperatures, concatenated — softer taus keep the vector dense while a
108
+ hard tau supplies the sign-code sharpness. v2 refinement (b)."""
109
+ A = F.normalize(self.codebook, dim=-1)
110
+ cos = F.normalize(x, dim=-1) @ A.transpose(-1, -2)
111
+ outs = []
112
+ for t in taus:
113
+ u = cos / t
114
+ m = u.abs().amax(dim=-1, keepdim=True)
115
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
116
+ outs.append((ep - en) / (ep + en).sum(dim=-1, keepdim=True))
117
+ return torch.cat(outs, dim=-1)
118
+
119
+ def m_hat(self, x):
120
+ """Closed-form soft read (decoders read M_hat, never M). v2 control (c)."""
121
+ u = self._u(x)
122
+ m = u.abs().amax(dim=-1, keepdim=True)
123
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
124
+ A = F.normalize(self.codebook, dim=-1)
125
+ return ((ep - en) @ A) / (ep + en).sum(dim=-1, keepdim=True)
126
+
127
+ def m_hard_ste(self, x):
128
+ """Hard mode (aleph-void article): M_hard = sign(cos_win) * A[win], straight-through to
129
+ the soft read — forward fully discrete SIGN CODE, backward soft gradient.
130
+ Legal per theme A (reconstructive sign code, not a one-hot roster pick)."""
131
+ u = self._u(x)
132
+ soft = self.m_hat(x)
133
+ win = u.abs().argmax(dim=-1)
134
+ A = F.normalize(self.codebook, dim=-1)
135
+ sign = torch.sign(torch.gather(u, -1, win.unsqueeze(-1))).squeeze(-1)
136
+ hard = sign.unsqueeze(-1) * A[win]
137
+ return hard + soft - soft.detach()
138
+
139
+ @torch.no_grad()
140
+ def vitals(self, x_sample) -> dict:
141
+ u = self._u(x_sample.reshape(-1, x_sample.shape[-1]))
142
+ p, n = self.oriented(x_sample.reshape(-1, x_sample.shape[-1]))
143
+ two_k = torch.cat([p, n], dim=-1)
144
+ win = two_k.argmax(dim=-1)
145
+ cos_win = (u.abs().amax(dim=-1) * self.tau) # winner |cos| — sign-code sat.
146
+ d = anchor_drift(self.codebook, self.home)
147
+ return {"drift": round(d["mean"], 4),
148
+ "binding_frac": round(d["binding_fraction"], 4),
149
+ "aliveness": axis_aliveness(two_k),
150
+ "win_cos_mean": round(cos_win.mean().item(), 4),
151
+ "paths": path_diversity(win)}
152
+
153
+
154
+ # ------------------------------------------------------------------------- blocks
155
+ class CausalSDPA(nn.Module):
156
+ def __init__(self, d: int, heads: int = 4):
157
+ super().__init__()
158
+ self.h = heads
159
+ self.qkv = nn.Linear(d, 3 * d, bias=False)
160
+ self.o = nn.Linear(d, d, bias=False)
161
+ nn.init.orthogonal_(self.qkv.weight); nn.init.orthogonal_(self.o.weight)
162
+
163
+ def forward(self, x):
164
+ B, n, d = x.shape
165
+ q, k, v = self.qkv(x).chunk(3, dim=-1)
166
+ q, k, v = (t.view(B, n, self.h, d // self.h).transpose(1, 2) for t in (q, k, v))
167
+ y = F.scaled_dot_product_attention(q, k, v, is_causal=True)
168
+ return self.o(y.transpose(1, 2).reshape(B, n, d))
169
+
170
+
171
+ class CausalHUB(nn.Module):
172
+ """Causal aleph linear attention: prefix-sum memories over the two K-wide
173
+ halves of the oriented address; 2K never materialized; no selection event."""
174
+
175
+ def __init__(self, d: int, K: int = 32, D: int = 4, tau: float = 0.1):
176
+ super().__init__()
177
+ self.addr = AlephAddress(K, D, tau)
178
+ self.q = nn.Linear(d, D, bias=False)
179
+ self.k = nn.Linear(d, D, bias=False)
180
+ self.v = nn.Linear(d, d, bias=False)
181
+ self.o = nn.Linear(d, d, bias=False)
182
+ for m in (self.q, self.k, self.v, self.o):
183
+ nn.init.orthogonal_(m.weight)
184
+
185
+ def forward(self, x):
186
+ qp, qn = self.addr.oriented(self.q(x)) # (B, n, K)
187
+ kp, kn = self.addr.oriented(self.k(x))
188
+ v = self.v(x) # (B, n, d)
189
+ Sp = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kp, v), dim=1)
190
+ Sn = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kn, v), dim=1)
191
+ zp = torch.cumsum(kp, dim=1)
192
+ zn = torch.cumsum(kn, dim=1)
193
+ num = torch.einsum("bnk,bnkd->bnd", qp, Sp) + torch.einsum("bnk,bnkd->bnd", qn, Sn)
194
+ den = (qp * zp).sum(-1, keepdim=True) + (qn * zn).sum(-1, keepdim=True)
195
+ return self.o(num / den.clamp_min(1e-12))
196
+
197
+
198
+ class MslRelay(nn.Module):
199
+ """Depth-composition unit (chain-rule probe): multi-slot M_hat read entering
200
+ the trunk as a NEAR-ZERO gated residual (gate init -3.0, sigma~0.047 — theme D:
201
+ geometry enters as a nudge and grows only if it earns gradient)."""
202
+
203
+ def __init__(self, d: int, n_slots: int = 16, K: int = 64):
204
+ super().__init__()
205
+ self.n_slots = n_slots
206
+ self.proj = nn.Linear(d, n_slots * 4, bias=False)
207
+ self.out = nn.Linear(n_slots * 4, d, bias=False)
208
+ nn.init.orthogonal_(self.proj.weight)
209
+ nn.init.orthogonal_(self.out.weight)
210
+ self.addr = AlephAddress(K, 4)
211
+ self.gate = nn.Parameter(torch.tensor(-3.0))
212
+
213
+ def forward(self, x):
214
+ B, n, _ = x.shape
215
+ slots = self.proj(x).view(B, n, self.n_slots, 4)
216
+ m = self.addr.m_hat(slots).reshape(B, n, -1)
217
+ return x + self.gate.sigmoid() * self.out(m)
218
+
219
+
220
+ class Block(nn.Module):
221
+ def __init__(self, d: int, attn: nn.Module):
222
+ super().__init__()
223
+ self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)
224
+ self.attn = attn
225
+ self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
226
+
227
+ def forward(self, x):
228
+ x = x + self.attn(self.n1(x))
229
+ return x + self.mlp(self.n2(x))
230
+
231
+
232
+ class ByteLM(nn.Module):
233
+ def __init__(self, arm: str, d: int = 192, layers: int = 4, block: int = 256,
234
+ K: int = 32, D: int = 4):
235
+ super().__init__()
236
+ # "<arm>_tri" suffix = trigram byte embedding (AlephLM byte_emb x3 lineage):
237
+ # token embedding is the sum of embeddings of bytes t, t-1, t-2.
238
+ self.trigram = arm.endswith("_tri")
239
+ if self.trigram:
240
+ arm = arm[:-4]
241
+ # "_fib" = super-Fibonacci S^3 codebook init (basin test: starts INSIDE
242
+ # the RP^3 attractor; primary observable is init->final geodesic drift).
243
+ self.fib = arm.endswith("_fib")
244
+ if self.fib:
245
+ arm = arm[:-4]
246
+ # "relay*" = stacked addresses in depth: MslRelay after every block.
247
+ # relay -> sdpa trunk + standard head; relay_msl64 -> + addressed head.
248
+ self.use_relay = arm.startswith("relay")
249
+ if arm == "relay":
250
+ arm = "sdpa"
251
+ elif arm == "relay_msl64":
252
+ arm = "addr_msl64"
253
+ self.arm, self.block = arm, block
254
+ self.emb = nn.Embedding(VOCAB, d)
255
+ if self.trigram:
256
+ self.emb1 = nn.Embedding(VOCAB, d)
257
+ self.emb2 = nn.Embedding(VOCAB, d)
258
+ self.pos = nn.Parameter(torch.zeros(1, block, d) + 0.01 * torch.randn(1, block, d))
259
+ mk_attn = (lambda: CausalHUB(d, K, D)) if arm == "hub" else (lambda: CausalSDPA(d))
260
+ self.blocks = nn.ModuleList([Block(d, mk_attn()) for _ in range(layers)])
261
+ if self.use_relay:
262
+ self.relays = nn.ModuleList([MslRelay(d) for _ in range(layers)])
263
+ self.nf = nn.LayerNorm(d)
264
+ if arm == "addr_head":
265
+ self.head_addr = AlephAddress(K, d) # v1: codebook in model dim — COLLAPSED
266
+ self.head = nn.Linear(K, VOCAB, bias=True)
267
+ elif arm in ("addr_d4", "addr_3tau", "addr_mhat"):
268
+ # v2 refinements: LOW-D HOME — learned projection to the native D=4 home
269
+ # before addressing (mirrors the healthy HUB arms), K=64.
270
+ self.head_proj = nn.Linear(d, 4, bias=False)
271
+ nn.init.orthogonal_(self.head_proj.weight)
272
+ self.head_addr = AlephAddress(64, 4)
273
+ if arm == "addr_d4":
274
+ self.head = nn.Linear(64, VOCAB, bias=True) # w alone, D=4 home
275
+ elif arm == "addr_3tau":
276
+ self.taus = (0.05, 0.1, 0.3) # rule-of-3 strobe
277
+ self.head = nn.Linear(64 * 3, VOCAB, bias=True)
278
+ else: # addr_mhat
279
+ self.head = nn.Linear(4, VOCAB, bias=True) # tightest: M_hat
280
+ elif arm.startswith("addr_msl"):
281
+ # v3: MULTI-SLOT heads — the 16s funnel widening: P parallel D=4 slots
282
+ # over a SHARED codebook. addr_msl consumes the reconstructive M_hat per
283
+ # slot (Px4 dims); addr_msl_w consumes signed w per slot (Px64) — tests
284
+ # whether slot-parallel consumption alone rescues the coefficient path.
285
+ # addr_msl<P> = slot-count dose-response. addr_mslh<P> = HARD sign-code
286
+ # consumption (straight-through M_hard per slot).
287
+ self.hard = arm.startswith("addr_mslh")
288
+ if arm in ("addr_msl", "addr_msl_w"):
289
+ self.n_slots = 16
290
+ else:
291
+ self.n_slots = int(arm[len("addr_mslh" if self.hard else "addr_msl"):])
292
+ self.head_proj = nn.Linear(d, self.n_slots * 4, bias=False)
293
+ nn.init.orthogonal_(self.head_proj.weight)
294
+ self.head_addr = AlephAddress(
295
+ 64, 4, init="fibonacci" if self.fib else "random")
296
+ width = self.n_slots * (64 if arm == "addr_msl_w" else 4)
297
+ self.head = nn.Linear(width, VOCAB, bias=True)
298
+ elif arm == "addr_3tau_mhat":
299
+ # v3: combine the two v2 winners — 3-tau stroboscope + reconstructive read.
300
+ self.head_proj = nn.Linear(d, 4, bias=False)
301
+ nn.init.orthogonal_(self.head_proj.weight)
302
+ self.head_addr = AlephAddress(64, 4)
303
+ self.taus = (0.05, 0.1, 0.3)
304
+ self.head = nn.Linear(64 * 3 + 4, VOCAB, bias=True)
305
+ else:
306
+ self.head = nn.Linear(d, VOCAB, bias=True)
307
+ self._last_h = None
308
+
309
+ def forward(self, idx):
310
+ x = self.emb(idx)
311
+ if self.trigram: # past-only shifts — causality preserved
312
+ x = x + self.emb1(F.pad(idx, (1, 0), value=0)[:, :-1]) \
313
+ + self.emb2(F.pad(idx, (2, 0), value=0)[:, :-2])
314
+ x = x + self.pos[:, : idx.shape[1]]
315
+ if self.use_relay:
316
+ for b, r in zip(self.blocks, self.relays):
317
+ x = r(b(x))
318
+ else:
319
+ for b in self.blocks:
320
+ x = b(x)
321
+ h = self.nf(x)
322
+ self._last_h = h.detach()
323
+ if self.arm == "addr_head":
324
+ return self.head(self.head_addr.signed(h))
325
+ if self.arm == "addr_d4":
326
+ return self.head(self.head_addr.signed(self.head_proj(h)))
327
+ if self.arm == "addr_3tau":
328
+ return self.head(self.head_addr.signed_at(self.head_proj(h), self.taus))
329
+ if self.arm == "addr_mhat":
330
+ return self.head(self.head_addr.m_hat(self.head_proj(h)))
331
+ if self.arm.startswith("addr_msl"):
332
+ B, n, _ = h.shape
333
+ slots = self.head_proj(h).view(B, n, self.n_slots, 4)
334
+ if self.arm == "addr_msl_w":
335
+ feats = self.head_addr.signed(slots).reshape(B, n, -1)
336
+ elif getattr(self, "hard", False):
337
+ feats = self.head_addr.m_hard_ste(slots).reshape(B, n, -1)
338
+ else:
339
+ feats = self.head_addr.m_hat(slots).reshape(B, n, -1)
340
+ return self.head(feats)
341
+ if self.arm == "addr_3tau_mhat":
342
+ p = self.head_proj(h)
343
+ feats = torch.cat([self.head_addr.signed_at(p, self.taus),
344
+ self.head_addr.m_hat(p)], dim=-1)
345
+ return self.head(feats)
346
+ return self.head(h)
347
+
348
+ @torch.no_grad()
349
+ def vitals(self) -> dict:
350
+ out = {}
351
+ if self.arm == "hub":
352
+ for i, b in enumerate(self.blocks):
353
+ if self._last_h is not None:
354
+ out[f"L{i}"] = b.attn.addr.vitals(b.attn.q(self._last_h[:2]))
355
+ elif self.arm == "addr_head" and self._last_h is not None:
356
+ out["head"] = self.head_addr.vitals(self._last_h[:2])
357
+ elif self.arm in ("addr_d4", "addr_3tau", "addr_mhat",
358
+ "addr_3tau_mhat") and self._last_h is not None:
359
+ out["head"] = self.head_addr.vitals(self.head_proj(self._last_h[:2]))
360
+ elif self.arm.startswith("addr_msl") and self._last_h is not None:
361
+ slots = self.head_proj(self._last_h[:2])
362
+ out["head"] = self.head_addr.vitals(
363
+ slots.reshape(*slots.shape[:-1], self.n_slots, 4))
364
+ if self.use_relay and self._last_h is not None:
365
+ for i, r in enumerate(self.relays):
366
+ s = r.proj(self._last_h[:2])
367
+ v = r.addr.vitals(s.reshape(*s.shape[:-1], r.n_slots, 4))
368
+ out[f"relay{i}"] = {"gate": round(r.gate.sigmoid().item(), 4),
369
+ "drift": v["drift"],
370
+ "binding_frac": v["binding_frac"],
371
+ "ppl": round(v["aliveness"]["usage_ppl"], 1)}
372
+ return out
373
+
374
+
375
+ # --------------------------------------------------------------------------- data
376
+ def _wikitext_bytes(data_root: str):
377
+ """wikitext-2-raw as flat uint8 tensors via the HF parquet CDN."""
378
+ from huggingface_hub import hf_hub_download
379
+ import pyarrow.parquet as pq
380
+
381
+ def load(split):
382
+ p = hf_hub_download("Salesforce/wikitext",
383
+ f"wikitext-2-raw-v1/{split}-00000-of-00001.parquet",
384
+ repo_type="dataset", local_dir=data_root)
385
+ text = "".join(pq.read_table(p).column("text").to_pylist())
386
+ return torch.frombuffer(bytearray(text.encode("utf-8")), dtype=torch.uint8).clone()
387
+
388
+ return load("train"), load("validation")
389
+
390
+
391
+ def _batch(data: torch.Tensor, batch: int, block: int, device, g: torch.Generator):
392
+ ix = torch.randint(0, data.numel() - block - 1, (batch,), generator=g)
393
+ x = torch.stack([data[i:i + block] for i in ix]).long().to(device)
394
+ y = torch.stack([data[i + 1:i + block + 1] for i in ix]).long().to(device)
395
+ return x, y
396
+
397
+
398
+ # -------------------------------------------------------------------- train/smoke
399
+ def train(arms=("sdpa", "hub", "addr_head"), steps: int = 2000, batch: int = 32,
400
+ block: int = 256, device: str = "cuda", data_root: str = "./data",
401
+ seed: int = 0, eval_every: int = 500, save: bool = True):
402
+ """Verdict run — GPU only. Pure Adam wd=0. Reports val bits-per-byte + vitals.
403
+ save=True writes {data_root}/ar_ckpts/{arm}_s{seed}_t{steps}.pt per arm —
404
+ the cultivated codebooks are SPECIMENS for the projective reading instruments."""
405
+ import os
406
+ if device == "cuda" and not torch.cuda.is_available():
407
+ raise RuntimeError("Verdict runs are GPU-only (never CPU-train for accuracy).")
408
+ ckpt_dir = os.path.join(data_root, "ar_ckpts")
409
+ os.makedirs(ckpt_dir, exist_ok=True)
410
+ tr, va = _wikitext_bytes(data_root)
411
+ print(f"data ready: train {tr.numel():,} bytes, val {va.numel():,} bytes", flush=True)
412
+ results = {}
413
+ for arm in arms:
414
+ torch.manual_seed(seed)
415
+ g = torch.Generator().manual_seed(seed)
416
+ model = ByteLM(arm, block=block).to(device)
417
+ n_params = sum(p.numel() for p in model.parameters())
418
+ opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
419
+ for step in range(1, steps + 1):
420
+ x, y = _batch(tr, batch, block, device, g)
421
+ logits = model(x)
422
+ loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
423
+ opt.zero_grad(set_to_none=True)
424
+ loss.backward()
425
+ opt.step()
426
+ if step % eval_every == 0 or step == steps:
427
+ model.eval()
428
+ with torch.no_grad():
429
+ losses = []
430
+ for _ in range(20):
431
+ xv, yv = _batch(va, batch, block, device, g)
432
+ lv = F.cross_entropy(model(xv).reshape(-1, VOCAB),
433
+ yv.reshape(-1))
434
+ losses.append(lv.item())
435
+ bpb = sum(losses) / len(losses) / math.log(2)
436
+ print(f"[{arm}] step {step} val_bpb={bpb:.4f} vitals={model.vitals()}",
437
+ flush=True)
438
+ model.train()
439
+ results[arm] = {"val_bpb": bpb, "params": n_params, "vitals": model.vitals()}
440
+ if save:
441
+ path = os.path.join(ckpt_dir, f"{arm}_s{seed}_t{steps}.pt")
442
+ torch.save({"arm": arm, "seed": seed, "steps": steps, "val_bpb": bpb,
443
+ "state_dict": {k: v.cpu() for k, v in
444
+ model.state_dict().items()}}, path)
445
+ print(f"saved specimen: {path}", flush=True)
446
+ print(results, flush=True)
447
+ return results
448
+
449
+
450
+ def smoke():
451
+ """Shapes/parse only — no accuracy claims."""
452
+ x = torch.randint(0, VOCAB, (2, 64))
453
+ for arm in ("sdpa", "hub", "addr_head"):
454
+ m = ByteLM(arm, d=96, layers=2, block=64, K=16)
455
+ logits = m(x)
456
+ assert logits.shape == (2, 64, VOCAB)
457
+ logits.sum().backward()
458
+ # causality check: future byte must not affect past logits
459
+ with torch.no_grad():
460
+ a = m(x)[0, 10]
461
+ x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
462
+ b = m(x2)[0, 10]
463
+ assert torch.allclose(a, b, atol=1e-4), f"{arm} leaks future context"
464
+ print(f"{arm}: OK params={sum(p.numel() for p in m.parameters()):,} "
465
+ f"vitals={m.vitals()}", flush=True)
466
+ print("OK — AR bed smoke passed (verdict run: train() on GPU)", flush=True)
467
+
468
+
469
+ def _in_notebook() -> bool:
470
+ try:
471
+ get_ipython() # type: ignore[name-defined] # noqa: F821
472
+ return True
473
+ except NameError:
474
+ return False
475
+
476
+
477
+ if __name__ == "__main__":
478
+ if _in_notebook():
479
+ smoke()
480
+ print("Notebook mode: call train(steps=2000) in the next cell (GPU).")
481
+ else:
482
+ import argparse
483
+ ap = argparse.ArgumentParser()
484
+ ap.add_argument("--train", action="store_true")
485
+ ap.add_argument("--steps", type=int, default=2000)
486
+ a, _ = ap.parse_known_args()
487
+ train(steps=a.steps) if a.train else smoke()
exp019_cr/build_results.py ADDED
@@ -0,0 +1,81 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """build_results.py — exp019_cr: read results/ledger.jsonl and RE-ASSERT every
2
+ claim in the README (retention sweep + 19b generalization block).
3
+ Run from inside this folder: python build_results.py
4
+ """
5
+ import json
6
+ import os
7
+
8
+ HERE = os.path.dirname(os.path.abspath(__file__))
9
+ rows = [json.loads(l) for l in
10
+ open(os.path.join(HERE, "results", "ledger.jsonl"), encoding="utf-8")]
11
+ a = [r for r in rows if r["exp"] == "19"]
12
+ b = [r for r in rows if r["exp"] == "19b"]
13
+ assert len(a) == 28 and len(b) == 8
14
+
15
+ def cellA(ch, n, s):
16
+ return next(r for r in a if r["channel"] == ch and r["n"] == n
17
+ and r["seed"] == s)
18
+
19
+ def cellB(ch, s):
20
+ return next(r for r in b if r["channel"] == ch and r["seed"] == s)
21
+
22
+ # retention claim 1: capacity curve — teachers near-perfect through N=256,
23
+ # cliff before 1024
24
+ for s in (0, 1):
25
+ assert cellA("teacher", 64, s)["recall"]["exact"] == 1.0
26
+ assert cellA("teacher", 256, s)["recall"]["exact"] > 0.98
27
+ assert cellA("teacher", 1024, s)["recall"]["exact"] < 0.20
28
+
29
+ # retention claim 2: distillation tax zero-to-negative where the teacher knows
30
+ # the content (kd_facts >= direct - 0.01 at N=64/256, both seeds)
31
+ for s in (0, 1):
32
+ for n in (64, 256):
33
+ kd = cellA("kd_facts", n, s)["recall"]["exact"]
34
+ di = cellA("direct", n, s)["recall"]["exact"]
35
+ assert kd >= di - 0.01, (n, s, kd, di)
36
+
37
+ # retention claim 3: no logit leakage on clean text (kd_general exact 0.0 all)
38
+ for r in a:
39
+ if r["channel"] == "kd_general":
40
+ assert r["recall"]["exact"] == 0.0
41
+
42
+ # retention claim 4: anchors carry no bytes (book_implant exact 0.0, both seeds)
43
+ for s in (0, 1):
44
+ assert cellA("book_implant", 256, s)["recall"]["exact"] == 0.0
45
+
46
+ # retention claim 5: universal catastrophic forgetting (every student channel
47
+ # -> 0.0 exact after 1k clean steps)
48
+ for r in a:
49
+ if "recall_interf" in r and r["recall_interf"] is not None:
50
+ assert r["recall_interf"]["exact"] == 0.0, r
51
+
52
+ # 19b claim 1: partial rule learning — teachers memorize, held-out byte acc
53
+ # ~10x chance but exact 0
54
+ for s in (0, 1):
55
+ t = cellB("teacher", s)["gauges"]
56
+ assert t["train"]["exact"] > 0.98
57
+ assert 0.15 < t["heldout"]["byte_acc"] < 0.35
58
+ assert t["heldout"]["exact"] == 0.0
59
+ # format lock: variant-format recall collapses even on train keys
60
+ assert t["train_varfmt"]["byte_acc"] < 0.05
61
+
62
+ # 19b claim 2: the logit channel beats direct on rule content, both seeds,
63
+ # and transfers the partial rule at full fidelity
64
+ for s in (0, 1):
65
+ kd, di = cellB("kd_facts", s)["gauges"], cellB("direct", s)["gauges"]
66
+ assert kd["train"]["exact"] > di["train"]["exact"], s
67
+ assert kd["heldout"]["byte_acc"] > 0.24
68
+ # kd students are less format-locked than direct students
69
+ assert kd["train_varfmt"]["byte_acc"] > di["train_varfmt"]["byte_acc"], s
70
+
71
+ out = {"capacity_curve": {f"N{n}_s{s}": cellA("teacher", n, s)["recall"]["exact"]
72
+ for n in (64, 256, 1024) for s in (0, 1)},
73
+ "kd_vs_direct_rule_train_exact": {
74
+ f"s{s}": [cellB("kd_facts", s)["gauges"]["train"]["exact"],
75
+ cellB("direct", s)["gauges"]["train"]["exact"]]
76
+ for s in (0, 1)},
77
+ "n_rows": len(rows)}
78
+ json.dump(out, open(os.path.join(HERE, "results", "results.json"), "w",
79
+ encoding="utf-8"), indent=1)
80
+ print(f"{len(rows)} rows -> results/results.json")
81
+ print("all README claims asserted OK")
exp019_cr/exp014_genetic_distillation.py ADDED
@@ -0,0 +1,515 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """exp014_genetic_distillation.py — genetic distillation + memory substrate.
2
+ 14_a: multi-generational tournament (GM3 paradigm) where the aleph codebook is the
3
+ explicit heritable genome. Lineages: ALEPH-FLAT (consensus book + KD) |
4
+ ALEPH-TREE (structured genome: root book + branch books) | MLP-KD
5
+ (traditional: best-parent weights + KD) | NO-INHERIT (evolution floor).
6
+ Both sides intentionally inherit logits (KD); only ours inherits geometry.
7
+ Consensus = Procrustes/GPA alignment of parents' books to mean shape
8
+ (placement by construction — replaces GM3's k-means-on-consensus init).
9
+ 14_b: memory substrate — the D=4 home makes books size-agnostic. Implant books
10
+ cultivated in a small organism into a larger one (frozen / trainable), and
11
+ into GPT-2 relay adapters (cross-architecture frozen distillation).
12
+
13
+ Riders: pure Adam wd=0; KD = KL to detached teacher probs (predictive pressure, no
14
+ contrastive); tree routing is DENSE SOFT (oriented weights; collapse monitor on the
15
+ root); drift-check precedes every freeze claim; GPU-only verdict runs; Colab-safe.
16
+ Founders share a COMMON-ANCESTOR book so GPA row correspondence is inherited.
17
+
18
+ Colab paste order: geolip_vitals.py -> ar_differentiation_bed.py ->
19
+ exp013_augmentation_bed.py (only for run_b2) -> this file.
20
+ """
21
+ from __future__ import annotations
22
+ import copy
23
+ import json
24
+ import math
25
+ import os
26
+ import torch
27
+ import torch.nn as nn
28
+ import torch.nn.functional as F
29
+
30
+ if "anchor_drift" not in globals():
31
+ try:
32
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
33
+ except ImportError:
34
+ _here = globals().get("__file__")
35
+ if _here is None:
36
+ raise ImportError("paste/run geolip_vitals.py first")
37
+ import sys, pathlib
38
+ sys.path.insert(0, str(pathlib.Path(_here).parent))
39
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
40
+ if "ByteLM" not in globals():
41
+ try:
42
+ from ar_differentiation_bed import (ByteLM, AlephAddress, _wikitext_bytes,
43
+ _batch, VOCAB)
44
+ except ImportError:
45
+ raise ImportError("paste/run ar_differentiation_bed.py first")
46
+
47
+ DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
48
+ EXP_DIR = os.path.join(DATA_ROOT, "exp014")
49
+
50
+
51
+ class SquaredReLU(nn.Module):
52
+ def forward(self, x):
53
+ return F.relu(x) ** 2
54
+
55
+
56
+ # ===================================================== consensus (the germline) ===
57
+ @torch.no_grad()
58
+ def procrustes_rotation(A: torch.Tensor, M: torch.Tensor) -> torch.Tensor:
59
+ """Orthogonal R minimizing ||A R - M||_F (rows correspond)."""
60
+ U, _, Vt = torch.linalg.svd(A.T.double() @ M.double())
61
+ return (U @ Vt).float()
62
+
63
+
64
+ @torch.no_grad()
65
+ def align_to(A: torch.Tensor, ref: torch.Tensor, iters: int = 20) -> torch.Tensor:
66
+ """Projective Procrustes (rows correspond, signs free): alternate the
67
+ orthogonal rotation and per-row sign flips (books live on RP^(D-1))."""
68
+ s = torch.ones(A.shape[0], 1)
69
+ for _ in range(iters):
70
+ R = procrustes_rotation(s * A, ref)
71
+ AR = (s * A) @ R
72
+ s_upd = torch.where((AR * ref).sum(-1, keepdim=True) < 0, -s, s)
73
+ if torch.equal(s_upd, s):
74
+ return AR
75
+ s = s_upd
76
+ return (s * A) @ procrustes_rotation(s * A, ref)
77
+
78
+
79
+ @torch.no_grad()
80
+ def consensus_codebook(books: list, iters: int = 50, tol: float = 1e-8):
81
+ """GPA to mean shape (GM3 machinery, applied to aleph books), anchored to the
82
+ FIRST parent's frame. Rows must correspond (common-ancestor convention); signs
83
+ are projective. Returns (consensus, n_iters, delta)."""
84
+ # device-pin to CPU: parent models may live on CUDA after KD teacher moves
85
+ Bs = [F.normalize(b.detach().float().cpu(), dim=-1).clone() for b in books]
86
+ # pairwise projective alignment to parent-0's frame, THEN GPA refinement
87
+ aligned = [Bs[0]] + [align_to(b, Bs[0]) for b in Bs[1:]]
88
+ M = F.normalize(torch.stack(aligned).mean(0), dim=-1)
89
+ delta, it = 0.0, 0
90
+ for it in range(1, iters + 1):
91
+ aligned = [align_to(b, M) for b in Bs]
92
+ M_new = F.normalize(torch.stack(aligned).mean(0), dim=-1)
93
+ delta = (M_new - M).norm().item()
94
+ M = M_new
95
+ if delta < tol:
96
+ break
97
+ # re-anchor to parent-0 (GPA drift of the global frame stays measurable)
98
+ M = align_to(M, Bs[0])
99
+ return M, it, delta
100
+
101
+
102
+ @torch.no_grad()
103
+ def implant_book(addr: "AlephAddress", book: torch.Tensor, trainable: bool = True):
104
+ """Load a book into an AlephAddress: codebook + home (drift measured from the
105
+ implant). Freeze only via trainable=False AFTER a drift-check justifies it."""
106
+ b = F.normalize(book.float(), dim=-1).to(addr.codebook.device)
107
+ assert b.shape == addr.codebook.shape, (b.shape, addr.codebook.shape)
108
+ addr.codebook.data.copy_(b)
109
+ addr.home.copy_(b)
110
+ addr.codebook.requires_grad_(trainable)
111
+
112
+
113
+ # ============================================================= tree head ==========
114
+ class TreeHead(nn.Module):
115
+ """Autoregressive tree (the constellation-anchor analogue, exp011 TREE operator
116
+ in the healthy consumption regime): a ROOT aleph (K=2, D=4) yields the 4
117
+ oriented weights (2K half-axes = the 4 branches, dense soft, sums to 1);
118
+ each BRANCH is a 64-slot... shared slot projection read by a branch-specific
119
+ book (K=64, D=4); output = branch-weighted mixture of branch reads -> vocab.
120
+ Heritable genome: root book (2,4) + 4 branch books (64,4)."""
121
+
122
+ ROOT_SLOTS = 4 # slot-parallel root consumption (the collapse cure)
123
+ ROOT_TAU = 0.3 # softer root temperature (wave-1 fix: single hard-tau
124
+ # root partially collapsed, usage [.85,.12,.01,.02])
125
+
126
+ def __init__(self, d: int, vocab: int = 256, n_slots: int = 16):
127
+ super().__init__()
128
+ self.n_slots = n_slots
129
+ self.root_proj = nn.Linear(d, self.ROOT_SLOTS * 4, bias=False)
130
+ self.slot_proj = nn.Linear(d, n_slots * 4, bias=False)
131
+ nn.init.orthogonal_(self.root_proj.weight)
132
+ nn.init.orthogonal_(self.slot_proj.weight)
133
+ self.root = AlephAddress(2, 4, tau=self.ROOT_TAU)
134
+ self.branches = nn.ModuleList([AlephAddress(64, 4) for _ in range(4)])
135
+ self.out = nn.Linear(n_slots * 4, vocab, bias=True)
136
+ self._last_root = None
137
+
138
+ def forward(self, h):
139
+ B, n, _ = h.shape
140
+ rs = self.root_proj(h).view(B, n, self.ROOT_SLOTS, 4)
141
+ p, m = self.root.oriented(rs) # (B,n,S,2) x2
142
+ w = torch.cat([p, m], dim=-1).mean(dim=-2) # slot-avg -> (B,n,4)
143
+ self._last_root = w.detach()
144
+ slots = self.slot_proj(h).view(B, n, self.n_slots, 4)
145
+ mix = 0
146
+ for b, br in enumerate(self.branches):
147
+ mix = mix + w[..., b:b + 1] * br.m_hat(slots).reshape(B, n, -1)
148
+ return self.out(mix)
149
+
150
+ def genome(self):
151
+ return {"root": self.root.codebook.detach().clone(),
152
+ **{f"branch{i}": br.codebook.detach().clone()
153
+ for i, br in enumerate(self.branches)}}
154
+
155
+ @torch.no_grad()
156
+ def inherit(self, genomes: list):
157
+ c, it, dl = consensus_codebook([g["root"] for g in genomes])
158
+ implant_book(self.root, c)
159
+ for i, br in enumerate(self.branches):
160
+ c, _, _ = consensus_codebook([g[f"branch{i}"] for g in genomes])
161
+ implant_book(br, c)
162
+
163
+ @torch.no_grad()
164
+ def vitals(self):
165
+ out = {"root_drift": round(anchor_drift(self.root.codebook,
166
+ self.root.home)["mean"], 4)}
167
+ if self._last_root is not None:
168
+ w = self._last_root.reshape(-1, 4)
169
+ usage = w.mean(0)
170
+ usage = usage / usage.sum()
171
+ out["root_usage"] = [round(float(u), 3) for u in usage]
172
+ ent = -(usage.clamp_min(1e-9) * usage.clamp_min(1e-9).log()).sum()
173
+ out["root_ppl4"] = round(float(ent.exp()), 3)
174
+ d = [anchor_drift(br.codebook, br.home)["mean"] for br in self.branches]
175
+ out["branch_drift"] = [round(x, 3) for x in d]
176
+ return out
177
+
178
+
179
+ # ============================================================ organisms ===========
180
+ def make_organism(lineage: str, d: int = 192, layers: int = 4, block: int = 256,
181
+ seed: int = 0):
182
+ """lineage in {aleph_flat, aleph_tree, mlp_kd, no_inherit}. no_inherit uses the
183
+ aleph_flat architecture (the control isolates INHERITANCE, not architecture)."""
184
+ torch.manual_seed(seed)
185
+ if lineage in ("aleph_flat", "no_inherit", "aleph_full", "aleph_weights"):
186
+ return ByteLM("addr_msl64", d=d, layers=layers, block=block)
187
+ if lineage == "aleph_tree":
188
+ m = ByteLM("sdpa", d=d, layers=layers, block=block)
189
+ m.head = TreeHead(d)
190
+ return m
191
+ if lineage == "mlp_kd":
192
+ m = ByteLM("sdpa", d=d, layers=layers, block=block)
193
+ m.head = nn.Sequential(nn.Linear(d, 224), SquaredReLU(),
194
+ nn.LayerNorm(224), nn.Linear(224, VOCAB))
195
+ return m
196
+ raise ValueError(lineage)
197
+
198
+
199
+ def genome_of(model):
200
+ """The heritable organ = co-adapted (projection, book) pair(s). Books inherit
201
+ by CONSENSUS (the geometric germline); projections inherit from the BEST
202
+ parent (weight copy) — implanting a book against a random projection puts the
203
+ child below random init (campaign-v2 lesson)."""
204
+ if isinstance(model.head, TreeHead):
205
+ g = model.head.genome()
206
+ g["root_proj"] = model.head.root_proj.weight.detach().cpu().clone()
207
+ g["slot_proj"] = model.head.slot_proj.weight.detach().cpu().clone()
208
+ return g
209
+ if hasattr(model, "head_addr"):
210
+ return {"flat": model.head_addr.codebook.detach().cpu().clone(),
211
+ "proj": model.head_proj.weight.detach().cpu().clone()}
212
+ return None
213
+
214
+
215
+ @torch.no_grad()
216
+ def inherit_genome(model, genomes: list):
217
+ """genomes[0] = the BEST parent (selection order matters)."""
218
+ if isinstance(model.head, TreeHead):
219
+ model.head.inherit(genomes)
220
+ model.head.root_proj.weight.copy_(genomes[0]["root_proj"].to(
221
+ model.head.root_proj.weight.device))
222
+ model.head.slot_proj.weight.copy_(genomes[0]["slot_proj"].to(
223
+ model.head.slot_proj.weight.device))
224
+ elif hasattr(model, "head_addr"):
225
+ c, it, dl = consensus_codebook([g["flat"] for g in genomes])
226
+ implant_book(model.head_addr, c)
227
+ model.head_proj.weight.copy_(genomes[0]["proj"].to(
228
+ model.head_proj.weight.device))
229
+
230
+
231
+ def organism_vitals(model):
232
+ if isinstance(model.head, TreeHead):
233
+ return model.head.vitals()
234
+ return model.vitals() if hasattr(model, "vitals") else {}
235
+
236
+
237
+ # ========================================================= train one member ======
238
+ def train_member(model, tr, va, steps=2000, batch=32, block=256, device="cuda",
239
+ seed=0, teachers=None, kd_alpha=1.0):
240
+ """CE (+ KL to detached mean teacher probs when teachers given). Pure Adam."""
241
+ g = torch.Generator().manual_seed(seed)
242
+ model = model.to(device)
243
+ opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
244
+ if teachers:
245
+ teachers = [t.to(device).eval() for t in teachers]
246
+ for step in range(1, steps + 1):
247
+ x, y = _batch(tr, batch, block, device, g)
248
+ logits = model(x)
249
+ loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
250
+ if teachers:
251
+ with torch.no_grad():
252
+ tp = torch.stack([F.softmax(t(x), -1) for t in teachers]).mean(0)
253
+ loss = loss + kd_alpha * F.kl_div(
254
+ F.log_softmax(logits, -1), tp, reduction="batchmean")
255
+ opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
256
+ model.eval()
257
+ with torch.no_grad():
258
+ ls = []
259
+ for _ in range(20):
260
+ xv, yv = _batch(va, batch, block, device, g)
261
+ ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
262
+ yv.reshape(-1)).item())
263
+ return sum(ls) / len(ls) / math.log(2) # bpb
264
+
265
+
266
+ # ============================================================ the tournament ======
267
+ def run_tournament(lineage: str, gens: int = 4, pop: int = 4, steps: int = 2000,
268
+ seed: int = 0, device: str = "cuda",
269
+ catastrophic_at: int | None = None):
270
+ """One lineage, one tournament seed. Logs per-gen to the ledger; saves the
271
+ champion genome per generation. catastrophic_at=G injects a 0-step random
272
+ parent into the consensus at generation G (the GM3 robustness probe)."""
273
+ if not torch.cuda.is_available():
274
+ raise RuntimeError("verdict runs are GPU-only")
275
+ os.makedirs(EXP_DIR, exist_ok=True)
276
+ tr, va = _wikitext_bytes(DATA_ROOT)
277
+ ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
278
+ # common ancestor: every founder book starts identical within a tournament
279
+ torch.manual_seed(9000 + seed)
280
+ ancestor = make_organism(lineage, seed=9000 + seed)
281
+ anc_genome = genome_of(ancestor)
282
+ parents, parent_models, champion_genomes = [], [], []
283
+ for gen in range(gens):
284
+ members = []
285
+ for i in range(pop):
286
+ mseed = seed * 1000 + gen * 100 + i
287
+ m = make_organism(lineage, seed=mseed)
288
+ if anc_genome and gen == 0:
289
+ inherit_genome(m, [anc_genome]) # common ancestor
290
+ if gen > 0:
291
+ is_fresh = (i == pop - 1) # gene flow founder
292
+ if not is_fresh:
293
+ if lineage in ("aleph_flat", "aleph_tree"):
294
+ gs = [genome_of(pm) for pm in parent_models]
295
+ if catastrophic_at == gen:
296
+ bad = make_organism(lineage, seed=666 + i)
297
+ gs = gs + [genome_of(bad)]
298
+ inherit_genome(m, gs)
299
+ elif lineage == "aleph_full":
300
+ # v4 arm (v3 lesson: continuity is what pays) — inherit the
301
+ # WHOLE best parent, then overwrite the book with the
302
+ # two-parent consensus: germline ON TOP of continuity.
303
+ m.load_state_dict(copy.deepcopy(
304
+ parent_models[0].state_dict()))
305
+ gs = [genome_of(pm) for pm in parent_models]
306
+ if catastrophic_at == gen:
307
+ bad = make_organism(lineage, seed=666 + i)
308
+ gs = gs + [genome_of(bad)]
309
+ c, _, _ = consensus_codebook([g["flat"] for g in gs])
310
+ implant_book(m.head_addr, c)
311
+ elif lineage in ("mlp_kd", "aleph_weights"):
312
+ # pure continuity (no germline op) — aleph_weights is the
313
+ # within-architecture control for aleph_full
314
+ m.load_state_dict(copy.deepcopy(
315
+ parent_models[0].state_dict()))
316
+ # no_inherit: nothing
317
+ # KD: alpha 0.25 (campaign-v2 lesson: alpha=1.0 from near-parity
318
+ # teachers COMPOUNDS DOWNWARD — inverse evolution; fresh-founder
319
+ # control isolated it). Fresh founders get NO KD (clean gene flow).
320
+ is_fresh_now = (gen > 0 and i == pop - 1)
321
+ teachers = parent_models if (gen > 0 and not is_fresh_now
322
+ and lineage != "no_inherit") else None
323
+ bpb = train_member(m, tr, va, steps=steps, device=device,
324
+ seed=mseed, teachers=teachers, kd_alpha=0.25)
325
+ vit = organism_vitals(m)
326
+ members.append((bpb, m))
327
+ rec = {"exp": "14a", "lineage": lineage, "tseed": seed, "gen": gen,
328
+ "member": i, "fresh": gen > 0 and i == pop - 1,
329
+ "catastrophic": catastrophic_at == gen, "bpb": round(bpb, 4),
330
+ "vitals": vit}
331
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
332
+ print(f"[14a {lineage} t{seed} g{gen} m{i}] bpb={bpb:.4f} {vit}",
333
+ flush=True)
334
+ members.sort(key=lambda t: t[0])
335
+ parent_models = [members[0][1].cpu(), members[1][1].cpu()]
336
+ best = members[0][0]
337
+ gene = genome_of(members[0][1])
338
+ if gene:
339
+ torch.save(gene, os.path.join(
340
+ EXP_DIR, f"champion_{lineage}_t{seed}_g{gen}.pt"))
341
+ champion_genomes.append(gene)
342
+ print(f"[14a {lineage} t{seed} g{gen}] BEST={best:.4f} "
343
+ f"mean={sum(b for b, _ in members)/pop:.4f}", flush=True)
344
+ for _, mm in members[2:]:
345
+ del mm
346
+ torch.cuda.empty_cache()
347
+ ledger.close()
348
+ return best
349
+
350
+
351
+ # ============================================================ 14_b implants ======
352
+ def run_b1(steps: int = 2000, seed: int = 0, device: str = "cuda",
353
+ donor_book: torch.Tensor | None = None, tag: str = "small_cultivated"):
354
+ """Cross-size: donor book (default: cultivate in a small organism) implanted
355
+ into a LARGE organism. Arms: fresh | implant-trainable | implant-frozen | mlp."""
356
+ if not torch.cuda.is_available():
357
+ raise RuntimeError("verdict runs are GPU-only")
358
+ os.makedirs(EXP_DIR, exist_ok=True)
359
+ tr, va = _wikitext_bytes(DATA_ROOT)
360
+ ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
361
+ if donor_book is None:
362
+ small = make_organism("aleph_flat", d=128, layers=2, seed=seed)
363
+ bpb_small = train_member(small, tr, va, steps=steps, device=device, seed=seed)
364
+ donor_book = genome_of(small)["flat"]
365
+ print(f"[14b donor small] bpb={bpb_small:.4f}", flush=True)
366
+ results = {}
367
+ for arm in ("fresh", "implant_train", "implant_frozen", "mlp"):
368
+ lineage = "mlp_kd" if arm == "mlp" else "aleph_flat"
369
+ m = make_organism(lineage, d=384, layers=6, seed=seed + 10)
370
+ if arm.startswith("implant"):
371
+ implant_book(m.head_addr, donor_book, trainable=(arm == "implant_train"))
372
+ bpb = train_member(m, tr, va, steps=steps, device=device, seed=seed + 10)
373
+ vit = organism_vitals(m)
374
+ results[arm] = {"bpb": round(bpb, 4), "vitals": vit}
375
+ rec = {"exp": "14b1", "arm": arm, "donor": tag, "seed": seed,
376
+ "bpb": round(bpb, 4), "vitals": vit}
377
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
378
+ print(f"[14b1 {arm} donor={tag}] bpb={bpb:.4f} {vit}", flush=True)
379
+ del m; torch.cuda.empty_cache()
380
+ ledger.close()
381
+ return results, donor_book
382
+
383
+
384
+ def run_b2(donor_book: torch.Tensor, steps: int = 1500, seed: int = 0,
385
+ device: str = "cuda", tag: str = "small_cultivated"):
386
+ """Cross-architecture: implant the donor book into every GPT-2 relay adapter
387
+ (exp013 Track C bed) vs random-init relays. Books are (64,4) — size-agnostic."""
388
+ from exp013_augmentation_bed import _wikitext_lines
389
+ from transformers import GPT2LMHeadModel, GPT2TokenizerFast
390
+ if "MslRelay" not in globals():
391
+ from ar_differentiation_bed import MslRelay
392
+ from exp013_augmentation_bed import _BlockWithAdapter
393
+ tok = GPT2TokenizerFast.from_pretrained("gpt2")
394
+ tr_lines, va_lines = _wikitext_lines(DATA_ROOT)
395
+ stream_tr = tok("\n\n".join(tr_lines[:8000]), return_tensors="pt").input_ids[0]
396
+ stream_va = tok("\n\n".join(va_lines[:1000]), return_tensors="pt").input_ids[0]
397
+ ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
398
+ out = {}
399
+ for arm in ("random_relays", "implanted_relays"):
400
+ torch.manual_seed(seed)
401
+ g = torch.Generator().manual_seed(seed)
402
+ model = GPT2LMHeadModel.from_pretrained("gpt2").to(device)
403
+ for p in model.parameters():
404
+ p.requires_grad_(False)
405
+ adapters = []
406
+ for i, blk in enumerate(model.transformer.h):
407
+ ad = MslRelay(model.config.n_embd).to(device)
408
+ if arm == "implanted_relays":
409
+ implant_book(ad.addr, donor_book, trainable=True)
410
+ model.transformer.h[i] = _BlockWithAdapter(blk, ad)
411
+ adapters.append(ad)
412
+ params = [p for ad in adapters for p in ad.parameters()
413
+ if p.requires_grad]
414
+ opt = torch.optim.Adam(params, lr=1e-3, weight_decay=0.0)
415
+ block = 256
416
+ for step in range(1, steps + 1):
417
+ ix = torch.randint(0, stream_tr.numel() - block - 1, (8,), generator=g)
418
+ x = torch.stack([stream_tr[i:i + block] for i in ix]).to(device)
419
+ loss = model(x, labels=x).loss
420
+ opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
421
+ model.eval()
422
+ with torch.no_grad():
423
+ ls = []
424
+ for j in range(0, stream_va.numel() - block - 1, block * 4):
425
+ x = stream_va[j:j + block].unsqueeze(0).to(device)
426
+ ls.append(model(x, labels=x).loss.item())
427
+ ppl = math.exp(sum(ls) / len(ls))
428
+ gates = [round(ad.gate.sigmoid().item(), 4) for ad in adapters]
429
+ drifts = [round(anchor_drift(ad.addr.codebook, ad.addr.home)["mean"], 3)
430
+ for ad in adapters]
431
+ out[arm] = {"ppl": round(ppl, 3), "gates": gates, "drift": drifts}
432
+ rec = {"exp": "14b2", "arm": arm, "donor": tag, "seed": seed,
433
+ "ppl": round(ppl, 3), "gates": gates, "drift": drifts}
434
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
435
+ print(f"[14b2 {arm} donor={tag}] ppl={ppl:.3f} gates={gates[:3]}.. "
436
+ f"drift={drifts[:3]}..", flush=True)
437
+ del model; torch.cuda.empty_cache()
438
+ ledger.close()
439
+ return out
440
+
441
+
442
+ # ================================================================ smoke ===========
443
+ def smoke():
444
+ """CPU shapes/parse only: GPA ground truth, tree causality, implant, KD path."""
445
+ g = torch.Generator().manual_seed(0)
446
+ # GPA: two rotated (+row-sign-flipped) copies of one book must align back to it
447
+ A = F.normalize(torch.randn(64, 4, generator=g), dim=-1)
448
+ q, _ = torch.linalg.qr(torch.randn(4, 4, generator=g))
449
+ B = A @ q
450
+ B[::3] = -B[::3]
451
+ C, it, dl = consensus_codebook([A, B])
452
+ cos = (F.normalize(C, dim=-1) * A).sum(-1).abs().mean()
453
+ assert cos > 0.999, cos
454
+ print(f"GPA OK (iters={it}, delta={dl:.2e}, |cos to truth|={cos:.5f})")
455
+ x = torch.randint(0, 256, (2, 64))
456
+ for lineage in ("aleph_flat", "aleph_tree", "mlp_kd", "no_inherit"):
457
+ m = make_organism(lineage, d=96, layers=2, block=64, seed=0)
458
+ lg = m(x); assert lg.shape == (2, 64, 256); lg.sum().backward()
459
+ with torch.no_grad():
460
+ a = m(x)[0, 10]; x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
461
+ b = m(x2)[0, 10]
462
+ assert torch.allclose(a, b, atol=1e-4), lineage + " leaks"
463
+ gnm = genome_of(m)
464
+ if gnm:
465
+ inherit_genome(m, [gnm, gnm]) # self-consensus = identity-ish
466
+ print(lineage, "OK params",
467
+ f"{sum(p.numel() for p in m.parameters()):,}",
468
+ organism_vitals(m) if lineage != "mlp_kd" else {})
469
+ # KD path: teacher forward + KL backward
470
+ t = make_organism("mlp_kd", d=96, layers=2, block=64, seed=1)
471
+ s = make_organism("aleph_flat", d=96, layers=2, block=64, seed=2)
472
+ tp = F.softmax(t(x), -1).detach()
473
+ loss = F.kl_div(F.log_softmax(s(x), -1), tp, reduction="batchmean")
474
+ loss.backward()
475
+ print("KD OK — exp014 smoke passed (tournament on GPU: run_tournament(...))")
476
+
477
+
478
+ def _in_notebook():
479
+ try:
480
+ get_ipython() # type: ignore[name-defined] # noqa: F821
481
+ return True
482
+ except NameError:
483
+ return False
484
+
485
+
486
+ if __name__ == "__main__":
487
+ if _in_notebook():
488
+ smoke()
489
+ print("Notebook: run_tournament('aleph_flat'), run_b1(), run_b2(book).")
490
+ else:
491
+ import argparse
492
+ ap = argparse.ArgumentParser()
493
+ ap.add_argument("--mode", default="smoke",
494
+ choices=["smoke", "tournament", "b1", "b2"])
495
+ ap.add_argument("--lineage", default="aleph_full",
496
+ help="tournament lineage: aleph_flat|aleph_full|"
497
+ "aleph_weights|aleph_tree|mlp_kd|no_inherit")
498
+ ap.add_argument("--seed", type=int, default=0)
499
+ ap.add_argument("--steps", type=int, default=2000)
500
+ ap.add_argument("--genome", default="genomes/champion_aleph_full_t0_g3.pt",
501
+ help="donor genome .pt for --mode b1/b2 (uses its 'flat' book)")
502
+ a, _ = ap.parse_known_args()
503
+ if a.mode == "smoke":
504
+ smoke()
505
+ elif a.mode == "tournament":
506
+ run_tournament(a.lineage, steps=a.steps, seed=a.seed)
507
+ elif a.mode == "b1":
508
+ donor = (torch.load(a.genome, map_location="cpu")["flat"]
509
+ if os.path.exists(a.genome) else None)
510
+ run_b1(steps=a.steps, seed=a.seed, donor_book=donor,
511
+ tag=os.path.basename(a.genome) if donor is not None
512
+ else "small_cultivated")
513
+ elif a.mode == "b2":
514
+ donor = torch.load(a.genome, map_location="cpu")["flat"]
515
+ run_b2(donor, seed=a.seed, tag=os.path.basename(a.genome))
exp019_cr/exp019_content_retention.py ADDED
@@ -0,0 +1,402 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """exp019_content_retention.py — CAPACITY FOR DISTILLED CONTENT RETENTION.
2
+ When content is DISTILLED rather than directly learned, how much is retained,
3
+ through which channel, and for how long? The memory-substrate compass made a
4
+ measurable instrument — and the weak-to-strong distillation cell made real: the
5
+ teacher holds content the student lacks (true headroom on the content axis).
6
+
7
+ THE CONTENT: N fact records "\n@<key6>=<value12>\n" with random alphanumeric
8
+ keys/values — uncompletable from language statistics, so exact-match completion
9
+ IS retention. Facts are mixed into the byte stream (fact-packed blocks at
10
+ FACT_RATE against wikitext blocks).
11
+
12
+ THE TEACHER: the certified bed model (addr_msl64 aleph substrate) trained
13
+ TEACHER_STEPS on the mix; gated on its own recall (the gate doubles as the
14
+ substrate's DIRECT capacity datum at each N).
15
+
16
+ THE CHANNELS (fresh student each, STUDENT_STEPS budget):
17
+ direct — ground-truth CE on the mix (ceiling: learning, not distillation)
18
+ kd_facts — CE on wikitext blocks; on fact blocks the ONLY signal is the
19
+ teacher's logits (KL) — pure distilled content
20
+ kd_general — CE + KL to teacher on CLEAN wikitext only; facts never shown —
21
+ does content leak through logits without exposure?
22
+ book_implant — teacher's farmed codebook implanted (trainable) into a fresh
23
+ student, clean-stream training — do the anchors carry
24
+ byte-content? (the open question from the exp014 implant studies, asked directly)
25
+ none — clean-stream only (floor)
26
+ THE AXES: capacity N in {64, 256, 1024} (main channels); retention = recall
27
+ right after training AND after INTERFERE_STEPS further clean-stream steps
28
+ (the forgetting measurement). KD alpha 1.0 here is LEGAL: the teacher has real
29
+ headroom on the content axis (the inverse-evolution failure was alpha 1.0 at
30
+ NEAR-PARITY — regime, not constant).
31
+
32
+ Preregistered forks:
33
+ F1 capacity curve: direct recall vs N = the substrate's raw content capacity.
34
+ F2 distillation tax: kd_facts vs direct at each N (what survives the logit
35
+ channel).
36
+ F3 leakage: kd_general recall > floor => content crosses on clean text alone.
37
+ F4 anchors: book_implant recall ~ floor => codebooks do not carry byte
38
+ content (mean-shape/content question closed in the direct sense).
39
+ F5 half-life: post-interference retention per channel (does distilled content
40
+ decay faster than learned content?).
41
+ Riders: pure Adam wd=0; GPU-only verdicts; >=2 seeds; Colab-safe. Paste order:
42
+ geolip_vitals -> ar_differentiation_bed -> exp014_genetic_distillation ->
43
+ this file.
44
+ """
45
+ from __future__ import annotations
46
+ import json
47
+ import math
48
+ import os
49
+ import string
50
+ import torch
51
+ import torch.nn.functional as F
52
+
53
+ if "ByteLM" not in globals():
54
+ try:
55
+ from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
56
+ from exp014_genetic_distillation import implant_book
57
+ from geolip_vitals import anchor_drift
58
+ except ImportError:
59
+ _here = globals().get("__file__")
60
+ if _here is None:
61
+ raise ImportError("paste geolip_vitals + ar_differentiation_bed + "
62
+ "exp014_genetic_distillation first")
63
+ import sys, pathlib
64
+ sys.path.insert(0, str(pathlib.Path(_here).parent))
65
+ from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
66
+ from exp014_genetic_distillation import implant_book
67
+ from geolip_vitals import anchor_drift
68
+
69
+ DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
70
+ EXP19_DIR = os.path.join(DATA_ROOT, "exp019")
71
+
72
+ KEY_LEN, VAL_LEN = 6, 12
73
+ FACT_RATE = 0.5 # fraction of training blocks drawn from fact stream
74
+ TEACHER_STEPS = 4000
75
+ STUDENT_STEPS = 2000
76
+ INTERFERE_STEPS = 1000
77
+ ALNUM = (string.ascii_lowercase + string.digits).encode()
78
+
79
+
80
+ def make_facts(n: int, seed: int = 0):
81
+ """N records '\\n@<key>=<value>\\n'; returns (records list, fact byte stream)."""
82
+ g = torch.Generator().manual_seed(4000 + seed)
83
+ recs = []
84
+ seen = set()
85
+ while len(recs) < n:
86
+ k = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (KEY_LEN,),
87
+ generator=g))
88
+ if k in seen:
89
+ continue
90
+ seen.add(k)
91
+ v = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (VAL_LEN,),
92
+ generator=g))
93
+ recs.append((k, v))
94
+ return recs
95
+
96
+
97
+ def fact_stream(recs, copies: int = 50, seed: int = 0) -> torch.Tensor:
98
+ """Byte stream of shuffled fact records (each record appears `copies` times)."""
99
+ g = torch.Generator().manual_seed(5000 + seed)
100
+ order = torch.cat([torch.randperm(len(recs), generator=g)
101
+ for _ in range(copies)])
102
+ blob = b"".join(b"\n@" + recs[i][0] + b"=" + recs[i][1] + b"\n"
103
+ for i in order.tolist())
104
+ return torch.frombuffer(bytearray(blob), dtype=torch.uint8).clone()
105
+
106
+
107
+ def _mix_batch(tr, fs, batch, block, device, g):
108
+ """Blocks drawn from the fact stream with prob FACT_RATE, else wikitext.
109
+ Returns (x, y, fact_mask (B,)) — mask marks fact-sourced rows."""
110
+ xw, yw = _batch(tr, batch, block, device, g)
111
+ xf, yf = _batch(fs, batch, block, device, g)
112
+ m = (torch.rand(batch, generator=g) < FACT_RATE).to(device)
113
+ x = torch.where(m[:, None], xf, xw)
114
+ y = torch.where(m[:, None], yf, yw)
115
+ return x, y, m
116
+
117
+
118
+ @torch.no_grad()
119
+ def recall(model, recs, device="cuda", max_eval: int = 256,
120
+ batch: int = 64) -> dict:
121
+ """Exact-match greedy completion: prompt '\\n@<key>=' -> VAL_LEN bytes."""
122
+ model = model.to(device).eval()
123
+ recs = recs[:max_eval]
124
+ prompts = torch.stack([torch.frombuffer(
125
+ bytearray(b"\n@" + k + b"="), dtype=torch.uint8).long()
126
+ for k, _ in recs]).to(device)
127
+ outs = []
128
+ for i in range(0, len(recs), batch):
129
+ x = prompts[i:i + batch]
130
+ for _ in range(VAL_LEN):
131
+ nxt = model(x)[:, -1].argmax(-1, keepdim=True)
132
+ x = torch.cat([x, nxt], dim=1)
133
+ outs.append(x[:, -VAL_LEN:].cpu())
134
+ got = torch.cat(outs)
135
+ tgt = torch.stack([torch.frombuffer(bytearray(v), dtype=torch.uint8).long()
136
+ for _, v in recs])
137
+ byte_acc = (got == tgt).float().mean().item()
138
+ exact = (got == tgt).all(dim=1).float().mean().item()
139
+ return {"exact": round(exact, 4), "byte_acc": round(byte_acc, 4)}
140
+
141
+
142
+ def train_stream(model, tr, va, fs=None, teacher=None, channel="direct",
143
+ steps=2000, batch=32, block=256, device="cuda", seed=0):
144
+ """One training run under a channel's signal routing (docstring above)."""
145
+ g = torch.Generator().manual_seed(seed)
146
+ model = model.to(device)
147
+ if teacher is not None:
148
+ teacher = teacher.to(device).eval()
149
+ opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
150
+ for step in range(1, steps + 1):
151
+ if channel in ("direct", "kd_facts") and fs is not None:
152
+ x, y, m = _mix_batch(tr, fs, batch, block, device, g)
153
+ else: # clean wikitext stream
154
+ x, y = _batch(tr, batch, block, device, g)
155
+ m = torch.zeros(batch, dtype=torch.bool, device=device)
156
+ logits = model(x)
157
+ if channel == "kd_facts":
158
+ # ground truth on wiki rows only; teacher logits are the ONLY
159
+ # signal on fact rows (pure distilled content)
160
+ ce_rows = ~m
161
+ loss = torch.tensor(0.0, device=device)
162
+ if ce_rows.any():
163
+ loss = F.cross_entropy(logits[ce_rows].reshape(-1, VOCAB),
164
+ y[ce_rows].reshape(-1))
165
+ if m.any():
166
+ with torch.no_grad():
167
+ tp = F.softmax(teacher(x[m]), -1)
168
+ loss = loss + F.kl_div(F.log_softmax(logits[m], -1), tp,
169
+ reduction="batchmean")
170
+ elif channel == "kd_general":
171
+ loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
172
+ with torch.no_grad():
173
+ tp = F.softmax(teacher(x), -1)
174
+ loss = loss + F.kl_div(F.log_softmax(logits, -1), tp,
175
+ reduction="batchmean")
176
+ else: # direct / book_implant / none
177
+ loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
178
+ opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
179
+ model.eval()
180
+ with torch.no_grad():
181
+ ls = []
182
+ for _ in range(10):
183
+ xv, yv = _batch(va, batch, block, device, g)
184
+ ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
185
+ yv.reshape(-1)).item())
186
+ return sum(ls) / len(ls) / math.log(2)
187
+
188
+
189
+ def run_retention(n_facts=(64, 256, 1024), seeds=(0, 1), device="cuda"):
190
+ if not torch.cuda.is_available():
191
+ raise RuntimeError("verdict runs are GPU-only")
192
+ os.makedirs(EXP19_DIR, exist_ok=True)
193
+ tr, va = _wikitext_bytes(DATA_ROOT)
194
+ ledger = open(os.path.join(EXP19_DIR, "ledger.jsonl"), "a", encoding="utf-8")
195
+
196
+ def log(rec):
197
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
198
+ print(f"[19 {rec['channel']} N={rec['n']} s{rec['seed']}] "
199
+ f"recall={rec['recall']} after_interf={rec.get('recall_interf')} "
200
+ f"bpb={rec['bpb']}", flush=True)
201
+
202
+ for seed in seeds:
203
+ for n in n_facts:
204
+ recs = make_facts(n, seed=seed)
205
+ fs = fact_stream(recs, seed=seed)
206
+ # ---- teacher (also the DIRECT capacity datum at TEACHER_STEPS)
207
+ torch.manual_seed(seed)
208
+ teacher = ByteLM("addr_msl64")
209
+ t_bpb = train_stream(teacher, tr, va, fs=fs, channel="direct",
210
+ steps=TEACHER_STEPS, device=device, seed=seed)
211
+ t_rec = recall(teacher, recs, device=device)
212
+ log({"exp": "19", "channel": "teacher", "n": n, "seed": seed,
213
+ "steps": TEACHER_STEPS, "recall": t_rec, "bpb": round(t_bpb, 4)})
214
+ torch.save({"n": n, "seed": seed,
215
+ "state_dict": {k: v.cpu() for k, v in
216
+ teacher.state_dict().items()}},
217
+ os.path.join(EXP19_DIR, f"teacher_N{n}_s{seed}.pt"))
218
+ # ---- channels
219
+ chans = ["direct", "kd_facts", "kd_general"]
220
+ if n == 256:
221
+ chans += ["book_implant", "none"]
222
+ for ch in chans:
223
+ torch.manual_seed(1000 + seed)
224
+ m = ByteLM("addr_msl64")
225
+ if ch == "book_implant":
226
+ implant_book(m.head_addr,
227
+ teacher.head_addr.codebook.detach().cpu())
228
+ bpb = train_stream(
229
+ m, tr, va, fs=fs if ch in ("direct", "kd_facts") else None,
230
+ teacher=teacher if ch.startswith("kd") else None,
231
+ channel=ch, steps=STUDENT_STEPS, device=device,
232
+ seed=1000 + seed)
233
+ r0 = recall(m, recs, device=device)
234
+ # retention under interference: further CLEAN-stream training
235
+ bpb2 = train_stream(m, tr, va, channel="none",
236
+ steps=INTERFERE_STEPS, device=device,
237
+ seed=2000 + seed)
238
+ r1 = recall(m, recs, device=device)
239
+ log({"exp": "19", "channel": ch, "n": n, "seed": seed,
240
+ "steps": STUDENT_STEPS, "recall": r0, "recall_interf": r1,
241
+ "bpb": round(bpb, 4), "bpb_after_interf": round(bpb2, 4)})
242
+ del m
243
+ torch.cuda.empty_cache()
244
+ del teacher
245
+ torch.cuda.empty_cache()
246
+ ledger.close()
247
+
248
+
249
+ # ==================== exp019b — GENERALIZATION block =========================
250
+ # Rule-bearing content: value = fixed random substitution cipher applied to the
251
+ # key, extended to VAL_LEN (v[i] = subst(k[i % KEY_LEN])). Teacher sees
252
+ # N_TRAIN rule-keys; N_TEST keys are HELD OUT. Held-out recall = the RULE
253
+ # generalizing, not the list. Sharp question: does the logit channel transfer
254
+ # the rule better than it transfers the rote list? Plus prompt-format variants
255
+ # (content vs surface form disentangled).
256
+
257
+ def make_rule_facts(n_train: int = 256, n_test: int = 128, seed: int = 0):
258
+ g = torch.Generator().manual_seed(6000 + seed)
259
+ subst = {ALNUM[i]: ALNUM[j] for i, j in
260
+ enumerate(torch.randperm(len(ALNUM), generator=g).tolist())}
261
+ keys, seen = [], set()
262
+ while len(keys) < n_train + n_test:
263
+ k = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (KEY_LEN,),
264
+ generator=g))
265
+ if k not in seen:
266
+ seen.add(k)
267
+ keys.append(k)
268
+ def val(k):
269
+ return bytes(subst[k[i % KEY_LEN]] for i in range(VAL_LEN))
270
+ train = [(k, val(k)) for k in keys[:n_train]]
271
+ test = [(k, val(k)) for k in keys[n_train:]]
272
+ return train, test
273
+
274
+
275
+ @torch.no_grad()
276
+ def recall_fmt(model, recs, fmt: bytes = b"\n@%s=", device="cuda",
277
+ max_eval: int = 256, batch: int = 64) -> dict:
278
+ """recall() under an arbitrary prompt format (b'\\n@%s=' = the training
279
+ format; variants probe surface-form generalization)."""
280
+ model = model.to(device).eval()
281
+ recs = recs[:max_eval]
282
+ proms = [torch.frombuffer(bytearray(fmt.replace(b"%s", k)),
283
+ dtype=torch.uint8).long() for k, _ in recs]
284
+ L = max(p.numel() for p in proms)
285
+ # left-pad with newlines to equal length (causal — padding is prefix noise)
286
+ prompts = torch.stack([torch.cat([torch.full((L - p.numel(),), 10,
287
+ dtype=torch.long), p])
288
+ for p in proms]).to(device)
289
+ outs = []
290
+ for i in range(0, len(recs), batch):
291
+ x = prompts[i:i + batch]
292
+ for _ in range(VAL_LEN):
293
+ nxt = model(x)[:, -1].argmax(-1, keepdim=True)
294
+ x = torch.cat([x, nxt], dim=1)
295
+ outs.append(x[:, -VAL_LEN:].cpu())
296
+ got = torch.cat(outs)
297
+ tgt = torch.stack([torch.frombuffer(bytearray(v), dtype=torch.uint8).long()
298
+ for _, v in recs])
299
+ return {"exact": round((got == tgt).all(dim=1).float().mean().item(), 4),
300
+ "byte_acc": round((got == tgt).float().mean().item(), 4)}
301
+
302
+
303
+ FMT_TRAIN = b"\n@%s="
304
+ FMT_VARIANT = b" @%s= " # never seen in training: pure format shift
305
+
306
+
307
+ def run_generalization(n_train: int = 256, n_test: int = 128, seeds=(0, 1),
308
+ device="cuda"):
309
+ if not torch.cuda.is_available():
310
+ raise RuntimeError("verdict runs are GPU-only")
311
+ os.makedirs(EXP19_DIR, exist_ok=True)
312
+ tr, va = _wikitext_bytes(DATA_ROOT)
313
+ ledger = open(os.path.join(EXP19_DIR, "ledger.jsonl"), "a", encoding="utf-8")
314
+
315
+ def gauges(model, train_recs, test_recs):
316
+ return {"train": recall_fmt(model, train_recs, FMT_TRAIN, device=device),
317
+ "heldout": recall_fmt(model, test_recs, FMT_TRAIN, device=device),
318
+ "train_varfmt": recall_fmt(model, train_recs, FMT_VARIANT,
319
+ device=device)}
320
+
321
+ for seed in seeds:
322
+ train_recs, test_recs = make_rule_facts(n_train, n_test, seed=seed)
323
+ fs = fact_stream(train_recs, seed=seed) # held-out NEVER streamed
324
+ torch.manual_seed(seed)
325
+ teacher = ByteLM("addr_msl64")
326
+ t_bpb = train_stream(teacher, tr, va, fs=fs, channel="direct",
327
+ steps=TEACHER_STEPS, device=device, seed=seed)
328
+ gt = gauges(teacher, train_recs, test_recs)
329
+ rec = {"exp": "19b", "channel": "teacher", "n": n_train, "seed": seed,
330
+ "steps": TEACHER_STEPS, "gauges": gt, "bpb": round(t_bpb, 4)}
331
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
332
+ print(f"[19b teacher s{seed}] {gt} bpb={t_bpb:.4f}", flush=True)
333
+ torch.save({"seed": seed, "state_dict": {k: v.cpu() for k, v in
334
+ teacher.state_dict().items()}},
335
+ os.path.join(EXP19_DIR, f"rule_teacher_s{seed}.pt"))
336
+ for ch in ("direct", "kd_facts", "kd_general"):
337
+ torch.manual_seed(1000 + seed)
338
+ m = ByteLM("addr_msl64")
339
+ bpb = train_stream(
340
+ m, tr, va, fs=fs if ch in ("direct", "kd_facts") else None,
341
+ teacher=teacher if ch.startswith("kd") else None,
342
+ channel=ch, steps=STUDENT_STEPS, device=device, seed=1000 + seed)
343
+ g0 = gauges(m, train_recs, test_recs)
344
+ bpb2 = train_stream(m, tr, va, channel="none",
345
+ steps=INTERFERE_STEPS, device=device,
346
+ seed=2000 + seed)
347
+ g1 = gauges(m, train_recs, test_recs)
348
+ rec = {"exp": "19b", "channel": ch, "n": n_train, "seed": seed,
349
+ "steps": STUDENT_STEPS, "gauges": g0, "gauges_interf": g1,
350
+ "bpb": round(bpb, 4), "bpb_after_interf": round(bpb2, 4)}
351
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
352
+ print(f"[19b {ch} s{seed}] {g0} interf_heldout="
353
+ f"{g1['heldout']} bpb={bpb:.4f}", flush=True)
354
+ torch.save({"channel": ch, "seed": seed,
355
+ "state_dict": {k: v.cpu() for k, v in
356
+ m.state_dict().items()}},
357
+ os.path.join(EXP19_DIR, f"rule_{ch}_s{seed}.pt"))
358
+ del m
359
+ torch.cuda.empty_cache()
360
+ del teacher
361
+ torch.cuda.empty_cache()
362
+ ledger.close()
363
+
364
+
365
+ def smoke():
366
+ recs = make_facts(8, seed=0)
367
+ assert len(recs) == 8 and all(len(k) == KEY_LEN and len(v) == VAL_LEN
368
+ for k, v in recs)
369
+ fs = fact_stream(recs, copies=3, seed=0)
370
+ assert fs.dtype == torch.uint8 and fs.numel() == 3 * 8 * (KEY_LEN + VAL_LEN + 4)
371
+ m = ByteLM("addr_msl64", d=96, layers=2, block=64)
372
+ r = recall(m, recs, device="cpu", max_eval=8, batch=4)
373
+ assert 0.0 <= r["exact"] <= 1.0 and 0.0 <= r["byte_acc"] <= 1.0
374
+ g = torch.Generator().manual_seed(0)
375
+ x, y, mask = _mix_batch(torch.randint(0, 256, (50000,),
376
+ dtype=torch.uint8, generator=g),
377
+ fs, 8, 64, "cpu", g)
378
+ assert x.shape == (8, 64) and mask.shape == (8,)
379
+ assert (x[:, 1:] == y[:, :-1]).all() # stream alignment
380
+ # 19b: rule facts are rule-consistent + disjoint; variant recall runs
381
+ tr8, te4 = make_rule_facts(8, 4, seed=0)
382
+ assert len(tr8) == 8 and len(te4) == 4
383
+ assert not set(k for k, _ in tr8) & set(k for k, _ in te4)
384
+ k0, v0 = tr8[0]
385
+ assert len(v0) == VAL_LEN and v0[:KEY_LEN] == v0[KEY_LEN:2 * KEY_LEN]
386
+ rv = recall_fmt(m, tr8, FMT_VARIANT, device="cpu", max_eval=8, batch=4)
387
+ assert 0.0 <= rv["exact"] <= 1.0
388
+ print(f"exp019 smoke passed (untrained recall exact={r['exact']} "
389
+ f"byte={r['byte_acc']} ~ chance; 19b rule+variant OK)")
390
+
391
+
392
+ def _in_notebook():
393
+ try:
394
+ get_ipython() # type: ignore[name-defined] # noqa: F821
395
+ return True
396
+ except NameError:
397
+ return False
398
+
399
+
400
+ if __name__ == "__main__":
401
+ smoke() if not _in_notebook() else (smoke(),
402
+ print("Notebook: run_retention() on GPU."))
exp019_cr/geolip_vitals.py ADDED
@@ -0,0 +1,219 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """geolip_vitals.py — shared diagnostic harness for the GeoLIP aleph experiments.
2
+ ALL functions are READOUTS: no gradients, no losses. CV is a readout, never a
3
+ force. Addressing is judged by drift->0.29154 and CV->0.20, never by recon cosine
4
+ (judgment criteria per the aleph-void article: https://huggingface.co/blog/AbstractPhil/geometric-vocabulary-patchwork-aleph-void).
5
+
6
+ Vitals provided:
7
+ anchor_drift — geodesic drift of anchors from init; binding fraction @0.29154
8
+ pentachoron_cv — CM 4-volume CV over random 5-row subsets (geovocab2 import)
9
+ axis_aliveness — oriented-address usage: axes alive, hppl, collapse flag
10
+ gate_stats — gate means vs the 0.012-0.03 band
11
+ path_diversity — unique-path counting, FIXED high-bits hash (low-16 bug is the
12
+ retracted artifact — never use the low bits)
13
+ grad_norm_spread — gradient democracy monitor (orders-of-magnitude spread)
14
+ CVScreen — CV@1000-batch early band screen (<0.30 LOW / .35-.50 MID / >.80 HIGH)
15
+
16
+ Smoke on a torch-capable env: python geolip_vitals.py
17
+ """
18
+ from __future__ import annotations
19
+ import math
20
+ import torch
21
+
22
+ BINDING = 0.29154 # radians; the binding/separation constant
23
+ CV_BAND = (0.13, 0.30) # CM CV band (discovery_catalog #4)
24
+ GATE_BAND = (0.012, 0.03) # live invariant candidate (acd_campaign)
25
+ KNUTH32 = 2654435761
26
+
27
+
28
+ # ----------------------------------------------------------------------------- drift
29
+ @torch.no_grad()
30
+ def anchor_drift(current: torch.Tensor, init: torch.Tensor, tol: float = 0.05) -> dict:
31
+ """Geodesic drift (radians) of each row of `current` from its row in `init`,
32
+ both row-normalized. Returns mean/std/per-row drift and the fraction of rows
33
+ within +/-tol of BINDING (the GLFM '46%' readout)."""
34
+ a = torch.nn.functional.normalize(current.float(), dim=-1)
35
+ b = torch.nn.functional.normalize(init.float(), dim=-1)
36
+ cos = (a * b).sum(-1).clamp(-1.0, 1.0)
37
+ drift = torch.arccos(cos)
38
+ frac = ((drift - BINDING).abs() <= tol).float().mean()
39
+ return {"mean": drift.mean().item(), "std": drift.std().item(),
40
+ "per_row": drift, "binding_fraction": frac.item()}
41
+
42
+
43
+ # -------------------------------------------------------------------------------- cv
44
+ @torch.no_grad()
45
+ def _pentachoron_volumes(pts: torch.Tensor) -> torch.Tensor:
46
+ """Batched Cayley-Menger 4-simplex volumes. pts: (B, 5, D) -> (B,) volumes.
47
+ One float64 det over all samples (vol^2 = -det(CM)/9216 for n=4). Built-in
48
+ for speed (the per-sample reference path is ~260x slower in a vitals loop);
49
+ geovocab2 remains the formula's reference implementation, parity-checked
50
+ via cv_reference_check()."""
51
+ B = pts.shape[0]
52
+ d2 = torch.cdist(pts.double(), pts.double()).pow(2) # (B,5,5)
53
+ cm = torch.ones(B, 6, 6, dtype=torch.float64, device=pts.device)
54
+ cm[:, 0, 0] = 0.0
55
+ cm[:, 1:, 1:] = d2
56
+ det = torch.linalg.det(cm)
57
+ return (-det / 9216.0).clamp_min(0.0).sqrt().float()
58
+
59
+
60
+ @torch.no_grad()
61
+ def pentachoron_cv(rows: torch.Tensor, n_samples: int = 200,
62
+ generator: torch.Generator | None = None) -> float:
63
+ """CV (std/mean) of Cayley-Menger 4-simplex volumes over n_samples random
64
+ 5-row subsets. Rows are row-normalized before measurement. Uses the built-in
65
+ batched CM (float64 det); validate against geovocab2 with
66
+ cv_reference_check() after any change to the volume math."""
67
+ x = torch.nn.functional.normalize(rows.float(), dim=-1)
68
+ n = x.shape[0]
69
+ if n < 5:
70
+ raise ValueError(f"pentachoron_cv needs >=5 rows, got {n}")
71
+ g = generator or torch.Generator(device="cpu").manual_seed(0)
72
+ idx = torch.stack([torch.randperm(n, generator=g)[:5]
73
+ for _ in range(n_samples)]) # (B,5)
74
+ v = _pentachoron_volumes(x[idx].cpu())
75
+ return (v.std() / v.mean().clamp_min(1e-12)).item()
76
+
77
+
78
+ @torch.no_grad()
79
+ def cv_reference_check(n_trials: int = 50, tol: float = 1e-5) -> float:
80
+ """Parity check of the built-in batched CM against geovocab2's reference
81
+ implementation (the formula's source of truth). Returns max |rel diff|;
82
+ raises if geovocab2 is absent or parity fails. Run after touching
83
+ _pentachoron_volumes."""
84
+ try:
85
+ from geovocab2.shapes.formula.symbolic.cayley_menger import (
86
+ CayleyMengerFromSimplex)
87
+ except Exception as e: # pragma: no cover
88
+ raise ImportError(
89
+ "cv_reference_check requires geovocab2 (install via the geolip-svae "
90
+ "umbrella: pip install git+https://github.com/AbstractEyes/"
91
+ "geolip-svae).") from e
92
+ ref = CayleyMengerFromSimplex()
93
+ g = torch.Generator().manual_seed(0)
94
+ pts = torch.nn.functional.normalize(
95
+ torch.randn(n_trials, 5, 4, generator=g), dim=-1)
96
+ mine = _pentachoron_volumes(pts)
97
+ # compare at float64: the reference computes in the INPUT dtype, and fp32
98
+ # dets lose up to ~4% on near-degenerate pentachora (measured 2026-07-11)
99
+ theirs = torch.stack([ref.forward(p.double())["volume"].float() for p in pts])
100
+ rel = ((mine - theirs).abs() / theirs.abs().clamp_min(1e-12)).max().item()
101
+ if rel > tol:
102
+ raise AssertionError(f"CM parity vs geovocab2 failed: max rel {rel}")
103
+ return rel
104
+
105
+
106
+ # ------------------------------------------------------------------------- aliveness
107
+ @torch.no_grad()
108
+ def axis_aliveness(oriented_weights: torch.Tensor, alive_thresh: float = 1e-3) -> dict:
109
+ """`oriented_weights`: (..., 2K) nonnegative oriented-softmax address rows
110
+ (sum to 1 on the last dim). Returns axes-alive count, mean-usage perplexity
111
+ (hppl analogue; healthy hosted reference 125-126/128), and a collapse flag.
112
+ Reference behavior: near-uniform aliveness at div_weight=0 (discovery #22)."""
113
+ w = oriented_weights.reshape(-1, oriented_weights.shape[-1]).float()
114
+ usage = w.mean(0)
115
+ usage = usage / usage.sum().clamp_min(1e-12)
116
+ # an axis is alive if its mean usage exceeds alive_thresh x the uniform share
117
+ alive = int((usage > alive_thresh * (1.0 / usage.numel())).sum())
118
+ ent = -(usage.clamp_min(1e-12) * usage.clamp_min(1e-12).log()).sum()
119
+ ppl = float(ent.exp())
120
+ return {"axes_total": usage.numel(), "axes_alive": alive, "usage_ppl": ppl,
121
+ "collapsed": ppl < 0.05 * usage.numel()}
122
+
123
+
124
+ # ------------------------------------------------------------------------------ gates
125
+ @torch.no_grad()
126
+ def gate_stats(gates: torch.Tensor) -> dict:
127
+ """Gate values (post-sigmoid/clamp). Reports mean and whether it sits in the
128
+ 0.012-0.03 band (read-only — the band is a candidate invariant, never a target)."""
129
+ g = gates.float().flatten()
130
+ m = g.mean().item()
131
+ return {"mean": m, "std": g.std().item(),
132
+ "in_band": GATE_BAND[0] <= m <= GATE_BAND[1]}
133
+
134
+
135
+ # ------------------------------------------------------------------------------ paths
136
+ @torch.no_grad()
137
+ def path_diversity(ids: torch.Tensor) -> dict:
138
+ """Unique-path counting with the FIXED multiplicative hash:
139
+ ((ids * 2654435761) % 2^32) >> 16 — Knuth needs the HIGH bits; the low-16
140
+ variant produced a retracted ~1,500 path ceiling in a prior campaign.
141
+ `ids`: integer tensor, one composed path id per row (any shape)."""
142
+ x = ids.reshape(-1).to(torch.int64)
143
+ hashed = ((x * KNUTH32) % (1 << 32)) >> 16
144
+ return {"n": int(x.numel()),
145
+ "unique_raw": int(torch.unique(x).numel()),
146
+ "unique_hashed": int(torch.unique(hashed).numel())}
147
+
148
+
149
+ @torch.no_grad()
150
+ def compose_path_ids(stage_indices: list[torch.Tensor], radix: int) -> torch.Tensor:
151
+ """Compose per-stage discrete indices (each (...,) int in [0, radix)) into a
152
+ single path id, positional base-`radix` — construction, not hashing."""
153
+ out = torch.zeros_like(stage_indices[0], dtype=torch.int64)
154
+ for s in stage_indices:
155
+ out = out * radix + s.to(torch.int64)
156
+ return out
157
+
158
+
159
+ # --------------------------------------------------------------------- grad democracy
160
+ @torch.no_grad()
161
+ def grad_norm_spread(groups: dict[str, list[torch.nn.Parameter]]) -> dict:
162
+ """Gradient-democracy monitor. `groups`: name -> params of one parallel member
163
+ (tower/expert). Reports per-group grad norms and the orders-of-magnitude spread.
164
+ Reference: unequalized heterogeneous towers spread ~20 orders (fibonacci dead at
165
+ 2.25e-21 under helix); equalized ~0.0 (geofractal gradient-democracy result)."""
166
+ norms = {}
167
+ for name, params in groups.items():
168
+ gs = [p.grad for p in params if p.grad is not None]
169
+ norms[name] = float(torch.sqrt(sum((g.float() ** 2).sum() for g in gs)).item()) \
170
+ if gs else 0.0
171
+ vals = [v for v in norms.values() if v > 0]
172
+ spread = (math.log10(max(vals)) - math.log10(min(vals))) if len(vals) >= 2 else 0.0
173
+ return {"norms": norms, "spread_orders": spread, "dead": [k for k, v in norms.items() if v == 0.0]}
174
+
175
+
176
+ # ----------------------------------------------------------------------------- screen
177
+ class CVScreen:
178
+ """CV@N early band screen (tri-band ft1): record pentachoron CV at `step_mark`
179
+ batches; classify <0.30 LOW / 0.35-0.50 MID / >0.80 HIGH. Turns ~2h/config
180
+ into ~7min. Readout only."""
181
+ def __init__(self, step_mark: int = 1000):
182
+ self.step_mark = step_mark
183
+ self.recorded: float | None = None
184
+
185
+ def maybe_record(self, step: int, rows: torch.Tensor) -> float | None:
186
+ if self.recorded is None and step >= self.step_mark:
187
+ self.recorded = pentachoron_cv(rows)
188
+ return self.recorded
189
+
190
+ @property
191
+ def band(self) -> str | None:
192
+ c = self.recorded
193
+ if c is None:
194
+ return None
195
+ if c < 0.30:
196
+ return "LOW"
197
+ if 0.35 <= c <= 0.50:
198
+ return "MID"
199
+ if c > 0.80:
200
+ return "HIGH"
201
+ return "BETWEEN"
202
+
203
+
204
+ # ------------------------------------------------------------------------------ smoke
205
+ if __name__ == "__main__": # shapes/parse smoke ONLY — no training, ever.
206
+ g = torch.Generator().manual_seed(0)
207
+ K, D = 64, 4
208
+ init = torch.nn.functional.normalize(torch.randn(K, D, generator=g), dim=-1)
209
+ cur = torch.nn.functional.normalize(init + 0.29 * torch.randn(K, D, generator=g), dim=-1)
210
+ print("drift:", {k: v for k, v in anchor_drift(cur, init).items() if k != "per_row"})
211
+ w = torch.softmax(torch.randn(32, 2 * K, generator=g), dim=-1)
212
+ print("aliveness:", axis_aliveness(w))
213
+ print("gates:", gate_stats(torch.full((8,), 0.024)))
214
+ ids = compose_path_ids([torch.randint(0, 16, (4096,), generator=g) for _ in range(4)], 16)
215
+ print("paths:", path_diversity(ids))
216
+ lin = torch.nn.Linear(8, 8)
217
+ lin(torch.randn(4, 8)).sum().backward()
218
+ print("democracy:", grad_norm_spread({"a": list(lin.parameters())}))
219
+ print("OK — vitals smoke passed (pentachoron_cv needs geovocab2; run on GPU env)")
exp019_cr/read_codebook.py ADDED
@@ -0,0 +1,159 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """read_codebook.py — projective reading of cultivated aleph codebooks.
2
+ Antipodal-collapse extraction on trained codebooks + projective statistics on
3
+ RP^(D-1) — applied to the exp012 AR-bed specimens.
4
+
5
+ Recipe per the Polygonal Omega article (geometric-tri-band-ft2): collapse = (row_i - row_j)/2 normalized for each MUTUAL-STRONGEST
6
+ pair with cos < -0.9 — "a deterministic tensor operation," not clustering.
7
+ Projective metric ALWAYS arccos|<a,b>| (metric-alignment rule, reading-voids-ft1).
8
+ D=4 scope is the validated regime (D=5 walked back; axis count grows with D).
9
+
10
+ Readouts per specimen:
11
+ pairs / n_axes / unpaired — antipodal structure
12
+ proj_angle mean vs uniform baseline, deviation — near-uniform RP^(D-1)?
13
+ drift from home + binding fraction @0.29154 — cultivation record
14
+ erank of the axis set — spectral occupancy
15
+ verdict: PROJECTIVE-CLEAN (|dev|<0.05, util>0.95, secondary pairs<=3) /
16
+ -MOSTLY / STRUCTURED / DEGENERATE (per Polygonal Omega thresholds)
17
+
18
+ Usage (terminal): python read_codebook.py <ckpt_or_dir> [more paths...]
19
+ Colab: paste geolip_vitals.py cell first (optional), then this file, then
20
+ read_all(r"/content/data/ar_ckpts").
21
+ """
22
+ from __future__ import annotations
23
+ import math
24
+ import sys
25
+ import torch
26
+ import torch.nn.functional as F
27
+
28
+ BINDING = 0.29154
29
+
30
+
31
+ @torch.no_grad()
32
+ def antipodal_collapse(codebook: torch.Tensor, thresh: float = -0.9) -> dict:
33
+ """Mutual-strongest antipodal pairing + collapse to axes on RP^(D-1)."""
34
+ A = F.normalize(codebook.float(), dim=-1)
35
+ K = A.shape[0]
36
+ cos = A @ A.T
37
+ cos.fill_diagonal_(2.0) # exclude self from minima
38
+ nearest_neg = cos.argmin(dim=-1) # most-antipodal partner
39
+ pairs = []
40
+ used = set()
41
+ for i in range(K):
42
+ j = int(nearest_neg[i])
43
+ if i < j and int(nearest_neg[j]) == i and cos[i, j] < thresh:
44
+ pairs.append((i, j))
45
+ used.update((i, j))
46
+ axes = [F.normalize((A[i] - A[j]) / 2.0, dim=-1) for i, j in pairs]
47
+ axes += [A[i] for i in range(K) if i not in used] # unpaired rows as axes
48
+ axes = torch.stack(axes) if axes else A[:0]
49
+ # sign-canon onto RP: first nonzero coordinate positive
50
+ for r in range(axes.shape[0]):
51
+ nz = torch.nonzero(axes[r].abs() > 1e-8)
52
+ if nz.numel() and axes[r, nz[0, 0]] < 0:
53
+ axes[r] = -axes[r]
54
+ return {"pairs": len(pairs), "n_axes": axes.shape[0],
55
+ "unpaired": K - 2 * len(pairs), "axes": axes}
56
+
57
+
58
+ @torch.no_grad()
59
+ def projective_stats(axes: torch.Tensor, n_baseline: int = 20000,
60
+ seed: int = 0) -> dict:
61
+ """Mean projective angle arccos|<a,b>| vs a uniform-RP baseline at same (n, D)."""
62
+ n, D = axes.shape
63
+ if n < 2:
64
+ return {"proj_angle_mean": None, "uniform_baseline": None,
65
+ "deviation": None, "erank": None}
66
+ def mean_angle(rows):
67
+ c = (rows @ rows.T).abs().clamp(max=1.0)
68
+ iu = torch.triu_indices(rows.shape[0], rows.shape[0], offset=1)
69
+ return torch.arccos(c[iu[0], iu[1]]).mean().item()
70
+ obs = mean_angle(axes)
71
+ g = torch.Generator().manual_seed(seed)
72
+ base_angles = []
73
+ m = max(2, n)
74
+ for _ in range(max(1, n_baseline // max(1, m * (m - 1) // 2))):
75
+ r = F.normalize(torch.randn(m, D, generator=g), dim=-1)
76
+ base_angles.append(mean_angle(r))
77
+ base = sum(base_angles) / len(base_angles)
78
+ s = torch.linalg.svdvals(axes)
79
+ p = (s / s.sum().clamp_min(1e-12))
80
+ erank = float(torch.exp(-(p.clamp_min(1e-12) * p.clamp_min(1e-12).log()).sum()))
81
+ return {"proj_angle_mean": round(obs, 4), "uniform_baseline": round(base, 4),
82
+ "deviation": round(obs - base, 4), "erank": round(erank, 3)}
83
+
84
+
85
+ @torch.no_grad()
86
+ def read_specimen(path: str) -> dict:
87
+ ck = torch.load(path, map_location="cpu", weights_only=True)
88
+ out = {"file": path.split("\\")[-1].split("/")[-1],
89
+ "arm": ck.get("arm"), "seed": ck.get("seed"),
90
+ "steps": ck.get("steps"), "val_bpb": round(ck.get("val_bpb", -1), 4)}
91
+ if "state_dict" in ck: # full specimen checkpoint
92
+ sd = ck["state_dict"]
93
+ books = {k[:-len(".codebook")]: sd[k] for k in sd
94
+ if k.endswith("addr.codebook") or k.endswith("head_addr.codebook")}
95
+ homes = {k[:-len(".home")]: sd[k] for k in sd if k.endswith(".home")}
96
+ else: # bare genome dict (exp014+ champion files):
97
+ # books under flat/root/branch* keys; *_proj entries are projections
98
+ books = {k: v for k, v in ck.items()
99
+ if torch.is_tensor(v) and v.ndim == 2
100
+ and (k in ("flat", "root") or k.startswith("branch"))}
101
+ homes = {}
102
+ reads = {}
103
+ for name, cb in books.items():
104
+ col = antipodal_collapse(cb)
105
+ stats = projective_stats(col["axes"])
106
+ home = homes.get(name)
107
+ drift = None
108
+ binding = None
109
+ if home is not None and home.shape == cb.shape:
110
+ a = F.normalize(cb.float(), dim=-1)
111
+ b = F.normalize(home.float(), dim=-1)
112
+ dr = torch.arccos((a * b).sum(-1).clamp(-1, 1))
113
+ drift = round(dr.mean().item(), 4)
114
+ binding = round(((dr - BINDING).abs() <= 0.05).float().mean().item(), 4)
115
+ util = col["n_axes"] / cb.shape[0]
116
+ dev = stats["deviation"]
117
+ if dev is not None and abs(dev) < 0.05 and util > 0.95 and col["pairs"] <= 3:
118
+ verdict = "PROJECTIVE-CLEAN"
119
+ elif dev is not None and abs(dev) < 0.05:
120
+ verdict = "PROJECTIVE-MOSTLY"
121
+ elif dev is not None and dev > 0.05:
122
+ verdict = "STRUCTURED(repulsive)"
123
+ else:
124
+ verdict = "DEGENERATE/CLUMPED" if dev is not None else "TOO-FEW-AXES"
125
+ reads[name] = {
126
+ "pairs": col["pairs"], "n_axes": col["n_axes"], **stats,
127
+ "drift": drift, "binding_frac": binding, "verdict": verdict}
128
+ out["codebooks"] = reads
129
+ return out
130
+
131
+
132
+ def read_all(root: str) -> list:
133
+ import glob, os
134
+ results = []
135
+ for p in sorted(glob.glob(os.path.join(root, "*.pt"))):
136
+ r = read_specimen(p)
137
+ print(r, flush=True)
138
+ results.append(r)
139
+ return results
140
+
141
+
142
+ def _in_notebook() -> bool:
143
+ try:
144
+ get_ipython() # type: ignore[name-defined] # noqa: F821
145
+ return True
146
+ except NameError:
147
+ return False
148
+
149
+
150
+ if __name__ == "__main__":
151
+ if _in_notebook():
152
+ print("Notebook mode: call read_all(r'<data_root>/ar_ckpts') in the next cell.")
153
+ else:
154
+ args = [a for a in sys.argv[1:] if not a.startswith("-")]
155
+ if not args:
156
+ print("usage: python read_codebook.py <ckpt_or_dir> [...]")
157
+ for a in args:
158
+ import os
159
+ read_all(a) if os.path.isdir(a) else print(read_specimen(a))
exp019_cr/repro.py ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """repro.py — standalone loader/runner for exp019_cr. Code dependencies live
2
+ in THIS folder (geolip_vitals.py, ar_differentiation_bed.py,
3
+ exp014_genetic_distillation.py, exp019_content_retention.py, read_codebook.py).
4
+
5
+ python repro.py # CPU smoke: facts, streams, recall, rule block
6
+ python repro.py --run # retention sweep (2 seeds x 3 N x channels, ~3h GPU)
7
+ python repro.py --run19b # generalization block (rule content, ~1h GPU)
8
+
9
+ Data lands in ./data (override with GEOLIP_DATA); ledger + checkpoints in
10
+ ./data/exp019.
11
+ """
12
+ import os
13
+ import sys
14
+
15
+ HERE = os.path.dirname(os.path.abspath(__file__))
16
+ sys.path.insert(0, HERE)
17
+
18
+ if __name__ == "__main__":
19
+ import geolip_vitals # noqa: F401 (paste order)
20
+ import ar_differentiation_bed # noqa: F401
21
+ import exp014_genetic_distillation # noqa: F401
22
+ import exp019_content_retention as cr
23
+ if "--run" in sys.argv[1:]:
24
+ cr.run_retention()
25
+ elif "--run19b" in sys.argv[1:]:
26
+ cr.run_generalization()
27
+ else:
28
+ cr.smoke()
29
+ print("repro smoke passed — --run (retention) / --run19b (rule block)")
exp019_cr/results/ledger.jsonl ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {"exp": "19", "channel": "teacher", "n": 64, "seed": 0, "steps": 4000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "bpb": 2.3987}
2
+ {"exp": "19", "channel": "direct", "n": 64, "seed": 0, "steps": 2000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0221}, "bpb": 2.7867, "bpb_after_interf": 2.3879}
3
+ {"exp": "19", "channel": "kd_facts", "n": 64, "seed": 0, "steps": 2000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0013}, "bpb": 4.0938, "bpb_after_interf": 2.5318}
4
+ {"exp": "19", "channel": "kd_general", "n": 64, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0065}, "bpb": 2.5001, "bpb_after_interf": 2.2824}
5
+ {"exp": "19", "channel": "teacher", "n": 256, "seed": 0, "steps": 4000, "recall": {"exact": 0.9883, "byte_acc": 0.9938}, "bpb": 2.4127}
6
+ {"exp": "19", "channel": "direct", "n": 256, "seed": 0, "steps": 2000, "recall": {"exact": 0.8711, "byte_acc": 0.9095}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0072}, "bpb": 2.8685, "bpb_after_interf": 2.3991}
7
+ {"exp": "19", "channel": "kd_facts", "n": 256, "seed": 0, "steps": 2000, "recall": {"exact": 0.9531, "byte_acc": 0.9694}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0186}, "bpb": 4.1136, "bpb_after_interf": 2.5174}
8
+ {"exp": "19", "channel": "kd_general", "n": 256, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0091}, "bpb": 2.5322, "bpb_after_interf": 2.3352}
9
+ {"exp": "19", "channel": "book_implant", "n": 256, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0085}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0075}, "bpb": 2.4608, "bpb_after_interf": 2.2463}
10
+ {"exp": "19", "channel": "none", "n": 256, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0101}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0078}, "bpb": 2.4685, "bpb_after_interf": 2.2399}
11
+ {"exp": "19", "channel": "teacher", "n": 1024, "seed": 0, "steps": 4000, "recall": {"exact": 0.0117, "byte_acc": 0.0693}, "bpb": 2.5179}
12
+ {"exp": "19", "channel": "direct", "n": 1024, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0319}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0065}, "bpb": 3.0269, "bpb_after_interf": 2.489}
13
+ {"exp": "19", "channel": "kd_facts", "n": 1024, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0352}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0199}, "bpb": 4.1922, "bpb_after_interf": 2.6121}
14
+ {"exp": "19", "channel": "kd_general", "n": 1024, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0013}, "bpb": 2.5957, "bpb_after_interf": 2.3458}
15
+ {"exp": "19", "channel": "teacher", "n": 64, "seed": 1, "steps": 4000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "bpb": 2.3943}
16
+ {"exp": "19", "channel": "direct", "n": 64, "seed": 1, "steps": 2000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.013}, "bpb": 2.752, "bpb_after_interf": 2.3639}
17
+ {"exp": "19", "channel": "kd_facts", "n": 64, "seed": 1, "steps": 2000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0052}, "bpb": 3.9946, "bpb_after_interf": 2.5456}
18
+ {"exp": "19", "channel": "kd_general", "n": 64, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0}, "bpb": 2.5233, "bpb_after_interf": 2.3492}
19
+ {"exp": "19", "channel": "teacher", "n": 256, "seed": 1, "steps": 4000, "recall": {"exact": 0.9922, "byte_acc": 0.9961}, "bpb": 2.3974}
20
+ {"exp": "19", "channel": "direct", "n": 256, "seed": 1, "steps": 2000, "recall": {"exact": 0.8555, "byte_acc": 0.9108}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0179}, "bpb": 2.9598, "bpb_after_interf": 2.4577}
21
+ {"exp": "19", "channel": "kd_facts", "n": 256, "seed": 1, "steps": 2000, "recall": {"exact": 0.8594, "byte_acc": 0.891}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0104}, "bpb": 4.149, "bpb_after_interf": 2.5129}
22
+ {"exp": "19", "channel": "kd_general", "n": 256, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0}, "bpb": 2.5208, "bpb_after_interf": 2.3414}
23
+ {"exp": "19", "channel": "book_implant", "n": 256, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0052}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0065}, "bpb": 2.4782, "bpb_after_interf": 2.2825}
24
+ {"exp": "19", "channel": "none", "n": 256, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0016}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0}, "bpb": 2.4772, "bpb_after_interf": 2.274}
25
+ {"exp": "19", "channel": "teacher", "n": 1024, "seed": 1, "steps": 4000, "recall": {"exact": 0.1992, "byte_acc": 0.3747}, "bpb": 2.4916}
26
+ {"exp": "19", "channel": "direct", "n": 1024, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0286}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0026}, "bpb": 3.0763, "bpb_after_interf": 2.5039}
27
+ {"exp": "19", "channel": "kd_facts", "n": 1024, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0319}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0179}, "bpb": 4.2629, "bpb_after_interf": 2.6531}
28
+ {"exp": "19", "channel": "kd_general", "n": 1024, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0003}, "recall_interf": {"exact": 0.0, "byte_acc": 0.001}, "bpb": 2.5645, "bpb_after_interf": 2.3509}
29
+ {"exp": "19b", "channel": "teacher", "n": 256, "seed": 0, "steps": 4000, "gauges": {"train": {"exact": 1.0, "byte_acc": 1.0}, "heldout": {"exact": 0.0, "byte_acc": 0.2702}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.016}}, "bpb": 2.3906}
30
+ {"exp": "19b", "channel": "direct", "n": 256, "seed": 0, "steps": 2000, "gauges": {"train": {"exact": 0.6523, "byte_acc": 0.8337}, "heldout": {"exact": 0.0, "byte_acc": 0.2695}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0241}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0391}, "heldout": {"exact": 0.0, "byte_acc": 0.0228}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0033}}, "bpb": 2.8785, "bpb_after_interf": 2.4644}
31
+ {"exp": "19b", "channel": "kd_facts", "n": 256, "seed": 0, "steps": 2000, "gauges": {"train": {"exact": 0.7578, "byte_acc": 0.8981}, "heldout": {"exact": 0.0, "byte_acc": 0.2637}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0908}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0247}, "heldout": {"exact": 0.0, "byte_acc": 0.0176}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0225}}, "bpb": 4.1461, "bpb_after_interf": 2.5325}
32
+ {"exp": "19b", "channel": "kd_general", "n": 256, "seed": 0, "steps": 2000, "gauges": {"train": {"exact": 0.0, "byte_acc": 0.0}, "heldout": {"exact": 0.0, "byte_acc": 0.0}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0029}, "heldout": {"exact": 0.0, "byte_acc": 0.0039}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0094}}, "bpb": 2.5091, "bpb_after_interf": 2.3152}
33
+ {"exp": "19b", "channel": "teacher", "n": 256, "seed": 1, "steps": 4000, "gauges": {"train": {"exact": 0.9844, "byte_acc": 0.9948}, "heldout": {"exact": 0.0, "byte_acc": 0.2448}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.002}}, "bpb": 2.4099}
34
+ {"exp": "19b", "channel": "direct", "n": 256, "seed": 1, "steps": 2000, "gauges": {"train": {"exact": 0.7734, "byte_acc": 0.9027}, "heldout": {"exact": 0.0, "byte_acc": 0.2559}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0238}, "heldout": {"exact": 0.0, "byte_acc": 0.0221}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0029}}, "bpb": 2.928, "bpb_after_interf": 2.4109}
35
+ {"exp": "19b", "channel": "kd_facts", "n": 256, "seed": 1, "steps": 2000, "gauges": {"train": {"exact": 0.8438, "byte_acc": 0.932}, "heldout": {"exact": 0.0, "byte_acc": 0.2786}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0771}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0075}, "heldout": {"exact": 0.0, "byte_acc": 0.0085}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0163}}, "bpb": 4.0264, "bpb_after_interf": 2.5571}
36
+ {"exp": "19b", "channel": "kd_general", "n": 256, "seed": 1, "steps": 2000, "gauges": {"train": {"exact": 0.0, "byte_acc": 0.0}, "heldout": {"exact": 0.0, "byte_acc": 0.0}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0}, "heldout": {"exact": 0.0, "byte_acc": 0.0}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0013}}, "bpb": 2.5298, "bpb_after_interf": 2.3408}
exp019_cr/results/results.json ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "capacity_curve": {
3
+ "N64_s0": 1.0,
4
+ "N64_s1": 1.0,
5
+ "N256_s0": 0.9883,
6
+ "N256_s1": 0.9922,
7
+ "N1024_s0": 0.0117,
8
+ "N1024_s1": 0.1992
9
+ },
10
+ "kd_vs_direct_rule_train_exact": {
11
+ "s0": [
12
+ 0.7578,
13
+ 0.6523
14
+ ],
15
+ "s1": [
16
+ 0.8438,
17
+ 0.7734
18
+ ]
19
+ },
20
+ "n_rows": 36
21
+ }
exp019_cr/specimens/rule_direct_s0.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7cb0de713dc1277787f40717c4de271815ed5e8480500c9fc5aab8ed12b35e91
3
+ size 7977425
exp019_cr/specimens/rule_direct_s1.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b63fcab4464246727a8dff2fa2208c4e304e50557397dce45ee5962ef3acc7e6
3
+ size 7977425
exp019_cr/specimens/rule_kd_facts_s0.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:637f1a4de3270b34ed59b589fa083516faf1bcf93c7a96d6118ebcc14cdea7c0
3
+ size 7977535
exp019_cr/specimens/rule_kd_facts_s1.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:771d084ad26ddb878371e385ef91f7d433ac31cfb14412aef4bc171529bf423f
3
+ size 7977535
exp019_cr/specimens/rule_kd_general_s0.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c9ad649e1a8fb6fcacd67866757eb7f26c5e561dfd94f42b3e470f84b374b3dc
3
+ size 7977645
exp019_cr/specimens/rule_kd_general_s1.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:56acd7c7ef2291bfceae6106902b90a3376a7f4fbbd00ab1e81b8fb2731dd9fb
3
+ size 7977645
exp019_cr/specimens/rule_teacher_s0.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9e2fd2b00e87f8e380e45169830225adb15b7c7883b6e6fb027a4e5c973c2154
3
+ size 7977480
exp019_cr/specimens/rule_teacher_s1.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9394e030f940e21b8f4cb9ffc4637bc2db1961c90040d01dca94d4fa66543b1b
3
+ size 7977480
exp019_cr/specimens/teacher_N1024_s0.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:af3c50e397e9443ebbaa1f2e7b9a231ee48e05eb1a9cadd8576f2c7ae162ba51
3
+ size 7977535
exp019_cr/specimens/teacher_N1024_s1.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:41211d812d569c081998188e7ba209bf6e411592dc300919a943ba32a42356c4
3
+ size 7977535
exp019_cr/specimens/teacher_N256_s0.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f33ce09f77188f0275b71d35abb510e58e6855242c8a0faf75e459b4d2420ff0
3
+ size 7977480
exp019_cr/specimens/teacher_N256_s1.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7e6d6917c2c1a78172b5ccd7f73b8f5ba34674224f0dae0c548191b18a95ba6f
3
+ size 7977480
exp019_cr/specimens/teacher_N64_s0.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c171d410e06adb6eb5e9a4d078095d92ba54f91598f5b24e3e5e406dc72c165f
3
+ size 7977425
exp019_cr/specimens/teacher_N64_s1.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ead81b62b73c759c75af1a84b02c477380cec45b835062fad06b64f427524f6c
3
+ size 7977425
exp020_gen/README.md ADDED
@@ -0,0 +1,73 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # exp020_gen — generalization: structure vs capacity on a hidden rule
2
+
3
+ Sequel to [exp019](../exp019_cr/), whose 19b block found rules learned only
4
+ fragmentarily (held-out byte ~0.26, exact 0.000) and content locked to its
5
+ surface format. exp020 races the program's structural methodologies —
6
+ codebooks and constellations — on that same rule task, against the honest
7
+ capacity control, asking what (if anything) unlocks generalization.
8
+
9
+ **Task** (identical to 19b): value = fixed random substitution cipher of the
10
+ key; 256 train keys in-stream, 128 held out; 4000 steps direct training.
11
+ **Judge**: held-out recall (rule induction), variant-format recall (the
12
+ lock), clean bpb.
13
+
14
+ ## The bakeoff (held-out byte accuracy; exact = 0.000 in all 12 cells)
15
+
16
+ | arm | head params | s0 | s1 | clean bpb |
17
+ |---|---|---|---|---|
18
+ | **aleph** (addr_msl64 bottleneck) | 115,200 | **.2611** | **.2493** | 2.40 / 2.41 |
19
+ | relay (addresses in depth) | 214,532 | .2305 | .2493 | 2.36 / 2.41 |
20
+ | aleph_fmtdiv (3 fact formats in-stream) | 115,200 | .2148 | .2363 | 2.35 / 2.40 |
21
+ | mlp (param-matched free head) | 115,261 | .1992 | .2044 | 2.39 / 2.42 |
22
+ | aleph_tri (trigram byte embeddings) | 115,200 | .1634 | .1582 | 2.37 / 2.30 |
23
+ | const (constellation 768 address) | 1,978,432 | .1133 | .1309 | **2.26 / 2.23** |
24
+
25
+ `build_results.py` re-asserts every claim below from `results/ledger.jsonl`.
26
+
27
+ ## Findings
28
+
29
+ 1. **The plateau is not a structure problem.** No methodology in the toolkit
30
+ lifts held-out accuracy off the ~0.26 byte plateau, and no cell produces a
31
+ single exact held-out completion. Cracking rule induction here needs
32
+ scale, budget, or curriculum — not a different head.
33
+ 2. **The aleph bottleneck is the best generalizer — certified both seeds** —
34
+ beating its param-matched free head by ~25% relative at a head budget
35
+ matched to within 0.2%. Structure beats capacity for generalization even
36
+ while losing on bpb.
37
+ 3. **The inverse law**: generalization ordering roughly REVERSES modeling
38
+ strength, both seeds. The constellation head (best clean bpb, 17× larger)
39
+ generalizes worst; the trigram memorizer lineage second-worst; the
40
+ tightest bottleneck best. **Memorization ease substitutes for rule
41
+ induction** — surplus representational capacity soaks the instances and
42
+ removes the pressure to induce. (This measured memorization↔generalization
43
+ axis is the empirical foundation for slider/registry-style composites
44
+ where blocks are graded by this property and gated accordingly.)
45
+ 4. **Format diversity generalizes across trained surface forms, not to novel
46
+ ones**: the diverse-format arm completes held-out keys at full rule level
47
+ in its *seen* alternate format (byte .205/.260) while the *unseen* format
48
+ stays locked (.017/.041).
49
+ 5. **Depth relays ≈ the head bottleneck** (no lift, no cost) — distributing
50
+ the address through depth neither helps nor hurts rule induction.
51
+
52
+ ## Files
53
+ - `exp020_generalization.py` — the six arms, format-diverse fact streams,
54
+ matched-head construction, bakeoff runner, smoke.
55
+ - `geolip_vitals.py` / `ar_differentiation_bed.py` /
56
+ `exp014_genetic_distillation.py` / `exp017_aleph_constellation.py` /
57
+ `exp019_content_retention.py` / `read_codebook.py` — this package's own
58
+ harness copies (the bed, the constellation head, the rule-task machinery).
59
+ Standalone.
60
+ - `repro.py`, `build_results.py`, `results/ledger.jsonl` (12 rows),
61
+ `specimens/` (all 12 checkpoints).
62
+
63
+ ## Reproduce (from inside this folder)
64
+ ```bash
65
+ pip install torch --index-url https://download.pytorch.org/whl/cu128
66
+ pip install pyarrow huggingface_hub
67
+ python repro.py # CPU smoke
68
+ python repro.py --run # the bakeoff (GPU, ~2h)
69
+ python build_results.py # re-assert every claim from the ledger
70
+ ```
71
+ Data lands in `./data` (override with `GEOLIP_DATA`).
72
+
73
+ License: MIT · AbstractPhil + Claude Fable 5 · July 11, 2026
exp020_gen/ar_differentiation_bed.py ADDED
@@ -0,0 +1,487 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """ar_differentiation_bed.py — exp012: autoregressive differentiation of the aleph.
2
+
3
+ Differentiation is cultivated by PREDICTIVE pressure along the sequence — the
4
+ address parameterizing the next-byte distribution (Law 2: chain-rule advantage pays
5
+ ONLY where the composed address directly parameterizes the predictive distribution).
6
+ This bed puts the aleph in the autoregressive gradient path and measures what
7
+ differentiates. The head arms enforce the employment law at its maximum: the
8
+ ENTIRE next-byte distribution is parameterized by the address.
9
+
10
+ Byte-level causal LM on wikitext-2-raw (HF parquet, CDN-fast), block 256. ARMS:
11
+ sdpa — standard causal transformer control (matched trunk).
12
+ hub — attention replaced by CAUSAL HUB: linear attention whose feature map
13
+ is the 2K-oriented aleph address, prefix-sum memories (no selection
14
+ event; O(n*K*d)). Differentiation cultivated INSIDE attention.
15
+ addr_head — sdpa trunk, but the OUTPUT HEAD reads ONLY the signed aleph
16
+ coefficient vector w_k = sinh(u_k)/sum_j cosh(u_j) of the final
17
+ hidden state (K -> 256 logits). The address MUST carry every bit of
18
+ next-byte information — the hardest Law-2 bottleneck.
19
+
20
+ JUDGED BY: val bits-per-byte per arm (task) + CULTIVATION VITALS on every aleph
21
+ codebook (readouts, never losses): axis aliveness/hppl, drift-from-init +
22
+ binding fraction @0.29154, winner-|cos| saturation (sign-code emergence), shadow
23
+ path diversity (fixed high-bits hash). Never by recon.
24
+
25
+ Riders: pure Adam wd=0; no BN/Dropout/GAP on geometric paths; orthogonal init;
26
+ Colab-cell-safe (paste-ahead imports, no bare argparse, no __file__ reliance);
27
+ GPU-only for verdict runs.
28
+
29
+ Terminal: python ar_differentiation_bed.py # shapes/parse smoke
30
+ python ar_differentiation_bed.py --train # verdict run
31
+ Colab: paste geolip_vitals.py cell, then this file (smoke auto-runs),
32
+ then train(steps=2000, data_root="/content/data") in the next cell.
33
+ """
34
+ from __future__ import annotations
35
+ import math
36
+ import torch
37
+ import torch.nn as nn
38
+ import torch.nn.functional as F
39
+
40
+ if "anchor_drift" not in globals():
41
+ try:
42
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
43
+ except ImportError:
44
+ _here = globals().get("__file__")
45
+ if _here is not None:
46
+ import sys, pathlib
47
+ sys.path.insert(0, str(pathlib.Path(_here).parent))
48
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
49
+ else:
50
+ raise ImportError(
51
+ "geolip_vitals not found — paste/run its cell first, or "
52
+ "hf_hub_download exp012_ar/geolip_vitals.py from "
53
+ "AbstractPhil/geolip-aleph-differentiation.")
54
+
55
+ VOCAB = 256 # bytes
56
+
57
+
58
+ # ------------------------------------------------------------------ aleph address
59
+ def _super_fibonacci_s3(n: int) -> torch.Tensor:
60
+ """Near-uniform unit quaternions (Alexa CVPR'22) —
61
+ starts the codebook INSIDE the RP^3 attractor basin. D=4 only."""
62
+ PHI, PSI = math.sqrt(2.0), 1.533751168755204288118041
63
+ i = torch.arange(n, dtype=torch.float64)
64
+ s = (i + 0.5) / n
65
+ r, R = torch.sqrt(s), torch.sqrt(1.0 - s)
66
+ a, b = 2 * math.pi * i / PHI, 2 * math.pi * i / PSI
67
+ q = torch.stack([r * torch.sin(a), r * torch.cos(a),
68
+ R * torch.sin(b), R * torch.cos(b)], dim=-1)
69
+ return F.normalize(q, dim=-1).float()
70
+
71
+
72
+ class AlephAddress(nn.Module):
73
+ """Closed-form aleph over 2K oriented half-axes (aleph-void article).
74
+ signed(x): (..., K) w_k = sinh(u_k)/sum_j cosh(u_j) — the Law-2 head feature.
75
+ oriented(x): ((..., K), (..., K)) positive halves of the 2K softmax — HUB map."""
76
+
77
+ def __init__(self, K: int, D: int, tau: float = 0.1, init: str = "random"):
78
+ super().__init__()
79
+ self.K, self.D, self.tau = K, D, tau
80
+ if init == "fibonacci":
81
+ assert D == 4, "fibonacci init lives on S^3 (D=4)"
82
+ A = _super_fibonacci_s3(K)
83
+ else:
84
+ A = F.normalize(torch.randn(K, D), dim=-1)
85
+ self.codebook = nn.Parameter(A)
86
+ self.register_buffer("home", self.codebook.detach().clone())
87
+
88
+ def _u(self, x):
89
+ A = F.normalize(self.codebook, dim=-1)
90
+ return (F.normalize(x, dim=-1) @ A.transpose(-1, -2)) / self.tau
91
+
92
+ def oriented(self, x):
93
+ u = self._u(x)
94
+ m = u.abs().amax(dim=-1, keepdim=True)
95
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
96
+ Z = (ep + en).sum(dim=-1, keepdim=True)
97
+ return ep / Z, en / Z
98
+
99
+ def signed(self, x):
100
+ u = self._u(x)
101
+ m = u.abs().amax(dim=-1, keepdim=True)
102
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
103
+ return (ep - en) / (ep + en).sum(dim=-1, keepdim=True)
104
+
105
+ def signed_at(self, x, taus):
106
+ """Multi-tau stroboscope (rule of 3): signed coefficients at several
107
+ temperatures, concatenated — softer taus keep the vector dense while a
108
+ hard tau supplies the sign-code sharpness. v2 refinement (b)."""
109
+ A = F.normalize(self.codebook, dim=-1)
110
+ cos = F.normalize(x, dim=-1) @ A.transpose(-1, -2)
111
+ outs = []
112
+ for t in taus:
113
+ u = cos / t
114
+ m = u.abs().amax(dim=-1, keepdim=True)
115
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
116
+ outs.append((ep - en) / (ep + en).sum(dim=-1, keepdim=True))
117
+ return torch.cat(outs, dim=-1)
118
+
119
+ def m_hat(self, x):
120
+ """Closed-form soft read (decoders read M_hat, never M). v2 control (c)."""
121
+ u = self._u(x)
122
+ m = u.abs().amax(dim=-1, keepdim=True)
123
+ ep, en = torch.exp(u - m), torch.exp(-u - m)
124
+ A = F.normalize(self.codebook, dim=-1)
125
+ return ((ep - en) @ A) / (ep + en).sum(dim=-1, keepdim=True)
126
+
127
+ def m_hard_ste(self, x):
128
+ """Hard mode (aleph-void article): M_hard = sign(cos_win) * A[win], straight-through to
129
+ the soft read — forward fully discrete SIGN CODE, backward soft gradient.
130
+ Legal per theme A (reconstructive sign code, not a one-hot roster pick)."""
131
+ u = self._u(x)
132
+ soft = self.m_hat(x)
133
+ win = u.abs().argmax(dim=-1)
134
+ A = F.normalize(self.codebook, dim=-1)
135
+ sign = torch.sign(torch.gather(u, -1, win.unsqueeze(-1))).squeeze(-1)
136
+ hard = sign.unsqueeze(-1) * A[win]
137
+ return hard + soft - soft.detach()
138
+
139
+ @torch.no_grad()
140
+ def vitals(self, x_sample) -> dict:
141
+ u = self._u(x_sample.reshape(-1, x_sample.shape[-1]))
142
+ p, n = self.oriented(x_sample.reshape(-1, x_sample.shape[-1]))
143
+ two_k = torch.cat([p, n], dim=-1)
144
+ win = two_k.argmax(dim=-1)
145
+ cos_win = (u.abs().amax(dim=-1) * self.tau) # winner |cos| — sign-code sat.
146
+ d = anchor_drift(self.codebook, self.home)
147
+ return {"drift": round(d["mean"], 4),
148
+ "binding_frac": round(d["binding_fraction"], 4),
149
+ "aliveness": axis_aliveness(two_k),
150
+ "win_cos_mean": round(cos_win.mean().item(), 4),
151
+ "paths": path_diversity(win)}
152
+
153
+
154
+ # ------------------------------------------------------------------------- blocks
155
+ class CausalSDPA(nn.Module):
156
+ def __init__(self, d: int, heads: int = 4):
157
+ super().__init__()
158
+ self.h = heads
159
+ self.qkv = nn.Linear(d, 3 * d, bias=False)
160
+ self.o = nn.Linear(d, d, bias=False)
161
+ nn.init.orthogonal_(self.qkv.weight); nn.init.orthogonal_(self.o.weight)
162
+
163
+ def forward(self, x):
164
+ B, n, d = x.shape
165
+ q, k, v = self.qkv(x).chunk(3, dim=-1)
166
+ q, k, v = (t.view(B, n, self.h, d // self.h).transpose(1, 2) for t in (q, k, v))
167
+ y = F.scaled_dot_product_attention(q, k, v, is_causal=True)
168
+ return self.o(y.transpose(1, 2).reshape(B, n, d))
169
+
170
+
171
+ class CausalHUB(nn.Module):
172
+ """Causal aleph linear attention: prefix-sum memories over the two K-wide
173
+ halves of the oriented address; 2K never materialized; no selection event."""
174
+
175
+ def __init__(self, d: int, K: int = 32, D: int = 4, tau: float = 0.1):
176
+ super().__init__()
177
+ self.addr = AlephAddress(K, D, tau)
178
+ self.q = nn.Linear(d, D, bias=False)
179
+ self.k = nn.Linear(d, D, bias=False)
180
+ self.v = nn.Linear(d, d, bias=False)
181
+ self.o = nn.Linear(d, d, bias=False)
182
+ for m in (self.q, self.k, self.v, self.o):
183
+ nn.init.orthogonal_(m.weight)
184
+
185
+ def forward(self, x):
186
+ qp, qn = self.addr.oriented(self.q(x)) # (B, n, K)
187
+ kp, kn = self.addr.oriented(self.k(x))
188
+ v = self.v(x) # (B, n, d)
189
+ Sp = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kp, v), dim=1)
190
+ Sn = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kn, v), dim=1)
191
+ zp = torch.cumsum(kp, dim=1)
192
+ zn = torch.cumsum(kn, dim=1)
193
+ num = torch.einsum("bnk,bnkd->bnd", qp, Sp) + torch.einsum("bnk,bnkd->bnd", qn, Sn)
194
+ den = (qp * zp).sum(-1, keepdim=True) + (qn * zn).sum(-1, keepdim=True)
195
+ return self.o(num / den.clamp_min(1e-12))
196
+
197
+
198
+ class MslRelay(nn.Module):
199
+ """Depth-composition unit (chain-rule probe): multi-slot M_hat read entering
200
+ the trunk as a NEAR-ZERO gated residual (gate init -3.0, sigma~0.047 — theme D:
201
+ geometry enters as a nudge and grows only if it earns gradient)."""
202
+
203
+ def __init__(self, d: int, n_slots: int = 16, K: int = 64):
204
+ super().__init__()
205
+ self.n_slots = n_slots
206
+ self.proj = nn.Linear(d, n_slots * 4, bias=False)
207
+ self.out = nn.Linear(n_slots * 4, d, bias=False)
208
+ nn.init.orthogonal_(self.proj.weight)
209
+ nn.init.orthogonal_(self.out.weight)
210
+ self.addr = AlephAddress(K, 4)
211
+ self.gate = nn.Parameter(torch.tensor(-3.0))
212
+
213
+ def forward(self, x):
214
+ B, n, _ = x.shape
215
+ slots = self.proj(x).view(B, n, self.n_slots, 4)
216
+ m = self.addr.m_hat(slots).reshape(B, n, -1)
217
+ return x + self.gate.sigmoid() * self.out(m)
218
+
219
+
220
+ class Block(nn.Module):
221
+ def __init__(self, d: int, attn: nn.Module):
222
+ super().__init__()
223
+ self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)
224
+ self.attn = attn
225
+ self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
226
+
227
+ def forward(self, x):
228
+ x = x + self.attn(self.n1(x))
229
+ return x + self.mlp(self.n2(x))
230
+
231
+
232
+ class ByteLM(nn.Module):
233
+ def __init__(self, arm: str, d: int = 192, layers: int = 4, block: int = 256,
234
+ K: int = 32, D: int = 4):
235
+ super().__init__()
236
+ # "<arm>_tri" suffix = trigram byte embedding (AlephLM byte_emb x3 lineage):
237
+ # token embedding is the sum of embeddings of bytes t, t-1, t-2.
238
+ self.trigram = arm.endswith("_tri")
239
+ if self.trigram:
240
+ arm = arm[:-4]
241
+ # "_fib" = super-Fibonacci S^3 codebook init (basin test: starts INSIDE
242
+ # the RP^3 attractor; primary observable is init->final geodesic drift).
243
+ self.fib = arm.endswith("_fib")
244
+ if self.fib:
245
+ arm = arm[:-4]
246
+ # "relay*" = stacked addresses in depth: MslRelay after every block.
247
+ # relay -> sdpa trunk + standard head; relay_msl64 -> + addressed head.
248
+ self.use_relay = arm.startswith("relay")
249
+ if arm == "relay":
250
+ arm = "sdpa"
251
+ elif arm == "relay_msl64":
252
+ arm = "addr_msl64"
253
+ self.arm, self.block = arm, block
254
+ self.emb = nn.Embedding(VOCAB, d)
255
+ if self.trigram:
256
+ self.emb1 = nn.Embedding(VOCAB, d)
257
+ self.emb2 = nn.Embedding(VOCAB, d)
258
+ self.pos = nn.Parameter(torch.zeros(1, block, d) + 0.01 * torch.randn(1, block, d))
259
+ mk_attn = (lambda: CausalHUB(d, K, D)) if arm == "hub" else (lambda: CausalSDPA(d))
260
+ self.blocks = nn.ModuleList([Block(d, mk_attn()) for _ in range(layers)])
261
+ if self.use_relay:
262
+ self.relays = nn.ModuleList([MslRelay(d) for _ in range(layers)])
263
+ self.nf = nn.LayerNorm(d)
264
+ if arm == "addr_head":
265
+ self.head_addr = AlephAddress(K, d) # v1: codebook in model dim — COLLAPSED
266
+ self.head = nn.Linear(K, VOCAB, bias=True)
267
+ elif arm in ("addr_d4", "addr_3tau", "addr_mhat"):
268
+ # v2 refinements: LOW-D HOME — learned projection to the native D=4 home
269
+ # before addressing (mirrors the healthy HUB arms), K=64.
270
+ self.head_proj = nn.Linear(d, 4, bias=False)
271
+ nn.init.orthogonal_(self.head_proj.weight)
272
+ self.head_addr = AlephAddress(64, 4)
273
+ if arm == "addr_d4":
274
+ self.head = nn.Linear(64, VOCAB, bias=True) # w alone, D=4 home
275
+ elif arm == "addr_3tau":
276
+ self.taus = (0.05, 0.1, 0.3) # rule-of-3 strobe
277
+ self.head = nn.Linear(64 * 3, VOCAB, bias=True)
278
+ else: # addr_mhat
279
+ self.head = nn.Linear(4, VOCAB, bias=True) # tightest: M_hat
280
+ elif arm.startswith("addr_msl"):
281
+ # v3: MULTI-SLOT heads — the 16s funnel widening: P parallel D=4 slots
282
+ # over a SHARED codebook. addr_msl consumes the reconstructive M_hat per
283
+ # slot (Px4 dims); addr_msl_w consumes signed w per slot (Px64) — tests
284
+ # whether slot-parallel consumption alone rescues the coefficient path.
285
+ # addr_msl<P> = slot-count dose-response. addr_mslh<P> = HARD sign-code
286
+ # consumption (straight-through M_hard per slot).
287
+ self.hard = arm.startswith("addr_mslh")
288
+ if arm in ("addr_msl", "addr_msl_w"):
289
+ self.n_slots = 16
290
+ else:
291
+ self.n_slots = int(arm[len("addr_mslh" if self.hard else "addr_msl"):])
292
+ self.head_proj = nn.Linear(d, self.n_slots * 4, bias=False)
293
+ nn.init.orthogonal_(self.head_proj.weight)
294
+ self.head_addr = AlephAddress(
295
+ 64, 4, init="fibonacci" if self.fib else "random")
296
+ width = self.n_slots * (64 if arm == "addr_msl_w" else 4)
297
+ self.head = nn.Linear(width, VOCAB, bias=True)
298
+ elif arm == "addr_3tau_mhat":
299
+ # v3: combine the two v2 winners — 3-tau stroboscope + reconstructive read.
300
+ self.head_proj = nn.Linear(d, 4, bias=False)
301
+ nn.init.orthogonal_(self.head_proj.weight)
302
+ self.head_addr = AlephAddress(64, 4)
303
+ self.taus = (0.05, 0.1, 0.3)
304
+ self.head = nn.Linear(64 * 3 + 4, VOCAB, bias=True)
305
+ else:
306
+ self.head = nn.Linear(d, VOCAB, bias=True)
307
+ self._last_h = None
308
+
309
+ def forward(self, idx):
310
+ x = self.emb(idx)
311
+ if self.trigram: # past-only shifts — causality preserved
312
+ x = x + self.emb1(F.pad(idx, (1, 0), value=0)[:, :-1]) \
313
+ + self.emb2(F.pad(idx, (2, 0), value=0)[:, :-2])
314
+ x = x + self.pos[:, : idx.shape[1]]
315
+ if self.use_relay:
316
+ for b, r in zip(self.blocks, self.relays):
317
+ x = r(b(x))
318
+ else:
319
+ for b in self.blocks:
320
+ x = b(x)
321
+ h = self.nf(x)
322
+ self._last_h = h.detach()
323
+ if self.arm == "addr_head":
324
+ return self.head(self.head_addr.signed(h))
325
+ if self.arm == "addr_d4":
326
+ return self.head(self.head_addr.signed(self.head_proj(h)))
327
+ if self.arm == "addr_3tau":
328
+ return self.head(self.head_addr.signed_at(self.head_proj(h), self.taus))
329
+ if self.arm == "addr_mhat":
330
+ return self.head(self.head_addr.m_hat(self.head_proj(h)))
331
+ if self.arm.startswith("addr_msl"):
332
+ B, n, _ = h.shape
333
+ slots = self.head_proj(h).view(B, n, self.n_slots, 4)
334
+ if self.arm == "addr_msl_w":
335
+ feats = self.head_addr.signed(slots).reshape(B, n, -1)
336
+ elif getattr(self, "hard", False):
337
+ feats = self.head_addr.m_hard_ste(slots).reshape(B, n, -1)
338
+ else:
339
+ feats = self.head_addr.m_hat(slots).reshape(B, n, -1)
340
+ return self.head(feats)
341
+ if self.arm == "addr_3tau_mhat":
342
+ p = self.head_proj(h)
343
+ feats = torch.cat([self.head_addr.signed_at(p, self.taus),
344
+ self.head_addr.m_hat(p)], dim=-1)
345
+ return self.head(feats)
346
+ return self.head(h)
347
+
348
+ @torch.no_grad()
349
+ def vitals(self) -> dict:
350
+ out = {}
351
+ if self.arm == "hub":
352
+ for i, b in enumerate(self.blocks):
353
+ if self._last_h is not None:
354
+ out[f"L{i}"] = b.attn.addr.vitals(b.attn.q(self._last_h[:2]))
355
+ elif self.arm == "addr_head" and self._last_h is not None:
356
+ out["head"] = self.head_addr.vitals(self._last_h[:2])
357
+ elif self.arm in ("addr_d4", "addr_3tau", "addr_mhat",
358
+ "addr_3tau_mhat") and self._last_h is not None:
359
+ out["head"] = self.head_addr.vitals(self.head_proj(self._last_h[:2]))
360
+ elif self.arm.startswith("addr_msl") and self._last_h is not None:
361
+ slots = self.head_proj(self._last_h[:2])
362
+ out["head"] = self.head_addr.vitals(
363
+ slots.reshape(*slots.shape[:-1], self.n_slots, 4))
364
+ if self.use_relay and self._last_h is not None:
365
+ for i, r in enumerate(self.relays):
366
+ s = r.proj(self._last_h[:2])
367
+ v = r.addr.vitals(s.reshape(*s.shape[:-1], r.n_slots, 4))
368
+ out[f"relay{i}"] = {"gate": round(r.gate.sigmoid().item(), 4),
369
+ "drift": v["drift"],
370
+ "binding_frac": v["binding_frac"],
371
+ "ppl": round(v["aliveness"]["usage_ppl"], 1)}
372
+ return out
373
+
374
+
375
+ # --------------------------------------------------------------------------- data
376
+ def _wikitext_bytes(data_root: str):
377
+ """wikitext-2-raw as flat uint8 tensors via the HF parquet CDN."""
378
+ from huggingface_hub import hf_hub_download
379
+ import pyarrow.parquet as pq
380
+
381
+ def load(split):
382
+ p = hf_hub_download("Salesforce/wikitext",
383
+ f"wikitext-2-raw-v1/{split}-00000-of-00001.parquet",
384
+ repo_type="dataset", local_dir=data_root)
385
+ text = "".join(pq.read_table(p).column("text").to_pylist())
386
+ return torch.frombuffer(bytearray(text.encode("utf-8")), dtype=torch.uint8).clone()
387
+
388
+ return load("train"), load("validation")
389
+
390
+
391
+ def _batch(data: torch.Tensor, batch: int, block: int, device, g: torch.Generator):
392
+ ix = torch.randint(0, data.numel() - block - 1, (batch,), generator=g)
393
+ x = torch.stack([data[i:i + block] for i in ix]).long().to(device)
394
+ y = torch.stack([data[i + 1:i + block + 1] for i in ix]).long().to(device)
395
+ return x, y
396
+
397
+
398
+ # -------------------------------------------------------------------- train/smoke
399
+ def train(arms=("sdpa", "hub", "addr_head"), steps: int = 2000, batch: int = 32,
400
+ block: int = 256, device: str = "cuda", data_root: str = "./data",
401
+ seed: int = 0, eval_every: int = 500, save: bool = True):
402
+ """Verdict run — GPU only. Pure Adam wd=0. Reports val bits-per-byte + vitals.
403
+ save=True writes {data_root}/ar_ckpts/{arm}_s{seed}_t{steps}.pt per arm —
404
+ the cultivated codebooks are SPECIMENS for the projective reading instruments."""
405
+ import os
406
+ if device == "cuda" and not torch.cuda.is_available():
407
+ raise RuntimeError("Verdict runs are GPU-only (never CPU-train for accuracy).")
408
+ ckpt_dir = os.path.join(data_root, "ar_ckpts")
409
+ os.makedirs(ckpt_dir, exist_ok=True)
410
+ tr, va = _wikitext_bytes(data_root)
411
+ print(f"data ready: train {tr.numel():,} bytes, val {va.numel():,} bytes", flush=True)
412
+ results = {}
413
+ for arm in arms:
414
+ torch.manual_seed(seed)
415
+ g = torch.Generator().manual_seed(seed)
416
+ model = ByteLM(arm, block=block).to(device)
417
+ n_params = sum(p.numel() for p in model.parameters())
418
+ opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
419
+ for step in range(1, steps + 1):
420
+ x, y = _batch(tr, batch, block, device, g)
421
+ logits = model(x)
422
+ loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
423
+ opt.zero_grad(set_to_none=True)
424
+ loss.backward()
425
+ opt.step()
426
+ if step % eval_every == 0 or step == steps:
427
+ model.eval()
428
+ with torch.no_grad():
429
+ losses = []
430
+ for _ in range(20):
431
+ xv, yv = _batch(va, batch, block, device, g)
432
+ lv = F.cross_entropy(model(xv).reshape(-1, VOCAB),
433
+ yv.reshape(-1))
434
+ losses.append(lv.item())
435
+ bpb = sum(losses) / len(losses) / math.log(2)
436
+ print(f"[{arm}] step {step} val_bpb={bpb:.4f} vitals={model.vitals()}",
437
+ flush=True)
438
+ model.train()
439
+ results[arm] = {"val_bpb": bpb, "params": n_params, "vitals": model.vitals()}
440
+ if save:
441
+ path = os.path.join(ckpt_dir, f"{arm}_s{seed}_t{steps}.pt")
442
+ torch.save({"arm": arm, "seed": seed, "steps": steps, "val_bpb": bpb,
443
+ "state_dict": {k: v.cpu() for k, v in
444
+ model.state_dict().items()}}, path)
445
+ print(f"saved specimen: {path}", flush=True)
446
+ print(results, flush=True)
447
+ return results
448
+
449
+
450
+ def smoke():
451
+ """Shapes/parse only — no accuracy claims."""
452
+ x = torch.randint(0, VOCAB, (2, 64))
453
+ for arm in ("sdpa", "hub", "addr_head"):
454
+ m = ByteLM(arm, d=96, layers=2, block=64, K=16)
455
+ logits = m(x)
456
+ assert logits.shape == (2, 64, VOCAB)
457
+ logits.sum().backward()
458
+ # causality check: future byte must not affect past logits
459
+ with torch.no_grad():
460
+ a = m(x)[0, 10]
461
+ x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
462
+ b = m(x2)[0, 10]
463
+ assert torch.allclose(a, b, atol=1e-4), f"{arm} leaks future context"
464
+ print(f"{arm}: OK params={sum(p.numel() for p in m.parameters()):,} "
465
+ f"vitals={m.vitals()}", flush=True)
466
+ print("OK — AR bed smoke passed (verdict run: train() on GPU)", flush=True)
467
+
468
+
469
+ def _in_notebook() -> bool:
470
+ try:
471
+ get_ipython() # type: ignore[name-defined] # noqa: F821
472
+ return True
473
+ except NameError:
474
+ return False
475
+
476
+
477
+ if __name__ == "__main__":
478
+ if _in_notebook():
479
+ smoke()
480
+ print("Notebook mode: call train(steps=2000) in the next cell (GPU).")
481
+ else:
482
+ import argparse
483
+ ap = argparse.ArgumentParser()
484
+ ap.add_argument("--train", action="store_true")
485
+ ap.add_argument("--steps", type=int, default=2000)
486
+ a, _ = ap.parse_known_args()
487
+ train(steps=a.steps) if a.train else smoke()
exp020_gen/build_results.py ADDED
@@ -0,0 +1,52 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """build_results.py — exp020_gen: read results/ledger.jsonl and RE-ASSERT every
2
+ claim in the README. Run from inside this folder: python build_results.py
3
+ """
4
+ import json
5
+ import os
6
+
7
+ HERE = os.path.dirname(os.path.abspath(__file__))
8
+ rows = [json.loads(l) for l in
9
+ open(os.path.join(HERE, "results", "ledger.jsonl"), encoding="utf-8")]
10
+ assert all(r["exp"] == "20" for r in rows) and len(rows) == 12
11
+ cell = {(r["arm"], r["seed"]): r for r in rows}
12
+
13
+ def held(arm, s):
14
+ return cell[(arm, s)]["gauges"]["heldout"]["byte_acc"]
15
+
16
+ # claim 1: no arm lifts rule induction off the plateau — held-out EXACT is
17
+ # 0.000 in all 12 cells, byte acc <= 0.27 everywhere
18
+ for r in rows:
19
+ assert r["gauges"]["heldout"]["exact"] == 0.0
20
+ assert r["gauges"]["heldout"]["byte_acc"] <= 0.27
21
+
22
+ # claim 2: the aleph bottleneck beats its param-matched free head, both seeds,
23
+ # at matched head budget (within 0.2%)
24
+ for s in (0, 1):
25
+ assert held("aleph", s) > held("mlp", s), s
26
+ p_al, p_ml = cell[("aleph", 0)]["head_params"], cell[("mlp", 0)]["head_params"]
27
+ assert abs(p_al - p_ml) / p_al < 0.002, (p_al, p_ml)
28
+
29
+ # claim 3 (the inverse law): the best clean-bpb arm (const) generalizes WORST,
30
+ # and the memorizer lineage (trigram) is second-worst — both seeds
31
+ for s in (0, 1):
32
+ bpbs = {a: cell[(a, s)]["bpb"] for a, _ in cell if _ == s}
33
+ assert min(bpbs, key=bpbs.get) == "const", s
34
+ order = sorted((held(a, s), a) for a, ss in cell if ss == s)
35
+ assert order[0][1] == "const" and order[1][1] == "aleph_tri", (s, order)
36
+ assert order[-1][1] in ("aleph", "relay"), (s, order)
37
+
38
+ # claim 4: format diversity teaches the SEEN alternate format at rule level
39
+ # but does NOT unlock the unseen variant
40
+ for s in (0, 1):
41
+ g = cell[("aleph_fmtdiv", s)]["gauges"]
42
+ assert g["heldout_r1"]["byte_acc"] > 0.20, s # seen format: rule-level
43
+ assert g["train_varfmt"]["byte_acc"] < 0.06, s # unseen: still locked
44
+
45
+ out = {"heldout_byte": {f"{a}_s{s}": held(a, s) for (a, s) in sorted(cell)},
46
+ "head_params": {a: cell[(a, 0)]["head_params"]
47
+ for a in {k[0] for k in cell}},
48
+ "n_rows": len(rows)}
49
+ json.dump(out, open(os.path.join(HERE, "results", "results.json"), "w",
50
+ encoding="utf-8"), indent=1)
51
+ print(f"{len(rows)} rows -> results/results.json")
52
+ print("all README claims asserted OK")
exp020_gen/exp014_genetic_distillation.py ADDED
@@ -0,0 +1,515 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """exp014_genetic_distillation.py — genetic distillation + memory substrate.
2
+ 14_a: multi-generational tournament (GM3 paradigm) where the aleph codebook is the
3
+ explicit heritable genome. Lineages: ALEPH-FLAT (consensus book + KD) |
4
+ ALEPH-TREE (structured genome: root book + branch books) | MLP-KD
5
+ (traditional: best-parent weights + KD) | NO-INHERIT (evolution floor).
6
+ Both sides intentionally inherit logits (KD); only ours inherits geometry.
7
+ Consensus = Procrustes/GPA alignment of parents' books to mean shape
8
+ (placement by construction — replaces GM3's k-means-on-consensus init).
9
+ 14_b: memory substrate — the D=4 home makes books size-agnostic. Implant books
10
+ cultivated in a small organism into a larger one (frozen / trainable), and
11
+ into GPT-2 relay adapters (cross-architecture frozen distillation).
12
+
13
+ Riders: pure Adam wd=0; KD = KL to detached teacher probs (predictive pressure, no
14
+ contrastive); tree routing is DENSE SOFT (oriented weights; collapse monitor on the
15
+ root); drift-check precedes every freeze claim; GPU-only verdict runs; Colab-safe.
16
+ Founders share a COMMON-ANCESTOR book so GPA row correspondence is inherited.
17
+
18
+ Colab paste order: geolip_vitals.py -> ar_differentiation_bed.py ->
19
+ exp013_augmentation_bed.py (only for run_b2) -> this file.
20
+ """
21
+ from __future__ import annotations
22
+ import copy
23
+ import json
24
+ import math
25
+ import os
26
+ import torch
27
+ import torch.nn as nn
28
+ import torch.nn.functional as F
29
+
30
+ if "anchor_drift" not in globals():
31
+ try:
32
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
33
+ except ImportError:
34
+ _here = globals().get("__file__")
35
+ if _here is None:
36
+ raise ImportError("paste/run geolip_vitals.py first")
37
+ import sys, pathlib
38
+ sys.path.insert(0, str(pathlib.Path(_here).parent))
39
+ from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
40
+ if "ByteLM" not in globals():
41
+ try:
42
+ from ar_differentiation_bed import (ByteLM, AlephAddress, _wikitext_bytes,
43
+ _batch, VOCAB)
44
+ except ImportError:
45
+ raise ImportError("paste/run ar_differentiation_bed.py first")
46
+
47
+ DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
48
+ EXP_DIR = os.path.join(DATA_ROOT, "exp014")
49
+
50
+
51
+ class SquaredReLU(nn.Module):
52
+ def forward(self, x):
53
+ return F.relu(x) ** 2
54
+
55
+
56
+ # ===================================================== consensus (the germline) ===
57
+ @torch.no_grad()
58
+ def procrustes_rotation(A: torch.Tensor, M: torch.Tensor) -> torch.Tensor:
59
+ """Orthogonal R minimizing ||A R - M||_F (rows correspond)."""
60
+ U, _, Vt = torch.linalg.svd(A.T.double() @ M.double())
61
+ return (U @ Vt).float()
62
+
63
+
64
+ @torch.no_grad()
65
+ def align_to(A: torch.Tensor, ref: torch.Tensor, iters: int = 20) -> torch.Tensor:
66
+ """Projective Procrustes (rows correspond, signs free): alternate the
67
+ orthogonal rotation and per-row sign flips (books live on RP^(D-1))."""
68
+ s = torch.ones(A.shape[0], 1)
69
+ for _ in range(iters):
70
+ R = procrustes_rotation(s * A, ref)
71
+ AR = (s * A) @ R
72
+ s_upd = torch.where((AR * ref).sum(-1, keepdim=True) < 0, -s, s)
73
+ if torch.equal(s_upd, s):
74
+ return AR
75
+ s = s_upd
76
+ return (s * A) @ procrustes_rotation(s * A, ref)
77
+
78
+
79
+ @torch.no_grad()
80
+ def consensus_codebook(books: list, iters: int = 50, tol: float = 1e-8):
81
+ """GPA to mean shape (GM3 machinery, applied to aleph books), anchored to the
82
+ FIRST parent's frame. Rows must correspond (common-ancestor convention); signs
83
+ are projective. Returns (consensus, n_iters, delta)."""
84
+ # device-pin to CPU: parent models may live on CUDA after KD teacher moves
85
+ Bs = [F.normalize(b.detach().float().cpu(), dim=-1).clone() for b in books]
86
+ # pairwise projective alignment to parent-0's frame, THEN GPA refinement
87
+ aligned = [Bs[0]] + [align_to(b, Bs[0]) for b in Bs[1:]]
88
+ M = F.normalize(torch.stack(aligned).mean(0), dim=-1)
89
+ delta, it = 0.0, 0
90
+ for it in range(1, iters + 1):
91
+ aligned = [align_to(b, M) for b in Bs]
92
+ M_new = F.normalize(torch.stack(aligned).mean(0), dim=-1)
93
+ delta = (M_new - M).norm().item()
94
+ M = M_new
95
+ if delta < tol:
96
+ break
97
+ # re-anchor to parent-0 (GPA drift of the global frame stays measurable)
98
+ M = align_to(M, Bs[0])
99
+ return M, it, delta
100
+
101
+
102
+ @torch.no_grad()
103
+ def implant_book(addr: "AlephAddress", book: torch.Tensor, trainable: bool = True):
104
+ """Load a book into an AlephAddress: codebook + home (drift measured from the
105
+ implant). Freeze only via trainable=False AFTER a drift-check justifies it."""
106
+ b = F.normalize(book.float(), dim=-1).to(addr.codebook.device)
107
+ assert b.shape == addr.codebook.shape, (b.shape, addr.codebook.shape)
108
+ addr.codebook.data.copy_(b)
109
+ addr.home.copy_(b)
110
+ addr.codebook.requires_grad_(trainable)
111
+
112
+
113
+ # ============================================================= tree head ==========
114
+ class TreeHead(nn.Module):
115
+ """Autoregressive tree (the constellation-anchor analogue, exp011 TREE operator
116
+ in the healthy consumption regime): a ROOT aleph (K=2, D=4) yields the 4
117
+ oriented weights (2K half-axes = the 4 branches, dense soft, sums to 1);
118
+ each BRANCH is a 64-slot... shared slot projection read by a branch-specific
119
+ book (K=64, D=4); output = branch-weighted mixture of branch reads -> vocab.
120
+ Heritable genome: root book (2,4) + 4 branch books (64,4)."""
121
+
122
+ ROOT_SLOTS = 4 # slot-parallel root consumption (the collapse cure)
123
+ ROOT_TAU = 0.3 # softer root temperature (wave-1 fix: single hard-tau
124
+ # root partially collapsed, usage [.85,.12,.01,.02])
125
+
126
+ def __init__(self, d: int, vocab: int = 256, n_slots: int = 16):
127
+ super().__init__()
128
+ self.n_slots = n_slots
129
+ self.root_proj = nn.Linear(d, self.ROOT_SLOTS * 4, bias=False)
130
+ self.slot_proj = nn.Linear(d, n_slots * 4, bias=False)
131
+ nn.init.orthogonal_(self.root_proj.weight)
132
+ nn.init.orthogonal_(self.slot_proj.weight)
133
+ self.root = AlephAddress(2, 4, tau=self.ROOT_TAU)
134
+ self.branches = nn.ModuleList([AlephAddress(64, 4) for _ in range(4)])
135
+ self.out = nn.Linear(n_slots * 4, vocab, bias=True)
136
+ self._last_root = None
137
+
138
+ def forward(self, h):
139
+ B, n, _ = h.shape
140
+ rs = self.root_proj(h).view(B, n, self.ROOT_SLOTS, 4)
141
+ p, m = self.root.oriented(rs) # (B,n,S,2) x2
142
+ w = torch.cat([p, m], dim=-1).mean(dim=-2) # slot-avg -> (B,n,4)
143
+ self._last_root = w.detach()
144
+ slots = self.slot_proj(h).view(B, n, self.n_slots, 4)
145
+ mix = 0
146
+ for b, br in enumerate(self.branches):
147
+ mix = mix + w[..., b:b + 1] * br.m_hat(slots).reshape(B, n, -1)
148
+ return self.out(mix)
149
+
150
+ def genome(self):
151
+ return {"root": self.root.codebook.detach().clone(),
152
+ **{f"branch{i}": br.codebook.detach().clone()
153
+ for i, br in enumerate(self.branches)}}
154
+
155
+ @torch.no_grad()
156
+ def inherit(self, genomes: list):
157
+ c, it, dl = consensus_codebook([g["root"] for g in genomes])
158
+ implant_book(self.root, c)
159
+ for i, br in enumerate(self.branches):
160
+ c, _, _ = consensus_codebook([g[f"branch{i}"] for g in genomes])
161
+ implant_book(br, c)
162
+
163
+ @torch.no_grad()
164
+ def vitals(self):
165
+ out = {"root_drift": round(anchor_drift(self.root.codebook,
166
+ self.root.home)["mean"], 4)}
167
+ if self._last_root is not None:
168
+ w = self._last_root.reshape(-1, 4)
169
+ usage = w.mean(0)
170
+ usage = usage / usage.sum()
171
+ out["root_usage"] = [round(float(u), 3) for u in usage]
172
+ ent = -(usage.clamp_min(1e-9) * usage.clamp_min(1e-9).log()).sum()
173
+ out["root_ppl4"] = round(float(ent.exp()), 3)
174
+ d = [anchor_drift(br.codebook, br.home)["mean"] for br in self.branches]
175
+ out["branch_drift"] = [round(x, 3) for x in d]
176
+ return out
177
+
178
+
179
+ # ============================================================ organisms ===========
180
+ def make_organism(lineage: str, d: int = 192, layers: int = 4, block: int = 256,
181
+ seed: int = 0):
182
+ """lineage in {aleph_flat, aleph_tree, mlp_kd, no_inherit}. no_inherit uses the
183
+ aleph_flat architecture (the control isolates INHERITANCE, not architecture)."""
184
+ torch.manual_seed(seed)
185
+ if lineage in ("aleph_flat", "no_inherit", "aleph_full", "aleph_weights"):
186
+ return ByteLM("addr_msl64", d=d, layers=layers, block=block)
187
+ if lineage == "aleph_tree":
188
+ m = ByteLM("sdpa", d=d, layers=layers, block=block)
189
+ m.head = TreeHead(d)
190
+ return m
191
+ if lineage == "mlp_kd":
192
+ m = ByteLM("sdpa", d=d, layers=layers, block=block)
193
+ m.head = nn.Sequential(nn.Linear(d, 224), SquaredReLU(),
194
+ nn.LayerNorm(224), nn.Linear(224, VOCAB))
195
+ return m
196
+ raise ValueError(lineage)
197
+
198
+
199
+ def genome_of(model):
200
+ """The heritable organ = co-adapted (projection, book) pair(s). Books inherit
201
+ by CONSENSUS (the geometric germline); projections inherit from the BEST
202
+ parent (weight copy) — implanting a book against a random projection puts the
203
+ child below random init (campaign-v2 lesson)."""
204
+ if isinstance(model.head, TreeHead):
205
+ g = model.head.genome()
206
+ g["root_proj"] = model.head.root_proj.weight.detach().cpu().clone()
207
+ g["slot_proj"] = model.head.slot_proj.weight.detach().cpu().clone()
208
+ return g
209
+ if hasattr(model, "head_addr"):
210
+ return {"flat": model.head_addr.codebook.detach().cpu().clone(),
211
+ "proj": model.head_proj.weight.detach().cpu().clone()}
212
+ return None
213
+
214
+
215
+ @torch.no_grad()
216
+ def inherit_genome(model, genomes: list):
217
+ """genomes[0] = the BEST parent (selection order matters)."""
218
+ if isinstance(model.head, TreeHead):
219
+ model.head.inherit(genomes)
220
+ model.head.root_proj.weight.copy_(genomes[0]["root_proj"].to(
221
+ model.head.root_proj.weight.device))
222
+ model.head.slot_proj.weight.copy_(genomes[0]["slot_proj"].to(
223
+ model.head.slot_proj.weight.device))
224
+ elif hasattr(model, "head_addr"):
225
+ c, it, dl = consensus_codebook([g["flat"] for g in genomes])
226
+ implant_book(model.head_addr, c)
227
+ model.head_proj.weight.copy_(genomes[0]["proj"].to(
228
+ model.head_proj.weight.device))
229
+
230
+
231
+ def organism_vitals(model):
232
+ if isinstance(model.head, TreeHead):
233
+ return model.head.vitals()
234
+ return model.vitals() if hasattr(model, "vitals") else {}
235
+
236
+
237
+ # ========================================================= train one member ======
238
+ def train_member(model, tr, va, steps=2000, batch=32, block=256, device="cuda",
239
+ seed=0, teachers=None, kd_alpha=1.0):
240
+ """CE (+ KL to detached mean teacher probs when teachers given). Pure Adam."""
241
+ g = torch.Generator().manual_seed(seed)
242
+ model = model.to(device)
243
+ opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
244
+ if teachers:
245
+ teachers = [t.to(device).eval() for t in teachers]
246
+ for step in range(1, steps + 1):
247
+ x, y = _batch(tr, batch, block, device, g)
248
+ logits = model(x)
249
+ loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
250
+ if teachers:
251
+ with torch.no_grad():
252
+ tp = torch.stack([F.softmax(t(x), -1) for t in teachers]).mean(0)
253
+ loss = loss + kd_alpha * F.kl_div(
254
+ F.log_softmax(logits, -1), tp, reduction="batchmean")
255
+ opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
256
+ model.eval()
257
+ with torch.no_grad():
258
+ ls = []
259
+ for _ in range(20):
260
+ xv, yv = _batch(va, batch, block, device, g)
261
+ ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
262
+ yv.reshape(-1)).item())
263
+ return sum(ls) / len(ls) / math.log(2) # bpb
264
+
265
+
266
+ # ============================================================ the tournament ======
267
+ def run_tournament(lineage: str, gens: int = 4, pop: int = 4, steps: int = 2000,
268
+ seed: int = 0, device: str = "cuda",
269
+ catastrophic_at: int | None = None):
270
+ """One lineage, one tournament seed. Logs per-gen to the ledger; saves the
271
+ champion genome per generation. catastrophic_at=G injects a 0-step random
272
+ parent into the consensus at generation G (the GM3 robustness probe)."""
273
+ if not torch.cuda.is_available():
274
+ raise RuntimeError("verdict runs are GPU-only")
275
+ os.makedirs(EXP_DIR, exist_ok=True)
276
+ tr, va = _wikitext_bytes(DATA_ROOT)
277
+ ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
278
+ # common ancestor: every founder book starts identical within a tournament
279
+ torch.manual_seed(9000 + seed)
280
+ ancestor = make_organism(lineage, seed=9000 + seed)
281
+ anc_genome = genome_of(ancestor)
282
+ parents, parent_models, champion_genomes = [], [], []
283
+ for gen in range(gens):
284
+ members = []
285
+ for i in range(pop):
286
+ mseed = seed * 1000 + gen * 100 + i
287
+ m = make_organism(lineage, seed=mseed)
288
+ if anc_genome and gen == 0:
289
+ inherit_genome(m, [anc_genome]) # common ancestor
290
+ if gen > 0:
291
+ is_fresh = (i == pop - 1) # gene flow founder
292
+ if not is_fresh:
293
+ if lineage in ("aleph_flat", "aleph_tree"):
294
+ gs = [genome_of(pm) for pm in parent_models]
295
+ if catastrophic_at == gen:
296
+ bad = make_organism(lineage, seed=666 + i)
297
+ gs = gs + [genome_of(bad)]
298
+ inherit_genome(m, gs)
299
+ elif lineage == "aleph_full":
300
+ # v4 arm (v3 lesson: continuity is what pays) — inherit the
301
+ # WHOLE best parent, then overwrite the book with the
302
+ # two-parent consensus: germline ON TOP of continuity.
303
+ m.load_state_dict(copy.deepcopy(
304
+ parent_models[0].state_dict()))
305
+ gs = [genome_of(pm) for pm in parent_models]
306
+ if catastrophic_at == gen:
307
+ bad = make_organism(lineage, seed=666 + i)
308
+ gs = gs + [genome_of(bad)]
309
+ c, _, _ = consensus_codebook([g["flat"] for g in gs])
310
+ implant_book(m.head_addr, c)
311
+ elif lineage in ("mlp_kd", "aleph_weights"):
312
+ # pure continuity (no germline op) — aleph_weights is the
313
+ # within-architecture control for aleph_full
314
+ m.load_state_dict(copy.deepcopy(
315
+ parent_models[0].state_dict()))
316
+ # no_inherit: nothing
317
+ # KD: alpha 0.25 (campaign-v2 lesson: alpha=1.0 from near-parity
318
+ # teachers COMPOUNDS DOWNWARD — inverse evolution; fresh-founder
319
+ # control isolated it). Fresh founders get NO KD (clean gene flow).
320
+ is_fresh_now = (gen > 0 and i == pop - 1)
321
+ teachers = parent_models if (gen > 0 and not is_fresh_now
322
+ and lineage != "no_inherit") else None
323
+ bpb = train_member(m, tr, va, steps=steps, device=device,
324
+ seed=mseed, teachers=teachers, kd_alpha=0.25)
325
+ vit = organism_vitals(m)
326
+ members.append((bpb, m))
327
+ rec = {"exp": "14a", "lineage": lineage, "tseed": seed, "gen": gen,
328
+ "member": i, "fresh": gen > 0 and i == pop - 1,
329
+ "catastrophic": catastrophic_at == gen, "bpb": round(bpb, 4),
330
+ "vitals": vit}
331
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
332
+ print(f"[14a {lineage} t{seed} g{gen} m{i}] bpb={bpb:.4f} {vit}",
333
+ flush=True)
334
+ members.sort(key=lambda t: t[0])
335
+ parent_models = [members[0][1].cpu(), members[1][1].cpu()]
336
+ best = members[0][0]
337
+ gene = genome_of(members[0][1])
338
+ if gene:
339
+ torch.save(gene, os.path.join(
340
+ EXP_DIR, f"champion_{lineage}_t{seed}_g{gen}.pt"))
341
+ champion_genomes.append(gene)
342
+ print(f"[14a {lineage} t{seed} g{gen}] BEST={best:.4f} "
343
+ f"mean={sum(b for b, _ in members)/pop:.4f}", flush=True)
344
+ for _, mm in members[2:]:
345
+ del mm
346
+ torch.cuda.empty_cache()
347
+ ledger.close()
348
+ return best
349
+
350
+
351
+ # ============================================================ 14_b implants ======
352
+ def run_b1(steps: int = 2000, seed: int = 0, device: str = "cuda",
353
+ donor_book: torch.Tensor | None = None, tag: str = "small_cultivated"):
354
+ """Cross-size: donor book (default: cultivate in a small organism) implanted
355
+ into a LARGE organism. Arms: fresh | implant-trainable | implant-frozen | mlp."""
356
+ if not torch.cuda.is_available():
357
+ raise RuntimeError("verdict runs are GPU-only")
358
+ os.makedirs(EXP_DIR, exist_ok=True)
359
+ tr, va = _wikitext_bytes(DATA_ROOT)
360
+ ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
361
+ if donor_book is None:
362
+ small = make_organism("aleph_flat", d=128, layers=2, seed=seed)
363
+ bpb_small = train_member(small, tr, va, steps=steps, device=device, seed=seed)
364
+ donor_book = genome_of(small)["flat"]
365
+ print(f"[14b donor small] bpb={bpb_small:.4f}", flush=True)
366
+ results = {}
367
+ for arm in ("fresh", "implant_train", "implant_frozen", "mlp"):
368
+ lineage = "mlp_kd" if arm == "mlp" else "aleph_flat"
369
+ m = make_organism(lineage, d=384, layers=6, seed=seed + 10)
370
+ if arm.startswith("implant"):
371
+ implant_book(m.head_addr, donor_book, trainable=(arm == "implant_train"))
372
+ bpb = train_member(m, tr, va, steps=steps, device=device, seed=seed + 10)
373
+ vit = organism_vitals(m)
374
+ results[arm] = {"bpb": round(bpb, 4), "vitals": vit}
375
+ rec = {"exp": "14b1", "arm": arm, "donor": tag, "seed": seed,
376
+ "bpb": round(bpb, 4), "vitals": vit}
377
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
378
+ print(f"[14b1 {arm} donor={tag}] bpb={bpb:.4f} {vit}", flush=True)
379
+ del m; torch.cuda.empty_cache()
380
+ ledger.close()
381
+ return results, donor_book
382
+
383
+
384
+ def run_b2(donor_book: torch.Tensor, steps: int = 1500, seed: int = 0,
385
+ device: str = "cuda", tag: str = "small_cultivated"):
386
+ """Cross-architecture: implant the donor book into every GPT-2 relay adapter
387
+ (exp013 Track C bed) vs random-init relays. Books are (64,4) — size-agnostic."""
388
+ from exp013_augmentation_bed import _wikitext_lines
389
+ from transformers import GPT2LMHeadModel, GPT2TokenizerFast
390
+ if "MslRelay" not in globals():
391
+ from ar_differentiation_bed import MslRelay
392
+ from exp013_augmentation_bed import _BlockWithAdapter
393
+ tok = GPT2TokenizerFast.from_pretrained("gpt2")
394
+ tr_lines, va_lines = _wikitext_lines(DATA_ROOT)
395
+ stream_tr = tok("\n\n".join(tr_lines[:8000]), return_tensors="pt").input_ids[0]
396
+ stream_va = tok("\n\n".join(va_lines[:1000]), return_tensors="pt").input_ids[0]
397
+ ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
398
+ out = {}
399
+ for arm in ("random_relays", "implanted_relays"):
400
+ torch.manual_seed(seed)
401
+ g = torch.Generator().manual_seed(seed)
402
+ model = GPT2LMHeadModel.from_pretrained("gpt2").to(device)
403
+ for p in model.parameters():
404
+ p.requires_grad_(False)
405
+ adapters = []
406
+ for i, blk in enumerate(model.transformer.h):
407
+ ad = MslRelay(model.config.n_embd).to(device)
408
+ if arm == "implanted_relays":
409
+ implant_book(ad.addr, donor_book, trainable=True)
410
+ model.transformer.h[i] = _BlockWithAdapter(blk, ad)
411
+ adapters.append(ad)
412
+ params = [p for ad in adapters for p in ad.parameters()
413
+ if p.requires_grad]
414
+ opt = torch.optim.Adam(params, lr=1e-3, weight_decay=0.0)
415
+ block = 256
416
+ for step in range(1, steps + 1):
417
+ ix = torch.randint(0, stream_tr.numel() - block - 1, (8,), generator=g)
418
+ x = torch.stack([stream_tr[i:i + block] for i in ix]).to(device)
419
+ loss = model(x, labels=x).loss
420
+ opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
421
+ model.eval()
422
+ with torch.no_grad():
423
+ ls = []
424
+ for j in range(0, stream_va.numel() - block - 1, block * 4):
425
+ x = stream_va[j:j + block].unsqueeze(0).to(device)
426
+ ls.append(model(x, labels=x).loss.item())
427
+ ppl = math.exp(sum(ls) / len(ls))
428
+ gates = [round(ad.gate.sigmoid().item(), 4) for ad in adapters]
429
+ drifts = [round(anchor_drift(ad.addr.codebook, ad.addr.home)["mean"], 3)
430
+ for ad in adapters]
431
+ out[arm] = {"ppl": round(ppl, 3), "gates": gates, "drift": drifts}
432
+ rec = {"exp": "14b2", "arm": arm, "donor": tag, "seed": seed,
433
+ "ppl": round(ppl, 3), "gates": gates, "drift": drifts}
434
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
435
+ print(f"[14b2 {arm} donor={tag}] ppl={ppl:.3f} gates={gates[:3]}.. "
436
+ f"drift={drifts[:3]}..", flush=True)
437
+ del model; torch.cuda.empty_cache()
438
+ ledger.close()
439
+ return out
440
+
441
+
442
+ # ================================================================ smoke ===========
443
+ def smoke():
444
+ """CPU shapes/parse only: GPA ground truth, tree causality, implant, KD path."""
445
+ g = torch.Generator().manual_seed(0)
446
+ # GPA: two rotated (+row-sign-flipped) copies of one book must align back to it
447
+ A = F.normalize(torch.randn(64, 4, generator=g), dim=-1)
448
+ q, _ = torch.linalg.qr(torch.randn(4, 4, generator=g))
449
+ B = A @ q
450
+ B[::3] = -B[::3]
451
+ C, it, dl = consensus_codebook([A, B])
452
+ cos = (F.normalize(C, dim=-1) * A).sum(-1).abs().mean()
453
+ assert cos > 0.999, cos
454
+ print(f"GPA OK (iters={it}, delta={dl:.2e}, |cos to truth|={cos:.5f})")
455
+ x = torch.randint(0, 256, (2, 64))
456
+ for lineage in ("aleph_flat", "aleph_tree", "mlp_kd", "no_inherit"):
457
+ m = make_organism(lineage, d=96, layers=2, block=64, seed=0)
458
+ lg = m(x); assert lg.shape == (2, 64, 256); lg.sum().backward()
459
+ with torch.no_grad():
460
+ a = m(x)[0, 10]; x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
461
+ b = m(x2)[0, 10]
462
+ assert torch.allclose(a, b, atol=1e-4), lineage + " leaks"
463
+ gnm = genome_of(m)
464
+ if gnm:
465
+ inherit_genome(m, [gnm, gnm]) # self-consensus = identity-ish
466
+ print(lineage, "OK params",
467
+ f"{sum(p.numel() for p in m.parameters()):,}",
468
+ organism_vitals(m) if lineage != "mlp_kd" else {})
469
+ # KD path: teacher forward + KL backward
470
+ t = make_organism("mlp_kd", d=96, layers=2, block=64, seed=1)
471
+ s = make_organism("aleph_flat", d=96, layers=2, block=64, seed=2)
472
+ tp = F.softmax(t(x), -1).detach()
473
+ loss = F.kl_div(F.log_softmax(s(x), -1), tp, reduction="batchmean")
474
+ loss.backward()
475
+ print("KD OK — exp014 smoke passed (tournament on GPU: run_tournament(...))")
476
+
477
+
478
+ def _in_notebook():
479
+ try:
480
+ get_ipython() # type: ignore[name-defined] # noqa: F821
481
+ return True
482
+ except NameError:
483
+ return False
484
+
485
+
486
+ if __name__ == "__main__":
487
+ if _in_notebook():
488
+ smoke()
489
+ print("Notebook: run_tournament('aleph_flat'), run_b1(), run_b2(book).")
490
+ else:
491
+ import argparse
492
+ ap = argparse.ArgumentParser()
493
+ ap.add_argument("--mode", default="smoke",
494
+ choices=["smoke", "tournament", "b1", "b2"])
495
+ ap.add_argument("--lineage", default="aleph_full",
496
+ help="tournament lineage: aleph_flat|aleph_full|"
497
+ "aleph_weights|aleph_tree|mlp_kd|no_inherit")
498
+ ap.add_argument("--seed", type=int, default=0)
499
+ ap.add_argument("--steps", type=int, default=2000)
500
+ ap.add_argument("--genome", default="genomes/champion_aleph_full_t0_g3.pt",
501
+ help="donor genome .pt for --mode b1/b2 (uses its 'flat' book)")
502
+ a, _ = ap.parse_known_args()
503
+ if a.mode == "smoke":
504
+ smoke()
505
+ elif a.mode == "tournament":
506
+ run_tournament(a.lineage, steps=a.steps, seed=a.seed)
507
+ elif a.mode == "b1":
508
+ donor = (torch.load(a.genome, map_location="cpu")["flat"]
509
+ if os.path.exists(a.genome) else None)
510
+ run_b1(steps=a.steps, seed=a.seed, donor_book=donor,
511
+ tag=os.path.basename(a.genome) if donor is not None
512
+ else "small_cultivated")
513
+ elif a.mode == "b2":
514
+ donor = torch.load(a.genome, map_location="cpu")["flat"]
515
+ run_b2(donor, seed=a.seed, tag=os.path.basename(a.genome))
exp020_gen/exp017_aleph_constellation.py ADDED
@@ -0,0 +1,266 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """exp017_aleph_constellation.py — PROTOTYPE ALEPH CONSTELLATION.
2
+ The constellation element for the aleph address: the 16s bridge (per the
3
+ constellation-diffusion-bottleneck article: 16 anchors x 16 patches x 3
4
+ stroboscope SLERP phases = the 768 address; aleph D=4 home lifted to the S^15
5
+ measurement sphere via 16 = 4^2) implemented as a byte-LM head — the first
6
+ configuration in the AR bed where the constellation gauges legitimately apply.
7
+
8
+ Funnel (per position, the 16s numerology intact on the aleph D=4 home):
9
+ h(d) -> proj -> 16 patches x 32 rows of D=4 (in_per_patch = 32*4 = 128)
10
+ -> AlephAddress M_hat read per row (farms the codebook)
11
+ -> crush MLP 128 -> 48 -> 16, SquaredReLU (8x squeeze, rule-of-3 hidden)
12
+ -> row-normalize = S^15 point per patch (the measurement sphere; 16 = 4^2)
13
+ -> triangulate 16 constellation anchors x 3 SLERP strobe phases
14
+ tri_k(t) = cos((1-t) * theta_k), t in {0, 1/3, 2/3} (acos in fp32)
15
+ -> address 16 patches x 16 anchors x 3 = 768
16
+ -> patchwork Linear(768,1536) -> SquaredReLU -> LN -> Linear(1536, VOCAB)
17
+ Constellation rules honored (per the Constellation Forms Catalogue,
18
+ AbstractPhil/geolip-constellation-activations): SquaredReLU in
19
+ all constellation paths (never GELU); patchwork spec; anchor dropout 30%;
20
+ Procrustes CALIBRATION of the anchor init (non-negotiable rule; ablated in
21
+ const_uncal); acos fp32; Adam wd=0. No gate: this is the whole output path, not
22
+ a residual entry.
23
+
24
+ JUDGED BY (16s law): constellation anchor drift -> 0.29154, crushed CV -> 0.20,
25
+ and task bpb — NEVER recon cosine. The aleph codebook underneath keeps its own
26
+ regime (CV~0.9 volatile home; recon/CE gradient the only pressure on it).
27
+
28
+ Arms:
29
+ const — CV strictly a readout (aleph-side discipline)
30
+ const_cv — + micro CV loss 1e-3 on the constellation bank (Form 3:
31
+ "CV ON THE BANK is load-bearing"; forward loss, never backward
32
+ injection) — tests whether the constellation rule transfers
33
+ const_uncal — calibration ablation (random anchors, no Procrustes calibration;
34
+ the "without: 1/256 anchors" rule test)
35
+ Baselines (exp012 certified ledger, cited not rerun): addr_msl64 2.4990,
36
+ sdpa 2.5182 @2k/seed0. Head is NOT param-matched to addr_msl64 (~2.0M vs ~114K)
37
+ — the prototype question is health + gauge placement, not a matched win.
38
+ Riders: pure Adam wd=0; GPU-only verdicts; >=2 seeds; Colab-safe. Paste order:
39
+ geolip_vitals -> ar_differentiation_bed -> this file.
40
+ """
41
+ from __future__ import annotations
42
+ import json
43
+ import math
44
+ import os
45
+ import torch
46
+ import torch.nn as nn
47
+ import torch.nn.functional as F
48
+
49
+ if "ByteLM" not in globals():
50
+ try:
51
+ from ar_differentiation_bed import (ByteLM, AlephAddress, _wikitext_bytes,
52
+ _batch, VOCAB)
53
+ from geolip_vitals import anchor_drift, pentachoron_cv, BINDING
54
+ except ImportError:
55
+ _here = globals().get("__file__")
56
+ if _here is None:
57
+ raise ImportError("paste geolip_vitals.py + ar_differentiation_bed.py first")
58
+ import sys, pathlib
59
+ sys.path.insert(0, str(pathlib.Path(_here).parent))
60
+ from ar_differentiation_bed import (ByteLM, AlephAddress, _wikitext_bytes,
61
+ _batch, VOCAB)
62
+ from geolip_vitals import anchor_drift, pentachoron_cv, BINDING
63
+
64
+ DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
65
+ EXP17_DIR = os.path.join(DATA_ROOT, "exp017")
66
+
67
+ N_PATCHES, N_ANCH, V_ROWS, GEO = 16, 16, 32, 16 # 16 = 4^2 = d^2 (16s law)
68
+ PHASES = (0.0, 1.0 / 3.0, 2.0 / 3.0) # stroboscope SLERP, rule of 3
69
+
70
+
71
+ class SquaredReLU(nn.Module):
72
+ def forward(self, x):
73
+ return F.relu(x) ** 2
74
+
75
+
76
+ class ConstellationHead(nn.Module):
77
+ """The 16s funnel over the aleph address (docstring above)."""
78
+
79
+ def __init__(self, d: int, vocab: int = 256, calibrated: bool = True,
80
+ anchor_dropout: float = 0.30):
81
+ super().__init__()
82
+ self.calibrated = calibrated
83
+ self.anchor_dropout = anchor_dropout
84
+ self.proj = nn.Linear(d, N_PATCHES * V_ROWS * 4, bias=False)
85
+ nn.init.orthogonal_(self.proj.weight)
86
+ self.addr = AlephAddress(64, 4) # the aleph layer beneath
87
+ self.crush = nn.Sequential( # 128 -> 48 -> 16 (8x squeeze)
88
+ nn.Linear(V_ROWS * 4, GEO * 3), SquaredReLU(),
89
+ nn.Linear(GEO * 3, GEO))
90
+ A = F.normalize(torch.randn(N_ANCH, GEO), dim=-1)
91
+ self.anchors = nn.Parameter(A) # S^15 constellation bank
92
+ self.register_buffer("anchors_home", A.clone())
93
+ tri = N_PATCHES * N_ANCH * len(PHASES) # 768, the address
94
+ self.patchwork = nn.Sequential(
95
+ nn.Linear(tri, tri * 2), SquaredReLU(),
96
+ nn.LayerNorm(tri * 2), nn.Linear(tri * 2, vocab))
97
+ self._calibrated_done = not calibrated
98
+
99
+ @torch.no_grad()
100
+ def calibrate(self, h_sample: torch.Tensor):
101
+ """Procrustes calibration of the anchor init (NON-NEGOTIABLE constellation
102
+ rule): rotate the anchor template into the principal frame of the actual
103
+ crushed S^15 activations, then re-home. Runs once, before training."""
104
+ x = self._crushed(h_sample.reshape(-1, h_sample.shape[-1])) # (N, P, GEO)
105
+ x = x.reshape(-1, GEO).float()
106
+ # data frame: eigenvectors of the crushed covariance (fp64)
107
+ C = (x.T.double() @ x.double()) / x.shape[0]
108
+ _, Vd = torch.linalg.eigh(C)
109
+ A = self.anchors.double()
110
+ Ca = (A.T @ A) / A.shape[0]
111
+ _, Va = torch.linalg.eigh(Ca)
112
+ R = Va @ Vd.T # template frame -> data frame
113
+ A_cal = F.normalize((A @ R).float(), dim=-1)
114
+ self.anchors.data.copy_(A_cal)
115
+ self.anchors_home.copy_(A_cal)
116
+ self._calibrated_done = True
117
+
118
+ def _crushed(self, h):
119
+ rows = self.proj(h).view(*h.shape[:-1], N_PATCHES, V_ROWS, 4)
120
+ m_hat = self.addr.m_hat(rows) # aleph read per row
121
+ flat = m_hat.reshape(*h.shape[:-1], N_PATCHES, V_ROWS * 4)
122
+ return F.normalize(self.crush(flat), dim=-1) # (..., P, GEO) on S^15
123
+
124
+ def forward(self, h):
125
+ x = self._crushed(h) # (..., P, GEO)
126
+ A = F.normalize(self.anchors, dim=-1)
127
+ cos = (x @ A.T).clamp(-1 + 1e-6, 1 - 1e-6) # (..., P, K)
128
+ theta = torch.acos(cos.float()) # acos in fp32
129
+ tri = torch.cat([torch.cos((1.0 - t) * theta) for t in PHASES],
130
+ dim=-1).to(h.dtype) # (..., P, K*3)
131
+ if self.training and self.anchor_dropout > 0:
132
+ keep = (torch.rand(N_ANCH, device=h.device)
133
+ > self.anchor_dropout).float()
134
+ tri = tri * keep.repeat(len(PHASES))
135
+ return self.patchwork(tri.reshape(*h.shape[:-1],
136
+ N_PATCHES * N_ANCH * len(PHASES)))
137
+
138
+ @torch.no_grad()
139
+ def constellation_vitals(self) -> dict:
140
+ d = anchor_drift(self.anchors, self.anchors_home)
141
+ return {"anchor_drift": round(d["mean"], 4),
142
+ "binding_frac": round(d["binding_fraction"], 4),
143
+ "crushed_cv": round(pentachoron_cv(self.anchors.detach()), 4),
144
+ "aleph": {"drift": round(anchor_drift(
145
+ self.addr.codebook, self.addr.home)["mean"], 4)}}
146
+
147
+
148
+ def make_const_model(arm: str, d: int = 192, layers: int = 4,
149
+ block: int = 256, seed: int = 0):
150
+ torch.manual_seed(seed)
151
+ m = ByteLM("sdpa", d=d, layers=layers, block=block)
152
+ m.head = ConstellationHead(d, calibrated=(arm != "const_uncal"))
153
+ return m
154
+
155
+
156
+ def cv_bank_loss(head: ConstellationHead) -> torch.Tensor:
157
+ """Micro CV loss on the CONSTELLATION bank only (const_cv arm; Form 3).
158
+ Forward loss (differentiable through the anchors); never touches the aleph
159
+ codebook, whose only pressure stays the CE gradient through M_hat."""
160
+ A = F.normalize(head.anchors, dim=-1)
161
+ n = A.shape[0]
162
+ g = torch.Generator(device="cpu").manual_seed(0)
163
+ idx = torch.stack([torch.randperm(n, generator=g)[:5] for _ in range(64)])
164
+ pts = A[idx] # (64, 5, GEO)
165
+ d2 = torch.cdist(pts.double(), pts.double()).pow(2)
166
+ cm = torch.ones(64, 6, 6, dtype=torch.float64, device=A.device)
167
+ cm[:, 0, 0] = 0.0
168
+ cm[:, 1:, 1:] = d2
169
+ v = (-torch.linalg.det(cm) / 9216.0).clamp_min(1e-24).sqrt()
170
+ return (v.std() / v.mean().clamp_min(1e-12)).float()
171
+
172
+
173
+ def train_const(model, tr, va, arm: str, steps=2000, batch=32, block=256,
174
+ device="cuda", seed=0, cv_weight=1e-3, log_every=500,
175
+ ledger=None, tag=None):
176
+ g = torch.Generator().manual_seed(seed)
177
+ model = model.to(device)
178
+ if not model.head._calibrated_done: # calibration pass (1 batch)
179
+ x, _ = _batch(tr, batch, block, device, g)
180
+ model(x)
181
+ model.head.calibrate(model._last_h)
182
+ opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
183
+ for step in range(1, steps + 1):
184
+ x, y = _batch(tr, batch, block, device, g)
185
+ logits = model(x)
186
+ loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
187
+ if arm == "const_cv":
188
+ loss = loss + cv_weight * cv_bank_loss(model.head)
189
+ opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
190
+ if step % log_every == 0:
191
+ cv = model.head.constellation_vitals()
192
+ print(f"[17 {tag or arm} step {step}] {cv}", flush=True)
193
+ model.eval()
194
+ with torch.no_grad():
195
+ ls = []
196
+ for _ in range(20):
197
+ xv, yv = _batch(va, batch, block, device, g)
198
+ ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
199
+ yv.reshape(-1)).item())
200
+ bpb = sum(ls) / len(ls) / math.log(2)
201
+ cv = model.head.constellation_vitals()
202
+ if ledger is not None:
203
+ rec = {"exp": "17", "arm": arm, "seed": seed, "steps": steps,
204
+ "bpb": round(bpb, 4), "constellation": cv}
205
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
206
+ print(f"[17 {tag or arm} s{seed}] FINAL bpb={bpb:.4f} {cv}", flush=True)
207
+ return bpb
208
+
209
+
210
+ def run_prototype(arms=("const", "const_cv", "const_uncal"), seeds=(0, 1),
211
+ steps=2000, device="cuda"):
212
+ if not torch.cuda.is_available():
213
+ raise RuntimeError("verdict runs are GPU-only")
214
+ os.makedirs(EXP17_DIR, exist_ok=True)
215
+ tr, va = _wikitext_bytes(DATA_ROOT)
216
+ ledger = open(os.path.join(EXP17_DIR, "ledger.jsonl"), "a", encoding="utf-8")
217
+ for seed in seeds:
218
+ for arm in arms:
219
+ m = make_const_model(arm, seed=seed)
220
+ train_const(m, tr, va, arm, steps=steps, device=device, seed=seed,
221
+ ledger=ledger, tag=f"{arm}")
222
+ torch.save({"arm": arm, "seed": seed, "steps": steps,
223
+ "state_dict": {k: v.cpu() for k, v in
224
+ m.state_dict().items()}},
225
+ os.path.join(EXP17_DIR, f"const_{arm}_s{seed}.pt"))
226
+ del m
227
+ torch.cuda.empty_cache()
228
+ ledger.close()
229
+
230
+
231
+ def smoke():
232
+ x = torch.randint(0, 256, (2, 64))
233
+ m = make_const_model("const", d=96, layers=2, block=64, seed=0)
234
+ lg = m(x)
235
+ assert lg.shape == (2, 64, 256)
236
+ lg.sum().backward()
237
+ assert m.head.addr.codebook.grad is not None # aleph farms through M_hat
238
+ assert m.head.anchors.grad is not None # constellation trains
239
+ m.zero_grad()
240
+ m(x)
241
+ m.head.calibrate(m._last_h) # calibration runs
242
+ cv = m.head.constellation_vitals()
243
+ assert "crushed_cv" in cv and cv["anchor_drift"] == 0.0
244
+ l = cv_bank_loss(m.head)
245
+ assert l.requires_grad and l.item() > 0
246
+ # causality: future byte must not affect past logits
247
+ x2 = x.clone(); x2[:, -1] = (x2[:, -1] + 1) % 256
248
+ m.eval()
249
+ with torch.no_grad():
250
+ a, b = m(x), m(x2)
251
+ assert torch.allclose(a[:, :-1], b[:, :-1], atol=1e-5)
252
+ print("exp017 smoke passed —", {k: cv[k] for k in ("crushed_cv",
253
+ "binding_frac")})
254
+
255
+
256
+ def _in_notebook():
257
+ try:
258
+ get_ipython() # type: ignore[name-defined] # noqa: F821
259
+ return True
260
+ except NameError:
261
+ return False
262
+
263
+
264
+ if __name__ == "__main__":
265
+ smoke() if not _in_notebook() else (smoke(),
266
+ print("Notebook: run_prototype() on GPU."))
exp020_gen/exp019_content_retention.py ADDED
@@ -0,0 +1,402 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """exp019_content_retention.py — CAPACITY FOR DISTILLED CONTENT RETENTION.
2
+ When content is DISTILLED rather than directly learned, how much is retained,
3
+ through which channel, and for how long? The memory-substrate compass made a
4
+ measurable instrument — and the weak-to-strong distillation cell made real: the
5
+ teacher holds content the student lacks (true headroom on the content axis).
6
+
7
+ THE CONTENT: N fact records "\n@<key6>=<value12>\n" with random alphanumeric
8
+ keys/values — uncompletable from language statistics, so exact-match completion
9
+ IS retention. Facts are mixed into the byte stream (fact-packed blocks at
10
+ FACT_RATE against wikitext blocks).
11
+
12
+ THE TEACHER: the certified bed model (addr_msl64 aleph substrate) trained
13
+ TEACHER_STEPS on the mix; gated on its own recall (the gate doubles as the
14
+ substrate's DIRECT capacity datum at each N).
15
+
16
+ THE CHANNELS (fresh student each, STUDENT_STEPS budget):
17
+ direct — ground-truth CE on the mix (ceiling: learning, not distillation)
18
+ kd_facts — CE on wikitext blocks; on fact blocks the ONLY signal is the
19
+ teacher's logits (KL) — pure distilled content
20
+ kd_general — CE + KL to teacher on CLEAN wikitext only; facts never shown —
21
+ does content leak through logits without exposure?
22
+ book_implant — teacher's farmed codebook implanted (trainable) into a fresh
23
+ student, clean-stream training — do the anchors carry
24
+ byte-content? (the open question from the exp014 implant studies, asked directly)
25
+ none — clean-stream only (floor)
26
+ THE AXES: capacity N in {64, 256, 1024} (main channels); retention = recall
27
+ right after training AND after INTERFERE_STEPS further clean-stream steps
28
+ (the forgetting measurement). KD alpha 1.0 here is LEGAL: the teacher has real
29
+ headroom on the content axis (the inverse-evolution failure was alpha 1.0 at
30
+ NEAR-PARITY — regime, not constant).
31
+
32
+ Preregistered forks:
33
+ F1 capacity curve: direct recall vs N = the substrate's raw content capacity.
34
+ F2 distillation tax: kd_facts vs direct at each N (what survives the logit
35
+ channel).
36
+ F3 leakage: kd_general recall > floor => content crosses on clean text alone.
37
+ F4 anchors: book_implant recall ~ floor => codebooks do not carry byte
38
+ content (mean-shape/content question closed in the direct sense).
39
+ F5 half-life: post-interference retention per channel (does distilled content
40
+ decay faster than learned content?).
41
+ Riders: pure Adam wd=0; GPU-only verdicts; >=2 seeds; Colab-safe. Paste order:
42
+ geolip_vitals -> ar_differentiation_bed -> exp014_genetic_distillation ->
43
+ this file.
44
+ """
45
+ from __future__ import annotations
46
+ import json
47
+ import math
48
+ import os
49
+ import string
50
+ import torch
51
+ import torch.nn.functional as F
52
+
53
+ if "ByteLM" not in globals():
54
+ try:
55
+ from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
56
+ from exp014_genetic_distillation import implant_book
57
+ from geolip_vitals import anchor_drift
58
+ except ImportError:
59
+ _here = globals().get("__file__")
60
+ if _here is None:
61
+ raise ImportError("paste geolip_vitals + ar_differentiation_bed + "
62
+ "exp014_genetic_distillation first")
63
+ import sys, pathlib
64
+ sys.path.insert(0, str(pathlib.Path(_here).parent))
65
+ from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
66
+ from exp014_genetic_distillation import implant_book
67
+ from geolip_vitals import anchor_drift
68
+
69
+ DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
70
+ EXP19_DIR = os.path.join(DATA_ROOT, "exp019")
71
+
72
+ KEY_LEN, VAL_LEN = 6, 12
73
+ FACT_RATE = 0.5 # fraction of training blocks drawn from fact stream
74
+ TEACHER_STEPS = 4000
75
+ STUDENT_STEPS = 2000
76
+ INTERFERE_STEPS = 1000
77
+ ALNUM = (string.ascii_lowercase + string.digits).encode()
78
+
79
+
80
+ def make_facts(n: int, seed: int = 0):
81
+ """N records '\\n@<key>=<value>\\n'; returns (records list, fact byte stream)."""
82
+ g = torch.Generator().manual_seed(4000 + seed)
83
+ recs = []
84
+ seen = set()
85
+ while len(recs) < n:
86
+ k = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (KEY_LEN,),
87
+ generator=g))
88
+ if k in seen:
89
+ continue
90
+ seen.add(k)
91
+ v = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (VAL_LEN,),
92
+ generator=g))
93
+ recs.append((k, v))
94
+ return recs
95
+
96
+
97
+ def fact_stream(recs, copies: int = 50, seed: int = 0) -> torch.Tensor:
98
+ """Byte stream of shuffled fact records (each record appears `copies` times)."""
99
+ g = torch.Generator().manual_seed(5000 + seed)
100
+ order = torch.cat([torch.randperm(len(recs), generator=g)
101
+ for _ in range(copies)])
102
+ blob = b"".join(b"\n@" + recs[i][0] + b"=" + recs[i][1] + b"\n"
103
+ for i in order.tolist())
104
+ return torch.frombuffer(bytearray(blob), dtype=torch.uint8).clone()
105
+
106
+
107
+ def _mix_batch(tr, fs, batch, block, device, g):
108
+ """Blocks drawn from the fact stream with prob FACT_RATE, else wikitext.
109
+ Returns (x, y, fact_mask (B,)) — mask marks fact-sourced rows."""
110
+ xw, yw = _batch(tr, batch, block, device, g)
111
+ xf, yf = _batch(fs, batch, block, device, g)
112
+ m = (torch.rand(batch, generator=g) < FACT_RATE).to(device)
113
+ x = torch.where(m[:, None], xf, xw)
114
+ y = torch.where(m[:, None], yf, yw)
115
+ return x, y, m
116
+
117
+
118
+ @torch.no_grad()
119
+ def recall(model, recs, device="cuda", max_eval: int = 256,
120
+ batch: int = 64) -> dict:
121
+ """Exact-match greedy completion: prompt '\\n@<key>=' -> VAL_LEN bytes."""
122
+ model = model.to(device).eval()
123
+ recs = recs[:max_eval]
124
+ prompts = torch.stack([torch.frombuffer(
125
+ bytearray(b"\n@" + k + b"="), dtype=torch.uint8).long()
126
+ for k, _ in recs]).to(device)
127
+ outs = []
128
+ for i in range(0, len(recs), batch):
129
+ x = prompts[i:i + batch]
130
+ for _ in range(VAL_LEN):
131
+ nxt = model(x)[:, -1].argmax(-1, keepdim=True)
132
+ x = torch.cat([x, nxt], dim=1)
133
+ outs.append(x[:, -VAL_LEN:].cpu())
134
+ got = torch.cat(outs)
135
+ tgt = torch.stack([torch.frombuffer(bytearray(v), dtype=torch.uint8).long()
136
+ for _, v in recs])
137
+ byte_acc = (got == tgt).float().mean().item()
138
+ exact = (got == tgt).all(dim=1).float().mean().item()
139
+ return {"exact": round(exact, 4), "byte_acc": round(byte_acc, 4)}
140
+
141
+
142
+ def train_stream(model, tr, va, fs=None, teacher=None, channel="direct",
143
+ steps=2000, batch=32, block=256, device="cuda", seed=0):
144
+ """One training run under a channel's signal routing (docstring above)."""
145
+ g = torch.Generator().manual_seed(seed)
146
+ model = model.to(device)
147
+ if teacher is not None:
148
+ teacher = teacher.to(device).eval()
149
+ opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
150
+ for step in range(1, steps + 1):
151
+ if channel in ("direct", "kd_facts") and fs is not None:
152
+ x, y, m = _mix_batch(tr, fs, batch, block, device, g)
153
+ else: # clean wikitext stream
154
+ x, y = _batch(tr, batch, block, device, g)
155
+ m = torch.zeros(batch, dtype=torch.bool, device=device)
156
+ logits = model(x)
157
+ if channel == "kd_facts":
158
+ # ground truth on wiki rows only; teacher logits are the ONLY
159
+ # signal on fact rows (pure distilled content)
160
+ ce_rows = ~m
161
+ loss = torch.tensor(0.0, device=device)
162
+ if ce_rows.any():
163
+ loss = F.cross_entropy(logits[ce_rows].reshape(-1, VOCAB),
164
+ y[ce_rows].reshape(-1))
165
+ if m.any():
166
+ with torch.no_grad():
167
+ tp = F.softmax(teacher(x[m]), -1)
168
+ loss = loss + F.kl_div(F.log_softmax(logits[m], -1), tp,
169
+ reduction="batchmean")
170
+ elif channel == "kd_general":
171
+ loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
172
+ with torch.no_grad():
173
+ tp = F.softmax(teacher(x), -1)
174
+ loss = loss + F.kl_div(F.log_softmax(logits, -1), tp,
175
+ reduction="batchmean")
176
+ else: # direct / book_implant / none
177
+ loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
178
+ opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
179
+ model.eval()
180
+ with torch.no_grad():
181
+ ls = []
182
+ for _ in range(10):
183
+ xv, yv = _batch(va, batch, block, device, g)
184
+ ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
185
+ yv.reshape(-1)).item())
186
+ return sum(ls) / len(ls) / math.log(2)
187
+
188
+
189
+ def run_retention(n_facts=(64, 256, 1024), seeds=(0, 1), device="cuda"):
190
+ if not torch.cuda.is_available():
191
+ raise RuntimeError("verdict runs are GPU-only")
192
+ os.makedirs(EXP19_DIR, exist_ok=True)
193
+ tr, va = _wikitext_bytes(DATA_ROOT)
194
+ ledger = open(os.path.join(EXP19_DIR, "ledger.jsonl"), "a", encoding="utf-8")
195
+
196
+ def log(rec):
197
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
198
+ print(f"[19 {rec['channel']} N={rec['n']} s{rec['seed']}] "
199
+ f"recall={rec['recall']} after_interf={rec.get('recall_interf')} "
200
+ f"bpb={rec['bpb']}", flush=True)
201
+
202
+ for seed in seeds:
203
+ for n in n_facts:
204
+ recs = make_facts(n, seed=seed)
205
+ fs = fact_stream(recs, seed=seed)
206
+ # ---- teacher (also the DIRECT capacity datum at TEACHER_STEPS)
207
+ torch.manual_seed(seed)
208
+ teacher = ByteLM("addr_msl64")
209
+ t_bpb = train_stream(teacher, tr, va, fs=fs, channel="direct",
210
+ steps=TEACHER_STEPS, device=device, seed=seed)
211
+ t_rec = recall(teacher, recs, device=device)
212
+ log({"exp": "19", "channel": "teacher", "n": n, "seed": seed,
213
+ "steps": TEACHER_STEPS, "recall": t_rec, "bpb": round(t_bpb, 4)})
214
+ torch.save({"n": n, "seed": seed,
215
+ "state_dict": {k: v.cpu() for k, v in
216
+ teacher.state_dict().items()}},
217
+ os.path.join(EXP19_DIR, f"teacher_N{n}_s{seed}.pt"))
218
+ # ---- channels
219
+ chans = ["direct", "kd_facts", "kd_general"]
220
+ if n == 256:
221
+ chans += ["book_implant", "none"]
222
+ for ch in chans:
223
+ torch.manual_seed(1000 + seed)
224
+ m = ByteLM("addr_msl64")
225
+ if ch == "book_implant":
226
+ implant_book(m.head_addr,
227
+ teacher.head_addr.codebook.detach().cpu())
228
+ bpb = train_stream(
229
+ m, tr, va, fs=fs if ch in ("direct", "kd_facts") else None,
230
+ teacher=teacher if ch.startswith("kd") else None,
231
+ channel=ch, steps=STUDENT_STEPS, device=device,
232
+ seed=1000 + seed)
233
+ r0 = recall(m, recs, device=device)
234
+ # retention under interference: further CLEAN-stream training
235
+ bpb2 = train_stream(m, tr, va, channel="none",
236
+ steps=INTERFERE_STEPS, device=device,
237
+ seed=2000 + seed)
238
+ r1 = recall(m, recs, device=device)
239
+ log({"exp": "19", "channel": ch, "n": n, "seed": seed,
240
+ "steps": STUDENT_STEPS, "recall": r0, "recall_interf": r1,
241
+ "bpb": round(bpb, 4), "bpb_after_interf": round(bpb2, 4)})
242
+ del m
243
+ torch.cuda.empty_cache()
244
+ del teacher
245
+ torch.cuda.empty_cache()
246
+ ledger.close()
247
+
248
+
249
+ # ==================== exp019b — GENERALIZATION block =========================
250
+ # Rule-bearing content: value = fixed random substitution cipher applied to the
251
+ # key, extended to VAL_LEN (v[i] = subst(k[i % KEY_LEN])). Teacher sees
252
+ # N_TRAIN rule-keys; N_TEST keys are HELD OUT. Held-out recall = the RULE
253
+ # generalizing, not the list. Sharp question: does the logit channel transfer
254
+ # the rule better than it transfers the rote list? Plus prompt-format variants
255
+ # (content vs surface form disentangled).
256
+
257
+ def make_rule_facts(n_train: int = 256, n_test: int = 128, seed: int = 0):
258
+ g = torch.Generator().manual_seed(6000 + seed)
259
+ subst = {ALNUM[i]: ALNUM[j] for i, j in
260
+ enumerate(torch.randperm(len(ALNUM), generator=g).tolist())}
261
+ keys, seen = [], set()
262
+ while len(keys) < n_train + n_test:
263
+ k = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (KEY_LEN,),
264
+ generator=g))
265
+ if k not in seen:
266
+ seen.add(k)
267
+ keys.append(k)
268
+ def val(k):
269
+ return bytes(subst[k[i % KEY_LEN]] for i in range(VAL_LEN))
270
+ train = [(k, val(k)) for k in keys[:n_train]]
271
+ test = [(k, val(k)) for k in keys[n_train:]]
272
+ return train, test
273
+
274
+
275
+ @torch.no_grad()
276
+ def recall_fmt(model, recs, fmt: bytes = b"\n@%s=", device="cuda",
277
+ max_eval: int = 256, batch: int = 64) -> dict:
278
+ """recall() under an arbitrary prompt format (b'\\n@%s=' = the training
279
+ format; variants probe surface-form generalization)."""
280
+ model = model.to(device).eval()
281
+ recs = recs[:max_eval]
282
+ proms = [torch.frombuffer(bytearray(fmt.replace(b"%s", k)),
283
+ dtype=torch.uint8).long() for k, _ in recs]
284
+ L = max(p.numel() for p in proms)
285
+ # left-pad with newlines to equal length (causal — padding is prefix noise)
286
+ prompts = torch.stack([torch.cat([torch.full((L - p.numel(),), 10,
287
+ dtype=torch.long), p])
288
+ for p in proms]).to(device)
289
+ outs = []
290
+ for i in range(0, len(recs), batch):
291
+ x = prompts[i:i + batch]
292
+ for _ in range(VAL_LEN):
293
+ nxt = model(x)[:, -1].argmax(-1, keepdim=True)
294
+ x = torch.cat([x, nxt], dim=1)
295
+ outs.append(x[:, -VAL_LEN:].cpu())
296
+ got = torch.cat(outs)
297
+ tgt = torch.stack([torch.frombuffer(bytearray(v), dtype=torch.uint8).long()
298
+ for _, v in recs])
299
+ return {"exact": round((got == tgt).all(dim=1).float().mean().item(), 4),
300
+ "byte_acc": round((got == tgt).float().mean().item(), 4)}
301
+
302
+
303
+ FMT_TRAIN = b"\n@%s="
304
+ FMT_VARIANT = b" @%s= " # never seen in training: pure format shift
305
+
306
+
307
+ def run_generalization(n_train: int = 256, n_test: int = 128, seeds=(0, 1),
308
+ device="cuda"):
309
+ if not torch.cuda.is_available():
310
+ raise RuntimeError("verdict runs are GPU-only")
311
+ os.makedirs(EXP19_DIR, exist_ok=True)
312
+ tr, va = _wikitext_bytes(DATA_ROOT)
313
+ ledger = open(os.path.join(EXP19_DIR, "ledger.jsonl"), "a", encoding="utf-8")
314
+
315
+ def gauges(model, train_recs, test_recs):
316
+ return {"train": recall_fmt(model, train_recs, FMT_TRAIN, device=device),
317
+ "heldout": recall_fmt(model, test_recs, FMT_TRAIN, device=device),
318
+ "train_varfmt": recall_fmt(model, train_recs, FMT_VARIANT,
319
+ device=device)}
320
+
321
+ for seed in seeds:
322
+ train_recs, test_recs = make_rule_facts(n_train, n_test, seed=seed)
323
+ fs = fact_stream(train_recs, seed=seed) # held-out NEVER streamed
324
+ torch.manual_seed(seed)
325
+ teacher = ByteLM("addr_msl64")
326
+ t_bpb = train_stream(teacher, tr, va, fs=fs, channel="direct",
327
+ steps=TEACHER_STEPS, device=device, seed=seed)
328
+ gt = gauges(teacher, train_recs, test_recs)
329
+ rec = {"exp": "19b", "channel": "teacher", "n": n_train, "seed": seed,
330
+ "steps": TEACHER_STEPS, "gauges": gt, "bpb": round(t_bpb, 4)}
331
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
332
+ print(f"[19b teacher s{seed}] {gt} bpb={t_bpb:.4f}", flush=True)
333
+ torch.save({"seed": seed, "state_dict": {k: v.cpu() for k, v in
334
+ teacher.state_dict().items()}},
335
+ os.path.join(EXP19_DIR, f"rule_teacher_s{seed}.pt"))
336
+ for ch in ("direct", "kd_facts", "kd_general"):
337
+ torch.manual_seed(1000 + seed)
338
+ m = ByteLM("addr_msl64")
339
+ bpb = train_stream(
340
+ m, tr, va, fs=fs if ch in ("direct", "kd_facts") else None,
341
+ teacher=teacher if ch.startswith("kd") else None,
342
+ channel=ch, steps=STUDENT_STEPS, device=device, seed=1000 + seed)
343
+ g0 = gauges(m, train_recs, test_recs)
344
+ bpb2 = train_stream(m, tr, va, channel="none",
345
+ steps=INTERFERE_STEPS, device=device,
346
+ seed=2000 + seed)
347
+ g1 = gauges(m, train_recs, test_recs)
348
+ rec = {"exp": "19b", "channel": ch, "n": n_train, "seed": seed,
349
+ "steps": STUDENT_STEPS, "gauges": g0, "gauges_interf": g1,
350
+ "bpb": round(bpb, 4), "bpb_after_interf": round(bpb2, 4)}
351
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
352
+ print(f"[19b {ch} s{seed}] {g0} interf_heldout="
353
+ f"{g1['heldout']} bpb={bpb:.4f}", flush=True)
354
+ torch.save({"channel": ch, "seed": seed,
355
+ "state_dict": {k: v.cpu() for k, v in
356
+ m.state_dict().items()}},
357
+ os.path.join(EXP19_DIR, f"rule_{ch}_s{seed}.pt"))
358
+ del m
359
+ torch.cuda.empty_cache()
360
+ del teacher
361
+ torch.cuda.empty_cache()
362
+ ledger.close()
363
+
364
+
365
+ def smoke():
366
+ recs = make_facts(8, seed=0)
367
+ assert len(recs) == 8 and all(len(k) == KEY_LEN and len(v) == VAL_LEN
368
+ for k, v in recs)
369
+ fs = fact_stream(recs, copies=3, seed=0)
370
+ assert fs.dtype == torch.uint8 and fs.numel() == 3 * 8 * (KEY_LEN + VAL_LEN + 4)
371
+ m = ByteLM("addr_msl64", d=96, layers=2, block=64)
372
+ r = recall(m, recs, device="cpu", max_eval=8, batch=4)
373
+ assert 0.0 <= r["exact"] <= 1.0 and 0.0 <= r["byte_acc"] <= 1.0
374
+ g = torch.Generator().manual_seed(0)
375
+ x, y, mask = _mix_batch(torch.randint(0, 256, (50000,),
376
+ dtype=torch.uint8, generator=g),
377
+ fs, 8, 64, "cpu", g)
378
+ assert x.shape == (8, 64) and mask.shape == (8,)
379
+ assert (x[:, 1:] == y[:, :-1]).all() # stream alignment
380
+ # 19b: rule facts are rule-consistent + disjoint; variant recall runs
381
+ tr8, te4 = make_rule_facts(8, 4, seed=0)
382
+ assert len(tr8) == 8 and len(te4) == 4
383
+ assert not set(k for k, _ in tr8) & set(k for k, _ in te4)
384
+ k0, v0 = tr8[0]
385
+ assert len(v0) == VAL_LEN and v0[:KEY_LEN] == v0[KEY_LEN:2 * KEY_LEN]
386
+ rv = recall_fmt(m, tr8, FMT_VARIANT, device="cpu", max_eval=8, batch=4)
387
+ assert 0.0 <= rv["exact"] <= 1.0
388
+ print(f"exp019 smoke passed (untrained recall exact={r['exact']} "
389
+ f"byte={r['byte_acc']} ~ chance; 19b rule+variant OK)")
390
+
391
+
392
+ def _in_notebook():
393
+ try:
394
+ get_ipython() # type: ignore[name-defined] # noqa: F821
395
+ return True
396
+ except NameError:
397
+ return False
398
+
399
+
400
+ if __name__ == "__main__":
401
+ smoke() if not _in_notebook() else (smoke(),
402
+ print("Notebook: run_retention() on GPU."))
exp020_gen/exp020_generalization.py ADDED
@@ -0,0 +1,231 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """exp020_generalization.py — GENERALIZATION via the program's best structural
2
+ methodologies (codebooks + constellations), raced on the exp019b rule task.
3
+ exp019b established the failure modes: teachers learn a substitution cipher
4
+ only fragmentarily (held-out byte ~0.26, exact 0.000) and bind content to its
5
+ surface format. exp020 asks what UNLOCKS generalization: geometric structure
6
+ (aleph codebook reads, trigram character embeddings, addresses in depth, the
7
+ constellation 768 address) and data-side format diversity — against the honest
8
+ capacity control (the exp017 lesson: always run the matched free head).
9
+
10
+ TASK (identical to exp019b for comparability): value = fixed random
11
+ substitution cipher of the key (v[i] = subst(k[i % 6]), VAL_LEN 12); 256
12
+ train keys in the stream at FACT_RATE 0.5, 128 keys HELD OUT; 4000 steps
13
+ direct training. JUDGE: held-out exact/byte (rule induction), train recall,
14
+ variant-format recall (the format lock), clean bpb.
15
+
16
+ ARMS (2 seeds each):
17
+ aleph — ByteLM('addr_msl64') (the exp019b reference, in-harness)
18
+ aleph_tri — ByteLM('addr_msl64_tri') (trigram byte embeddings: certified
19
+ -10% bpb; character structure is
20
+ the cipher's own factorization)
21
+ relay — ByteLM('relay_msl64') (addresses in depth, GPT-2-certified
22
+ methodology on the co-trained bed)
23
+ const — ByteLM('sdpa') + ConstellationHead (exp017: the 768 address;
24
+ params NOT matched — recorded)
25
+ mlp — matched free head (params == aleph head budget; capacity control)
26
+ aleph_fmtdiv — addr_msl64 with FORMAT-DIVERSE fact rendering (3 formats in the
27
+ stream; the exp019b format-lock finding turned into a training
28
+ methodology) — judged on the base format + the UNSEEN variant
29
+ Preregistered forks:
30
+ F1 does ANY structural arm lift held-out byte acc off the flat ~0.26 plateau
31
+ (rule induction unlocked by geometry)?
32
+ F2 structure vs capacity: const/aleph vs mlp on held-out at recorded budgets.
33
+ F3 trigram: strongest prior for a per-character cipher.
34
+ F4 format diversity: does it break the format lock (variant recall >> 0) and
35
+ does breaking the lock ALSO lift rule induction?
36
+ F5 relays in depth: does distributed addressing help where the head alone
37
+ plateaus?
38
+ Riders: pure Adam wd=0; GPU-only verdicts; >=2 seeds; Colab-safe. Paste order:
39
+ geolip_vitals -> ar_differentiation_bed -> exp014_genetic_distillation ->
40
+ exp017_aleph_constellation -> exp019_content_retention -> this file.
41
+ """
42
+ from __future__ import annotations
43
+ import json
44
+ import math
45
+ import os
46
+ import torch
47
+ import torch.nn as nn
48
+ import torch.nn.functional as F
49
+
50
+ if "make_rule_facts" not in globals():
51
+ try:
52
+ from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
53
+ from exp017_aleph_constellation import ConstellationHead, SquaredReLU
54
+ from exp019_content_retention import (
55
+ make_rule_facts, recall_fmt, FMT_TRAIN, FMT_VARIANT,
56
+ KEY_LEN, VAL_LEN, FACT_RATE, ALNUM)
57
+ except ImportError:
58
+ _here = globals().get("__file__")
59
+ if _here is None:
60
+ raise ImportError("paste the stack (vitals, bed, exp017, exp019) first")
61
+ import sys, pathlib
62
+ sys.path.insert(0, str(pathlib.Path(_here).parent))
63
+ from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
64
+ from exp017_aleph_constellation import ConstellationHead, SquaredReLU
65
+ from exp019_content_retention import (
66
+ make_rule_facts, recall_fmt, FMT_TRAIN, FMT_VARIANT,
67
+ KEY_LEN, VAL_LEN, FACT_RATE, ALNUM)
68
+
69
+ DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
70
+ EXP20_DIR = os.path.join(DATA_ROOT, "exp020")
71
+ STEPS = 4000
72
+
73
+ # fact renderers: index 0 is the base format (matches FMT_TRAIN prompts);
74
+ # fmtdiv streams all three. FMT_VARIANT (" @%s= ") stays UNSEEN by every arm.
75
+ RENDERERS = (lambda k, v: b"\n@" + k + b"=" + v + b"\n",
76
+ lambda k, v: b"\n" + k + b" -> " + v + b"\n",
77
+ lambda k, v: b"\n<" + k + b"|" + v + b">\n")
78
+ FMT_R1 = b"\n%s -> " # prompt form of renderer 1 (fmtdiv-seen)
79
+
80
+
81
+ def fact_stream_fmt(recs, renderers, copies: int = 50, seed: int = 0):
82
+ g = torch.Generator().manual_seed(5000 + seed)
83
+ order = torch.cat([torch.randperm(len(recs), generator=g)
84
+ for _ in range(copies)])
85
+ rsel = torch.randint(len(renderers), (order.numel(),), generator=g)
86
+ blob = b"".join(renderers[rsel[i].item()](*recs[order[i].item()])
87
+ for i in range(order.numel()))
88
+ return torch.frombuffer(bytearray(blob), dtype=torch.uint8).clone()
89
+
90
+
91
+ def make_arm(arm: str, seed: int, d: int = 192, layers: int = 4,
92
+ block: int = 256):
93
+ torch.manual_seed(seed)
94
+ if arm in ("aleph", "aleph_fmtdiv"):
95
+ return ByteLM("addr_msl64", d=d, layers=layers, block=block)
96
+ if arm == "aleph_tri":
97
+ return ByteLM("addr_msl64_tri", d=d, layers=layers, block=block)
98
+ if arm == "relay":
99
+ return ByteLM("relay_msl64", d=d, layers=layers, block=block)
100
+ if arm == "const":
101
+ m = ByteLM("sdpa", d=d, layers=layers, block=block)
102
+ m.head = ConstellationHead(d)
103
+ return m
104
+ if arm == "mlp":
105
+ m = ByteLM("sdpa", d=d, layers=layers, block=block)
106
+ ref = ByteLM("addr_msl64", d=d, layers=layers, block=block)
107
+ target = (sum(p.numel() for p in ref.head_proj.parameters())
108
+ + sum(p.numel() for p in ref.head_addr.parameters())
109
+ + sum(p.numel() for p in ref.head.parameters()))
110
+ h = max(8, round((target - VOCAB) / (d + 1 + 2 + VOCAB)))
111
+ m.head = nn.Sequential(nn.Linear(d, h), SquaredReLU(),
112
+ nn.LayerNorm(h), nn.Linear(h, VOCAB))
113
+ del ref
114
+ return m
115
+ raise ValueError(arm)
116
+
117
+
118
+ def head_params(m) -> int:
119
+ tot = 0
120
+ for name in ("head", "head_proj", "head_addr"):
121
+ mod = getattr(m, name, None)
122
+ if mod is not None and isinstance(mod, nn.Module):
123
+ tot += sum(p.numel() for p in mod.parameters())
124
+ if hasattr(m, "relays"):
125
+ tot += sum(p.numel() for p in m.relays.parameters())
126
+ return tot
127
+
128
+
129
+ def train_direct(model, tr, va, fs, steps=STEPS, batch=32, block=256,
130
+ device="cuda", seed=0, needs_calibration=False):
131
+ g = torch.Generator().manual_seed(seed)
132
+ model = model.to(device)
133
+ if needs_calibration:
134
+ xw, _ = _batch(tr, batch, block, device, g)
135
+ model(xw)
136
+ model.head.calibrate(model._last_h)
137
+ opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
138
+ for step in range(1, steps + 1):
139
+ xw, yw = _batch(tr, batch, block, device, g)
140
+ xf, yf = _batch(fs, batch, block, device, g)
141
+ mk = (torch.rand(batch, generator=g) < FACT_RATE).to(device)
142
+ x = torch.where(mk[:, None], xf, xw)
143
+ y = torch.where(mk[:, None], yf, yw)
144
+ loss = F.cross_entropy(model(x).reshape(-1, VOCAB), y.reshape(-1))
145
+ opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
146
+ model.eval()
147
+ with torch.no_grad():
148
+ ls = []
149
+ for _ in range(10):
150
+ xv, yv = _batch(va, batch, block, device, g)
151
+ ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
152
+ yv.reshape(-1)).item())
153
+ return sum(ls) / len(ls) / math.log(2)
154
+
155
+
156
+ ARMS = ("aleph", "aleph_tri", "relay", "const", "mlp", "aleph_fmtdiv")
157
+
158
+
159
+ def run_bakeoff(arms=ARMS, seeds=(0, 1), device="cuda"):
160
+ if not torch.cuda.is_available():
161
+ raise RuntimeError("verdict runs are GPU-only")
162
+ os.makedirs(EXP20_DIR, exist_ok=True)
163
+ tr, va = _wikitext_bytes(DATA_ROOT)
164
+ ledger = open(os.path.join(EXP20_DIR, "ledger.jsonl"), "a", encoding="utf-8")
165
+ for seed in seeds:
166
+ train_recs, test_recs = make_rule_facts(256, 128, seed=seed)
167
+ fs_base = fact_stream_fmt(train_recs, RENDERERS[:1], seed=seed)
168
+ fs_div = fact_stream_fmt(train_recs, RENDERERS, seed=seed)
169
+ for arm in arms:
170
+ m = make_arm(arm, seed=seed)
171
+ hp = head_params(m)
172
+ bpb = train_direct(m, tr, va,
173
+ fs_div if arm == "aleph_fmtdiv" else fs_base,
174
+ device=device, seed=seed,
175
+ needs_calibration=(arm == "const"))
176
+ gz = {"train": recall_fmt(m, train_recs, FMT_TRAIN, device=device),
177
+ "heldout": recall_fmt(m, test_recs, FMT_TRAIN, device=device),
178
+ "train_varfmt": recall_fmt(m, train_recs, FMT_VARIANT,
179
+ device=device)}
180
+ if arm == "aleph_fmtdiv":
181
+ gz["heldout_r1"] = recall_fmt(m, test_recs, FMT_R1,
182
+ device=device)
183
+ rec = {"exp": "20", "arm": arm, "seed": seed, "steps": STEPS,
184
+ "head_params": hp, "gauges": gz, "bpb": round(bpb, 4)}
185
+ ledger.write(json.dumps(rec) + "\n"); ledger.flush()
186
+ print(f"[20 {arm} s{seed}] heldout={gz['heldout']} "
187
+ f"train={gz['train']['exact']} varfmt="
188
+ f"{gz['train_varfmt']['byte_acc']} bpb={bpb:.4f} "
189
+ f"hp={hp}", flush=True)
190
+ torch.save({"arm": arm, "seed": seed,
191
+ "state_dict": {k: v.cpu() for k, v in
192
+ m.state_dict().items()}},
193
+ os.path.join(EXP20_DIR, f"gen_{arm}_s{seed}.pt"))
194
+ del m
195
+ torch.cuda.empty_cache()
196
+ ledger.close()
197
+
198
+
199
+ def smoke():
200
+ recs, test = make_rule_facts(8, 4, seed=0)
201
+ fs = fact_stream_fmt(recs, RENDERERS, copies=3, seed=0)
202
+ assert fs.dtype == torch.uint8 and fs.numel() > 0
203
+ blob = bytes(fs.tolist())
204
+ assert b"->" in blob and b"<" in blob and b"@" in blob # all renderers hit
205
+ x = torch.randint(0, 256, (2, 64))
206
+ for arm in ARMS:
207
+ m = make_arm(arm, seed=0, d=96, layers=2, block=64)
208
+ lg = m(x)
209
+ assert lg.shape == (2, 64, 256), arm
210
+ lg.sum().backward()
211
+ assert head_params(m) > 0, arm
212
+ m.zero_grad()
213
+ del m
214
+ ref = make_arm("aleph", seed=0, d=96, layers=2, block=64)
215
+ mm = make_arm("mlp", seed=0, d=96, layers=2, block=64)
216
+ rp, mp = head_params(ref), head_params(mm)
217
+ assert abs(rp - mp) / rp < 0.05, (rp, mp) # matched within 5%
218
+ print(f"exp020 smoke passed (mlp head {mp} ~ aleph head {rp})")
219
+
220
+
221
+ def _in_notebook():
222
+ try:
223
+ get_ipython() # type: ignore[name-defined] # noqa: F821
224
+ return True
225
+ except NameError:
226
+ return False
227
+
228
+
229
+ if __name__ == "__main__":
230
+ smoke() if not _in_notebook() else (smoke(),
231
+ print("Notebook: run_bakeoff() on GPU."))