exp018_r12 + exp019_cr + exp020_gen packages (protocol ship: sub-article READMEs, standalone scrubbed code, self-asserting builders, ledgers, all specimens). exp018: keystone init avenues priced (neutral) / tied readout scoped (reconstruction-regime). exp019: capacity cliff, zero-to-negative distillation tax, anchors carry no bytes, universal forgetting. exp020: rule-induction plateau structure-independent; aleph bottleneck best generalizer; inverse law (memorization ease substitutes for induction)
Browse filesThis view is limited to 50 files because it contains too many changes. See raw diff
- README.md +39 -0
- exp018_r12/README.md +64 -0
- exp018_r12/ar_differentiation_bed.py +487 -0
- exp018_r12/build_results.py +36 -0
- exp018_r12/exp014_genetic_distillation.py +515 -0
- exp018_r12/exp018_reran12.py +176 -0
- exp018_r12/geolip_vitals.py +219 -0
- exp018_r12/read_codebook.py +159 -0
- exp018_r12/repro.py +27 -0
- exp018_r12/results/ledger.jsonl +8 -0
- exp018_r12/results/results.json +34 -0
- exp018_r12/specimens/r12_farmed_init_s0.pt +3 -0
- exp018_r12/specimens/r12_farmed_init_s1.pt +3 -0
- exp018_r12/specimens/r12_keystone_s0.pt +3 -0
- exp018_r12/specimens/r12_keystone_s1.pt +3 -0
- exp018_r12/specimens/r12_penta_init_s0.pt +3 -0
- exp018_r12/specimens/r12_penta_init_s1.pt +3 -0
- exp018_r12/specimens/r12_tied_s0.pt +3 -0
- exp018_r12/specimens/r12_tied_s1.pt +3 -0
- exp019_cr/README.md +82 -0
- exp019_cr/ar_differentiation_bed.py +487 -0
- exp019_cr/build_results.py +81 -0
- exp019_cr/exp014_genetic_distillation.py +515 -0
- exp019_cr/exp019_content_retention.py +402 -0
- exp019_cr/geolip_vitals.py +219 -0
- exp019_cr/read_codebook.py +159 -0
- exp019_cr/repro.py +29 -0
- exp019_cr/results/ledger.jsonl +36 -0
- exp019_cr/results/results.json +21 -0
- exp019_cr/specimens/rule_direct_s0.pt +3 -0
- exp019_cr/specimens/rule_direct_s1.pt +3 -0
- exp019_cr/specimens/rule_kd_facts_s0.pt +3 -0
- exp019_cr/specimens/rule_kd_facts_s1.pt +3 -0
- exp019_cr/specimens/rule_kd_general_s0.pt +3 -0
- exp019_cr/specimens/rule_kd_general_s1.pt +3 -0
- exp019_cr/specimens/rule_teacher_s0.pt +3 -0
- exp019_cr/specimens/rule_teacher_s1.pt +3 -0
- exp019_cr/specimens/teacher_N1024_s0.pt +3 -0
- exp019_cr/specimens/teacher_N1024_s1.pt +3 -0
- exp019_cr/specimens/teacher_N256_s0.pt +3 -0
- exp019_cr/specimens/teacher_N256_s1.pt +3 -0
- exp019_cr/specimens/teacher_N64_s0.pt +3 -0
- exp019_cr/specimens/teacher_N64_s1.pt +3 -0
- exp020_gen/README.md +73 -0
- exp020_gen/ar_differentiation_bed.py +487 -0
- exp020_gen/build_results.py +52 -0
- exp020_gen/exp014_genetic_distillation.py +515 -0
- exp020_gen/exp017_aleph_constellation.py +266 -0
- exp020_gen/exp019_content_retention.py +402 -0
- exp020_gen/exp020_generalization.py +231 -0
README.md
CHANGED
|
@@ -198,6 +198,45 @@ trainable bank. Value case: the structured 768 address (conditioning/routing/
|
|
| 198 |
lookup), not raw bpb. 8-row ledger + 6 checkpoints;
|
| 199 |
write-up in [exp017_ac/README.md](./exp017_ac/README.md).
|
| 200 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 201 |
## Reproducibility
|
| 202 |
|
| 203 |
Every experiment package (`exp012_ar/`, `exp013_aug/`, `exp014_gd/`,
|
|
|
|
| 198 |
lookup), not raw bpb. 8-row ledger + 6 checkpoints;
|
| 199 |
write-up in [exp017_ac/README.md](./exp017_ac/README.md).
|
| 200 |
|
| 201 |
+
## exp018 — experiment 18: exp012 rerun with keystone parameters (July 11, 2026)
|
| 202 |
+
|
| 203 |
+
[`exp018_r12/`](./exp018_r12): a labeled rerun (exp012's record untouched)
|
| 204 |
+
pricing the two keystone avenues the original bed never exercised. Verdicts:
|
| 205 |
+
the init avenue is task-neutral (geovocab2 regular-pentachoron vertices and a
|
| 206 |
+
farmed recon-real codebook both land in the certified band; farmed anchors
|
| 207 |
+
drift less, accelerate nothing); the tied M̂ readout fails in autoregression
|
| 208 |
+
(+1.0 bpb, codebook starved) — a reconstruction-regime device. Net: exp012's
|
| 209 |
+
original parameters stand. Write-up: [exp018_r12/README.md](./exp018_r12/README.md).
|
| 210 |
+
|
| 211 |
+
## exp019 — the capacity for distilled content retention (July 11, 2026)
|
| 212 |
+
|
| 213 |
+
[`exp019_cr/`](./exp019_cr): fact corpora in the byte stream; exact-match
|
| 214 |
+
recall as the retention gauge; distillation channels vs direct learning across
|
| 215 |
+
N ∈ {64, 256, 1024}, plus a rule-content generalization block. Verdicts: a
|
| 216 |
+
capacity cliff between 256 and 1024 facts; **the distillation tax is
|
| 217 |
+
zero-to-negative** (teacher logits alone transfer rote content at parity+ and
|
| 218 |
+
rule content better than ground truth, 2/2 seeds — with a steep clean-bpb
|
| 219 |
+
cost); no logit leakage without exposure; **anchors carry no bytes** (a
|
| 220 |
+
trained codebook from a 256-fact teacher transfers none of them); **universal
|
| 221 |
+
catastrophic forgetting** (every channel → 0.000 exact after 1k clean steps) —
|
| 222 |
+
persistence, not transfer, is the unsolved axis. Rules are learned only
|
| 223 |
+
fragmentarily and content is format-locked (KD students less so). Write-up:
|
| 224 |
+
[exp019_cr/README.md](./exp019_cr/README.md).
|
| 225 |
+
|
| 226 |
+
## exp020 — generalization: structure vs capacity on a hidden rule (July 11, 2026)
|
| 227 |
+
|
| 228 |
+
[`exp020_gen/`](./exp020_gen): the exp019 rule task raced across six
|
| 229 |
+
structural arms with an exact param-matched control. Verdicts: no structure
|
| 230 |
+
lifts rule induction off the ~0.26 plateau (exact 0.000 in all 12 cells); the
|
| 231 |
+
**aleph bottleneck is the best generalizer** (+25% over its matched free head,
|
| 232 |
+
both seeds); **the inverse law** — generalization ordering reverses modeling
|
| 233 |
+
strength (the constellation head models best and generalizes worst;
|
| 234 |
+
memorization ease substitutes for rule induction); format diversity covers
|
| 235 |
+
seen formats at full rule level but not novel ones. The measured
|
| 236 |
+
memorization↔generalization axis doubles as a component registry for
|
| 237 |
+
slider/registry-style composites. Write-up:
|
| 238 |
+
[exp020_gen/README.md](./exp020_gen/README.md).
|
| 239 |
+
|
| 240 |
## Reproducibility
|
| 241 |
|
| 242 |
Every experiment package (`exp012_ar/`, `exp013_aug/`, `exp014_gd/`,
|
exp018_r12/README.md
ADDED
|
@@ -0,0 +1,64 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# exp018_r12 — experiment 18: exp012 rerun with the keystone parameters
|
| 2 |
+
|
| 3 |
+
A labeled rerun, not a replacement: [exp012](../exp012_ar/)'s certified record
|
| 4 |
+
stands untouched. A process review of the exp012 bed against the aleph
|
| 5 |
+
keystone ([aleph-void article](https://huggingface.co/blog/AbstractPhil/geometric-vocabulary-patchwork-aleph-void))
|
| 6 |
+
found the `AlephAddress` primitive faithful (K=64, D=4, τ=0.1, M̂ reads,
|
| 7 |
+
pure Adam) but two keystone avenues never exercised. exp018 prices both.
|
| 8 |
+
|
| 9 |
+
## Arms (2 seeds, 2000 steps, otherwise the certified bed unchanged)
|
| 10 |
+
|
| 11 |
+
| arm | s0 | s1 | drift s0/s1 |
|
| 12 |
+
|---|---|---|---|
|
| 13 |
+
| penta_init — geovocab2 regular pentachoron vertices as codebook init | 2.4957 | 2.4389 | .242 / .211 |
|
| 14 |
+
| farmed_init — an exp012 specimen's trained (recon-real) codebook as init | 2.4934 | 2.4614 | .179 / .199 |
|
| 15 |
+
| tied — zero-parameter M̂ readout (decode tied through the projection + byte embedding) | 3.5131 | 3.4997 | .019 / .020 |
|
| 16 |
+
| keystone — penta_init + tied | 3.5169 | 3.5006 | .020 / .021 |
|
| 17 |
+
|
| 18 |
+
exp012 certified reference: addr_msl64 3-seed mean 2.469; s0 bed 2.4990.
|
| 19 |
+
`build_results.py` re-asserts every claim from `results/ledger.jsonl`.
|
| 20 |
+
|
| 21 |
+
## Verdicts
|
| 22 |
+
|
| 23 |
+
1. **The init avenue is task-neutral.** Both the exact regular-pentachoron
|
| 24 |
+
init (the factory's 5-dim construction, centered and SVD-projected to R⁴ —
|
| 25 |
+
pairwise cos = −¼ asserted in the smoke) and the farmed recon-real
|
| 26 |
+
codebook land inside the certified band. Consistent with the
|
| 27 |
+
basin-set-at-init and interchangeable-scaffold results: codebook identity
|
| 28 |
+
does not price into bits-per-byte on this bed at this budget.
|
| 29 |
+
2. **Farmed codebooks drift less** (0.18–0.20 vs 0.21–0.24) — mature anchors
|
| 30 |
+
are more stationary — **and accelerate nothing**, echoing the exp014
|
| 31 |
+
implant studies.
|
| 32 |
+
3. **The tied M̂ readout fails in autoregression** (+1.0 bpb, both seeds, with
|
| 33 |
+
or without the pentachoron init) **and starves the codebook** (drift ~0.02,
|
| 34 |
+
binding fraction 0): it is a reconstruction-regime device; as a
|
| 35 |
+
zero-parameter AR head it cannot shape a 256-way distribution and passes
|
| 36 |
+
almost no cultivating gradient into the address.
|
| 37 |
+
4. **Net: exp012's original parameters stand.** Random init + free linear
|
| 38 |
+
head was not a wrong configuration — the two unexercised keystone avenues
|
| 39 |
+
are now priced (one neutral, one negative in this regime).
|
| 40 |
+
|
| 41 |
+
## Files
|
| 42 |
+
- `exp018_reran12.py` — pentachoron/farmed inits, the tied readout, runner,
|
| 43 |
+
smoke (asserts exact simplex regularity).
|
| 44 |
+
- `geolip_vitals.py` / `ar_differentiation_bed.py` /
|
| 45 |
+
`exp014_genetic_distillation.py` / `read_codebook.py` — this package's own
|
| 46 |
+
copies of the shared harness. Standalone.
|
| 47 |
+
- `repro.py` — loads the code files from this folder and runs them.
|
| 48 |
+
- `build_results.py` → `results/results.json` — re-asserts every claim above.
|
| 49 |
+
- `results/ledger.jsonl` — 8 rows. `specimens/` — all 8 checkpoints.
|
| 50 |
+
|
| 51 |
+
## Reproduce (from inside this folder)
|
| 52 |
+
```bash
|
| 53 |
+
pip install torch --index-url https://download.pytorch.org/whl/cu128
|
| 54 |
+
pip install pyarrow huggingface_hub
|
| 55 |
+
pip install "git+https://github.com/AbstractEyes/geolip-svae" # geovocab2 (penta init)
|
| 56 |
+
python repro.py # CPU smoke
|
| 57 |
+
python repro.py --run # 4 arms x 2 seeds (GPU, ~40 min)
|
| 58 |
+
python build_results.py # re-assert every claim from the ledger
|
| 59 |
+
```
|
| 60 |
+
Data lands in `./data` (override `GEOLIP_DATA`); the farmed donor is
|
| 61 |
+
`../exp012_ar/specimens/addr_msl64_s0_t2000.pt` (in this repo) or set
|
| 62 |
+
`GEOLIP_FARMED_SPECIMEN`.
|
| 63 |
+
|
| 64 |
+
License: MIT · AbstractPhil + Claude Fable 5 · July 11, 2026
|
exp018_r12/ar_differentiation_bed.py
ADDED
|
@@ -0,0 +1,487 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""ar_differentiation_bed.py — exp012: autoregressive differentiation of the aleph.
|
| 2 |
+
|
| 3 |
+
Differentiation is cultivated by PREDICTIVE pressure along the sequence — the
|
| 4 |
+
address parameterizing the next-byte distribution (Law 2: chain-rule advantage pays
|
| 5 |
+
ONLY where the composed address directly parameterizes the predictive distribution).
|
| 6 |
+
This bed puts the aleph in the autoregressive gradient path and measures what
|
| 7 |
+
differentiates. The head arms enforce the employment law at its maximum: the
|
| 8 |
+
ENTIRE next-byte distribution is parameterized by the address.
|
| 9 |
+
|
| 10 |
+
Byte-level causal LM on wikitext-2-raw (HF parquet, CDN-fast), block 256. ARMS:
|
| 11 |
+
sdpa — standard causal transformer control (matched trunk).
|
| 12 |
+
hub — attention replaced by CAUSAL HUB: linear attention whose feature map
|
| 13 |
+
is the 2K-oriented aleph address, prefix-sum memories (no selection
|
| 14 |
+
event; O(n*K*d)). Differentiation cultivated INSIDE attention.
|
| 15 |
+
addr_head — sdpa trunk, but the OUTPUT HEAD reads ONLY the signed aleph
|
| 16 |
+
coefficient vector w_k = sinh(u_k)/sum_j cosh(u_j) of the final
|
| 17 |
+
hidden state (K -> 256 logits). The address MUST carry every bit of
|
| 18 |
+
next-byte information — the hardest Law-2 bottleneck.
|
| 19 |
+
|
| 20 |
+
JUDGED BY: val bits-per-byte per arm (task) + CULTIVATION VITALS on every aleph
|
| 21 |
+
codebook (readouts, never losses): axis aliveness/hppl, drift-from-init +
|
| 22 |
+
binding fraction @0.29154, winner-|cos| saturation (sign-code emergence), shadow
|
| 23 |
+
path diversity (fixed high-bits hash). Never by recon.
|
| 24 |
+
|
| 25 |
+
Riders: pure Adam wd=0; no BN/Dropout/GAP on geometric paths; orthogonal init;
|
| 26 |
+
Colab-cell-safe (paste-ahead imports, no bare argparse, no __file__ reliance);
|
| 27 |
+
GPU-only for verdict runs.
|
| 28 |
+
|
| 29 |
+
Terminal: python ar_differentiation_bed.py # shapes/parse smoke
|
| 30 |
+
python ar_differentiation_bed.py --train # verdict run
|
| 31 |
+
Colab: paste geolip_vitals.py cell, then this file (smoke auto-runs),
|
| 32 |
+
then train(steps=2000, data_root="/content/data") in the next cell.
|
| 33 |
+
"""
|
| 34 |
+
from __future__ import annotations
|
| 35 |
+
import math
|
| 36 |
+
import torch
|
| 37 |
+
import torch.nn as nn
|
| 38 |
+
import torch.nn.functional as F
|
| 39 |
+
|
| 40 |
+
if "anchor_drift" not in globals():
|
| 41 |
+
try:
|
| 42 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 43 |
+
except ImportError:
|
| 44 |
+
_here = globals().get("__file__")
|
| 45 |
+
if _here is not None:
|
| 46 |
+
import sys, pathlib
|
| 47 |
+
sys.path.insert(0, str(pathlib.Path(_here).parent))
|
| 48 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 49 |
+
else:
|
| 50 |
+
raise ImportError(
|
| 51 |
+
"geolip_vitals not found — paste/run its cell first, or "
|
| 52 |
+
"hf_hub_download exp012_ar/geolip_vitals.py from "
|
| 53 |
+
"AbstractPhil/geolip-aleph-differentiation.")
|
| 54 |
+
|
| 55 |
+
VOCAB = 256 # bytes
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
# ------------------------------------------------------------------ aleph address
|
| 59 |
+
def _super_fibonacci_s3(n: int) -> torch.Tensor:
|
| 60 |
+
"""Near-uniform unit quaternions (Alexa CVPR'22) —
|
| 61 |
+
starts the codebook INSIDE the RP^3 attractor basin. D=4 only."""
|
| 62 |
+
PHI, PSI = math.sqrt(2.0), 1.533751168755204288118041
|
| 63 |
+
i = torch.arange(n, dtype=torch.float64)
|
| 64 |
+
s = (i + 0.5) / n
|
| 65 |
+
r, R = torch.sqrt(s), torch.sqrt(1.0 - s)
|
| 66 |
+
a, b = 2 * math.pi * i / PHI, 2 * math.pi * i / PSI
|
| 67 |
+
q = torch.stack([r * torch.sin(a), r * torch.cos(a),
|
| 68 |
+
R * torch.sin(b), R * torch.cos(b)], dim=-1)
|
| 69 |
+
return F.normalize(q, dim=-1).float()
|
| 70 |
+
|
| 71 |
+
|
| 72 |
+
class AlephAddress(nn.Module):
|
| 73 |
+
"""Closed-form aleph over 2K oriented half-axes (aleph-void article).
|
| 74 |
+
signed(x): (..., K) w_k = sinh(u_k)/sum_j cosh(u_j) — the Law-2 head feature.
|
| 75 |
+
oriented(x): ((..., K), (..., K)) positive halves of the 2K softmax — HUB map."""
|
| 76 |
+
|
| 77 |
+
def __init__(self, K: int, D: int, tau: float = 0.1, init: str = "random"):
|
| 78 |
+
super().__init__()
|
| 79 |
+
self.K, self.D, self.tau = K, D, tau
|
| 80 |
+
if init == "fibonacci":
|
| 81 |
+
assert D == 4, "fibonacci init lives on S^3 (D=4)"
|
| 82 |
+
A = _super_fibonacci_s3(K)
|
| 83 |
+
else:
|
| 84 |
+
A = F.normalize(torch.randn(K, D), dim=-1)
|
| 85 |
+
self.codebook = nn.Parameter(A)
|
| 86 |
+
self.register_buffer("home", self.codebook.detach().clone())
|
| 87 |
+
|
| 88 |
+
def _u(self, x):
|
| 89 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 90 |
+
return (F.normalize(x, dim=-1) @ A.transpose(-1, -2)) / self.tau
|
| 91 |
+
|
| 92 |
+
def oriented(self, x):
|
| 93 |
+
u = self._u(x)
|
| 94 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 95 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 96 |
+
Z = (ep + en).sum(dim=-1, keepdim=True)
|
| 97 |
+
return ep / Z, en / Z
|
| 98 |
+
|
| 99 |
+
def signed(self, x):
|
| 100 |
+
u = self._u(x)
|
| 101 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 102 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 103 |
+
return (ep - en) / (ep + en).sum(dim=-1, keepdim=True)
|
| 104 |
+
|
| 105 |
+
def signed_at(self, x, taus):
|
| 106 |
+
"""Multi-tau stroboscope (rule of 3): signed coefficients at several
|
| 107 |
+
temperatures, concatenated — softer taus keep the vector dense while a
|
| 108 |
+
hard tau supplies the sign-code sharpness. v2 refinement (b)."""
|
| 109 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 110 |
+
cos = F.normalize(x, dim=-1) @ A.transpose(-1, -2)
|
| 111 |
+
outs = []
|
| 112 |
+
for t in taus:
|
| 113 |
+
u = cos / t
|
| 114 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 115 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 116 |
+
outs.append((ep - en) / (ep + en).sum(dim=-1, keepdim=True))
|
| 117 |
+
return torch.cat(outs, dim=-1)
|
| 118 |
+
|
| 119 |
+
def m_hat(self, x):
|
| 120 |
+
"""Closed-form soft read (decoders read M_hat, never M). v2 control (c)."""
|
| 121 |
+
u = self._u(x)
|
| 122 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 123 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 124 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 125 |
+
return ((ep - en) @ A) / (ep + en).sum(dim=-1, keepdim=True)
|
| 126 |
+
|
| 127 |
+
def m_hard_ste(self, x):
|
| 128 |
+
"""Hard mode (aleph-void article): M_hard = sign(cos_win) * A[win], straight-through to
|
| 129 |
+
the soft read — forward fully discrete SIGN CODE, backward soft gradient.
|
| 130 |
+
Legal per theme A (reconstructive sign code, not a one-hot roster pick)."""
|
| 131 |
+
u = self._u(x)
|
| 132 |
+
soft = self.m_hat(x)
|
| 133 |
+
win = u.abs().argmax(dim=-1)
|
| 134 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 135 |
+
sign = torch.sign(torch.gather(u, -1, win.unsqueeze(-1))).squeeze(-1)
|
| 136 |
+
hard = sign.unsqueeze(-1) * A[win]
|
| 137 |
+
return hard + soft - soft.detach()
|
| 138 |
+
|
| 139 |
+
@torch.no_grad()
|
| 140 |
+
def vitals(self, x_sample) -> dict:
|
| 141 |
+
u = self._u(x_sample.reshape(-1, x_sample.shape[-1]))
|
| 142 |
+
p, n = self.oriented(x_sample.reshape(-1, x_sample.shape[-1]))
|
| 143 |
+
two_k = torch.cat([p, n], dim=-1)
|
| 144 |
+
win = two_k.argmax(dim=-1)
|
| 145 |
+
cos_win = (u.abs().amax(dim=-1) * self.tau) # winner |cos| — sign-code sat.
|
| 146 |
+
d = anchor_drift(self.codebook, self.home)
|
| 147 |
+
return {"drift": round(d["mean"], 4),
|
| 148 |
+
"binding_frac": round(d["binding_fraction"], 4),
|
| 149 |
+
"aliveness": axis_aliveness(two_k),
|
| 150 |
+
"win_cos_mean": round(cos_win.mean().item(), 4),
|
| 151 |
+
"paths": path_diversity(win)}
|
| 152 |
+
|
| 153 |
+
|
| 154 |
+
# ------------------------------------------------------------------------- blocks
|
| 155 |
+
class CausalSDPA(nn.Module):
|
| 156 |
+
def __init__(self, d: int, heads: int = 4):
|
| 157 |
+
super().__init__()
|
| 158 |
+
self.h = heads
|
| 159 |
+
self.qkv = nn.Linear(d, 3 * d, bias=False)
|
| 160 |
+
self.o = nn.Linear(d, d, bias=False)
|
| 161 |
+
nn.init.orthogonal_(self.qkv.weight); nn.init.orthogonal_(self.o.weight)
|
| 162 |
+
|
| 163 |
+
def forward(self, x):
|
| 164 |
+
B, n, d = x.shape
|
| 165 |
+
q, k, v = self.qkv(x).chunk(3, dim=-1)
|
| 166 |
+
q, k, v = (t.view(B, n, self.h, d // self.h).transpose(1, 2) for t in (q, k, v))
|
| 167 |
+
y = F.scaled_dot_product_attention(q, k, v, is_causal=True)
|
| 168 |
+
return self.o(y.transpose(1, 2).reshape(B, n, d))
|
| 169 |
+
|
| 170 |
+
|
| 171 |
+
class CausalHUB(nn.Module):
|
| 172 |
+
"""Causal aleph linear attention: prefix-sum memories over the two K-wide
|
| 173 |
+
halves of the oriented address; 2K never materialized; no selection event."""
|
| 174 |
+
|
| 175 |
+
def __init__(self, d: int, K: int = 32, D: int = 4, tau: float = 0.1):
|
| 176 |
+
super().__init__()
|
| 177 |
+
self.addr = AlephAddress(K, D, tau)
|
| 178 |
+
self.q = nn.Linear(d, D, bias=False)
|
| 179 |
+
self.k = nn.Linear(d, D, bias=False)
|
| 180 |
+
self.v = nn.Linear(d, d, bias=False)
|
| 181 |
+
self.o = nn.Linear(d, d, bias=False)
|
| 182 |
+
for m in (self.q, self.k, self.v, self.o):
|
| 183 |
+
nn.init.orthogonal_(m.weight)
|
| 184 |
+
|
| 185 |
+
def forward(self, x):
|
| 186 |
+
qp, qn = self.addr.oriented(self.q(x)) # (B, n, K)
|
| 187 |
+
kp, kn = self.addr.oriented(self.k(x))
|
| 188 |
+
v = self.v(x) # (B, n, d)
|
| 189 |
+
Sp = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kp, v), dim=1)
|
| 190 |
+
Sn = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kn, v), dim=1)
|
| 191 |
+
zp = torch.cumsum(kp, dim=1)
|
| 192 |
+
zn = torch.cumsum(kn, dim=1)
|
| 193 |
+
num = torch.einsum("bnk,bnkd->bnd", qp, Sp) + torch.einsum("bnk,bnkd->bnd", qn, Sn)
|
| 194 |
+
den = (qp * zp).sum(-1, keepdim=True) + (qn * zn).sum(-1, keepdim=True)
|
| 195 |
+
return self.o(num / den.clamp_min(1e-12))
|
| 196 |
+
|
| 197 |
+
|
| 198 |
+
class MslRelay(nn.Module):
|
| 199 |
+
"""Depth-composition unit (chain-rule probe): multi-slot M_hat read entering
|
| 200 |
+
the trunk as a NEAR-ZERO gated residual (gate init -3.0, sigma~0.047 — theme D:
|
| 201 |
+
geometry enters as a nudge and grows only if it earns gradient)."""
|
| 202 |
+
|
| 203 |
+
def __init__(self, d: int, n_slots: int = 16, K: int = 64):
|
| 204 |
+
super().__init__()
|
| 205 |
+
self.n_slots = n_slots
|
| 206 |
+
self.proj = nn.Linear(d, n_slots * 4, bias=False)
|
| 207 |
+
self.out = nn.Linear(n_slots * 4, d, bias=False)
|
| 208 |
+
nn.init.orthogonal_(self.proj.weight)
|
| 209 |
+
nn.init.orthogonal_(self.out.weight)
|
| 210 |
+
self.addr = AlephAddress(K, 4)
|
| 211 |
+
self.gate = nn.Parameter(torch.tensor(-3.0))
|
| 212 |
+
|
| 213 |
+
def forward(self, x):
|
| 214 |
+
B, n, _ = x.shape
|
| 215 |
+
slots = self.proj(x).view(B, n, self.n_slots, 4)
|
| 216 |
+
m = self.addr.m_hat(slots).reshape(B, n, -1)
|
| 217 |
+
return x + self.gate.sigmoid() * self.out(m)
|
| 218 |
+
|
| 219 |
+
|
| 220 |
+
class Block(nn.Module):
|
| 221 |
+
def __init__(self, d: int, attn: nn.Module):
|
| 222 |
+
super().__init__()
|
| 223 |
+
self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)
|
| 224 |
+
self.attn = attn
|
| 225 |
+
self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
|
| 226 |
+
|
| 227 |
+
def forward(self, x):
|
| 228 |
+
x = x + self.attn(self.n1(x))
|
| 229 |
+
return x + self.mlp(self.n2(x))
|
| 230 |
+
|
| 231 |
+
|
| 232 |
+
class ByteLM(nn.Module):
|
| 233 |
+
def __init__(self, arm: str, d: int = 192, layers: int = 4, block: int = 256,
|
| 234 |
+
K: int = 32, D: int = 4):
|
| 235 |
+
super().__init__()
|
| 236 |
+
# "<arm>_tri" suffix = trigram byte embedding (AlephLM byte_emb x3 lineage):
|
| 237 |
+
# token embedding is the sum of embeddings of bytes t, t-1, t-2.
|
| 238 |
+
self.trigram = arm.endswith("_tri")
|
| 239 |
+
if self.trigram:
|
| 240 |
+
arm = arm[:-4]
|
| 241 |
+
# "_fib" = super-Fibonacci S^3 codebook init (basin test: starts INSIDE
|
| 242 |
+
# the RP^3 attractor; primary observable is init->final geodesic drift).
|
| 243 |
+
self.fib = arm.endswith("_fib")
|
| 244 |
+
if self.fib:
|
| 245 |
+
arm = arm[:-4]
|
| 246 |
+
# "relay*" = stacked addresses in depth: MslRelay after every block.
|
| 247 |
+
# relay -> sdpa trunk + standard head; relay_msl64 -> + addressed head.
|
| 248 |
+
self.use_relay = arm.startswith("relay")
|
| 249 |
+
if arm == "relay":
|
| 250 |
+
arm = "sdpa"
|
| 251 |
+
elif arm == "relay_msl64":
|
| 252 |
+
arm = "addr_msl64"
|
| 253 |
+
self.arm, self.block = arm, block
|
| 254 |
+
self.emb = nn.Embedding(VOCAB, d)
|
| 255 |
+
if self.trigram:
|
| 256 |
+
self.emb1 = nn.Embedding(VOCAB, d)
|
| 257 |
+
self.emb2 = nn.Embedding(VOCAB, d)
|
| 258 |
+
self.pos = nn.Parameter(torch.zeros(1, block, d) + 0.01 * torch.randn(1, block, d))
|
| 259 |
+
mk_attn = (lambda: CausalHUB(d, K, D)) if arm == "hub" else (lambda: CausalSDPA(d))
|
| 260 |
+
self.blocks = nn.ModuleList([Block(d, mk_attn()) for _ in range(layers)])
|
| 261 |
+
if self.use_relay:
|
| 262 |
+
self.relays = nn.ModuleList([MslRelay(d) for _ in range(layers)])
|
| 263 |
+
self.nf = nn.LayerNorm(d)
|
| 264 |
+
if arm == "addr_head":
|
| 265 |
+
self.head_addr = AlephAddress(K, d) # v1: codebook in model dim — COLLAPSED
|
| 266 |
+
self.head = nn.Linear(K, VOCAB, bias=True)
|
| 267 |
+
elif arm in ("addr_d4", "addr_3tau", "addr_mhat"):
|
| 268 |
+
# v2 refinements: LOW-D HOME — learned projection to the native D=4 home
|
| 269 |
+
# before addressing (mirrors the healthy HUB arms), K=64.
|
| 270 |
+
self.head_proj = nn.Linear(d, 4, bias=False)
|
| 271 |
+
nn.init.orthogonal_(self.head_proj.weight)
|
| 272 |
+
self.head_addr = AlephAddress(64, 4)
|
| 273 |
+
if arm == "addr_d4":
|
| 274 |
+
self.head = nn.Linear(64, VOCAB, bias=True) # w alone, D=4 home
|
| 275 |
+
elif arm == "addr_3tau":
|
| 276 |
+
self.taus = (0.05, 0.1, 0.3) # rule-of-3 strobe
|
| 277 |
+
self.head = nn.Linear(64 * 3, VOCAB, bias=True)
|
| 278 |
+
else: # addr_mhat
|
| 279 |
+
self.head = nn.Linear(4, VOCAB, bias=True) # tightest: M_hat
|
| 280 |
+
elif arm.startswith("addr_msl"):
|
| 281 |
+
# v3: MULTI-SLOT heads — the 16s funnel widening: P parallel D=4 slots
|
| 282 |
+
# over a SHARED codebook. addr_msl consumes the reconstructive M_hat per
|
| 283 |
+
# slot (Px4 dims); addr_msl_w consumes signed w per slot (Px64) — tests
|
| 284 |
+
# whether slot-parallel consumption alone rescues the coefficient path.
|
| 285 |
+
# addr_msl<P> = slot-count dose-response. addr_mslh<P> = HARD sign-code
|
| 286 |
+
# consumption (straight-through M_hard per slot).
|
| 287 |
+
self.hard = arm.startswith("addr_mslh")
|
| 288 |
+
if arm in ("addr_msl", "addr_msl_w"):
|
| 289 |
+
self.n_slots = 16
|
| 290 |
+
else:
|
| 291 |
+
self.n_slots = int(arm[len("addr_mslh" if self.hard else "addr_msl"):])
|
| 292 |
+
self.head_proj = nn.Linear(d, self.n_slots * 4, bias=False)
|
| 293 |
+
nn.init.orthogonal_(self.head_proj.weight)
|
| 294 |
+
self.head_addr = AlephAddress(
|
| 295 |
+
64, 4, init="fibonacci" if self.fib else "random")
|
| 296 |
+
width = self.n_slots * (64 if arm == "addr_msl_w" else 4)
|
| 297 |
+
self.head = nn.Linear(width, VOCAB, bias=True)
|
| 298 |
+
elif arm == "addr_3tau_mhat":
|
| 299 |
+
# v3: combine the two v2 winners — 3-tau stroboscope + reconstructive read.
|
| 300 |
+
self.head_proj = nn.Linear(d, 4, bias=False)
|
| 301 |
+
nn.init.orthogonal_(self.head_proj.weight)
|
| 302 |
+
self.head_addr = AlephAddress(64, 4)
|
| 303 |
+
self.taus = (0.05, 0.1, 0.3)
|
| 304 |
+
self.head = nn.Linear(64 * 3 + 4, VOCAB, bias=True)
|
| 305 |
+
else:
|
| 306 |
+
self.head = nn.Linear(d, VOCAB, bias=True)
|
| 307 |
+
self._last_h = None
|
| 308 |
+
|
| 309 |
+
def forward(self, idx):
|
| 310 |
+
x = self.emb(idx)
|
| 311 |
+
if self.trigram: # past-only shifts — causality preserved
|
| 312 |
+
x = x + self.emb1(F.pad(idx, (1, 0), value=0)[:, :-1]) \
|
| 313 |
+
+ self.emb2(F.pad(idx, (2, 0), value=0)[:, :-2])
|
| 314 |
+
x = x + self.pos[:, : idx.shape[1]]
|
| 315 |
+
if self.use_relay:
|
| 316 |
+
for b, r in zip(self.blocks, self.relays):
|
| 317 |
+
x = r(b(x))
|
| 318 |
+
else:
|
| 319 |
+
for b in self.blocks:
|
| 320 |
+
x = b(x)
|
| 321 |
+
h = self.nf(x)
|
| 322 |
+
self._last_h = h.detach()
|
| 323 |
+
if self.arm == "addr_head":
|
| 324 |
+
return self.head(self.head_addr.signed(h))
|
| 325 |
+
if self.arm == "addr_d4":
|
| 326 |
+
return self.head(self.head_addr.signed(self.head_proj(h)))
|
| 327 |
+
if self.arm == "addr_3tau":
|
| 328 |
+
return self.head(self.head_addr.signed_at(self.head_proj(h), self.taus))
|
| 329 |
+
if self.arm == "addr_mhat":
|
| 330 |
+
return self.head(self.head_addr.m_hat(self.head_proj(h)))
|
| 331 |
+
if self.arm.startswith("addr_msl"):
|
| 332 |
+
B, n, _ = h.shape
|
| 333 |
+
slots = self.head_proj(h).view(B, n, self.n_slots, 4)
|
| 334 |
+
if self.arm == "addr_msl_w":
|
| 335 |
+
feats = self.head_addr.signed(slots).reshape(B, n, -1)
|
| 336 |
+
elif getattr(self, "hard", False):
|
| 337 |
+
feats = self.head_addr.m_hard_ste(slots).reshape(B, n, -1)
|
| 338 |
+
else:
|
| 339 |
+
feats = self.head_addr.m_hat(slots).reshape(B, n, -1)
|
| 340 |
+
return self.head(feats)
|
| 341 |
+
if self.arm == "addr_3tau_mhat":
|
| 342 |
+
p = self.head_proj(h)
|
| 343 |
+
feats = torch.cat([self.head_addr.signed_at(p, self.taus),
|
| 344 |
+
self.head_addr.m_hat(p)], dim=-1)
|
| 345 |
+
return self.head(feats)
|
| 346 |
+
return self.head(h)
|
| 347 |
+
|
| 348 |
+
@torch.no_grad()
|
| 349 |
+
def vitals(self) -> dict:
|
| 350 |
+
out = {}
|
| 351 |
+
if self.arm == "hub":
|
| 352 |
+
for i, b in enumerate(self.blocks):
|
| 353 |
+
if self._last_h is not None:
|
| 354 |
+
out[f"L{i}"] = b.attn.addr.vitals(b.attn.q(self._last_h[:2]))
|
| 355 |
+
elif self.arm == "addr_head" and self._last_h is not None:
|
| 356 |
+
out["head"] = self.head_addr.vitals(self._last_h[:2])
|
| 357 |
+
elif self.arm in ("addr_d4", "addr_3tau", "addr_mhat",
|
| 358 |
+
"addr_3tau_mhat") and self._last_h is not None:
|
| 359 |
+
out["head"] = self.head_addr.vitals(self.head_proj(self._last_h[:2]))
|
| 360 |
+
elif self.arm.startswith("addr_msl") and self._last_h is not None:
|
| 361 |
+
slots = self.head_proj(self._last_h[:2])
|
| 362 |
+
out["head"] = self.head_addr.vitals(
|
| 363 |
+
slots.reshape(*slots.shape[:-1], self.n_slots, 4))
|
| 364 |
+
if self.use_relay and self._last_h is not None:
|
| 365 |
+
for i, r in enumerate(self.relays):
|
| 366 |
+
s = r.proj(self._last_h[:2])
|
| 367 |
+
v = r.addr.vitals(s.reshape(*s.shape[:-1], r.n_slots, 4))
|
| 368 |
+
out[f"relay{i}"] = {"gate": round(r.gate.sigmoid().item(), 4),
|
| 369 |
+
"drift": v["drift"],
|
| 370 |
+
"binding_frac": v["binding_frac"],
|
| 371 |
+
"ppl": round(v["aliveness"]["usage_ppl"], 1)}
|
| 372 |
+
return out
|
| 373 |
+
|
| 374 |
+
|
| 375 |
+
# --------------------------------------------------------------------------- data
|
| 376 |
+
def _wikitext_bytes(data_root: str):
|
| 377 |
+
"""wikitext-2-raw as flat uint8 tensors via the HF parquet CDN."""
|
| 378 |
+
from huggingface_hub import hf_hub_download
|
| 379 |
+
import pyarrow.parquet as pq
|
| 380 |
+
|
| 381 |
+
def load(split):
|
| 382 |
+
p = hf_hub_download("Salesforce/wikitext",
|
| 383 |
+
f"wikitext-2-raw-v1/{split}-00000-of-00001.parquet",
|
| 384 |
+
repo_type="dataset", local_dir=data_root)
|
| 385 |
+
text = "".join(pq.read_table(p).column("text").to_pylist())
|
| 386 |
+
return torch.frombuffer(bytearray(text.encode("utf-8")), dtype=torch.uint8).clone()
|
| 387 |
+
|
| 388 |
+
return load("train"), load("validation")
|
| 389 |
+
|
| 390 |
+
|
| 391 |
+
def _batch(data: torch.Tensor, batch: int, block: int, device, g: torch.Generator):
|
| 392 |
+
ix = torch.randint(0, data.numel() - block - 1, (batch,), generator=g)
|
| 393 |
+
x = torch.stack([data[i:i + block] for i in ix]).long().to(device)
|
| 394 |
+
y = torch.stack([data[i + 1:i + block + 1] for i in ix]).long().to(device)
|
| 395 |
+
return x, y
|
| 396 |
+
|
| 397 |
+
|
| 398 |
+
# -------------------------------------------------------------------- train/smoke
|
| 399 |
+
def train(arms=("sdpa", "hub", "addr_head"), steps: int = 2000, batch: int = 32,
|
| 400 |
+
block: int = 256, device: str = "cuda", data_root: str = "./data",
|
| 401 |
+
seed: int = 0, eval_every: int = 500, save: bool = True):
|
| 402 |
+
"""Verdict run — GPU only. Pure Adam wd=0. Reports val bits-per-byte + vitals.
|
| 403 |
+
save=True writes {data_root}/ar_ckpts/{arm}_s{seed}_t{steps}.pt per arm —
|
| 404 |
+
the cultivated codebooks are SPECIMENS for the projective reading instruments."""
|
| 405 |
+
import os
|
| 406 |
+
if device == "cuda" and not torch.cuda.is_available():
|
| 407 |
+
raise RuntimeError("Verdict runs are GPU-only (never CPU-train for accuracy).")
|
| 408 |
+
ckpt_dir = os.path.join(data_root, "ar_ckpts")
|
| 409 |
+
os.makedirs(ckpt_dir, exist_ok=True)
|
| 410 |
+
tr, va = _wikitext_bytes(data_root)
|
| 411 |
+
print(f"data ready: train {tr.numel():,} bytes, val {va.numel():,} bytes", flush=True)
|
| 412 |
+
results = {}
|
| 413 |
+
for arm in arms:
|
| 414 |
+
torch.manual_seed(seed)
|
| 415 |
+
g = torch.Generator().manual_seed(seed)
|
| 416 |
+
model = ByteLM(arm, block=block).to(device)
|
| 417 |
+
n_params = sum(p.numel() for p in model.parameters())
|
| 418 |
+
opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
|
| 419 |
+
for step in range(1, steps + 1):
|
| 420 |
+
x, y = _batch(tr, batch, block, device, g)
|
| 421 |
+
logits = model(x)
|
| 422 |
+
loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
|
| 423 |
+
opt.zero_grad(set_to_none=True)
|
| 424 |
+
loss.backward()
|
| 425 |
+
opt.step()
|
| 426 |
+
if step % eval_every == 0 or step == steps:
|
| 427 |
+
model.eval()
|
| 428 |
+
with torch.no_grad():
|
| 429 |
+
losses = []
|
| 430 |
+
for _ in range(20):
|
| 431 |
+
xv, yv = _batch(va, batch, block, device, g)
|
| 432 |
+
lv = F.cross_entropy(model(xv).reshape(-1, VOCAB),
|
| 433 |
+
yv.reshape(-1))
|
| 434 |
+
losses.append(lv.item())
|
| 435 |
+
bpb = sum(losses) / len(losses) / math.log(2)
|
| 436 |
+
print(f"[{arm}] step {step} val_bpb={bpb:.4f} vitals={model.vitals()}",
|
| 437 |
+
flush=True)
|
| 438 |
+
model.train()
|
| 439 |
+
results[arm] = {"val_bpb": bpb, "params": n_params, "vitals": model.vitals()}
|
| 440 |
+
if save:
|
| 441 |
+
path = os.path.join(ckpt_dir, f"{arm}_s{seed}_t{steps}.pt")
|
| 442 |
+
torch.save({"arm": arm, "seed": seed, "steps": steps, "val_bpb": bpb,
|
| 443 |
+
"state_dict": {k: v.cpu() for k, v in
|
| 444 |
+
model.state_dict().items()}}, path)
|
| 445 |
+
print(f"saved specimen: {path}", flush=True)
|
| 446 |
+
print(results, flush=True)
|
| 447 |
+
return results
|
| 448 |
+
|
| 449 |
+
|
| 450 |
+
def smoke():
|
| 451 |
+
"""Shapes/parse only — no accuracy claims."""
|
| 452 |
+
x = torch.randint(0, VOCAB, (2, 64))
|
| 453 |
+
for arm in ("sdpa", "hub", "addr_head"):
|
| 454 |
+
m = ByteLM(arm, d=96, layers=2, block=64, K=16)
|
| 455 |
+
logits = m(x)
|
| 456 |
+
assert logits.shape == (2, 64, VOCAB)
|
| 457 |
+
logits.sum().backward()
|
| 458 |
+
# causality check: future byte must not affect past logits
|
| 459 |
+
with torch.no_grad():
|
| 460 |
+
a = m(x)[0, 10]
|
| 461 |
+
x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
|
| 462 |
+
b = m(x2)[0, 10]
|
| 463 |
+
assert torch.allclose(a, b, atol=1e-4), f"{arm} leaks future context"
|
| 464 |
+
print(f"{arm}: OK params={sum(p.numel() for p in m.parameters()):,} "
|
| 465 |
+
f"vitals={m.vitals()}", flush=True)
|
| 466 |
+
print("OK — AR bed smoke passed (verdict run: train() on GPU)", flush=True)
|
| 467 |
+
|
| 468 |
+
|
| 469 |
+
def _in_notebook() -> bool:
|
| 470 |
+
try:
|
| 471 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 472 |
+
return True
|
| 473 |
+
except NameError:
|
| 474 |
+
return False
|
| 475 |
+
|
| 476 |
+
|
| 477 |
+
if __name__ == "__main__":
|
| 478 |
+
if _in_notebook():
|
| 479 |
+
smoke()
|
| 480 |
+
print("Notebook mode: call train(steps=2000) in the next cell (GPU).")
|
| 481 |
+
else:
|
| 482 |
+
import argparse
|
| 483 |
+
ap = argparse.ArgumentParser()
|
| 484 |
+
ap.add_argument("--train", action="store_true")
|
| 485 |
+
ap.add_argument("--steps", type=int, default=2000)
|
| 486 |
+
a, _ = ap.parse_known_args()
|
| 487 |
+
train(steps=a.steps) if a.train else smoke()
|
exp018_r12/build_results.py
ADDED
|
@@ -0,0 +1,36 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""build_results.py — exp018_r12: read results/ledger.jsonl and RE-ASSERT every
|
| 2 |
+
claim in the README. Run from inside this folder: python build_results.py
|
| 3 |
+
"""
|
| 4 |
+
import json
|
| 5 |
+
import os
|
| 6 |
+
|
| 7 |
+
HERE = os.path.dirname(os.path.abspath(__file__))
|
| 8 |
+
rows = [json.loads(l) for l in
|
| 9 |
+
open(os.path.join(HERE, "results", "ledger.jsonl"), encoding="utf-8")]
|
| 10 |
+
assert all(r["exp"] == "18" for r in rows) and len(rows) == 8
|
| 11 |
+
cell = {(r["arm"], r["seed"]): r for r in rows}
|
| 12 |
+
CERT_BAND = (2.43, 2.50) # exp012 certified: 3-seed mean 2.469, s0 bed 2.4990
|
| 13 |
+
|
| 14 |
+
# claim 1: both init avenues land inside the certified band, both seeds
|
| 15 |
+
for arm in ("penta_init", "farmed_init"):
|
| 16 |
+
for s in (0, 1):
|
| 17 |
+
assert CERT_BAND[0] < cell[(arm, s)]["bpb"] < CERT_BAND[1], (arm, s)
|
| 18 |
+
|
| 19 |
+
# claim 2: the farmed codebook drifts LESS than the pentachoron init (maturity
|
| 20 |
+
# = stationarity), both seeds — and neither accelerates
|
| 21 |
+
for s in (0, 1):
|
| 22 |
+
assert cell[("farmed_init", s)]["drift"] < cell[("penta_init", s)]["drift"]
|
| 23 |
+
|
| 24 |
+
# claim 3: the tied M_hat readout fails in AR (+1.0 bpb) and starves the
|
| 25 |
+
# codebook (drift ~0.02, binding 0), both seeds, with or without penta init
|
| 26 |
+
for arm in ("tied", "keystone"):
|
| 27 |
+
for s in (0, 1):
|
| 28 |
+
r = cell[(arm, s)]
|
| 29 |
+
assert r["bpb"] > 3.4 and r["drift"] < 0.03 and r["binding_frac"] == 0.0
|
| 30 |
+
|
| 31 |
+
out = {f"{a}_s{s}": {"bpb": cell[(a, s)]["bpb"], "drift": cell[(a, s)]["drift"]}
|
| 32 |
+
for (a, s) in sorted(cell)}
|
| 33 |
+
json.dump(out, open(os.path.join(HERE, "results", "results.json"), "w",
|
| 34 |
+
encoding="utf-8"), indent=1)
|
| 35 |
+
print(f"{len(rows)} rows -> results/results.json")
|
| 36 |
+
print("all README claims asserted OK")
|
exp018_r12/exp014_genetic_distillation.py
ADDED
|
@@ -0,0 +1,515 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""exp014_genetic_distillation.py — genetic distillation + memory substrate.
|
| 2 |
+
14_a: multi-generational tournament (GM3 paradigm) where the aleph codebook is the
|
| 3 |
+
explicit heritable genome. Lineages: ALEPH-FLAT (consensus book + KD) |
|
| 4 |
+
ALEPH-TREE (structured genome: root book + branch books) | MLP-KD
|
| 5 |
+
(traditional: best-parent weights + KD) | NO-INHERIT (evolution floor).
|
| 6 |
+
Both sides intentionally inherit logits (KD); only ours inherits geometry.
|
| 7 |
+
Consensus = Procrustes/GPA alignment of parents' books to mean shape
|
| 8 |
+
(placement by construction — replaces GM3's k-means-on-consensus init).
|
| 9 |
+
14_b: memory substrate — the D=4 home makes books size-agnostic. Implant books
|
| 10 |
+
cultivated in a small organism into a larger one (frozen / trainable), and
|
| 11 |
+
into GPT-2 relay adapters (cross-architecture frozen distillation).
|
| 12 |
+
|
| 13 |
+
Riders: pure Adam wd=0; KD = KL to detached teacher probs (predictive pressure, no
|
| 14 |
+
contrastive); tree routing is DENSE SOFT (oriented weights; collapse monitor on the
|
| 15 |
+
root); drift-check precedes every freeze claim; GPU-only verdict runs; Colab-safe.
|
| 16 |
+
Founders share a COMMON-ANCESTOR book so GPA row correspondence is inherited.
|
| 17 |
+
|
| 18 |
+
Colab paste order: geolip_vitals.py -> ar_differentiation_bed.py ->
|
| 19 |
+
exp013_augmentation_bed.py (only for run_b2) -> this file.
|
| 20 |
+
"""
|
| 21 |
+
from __future__ import annotations
|
| 22 |
+
import copy
|
| 23 |
+
import json
|
| 24 |
+
import math
|
| 25 |
+
import os
|
| 26 |
+
import torch
|
| 27 |
+
import torch.nn as nn
|
| 28 |
+
import torch.nn.functional as F
|
| 29 |
+
|
| 30 |
+
if "anchor_drift" not in globals():
|
| 31 |
+
try:
|
| 32 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 33 |
+
except ImportError:
|
| 34 |
+
_here = globals().get("__file__")
|
| 35 |
+
if _here is None:
|
| 36 |
+
raise ImportError("paste/run geolip_vitals.py first")
|
| 37 |
+
import sys, pathlib
|
| 38 |
+
sys.path.insert(0, str(pathlib.Path(_here).parent))
|
| 39 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 40 |
+
if "ByteLM" not in globals():
|
| 41 |
+
try:
|
| 42 |
+
from ar_differentiation_bed import (ByteLM, AlephAddress, _wikitext_bytes,
|
| 43 |
+
_batch, VOCAB)
|
| 44 |
+
except ImportError:
|
| 45 |
+
raise ImportError("paste/run ar_differentiation_bed.py first")
|
| 46 |
+
|
| 47 |
+
DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
|
| 48 |
+
EXP_DIR = os.path.join(DATA_ROOT, "exp014")
|
| 49 |
+
|
| 50 |
+
|
| 51 |
+
class SquaredReLU(nn.Module):
|
| 52 |
+
def forward(self, x):
|
| 53 |
+
return F.relu(x) ** 2
|
| 54 |
+
|
| 55 |
+
|
| 56 |
+
# ===================================================== consensus (the germline) ===
|
| 57 |
+
@torch.no_grad()
|
| 58 |
+
def procrustes_rotation(A: torch.Tensor, M: torch.Tensor) -> torch.Tensor:
|
| 59 |
+
"""Orthogonal R minimizing ||A R - M||_F (rows correspond)."""
|
| 60 |
+
U, _, Vt = torch.linalg.svd(A.T.double() @ M.double())
|
| 61 |
+
return (U @ Vt).float()
|
| 62 |
+
|
| 63 |
+
|
| 64 |
+
@torch.no_grad()
|
| 65 |
+
def align_to(A: torch.Tensor, ref: torch.Tensor, iters: int = 20) -> torch.Tensor:
|
| 66 |
+
"""Projective Procrustes (rows correspond, signs free): alternate the
|
| 67 |
+
orthogonal rotation and per-row sign flips (books live on RP^(D-1))."""
|
| 68 |
+
s = torch.ones(A.shape[0], 1)
|
| 69 |
+
for _ in range(iters):
|
| 70 |
+
R = procrustes_rotation(s * A, ref)
|
| 71 |
+
AR = (s * A) @ R
|
| 72 |
+
s_upd = torch.where((AR * ref).sum(-1, keepdim=True) < 0, -s, s)
|
| 73 |
+
if torch.equal(s_upd, s):
|
| 74 |
+
return AR
|
| 75 |
+
s = s_upd
|
| 76 |
+
return (s * A) @ procrustes_rotation(s * A, ref)
|
| 77 |
+
|
| 78 |
+
|
| 79 |
+
@torch.no_grad()
|
| 80 |
+
def consensus_codebook(books: list, iters: int = 50, tol: float = 1e-8):
|
| 81 |
+
"""GPA to mean shape (GM3 machinery, applied to aleph books), anchored to the
|
| 82 |
+
FIRST parent's frame. Rows must correspond (common-ancestor convention); signs
|
| 83 |
+
are projective. Returns (consensus, n_iters, delta)."""
|
| 84 |
+
# device-pin to CPU: parent models may live on CUDA after KD teacher moves
|
| 85 |
+
Bs = [F.normalize(b.detach().float().cpu(), dim=-1).clone() for b in books]
|
| 86 |
+
# pairwise projective alignment to parent-0's frame, THEN GPA refinement
|
| 87 |
+
aligned = [Bs[0]] + [align_to(b, Bs[0]) for b in Bs[1:]]
|
| 88 |
+
M = F.normalize(torch.stack(aligned).mean(0), dim=-1)
|
| 89 |
+
delta, it = 0.0, 0
|
| 90 |
+
for it in range(1, iters + 1):
|
| 91 |
+
aligned = [align_to(b, M) for b in Bs]
|
| 92 |
+
M_new = F.normalize(torch.stack(aligned).mean(0), dim=-1)
|
| 93 |
+
delta = (M_new - M).norm().item()
|
| 94 |
+
M = M_new
|
| 95 |
+
if delta < tol:
|
| 96 |
+
break
|
| 97 |
+
# re-anchor to parent-0 (GPA drift of the global frame stays measurable)
|
| 98 |
+
M = align_to(M, Bs[0])
|
| 99 |
+
return M, it, delta
|
| 100 |
+
|
| 101 |
+
|
| 102 |
+
@torch.no_grad()
|
| 103 |
+
def implant_book(addr: "AlephAddress", book: torch.Tensor, trainable: bool = True):
|
| 104 |
+
"""Load a book into an AlephAddress: codebook + home (drift measured from the
|
| 105 |
+
implant). Freeze only via trainable=False AFTER a drift-check justifies it."""
|
| 106 |
+
b = F.normalize(book.float(), dim=-1).to(addr.codebook.device)
|
| 107 |
+
assert b.shape == addr.codebook.shape, (b.shape, addr.codebook.shape)
|
| 108 |
+
addr.codebook.data.copy_(b)
|
| 109 |
+
addr.home.copy_(b)
|
| 110 |
+
addr.codebook.requires_grad_(trainable)
|
| 111 |
+
|
| 112 |
+
|
| 113 |
+
# ============================================================= tree head ==========
|
| 114 |
+
class TreeHead(nn.Module):
|
| 115 |
+
"""Autoregressive tree (the constellation-anchor analogue, exp011 TREE operator
|
| 116 |
+
in the healthy consumption regime): a ROOT aleph (K=2, D=4) yields the 4
|
| 117 |
+
oriented weights (2K half-axes = the 4 branches, dense soft, sums to 1);
|
| 118 |
+
each BRANCH is a 64-slot... shared slot projection read by a branch-specific
|
| 119 |
+
book (K=64, D=4); output = branch-weighted mixture of branch reads -> vocab.
|
| 120 |
+
Heritable genome: root book (2,4) + 4 branch books (64,4)."""
|
| 121 |
+
|
| 122 |
+
ROOT_SLOTS = 4 # slot-parallel root consumption (the collapse cure)
|
| 123 |
+
ROOT_TAU = 0.3 # softer root temperature (wave-1 fix: single hard-tau
|
| 124 |
+
# root partially collapsed, usage [.85,.12,.01,.02])
|
| 125 |
+
|
| 126 |
+
def __init__(self, d: int, vocab: int = 256, n_slots: int = 16):
|
| 127 |
+
super().__init__()
|
| 128 |
+
self.n_slots = n_slots
|
| 129 |
+
self.root_proj = nn.Linear(d, self.ROOT_SLOTS * 4, bias=False)
|
| 130 |
+
self.slot_proj = nn.Linear(d, n_slots * 4, bias=False)
|
| 131 |
+
nn.init.orthogonal_(self.root_proj.weight)
|
| 132 |
+
nn.init.orthogonal_(self.slot_proj.weight)
|
| 133 |
+
self.root = AlephAddress(2, 4, tau=self.ROOT_TAU)
|
| 134 |
+
self.branches = nn.ModuleList([AlephAddress(64, 4) for _ in range(4)])
|
| 135 |
+
self.out = nn.Linear(n_slots * 4, vocab, bias=True)
|
| 136 |
+
self._last_root = None
|
| 137 |
+
|
| 138 |
+
def forward(self, h):
|
| 139 |
+
B, n, _ = h.shape
|
| 140 |
+
rs = self.root_proj(h).view(B, n, self.ROOT_SLOTS, 4)
|
| 141 |
+
p, m = self.root.oriented(rs) # (B,n,S,2) x2
|
| 142 |
+
w = torch.cat([p, m], dim=-1).mean(dim=-2) # slot-avg -> (B,n,4)
|
| 143 |
+
self._last_root = w.detach()
|
| 144 |
+
slots = self.slot_proj(h).view(B, n, self.n_slots, 4)
|
| 145 |
+
mix = 0
|
| 146 |
+
for b, br in enumerate(self.branches):
|
| 147 |
+
mix = mix + w[..., b:b + 1] * br.m_hat(slots).reshape(B, n, -1)
|
| 148 |
+
return self.out(mix)
|
| 149 |
+
|
| 150 |
+
def genome(self):
|
| 151 |
+
return {"root": self.root.codebook.detach().clone(),
|
| 152 |
+
**{f"branch{i}": br.codebook.detach().clone()
|
| 153 |
+
for i, br in enumerate(self.branches)}}
|
| 154 |
+
|
| 155 |
+
@torch.no_grad()
|
| 156 |
+
def inherit(self, genomes: list):
|
| 157 |
+
c, it, dl = consensus_codebook([g["root"] for g in genomes])
|
| 158 |
+
implant_book(self.root, c)
|
| 159 |
+
for i, br in enumerate(self.branches):
|
| 160 |
+
c, _, _ = consensus_codebook([g[f"branch{i}"] for g in genomes])
|
| 161 |
+
implant_book(br, c)
|
| 162 |
+
|
| 163 |
+
@torch.no_grad()
|
| 164 |
+
def vitals(self):
|
| 165 |
+
out = {"root_drift": round(anchor_drift(self.root.codebook,
|
| 166 |
+
self.root.home)["mean"], 4)}
|
| 167 |
+
if self._last_root is not None:
|
| 168 |
+
w = self._last_root.reshape(-1, 4)
|
| 169 |
+
usage = w.mean(0)
|
| 170 |
+
usage = usage / usage.sum()
|
| 171 |
+
out["root_usage"] = [round(float(u), 3) for u in usage]
|
| 172 |
+
ent = -(usage.clamp_min(1e-9) * usage.clamp_min(1e-9).log()).sum()
|
| 173 |
+
out["root_ppl4"] = round(float(ent.exp()), 3)
|
| 174 |
+
d = [anchor_drift(br.codebook, br.home)["mean"] for br in self.branches]
|
| 175 |
+
out["branch_drift"] = [round(x, 3) for x in d]
|
| 176 |
+
return out
|
| 177 |
+
|
| 178 |
+
|
| 179 |
+
# ============================================================ organisms ===========
|
| 180 |
+
def make_organism(lineage: str, d: int = 192, layers: int = 4, block: int = 256,
|
| 181 |
+
seed: int = 0):
|
| 182 |
+
"""lineage in {aleph_flat, aleph_tree, mlp_kd, no_inherit}. no_inherit uses the
|
| 183 |
+
aleph_flat architecture (the control isolates INHERITANCE, not architecture)."""
|
| 184 |
+
torch.manual_seed(seed)
|
| 185 |
+
if lineage in ("aleph_flat", "no_inherit", "aleph_full", "aleph_weights"):
|
| 186 |
+
return ByteLM("addr_msl64", d=d, layers=layers, block=block)
|
| 187 |
+
if lineage == "aleph_tree":
|
| 188 |
+
m = ByteLM("sdpa", d=d, layers=layers, block=block)
|
| 189 |
+
m.head = TreeHead(d)
|
| 190 |
+
return m
|
| 191 |
+
if lineage == "mlp_kd":
|
| 192 |
+
m = ByteLM("sdpa", d=d, layers=layers, block=block)
|
| 193 |
+
m.head = nn.Sequential(nn.Linear(d, 224), SquaredReLU(),
|
| 194 |
+
nn.LayerNorm(224), nn.Linear(224, VOCAB))
|
| 195 |
+
return m
|
| 196 |
+
raise ValueError(lineage)
|
| 197 |
+
|
| 198 |
+
|
| 199 |
+
def genome_of(model):
|
| 200 |
+
"""The heritable organ = co-adapted (projection, book) pair(s). Books inherit
|
| 201 |
+
by CONSENSUS (the geometric germline); projections inherit from the BEST
|
| 202 |
+
parent (weight copy) — implanting a book against a random projection puts the
|
| 203 |
+
child below random init (campaign-v2 lesson)."""
|
| 204 |
+
if isinstance(model.head, TreeHead):
|
| 205 |
+
g = model.head.genome()
|
| 206 |
+
g["root_proj"] = model.head.root_proj.weight.detach().cpu().clone()
|
| 207 |
+
g["slot_proj"] = model.head.slot_proj.weight.detach().cpu().clone()
|
| 208 |
+
return g
|
| 209 |
+
if hasattr(model, "head_addr"):
|
| 210 |
+
return {"flat": model.head_addr.codebook.detach().cpu().clone(),
|
| 211 |
+
"proj": model.head_proj.weight.detach().cpu().clone()}
|
| 212 |
+
return None
|
| 213 |
+
|
| 214 |
+
|
| 215 |
+
@torch.no_grad()
|
| 216 |
+
def inherit_genome(model, genomes: list):
|
| 217 |
+
"""genomes[0] = the BEST parent (selection order matters)."""
|
| 218 |
+
if isinstance(model.head, TreeHead):
|
| 219 |
+
model.head.inherit(genomes)
|
| 220 |
+
model.head.root_proj.weight.copy_(genomes[0]["root_proj"].to(
|
| 221 |
+
model.head.root_proj.weight.device))
|
| 222 |
+
model.head.slot_proj.weight.copy_(genomes[0]["slot_proj"].to(
|
| 223 |
+
model.head.slot_proj.weight.device))
|
| 224 |
+
elif hasattr(model, "head_addr"):
|
| 225 |
+
c, it, dl = consensus_codebook([g["flat"] for g in genomes])
|
| 226 |
+
implant_book(model.head_addr, c)
|
| 227 |
+
model.head_proj.weight.copy_(genomes[0]["proj"].to(
|
| 228 |
+
model.head_proj.weight.device))
|
| 229 |
+
|
| 230 |
+
|
| 231 |
+
def organism_vitals(model):
|
| 232 |
+
if isinstance(model.head, TreeHead):
|
| 233 |
+
return model.head.vitals()
|
| 234 |
+
return model.vitals() if hasattr(model, "vitals") else {}
|
| 235 |
+
|
| 236 |
+
|
| 237 |
+
# ========================================================= train one member ======
|
| 238 |
+
def train_member(model, tr, va, steps=2000, batch=32, block=256, device="cuda",
|
| 239 |
+
seed=0, teachers=None, kd_alpha=1.0):
|
| 240 |
+
"""CE (+ KL to detached mean teacher probs when teachers given). Pure Adam."""
|
| 241 |
+
g = torch.Generator().manual_seed(seed)
|
| 242 |
+
model = model.to(device)
|
| 243 |
+
opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
|
| 244 |
+
if teachers:
|
| 245 |
+
teachers = [t.to(device).eval() for t in teachers]
|
| 246 |
+
for step in range(1, steps + 1):
|
| 247 |
+
x, y = _batch(tr, batch, block, device, g)
|
| 248 |
+
logits = model(x)
|
| 249 |
+
loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
|
| 250 |
+
if teachers:
|
| 251 |
+
with torch.no_grad():
|
| 252 |
+
tp = torch.stack([F.softmax(t(x), -1) for t in teachers]).mean(0)
|
| 253 |
+
loss = loss + kd_alpha * F.kl_div(
|
| 254 |
+
F.log_softmax(logits, -1), tp, reduction="batchmean")
|
| 255 |
+
opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
|
| 256 |
+
model.eval()
|
| 257 |
+
with torch.no_grad():
|
| 258 |
+
ls = []
|
| 259 |
+
for _ in range(20):
|
| 260 |
+
xv, yv = _batch(va, batch, block, device, g)
|
| 261 |
+
ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
|
| 262 |
+
yv.reshape(-1)).item())
|
| 263 |
+
return sum(ls) / len(ls) / math.log(2) # bpb
|
| 264 |
+
|
| 265 |
+
|
| 266 |
+
# ============================================================ the tournament ======
|
| 267 |
+
def run_tournament(lineage: str, gens: int = 4, pop: int = 4, steps: int = 2000,
|
| 268 |
+
seed: int = 0, device: str = "cuda",
|
| 269 |
+
catastrophic_at: int | None = None):
|
| 270 |
+
"""One lineage, one tournament seed. Logs per-gen to the ledger; saves the
|
| 271 |
+
champion genome per generation. catastrophic_at=G injects a 0-step random
|
| 272 |
+
parent into the consensus at generation G (the GM3 robustness probe)."""
|
| 273 |
+
if not torch.cuda.is_available():
|
| 274 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 275 |
+
os.makedirs(EXP_DIR, exist_ok=True)
|
| 276 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 277 |
+
ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 278 |
+
# common ancestor: every founder book starts identical within a tournament
|
| 279 |
+
torch.manual_seed(9000 + seed)
|
| 280 |
+
ancestor = make_organism(lineage, seed=9000 + seed)
|
| 281 |
+
anc_genome = genome_of(ancestor)
|
| 282 |
+
parents, parent_models, champion_genomes = [], [], []
|
| 283 |
+
for gen in range(gens):
|
| 284 |
+
members = []
|
| 285 |
+
for i in range(pop):
|
| 286 |
+
mseed = seed * 1000 + gen * 100 + i
|
| 287 |
+
m = make_organism(lineage, seed=mseed)
|
| 288 |
+
if anc_genome and gen == 0:
|
| 289 |
+
inherit_genome(m, [anc_genome]) # common ancestor
|
| 290 |
+
if gen > 0:
|
| 291 |
+
is_fresh = (i == pop - 1) # gene flow founder
|
| 292 |
+
if not is_fresh:
|
| 293 |
+
if lineage in ("aleph_flat", "aleph_tree"):
|
| 294 |
+
gs = [genome_of(pm) for pm in parent_models]
|
| 295 |
+
if catastrophic_at == gen:
|
| 296 |
+
bad = make_organism(lineage, seed=666 + i)
|
| 297 |
+
gs = gs + [genome_of(bad)]
|
| 298 |
+
inherit_genome(m, gs)
|
| 299 |
+
elif lineage == "aleph_full":
|
| 300 |
+
# v4 arm (v3 lesson: continuity is what pays) — inherit the
|
| 301 |
+
# WHOLE best parent, then overwrite the book with the
|
| 302 |
+
# two-parent consensus: germline ON TOP of continuity.
|
| 303 |
+
m.load_state_dict(copy.deepcopy(
|
| 304 |
+
parent_models[0].state_dict()))
|
| 305 |
+
gs = [genome_of(pm) for pm in parent_models]
|
| 306 |
+
if catastrophic_at == gen:
|
| 307 |
+
bad = make_organism(lineage, seed=666 + i)
|
| 308 |
+
gs = gs + [genome_of(bad)]
|
| 309 |
+
c, _, _ = consensus_codebook([g["flat"] for g in gs])
|
| 310 |
+
implant_book(m.head_addr, c)
|
| 311 |
+
elif lineage in ("mlp_kd", "aleph_weights"):
|
| 312 |
+
# pure continuity (no germline op) — aleph_weights is the
|
| 313 |
+
# within-architecture control for aleph_full
|
| 314 |
+
m.load_state_dict(copy.deepcopy(
|
| 315 |
+
parent_models[0].state_dict()))
|
| 316 |
+
# no_inherit: nothing
|
| 317 |
+
# KD: alpha 0.25 (campaign-v2 lesson: alpha=1.0 from near-parity
|
| 318 |
+
# teachers COMPOUNDS DOWNWARD — inverse evolution; fresh-founder
|
| 319 |
+
# control isolated it). Fresh founders get NO KD (clean gene flow).
|
| 320 |
+
is_fresh_now = (gen > 0 and i == pop - 1)
|
| 321 |
+
teachers = parent_models if (gen > 0 and not is_fresh_now
|
| 322 |
+
and lineage != "no_inherit") else None
|
| 323 |
+
bpb = train_member(m, tr, va, steps=steps, device=device,
|
| 324 |
+
seed=mseed, teachers=teachers, kd_alpha=0.25)
|
| 325 |
+
vit = organism_vitals(m)
|
| 326 |
+
members.append((bpb, m))
|
| 327 |
+
rec = {"exp": "14a", "lineage": lineage, "tseed": seed, "gen": gen,
|
| 328 |
+
"member": i, "fresh": gen > 0 and i == pop - 1,
|
| 329 |
+
"catastrophic": catastrophic_at == gen, "bpb": round(bpb, 4),
|
| 330 |
+
"vitals": vit}
|
| 331 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 332 |
+
print(f"[14a {lineage} t{seed} g{gen} m{i}] bpb={bpb:.4f} {vit}",
|
| 333 |
+
flush=True)
|
| 334 |
+
members.sort(key=lambda t: t[0])
|
| 335 |
+
parent_models = [members[0][1].cpu(), members[1][1].cpu()]
|
| 336 |
+
best = members[0][0]
|
| 337 |
+
gene = genome_of(members[0][1])
|
| 338 |
+
if gene:
|
| 339 |
+
torch.save(gene, os.path.join(
|
| 340 |
+
EXP_DIR, f"champion_{lineage}_t{seed}_g{gen}.pt"))
|
| 341 |
+
champion_genomes.append(gene)
|
| 342 |
+
print(f"[14a {lineage} t{seed} g{gen}] BEST={best:.4f} "
|
| 343 |
+
f"mean={sum(b for b, _ in members)/pop:.4f}", flush=True)
|
| 344 |
+
for _, mm in members[2:]:
|
| 345 |
+
del mm
|
| 346 |
+
torch.cuda.empty_cache()
|
| 347 |
+
ledger.close()
|
| 348 |
+
return best
|
| 349 |
+
|
| 350 |
+
|
| 351 |
+
# ============================================================ 14_b implants ======
|
| 352 |
+
def run_b1(steps: int = 2000, seed: int = 0, device: str = "cuda",
|
| 353 |
+
donor_book: torch.Tensor | None = None, tag: str = "small_cultivated"):
|
| 354 |
+
"""Cross-size: donor book (default: cultivate in a small organism) implanted
|
| 355 |
+
into a LARGE organism. Arms: fresh | implant-trainable | implant-frozen | mlp."""
|
| 356 |
+
if not torch.cuda.is_available():
|
| 357 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 358 |
+
os.makedirs(EXP_DIR, exist_ok=True)
|
| 359 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 360 |
+
ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 361 |
+
if donor_book is None:
|
| 362 |
+
small = make_organism("aleph_flat", d=128, layers=2, seed=seed)
|
| 363 |
+
bpb_small = train_member(small, tr, va, steps=steps, device=device, seed=seed)
|
| 364 |
+
donor_book = genome_of(small)["flat"]
|
| 365 |
+
print(f"[14b donor small] bpb={bpb_small:.4f}", flush=True)
|
| 366 |
+
results = {}
|
| 367 |
+
for arm in ("fresh", "implant_train", "implant_frozen", "mlp"):
|
| 368 |
+
lineage = "mlp_kd" if arm == "mlp" else "aleph_flat"
|
| 369 |
+
m = make_organism(lineage, d=384, layers=6, seed=seed + 10)
|
| 370 |
+
if arm.startswith("implant"):
|
| 371 |
+
implant_book(m.head_addr, donor_book, trainable=(arm == "implant_train"))
|
| 372 |
+
bpb = train_member(m, tr, va, steps=steps, device=device, seed=seed + 10)
|
| 373 |
+
vit = organism_vitals(m)
|
| 374 |
+
results[arm] = {"bpb": round(bpb, 4), "vitals": vit}
|
| 375 |
+
rec = {"exp": "14b1", "arm": arm, "donor": tag, "seed": seed,
|
| 376 |
+
"bpb": round(bpb, 4), "vitals": vit}
|
| 377 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 378 |
+
print(f"[14b1 {arm} donor={tag}] bpb={bpb:.4f} {vit}", flush=True)
|
| 379 |
+
del m; torch.cuda.empty_cache()
|
| 380 |
+
ledger.close()
|
| 381 |
+
return results, donor_book
|
| 382 |
+
|
| 383 |
+
|
| 384 |
+
def run_b2(donor_book: torch.Tensor, steps: int = 1500, seed: int = 0,
|
| 385 |
+
device: str = "cuda", tag: str = "small_cultivated"):
|
| 386 |
+
"""Cross-architecture: implant the donor book into every GPT-2 relay adapter
|
| 387 |
+
(exp013 Track C bed) vs random-init relays. Books are (64,4) — size-agnostic."""
|
| 388 |
+
from exp013_augmentation_bed import _wikitext_lines
|
| 389 |
+
from transformers import GPT2LMHeadModel, GPT2TokenizerFast
|
| 390 |
+
if "MslRelay" not in globals():
|
| 391 |
+
from ar_differentiation_bed import MslRelay
|
| 392 |
+
from exp013_augmentation_bed import _BlockWithAdapter
|
| 393 |
+
tok = GPT2TokenizerFast.from_pretrained("gpt2")
|
| 394 |
+
tr_lines, va_lines = _wikitext_lines(DATA_ROOT)
|
| 395 |
+
stream_tr = tok("\n\n".join(tr_lines[:8000]), return_tensors="pt").input_ids[0]
|
| 396 |
+
stream_va = tok("\n\n".join(va_lines[:1000]), return_tensors="pt").input_ids[0]
|
| 397 |
+
ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 398 |
+
out = {}
|
| 399 |
+
for arm in ("random_relays", "implanted_relays"):
|
| 400 |
+
torch.manual_seed(seed)
|
| 401 |
+
g = torch.Generator().manual_seed(seed)
|
| 402 |
+
model = GPT2LMHeadModel.from_pretrained("gpt2").to(device)
|
| 403 |
+
for p in model.parameters():
|
| 404 |
+
p.requires_grad_(False)
|
| 405 |
+
adapters = []
|
| 406 |
+
for i, blk in enumerate(model.transformer.h):
|
| 407 |
+
ad = MslRelay(model.config.n_embd).to(device)
|
| 408 |
+
if arm == "implanted_relays":
|
| 409 |
+
implant_book(ad.addr, donor_book, trainable=True)
|
| 410 |
+
model.transformer.h[i] = _BlockWithAdapter(blk, ad)
|
| 411 |
+
adapters.append(ad)
|
| 412 |
+
params = [p for ad in adapters for p in ad.parameters()
|
| 413 |
+
if p.requires_grad]
|
| 414 |
+
opt = torch.optim.Adam(params, lr=1e-3, weight_decay=0.0)
|
| 415 |
+
block = 256
|
| 416 |
+
for step in range(1, steps + 1):
|
| 417 |
+
ix = torch.randint(0, stream_tr.numel() - block - 1, (8,), generator=g)
|
| 418 |
+
x = torch.stack([stream_tr[i:i + block] for i in ix]).to(device)
|
| 419 |
+
loss = model(x, labels=x).loss
|
| 420 |
+
opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
|
| 421 |
+
model.eval()
|
| 422 |
+
with torch.no_grad():
|
| 423 |
+
ls = []
|
| 424 |
+
for j in range(0, stream_va.numel() - block - 1, block * 4):
|
| 425 |
+
x = stream_va[j:j + block].unsqueeze(0).to(device)
|
| 426 |
+
ls.append(model(x, labels=x).loss.item())
|
| 427 |
+
ppl = math.exp(sum(ls) / len(ls))
|
| 428 |
+
gates = [round(ad.gate.sigmoid().item(), 4) for ad in adapters]
|
| 429 |
+
drifts = [round(anchor_drift(ad.addr.codebook, ad.addr.home)["mean"], 3)
|
| 430 |
+
for ad in adapters]
|
| 431 |
+
out[arm] = {"ppl": round(ppl, 3), "gates": gates, "drift": drifts}
|
| 432 |
+
rec = {"exp": "14b2", "arm": arm, "donor": tag, "seed": seed,
|
| 433 |
+
"ppl": round(ppl, 3), "gates": gates, "drift": drifts}
|
| 434 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 435 |
+
print(f"[14b2 {arm} donor={tag}] ppl={ppl:.3f} gates={gates[:3]}.. "
|
| 436 |
+
f"drift={drifts[:3]}..", flush=True)
|
| 437 |
+
del model; torch.cuda.empty_cache()
|
| 438 |
+
ledger.close()
|
| 439 |
+
return out
|
| 440 |
+
|
| 441 |
+
|
| 442 |
+
# ================================================================ smoke ===========
|
| 443 |
+
def smoke():
|
| 444 |
+
"""CPU shapes/parse only: GPA ground truth, tree causality, implant, KD path."""
|
| 445 |
+
g = torch.Generator().manual_seed(0)
|
| 446 |
+
# GPA: two rotated (+row-sign-flipped) copies of one book must align back to it
|
| 447 |
+
A = F.normalize(torch.randn(64, 4, generator=g), dim=-1)
|
| 448 |
+
q, _ = torch.linalg.qr(torch.randn(4, 4, generator=g))
|
| 449 |
+
B = A @ q
|
| 450 |
+
B[::3] = -B[::3]
|
| 451 |
+
C, it, dl = consensus_codebook([A, B])
|
| 452 |
+
cos = (F.normalize(C, dim=-1) * A).sum(-1).abs().mean()
|
| 453 |
+
assert cos > 0.999, cos
|
| 454 |
+
print(f"GPA OK (iters={it}, delta={dl:.2e}, |cos to truth|={cos:.5f})")
|
| 455 |
+
x = torch.randint(0, 256, (2, 64))
|
| 456 |
+
for lineage in ("aleph_flat", "aleph_tree", "mlp_kd", "no_inherit"):
|
| 457 |
+
m = make_organism(lineage, d=96, layers=2, block=64, seed=0)
|
| 458 |
+
lg = m(x); assert lg.shape == (2, 64, 256); lg.sum().backward()
|
| 459 |
+
with torch.no_grad():
|
| 460 |
+
a = m(x)[0, 10]; x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
|
| 461 |
+
b = m(x2)[0, 10]
|
| 462 |
+
assert torch.allclose(a, b, atol=1e-4), lineage + " leaks"
|
| 463 |
+
gnm = genome_of(m)
|
| 464 |
+
if gnm:
|
| 465 |
+
inherit_genome(m, [gnm, gnm]) # self-consensus = identity-ish
|
| 466 |
+
print(lineage, "OK params",
|
| 467 |
+
f"{sum(p.numel() for p in m.parameters()):,}",
|
| 468 |
+
organism_vitals(m) if lineage != "mlp_kd" else {})
|
| 469 |
+
# KD path: teacher forward + KL backward
|
| 470 |
+
t = make_organism("mlp_kd", d=96, layers=2, block=64, seed=1)
|
| 471 |
+
s = make_organism("aleph_flat", d=96, layers=2, block=64, seed=2)
|
| 472 |
+
tp = F.softmax(t(x), -1).detach()
|
| 473 |
+
loss = F.kl_div(F.log_softmax(s(x), -1), tp, reduction="batchmean")
|
| 474 |
+
loss.backward()
|
| 475 |
+
print("KD OK — exp014 smoke passed (tournament on GPU: run_tournament(...))")
|
| 476 |
+
|
| 477 |
+
|
| 478 |
+
def _in_notebook():
|
| 479 |
+
try:
|
| 480 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 481 |
+
return True
|
| 482 |
+
except NameError:
|
| 483 |
+
return False
|
| 484 |
+
|
| 485 |
+
|
| 486 |
+
if __name__ == "__main__":
|
| 487 |
+
if _in_notebook():
|
| 488 |
+
smoke()
|
| 489 |
+
print("Notebook: run_tournament('aleph_flat'), run_b1(), run_b2(book).")
|
| 490 |
+
else:
|
| 491 |
+
import argparse
|
| 492 |
+
ap = argparse.ArgumentParser()
|
| 493 |
+
ap.add_argument("--mode", default="smoke",
|
| 494 |
+
choices=["smoke", "tournament", "b1", "b2"])
|
| 495 |
+
ap.add_argument("--lineage", default="aleph_full",
|
| 496 |
+
help="tournament lineage: aleph_flat|aleph_full|"
|
| 497 |
+
"aleph_weights|aleph_tree|mlp_kd|no_inherit")
|
| 498 |
+
ap.add_argument("--seed", type=int, default=0)
|
| 499 |
+
ap.add_argument("--steps", type=int, default=2000)
|
| 500 |
+
ap.add_argument("--genome", default="genomes/champion_aleph_full_t0_g3.pt",
|
| 501 |
+
help="donor genome .pt for --mode b1/b2 (uses its 'flat' book)")
|
| 502 |
+
a, _ = ap.parse_known_args()
|
| 503 |
+
if a.mode == "smoke":
|
| 504 |
+
smoke()
|
| 505 |
+
elif a.mode == "tournament":
|
| 506 |
+
run_tournament(a.lineage, steps=a.steps, seed=a.seed)
|
| 507 |
+
elif a.mode == "b1":
|
| 508 |
+
donor = (torch.load(a.genome, map_location="cpu")["flat"]
|
| 509 |
+
if os.path.exists(a.genome) else None)
|
| 510 |
+
run_b1(steps=a.steps, seed=a.seed, donor_book=donor,
|
| 511 |
+
tag=os.path.basename(a.genome) if donor is not None
|
| 512 |
+
else "small_cultivated")
|
| 513 |
+
elif a.mode == "b2":
|
| 514 |
+
donor = torch.load(a.genome, map_location="cpu")["flat"]
|
| 515 |
+
run_b2(donor, seed=a.seed, tag=os.path.basename(a.genome))
|
exp018_r12/exp018_reran12.py
ADDED
|
@@ -0,0 +1,176 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""exp018_reran12.py — EXPERIMENT 18: exp012 RERUN with the keystone aleph model
|
| 2 |
+
parameters (exp012 and its certified record stay untouched — this is a labeled
|
| 3 |
+
rerun, not a replacement). A process review found the exp012 bed's AlephAddress
|
| 4 |
+
primitive faithful to the keystone (per the aleph-void article) but two keystone avenues
|
| 5 |
+
never exercised: the caller-supplied FARMED/PENTACHORON codebook init and the
|
| 6 |
+
TIED M_hat readout (U=M_hat, S=Omega-token, Vt=I). This rerun exercises them on the
|
| 7 |
+
certified bed, unchanged otherwise (d=192, 4 layers, wikitext-2 bytes, 2000
|
| 8 |
+
steps, pure Adam wd=0, addr_msl64 multi-slot base).
|
| 9 |
+
|
| 10 |
+
Arms (x2 seeds):
|
| 11 |
+
penta_init — codebook initialized from geovocab2 REGULAR PENTACHORON
|
| 12 |
+
vertices (13 randomly-rotated regular 4-simplices -> 64 rows on
|
| 13 |
+
S^3, row-normalized; SimplexFactory is the source per keystone
|
| 14 |
+
"geovocab pentachoron vertices"). Free linear head (isolates
|
| 15 |
+
the init).
|
| 16 |
+
farmed_init — codebook initialized from a FARMED recon-real book: the exp012
|
| 17 |
+
certified specimen addr_msl64_s0_t2000's trained codebook —
|
| 18 |
+
anchors the M_hat gradient itself produced. Free linear head.
|
| 19 |
+
(The direct "farm FROM the aleph logit structure" arm.)
|
| 20 |
+
tied — random init + TIED readout: zero-parameter head; logits =
|
| 21 |
+
M_hat_slots @ head_proj.weight @ emb.weight^T (decode tied back
|
| 22 |
+
through the encoder projection and the byte embedding — the
|
| 23 |
+
AR form of the keystone's tied linear off M_hat).
|
| 24 |
+
keystone — penta_init + tied (the keystone-faithful configuration).
|
| 25 |
+
Baselines (exp012 certified, cited not rerun): addr_msl64 2.469 / sdpa 2.518
|
| 26 |
+
(3-seed means @2k); certified bed reproduction 2.4990 (s0).
|
| 27 |
+
Riders: pure Adam wd=0; GPU-only verdicts; >=2 seeds; drift-check before any
|
| 28 |
+
freeze claim; Colab-safe. Paste order: geolip_vitals ->
|
| 29 |
+
ar_differentiation_bed -> exp014_genetic_distillation -> this file.
|
| 30 |
+
"""
|
| 31 |
+
from __future__ import annotations
|
| 32 |
+
import json
|
| 33 |
+
import os
|
| 34 |
+
import torch
|
| 35 |
+
import torch.nn as nn
|
| 36 |
+
import torch.nn.functional as F
|
| 37 |
+
|
| 38 |
+
if "ByteLM" not in globals():
|
| 39 |
+
try:
|
| 40 |
+
from ar_differentiation_bed import ByteLM, _wikitext_bytes
|
| 41 |
+
from exp014_genetic_distillation import train_member, implant_book
|
| 42 |
+
from geolip_vitals import anchor_drift
|
| 43 |
+
except ImportError:
|
| 44 |
+
_here = globals().get("__file__")
|
| 45 |
+
if _here is None:
|
| 46 |
+
raise ImportError("paste geolip_vitals + ar_differentiation_bed + "
|
| 47 |
+
"exp014_genetic_distillation first")
|
| 48 |
+
import sys, pathlib
|
| 49 |
+
sys.path.insert(0, str(pathlib.Path(_here).parent))
|
| 50 |
+
from ar_differentiation_bed import ByteLM, _wikitext_bytes
|
| 51 |
+
from exp014_genetic_distillation import train_member, implant_book
|
| 52 |
+
from geolip_vitals import anchor_drift
|
| 53 |
+
|
| 54 |
+
DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
|
| 55 |
+
RERUN_DIR = os.path.join(DATA_ROOT, "exp018")
|
| 56 |
+
FARMED_SPECIMEN = os.environ.get(
|
| 57 |
+
"GEOLIP_FARMED_SPECIMEN",
|
| 58 |
+
os.path.join(os.path.dirname(os.path.abspath(__file__)), "..", "exp012_ar",
|
| 59 |
+
"specimens", "addr_msl64_s0_t2000.pt"))
|
| 60 |
+
|
| 61 |
+
|
| 62 |
+
def pentachoron_book(K: int = 64, seed: int = 0) -> torch.Tensor:
|
| 63 |
+
"""K rows on S^3 from geovocab2 regular pentachoron vertices: ceil(K/5)
|
| 64 |
+
regular 4-simplices, each rotated by a seeded random SO(4), stacked and
|
| 65 |
+
row-normalized. The keystone's 'geovocab pentachoron vertices' init."""
|
| 66 |
+
from geovocab2.shapes.factory.simplex_factory import SimplexFactory
|
| 67 |
+
g = torch.Generator().manual_seed(seed)
|
| 68 |
+
# the factory's regular construction lives in k+1 = 5 dims; center it and
|
| 69 |
+
# project isometrically onto its rank-4 subspace (SVD) -> exact regular
|
| 70 |
+
# pentachoron in R^4 (pairwise cos = -1/4 on the sphere)
|
| 71 |
+
fac = SimplexFactory(k=4, embed_dim=5, method="regular")
|
| 72 |
+
base5 = fac.build_torch().double() # (5, 5)
|
| 73 |
+
base5 = base5 - base5.mean(dim=0, keepdim=True)
|
| 74 |
+
_, _, Vt = torch.linalg.svd(base5, full_matrices=False)
|
| 75 |
+
base = (base5 @ Vt[:4].T).float() # (5, 4) regular
|
| 76 |
+
rows = []
|
| 77 |
+
for _ in range((K + 4) // 5):
|
| 78 |
+
q, _ = torch.linalg.qr(torch.randn(4, 4, generator=g, dtype=torch.float64))
|
| 79 |
+
rows.append((base.double() @ q).float())
|
| 80 |
+
book = torch.cat(rows)[:K]
|
| 81 |
+
return F.normalize(book, dim=-1)
|
| 82 |
+
|
| 83 |
+
|
| 84 |
+
def farmed_book() -> torch.Tensor:
|
| 85 |
+
"""The exp012 certified specimen's trained codebook — anchors farmed by the
|
| 86 |
+
M_hat recon/CE gradient itself (recon-real by construction)."""
|
| 87 |
+
ck = torch.load(FARMED_SPECIMEN, map_location="cpu", weights_only=True)
|
| 88 |
+
return F.normalize(ck["state_dict"]["head_addr.codebook"].float(), dim=-1)
|
| 89 |
+
|
| 90 |
+
|
| 91 |
+
class TiedReadout(nn.Module):
|
| 92 |
+
"""Zero-parameter tied decode: M_hat slots -> back through the encoder
|
| 93 |
+
projection -> byte-embedding transpose. Gradients flow into head_proj and
|
| 94 |
+
emb through both encode and decode paths (the tie)."""
|
| 95 |
+
|
| 96 |
+
def __init__(self, head_proj: nn.Linear, emb: nn.Embedding):
|
| 97 |
+
super().__init__()
|
| 98 |
+
self._proj = [head_proj] # references, not submodules
|
| 99 |
+
self._emb = [emb]
|
| 100 |
+
|
| 101 |
+
def forward(self, feats): # feats (..., P*4) = M_hat per slot
|
| 102 |
+
h = feats @ self._proj[0].weight # (P*4, d) -> (..., d)
|
| 103 |
+
return h @ self._emb[0].weight.T # (..., VOCAB)
|
| 104 |
+
|
| 105 |
+
|
| 106 |
+
def make_rerun_model(arm: str, seed: int = 0, d: int = 192, layers: int = 4,
|
| 107 |
+
block: int = 256):
|
| 108 |
+
torch.manual_seed(seed)
|
| 109 |
+
m = ByteLM("addr_msl64", d=d, layers=layers, block=block)
|
| 110 |
+
if arm in ("penta_init", "keystone"):
|
| 111 |
+
implant_book(m.head_addr, pentachoron_book(64, seed=seed))
|
| 112 |
+
elif arm == "farmed_init":
|
| 113 |
+
implant_book(m.head_addr, farmed_book())
|
| 114 |
+
if arm in ("tied", "keystone"):
|
| 115 |
+
m.head = TiedReadout(m.head_proj, m.emb)
|
| 116 |
+
return m
|
| 117 |
+
|
| 118 |
+
|
| 119 |
+
def run_rerun(arms=("penta_init", "farmed_init", "tied", "keystone"),
|
| 120 |
+
seeds=(0, 1), steps=2000, device="cuda"):
|
| 121 |
+
if not torch.cuda.is_available():
|
| 122 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 123 |
+
os.makedirs(RERUN_DIR, exist_ok=True)
|
| 124 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 125 |
+
ledger = open(os.path.join(RERUN_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 126 |
+
for seed in seeds:
|
| 127 |
+
for arm in arms:
|
| 128 |
+
m = make_rerun_model(arm, seed=seed)
|
| 129 |
+
bpb = train_member(m, tr, va, steps=steps, device=device, seed=seed)
|
| 130 |
+
d = anchor_drift(m.head_addr.codebook, m.head_addr.home)
|
| 131 |
+
rec = {"exp": "18", "arm": arm, "seed": seed, "steps": steps,
|
| 132 |
+
"bpb": round(bpb, 4), "drift": round(d["mean"], 4),
|
| 133 |
+
"binding_frac": round(d["binding_fraction"], 4)}
|
| 134 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 135 |
+
print(f"[18r {arm} s{seed}] FINAL bpb={bpb:.4f} "
|
| 136 |
+
f"drift={rec['drift']} bind={rec['binding_frac']}", flush=True)
|
| 137 |
+
torch.save({"arm": arm, "seed": seed, "steps": steps,
|
| 138 |
+
"state_dict": {k: v.cpu() for k, v in
|
| 139 |
+
m.state_dict().items()}},
|
| 140 |
+
os.path.join(RERUN_DIR, f"r12_{arm}_s{seed}.pt"))
|
| 141 |
+
del m
|
| 142 |
+
torch.cuda.empty_cache()
|
| 143 |
+
ledger.close()
|
| 144 |
+
|
| 145 |
+
|
| 146 |
+
def smoke():
|
| 147 |
+
b = pentachoron_book(64, seed=0)
|
| 148 |
+
assert b.shape == (64, 4)
|
| 149 |
+
assert torch.allclose(b.norm(dim=-1), torch.ones(64), atol=1e-5)
|
| 150 |
+
# regular-simplex signature: within one pentachoron, pairwise cos ~ -1/4
|
| 151 |
+
c = (b[:5] @ b[:5].T)
|
| 152 |
+
off = c[~torch.eye(5, dtype=torch.bool)]
|
| 153 |
+
assert (off - (-0.25)).abs().max() < 1e-4, off
|
| 154 |
+
x = torch.randint(0, 256, (2, 64))
|
| 155 |
+
for arm in ("penta_init", "tied", "keystone"):
|
| 156 |
+
m = make_rerun_model(arm, seed=0, d=96, layers=2, block=64)
|
| 157 |
+
lg = m(x); assert lg.shape == (2, 64, 256); lg.sum().backward()
|
| 158 |
+
assert m.head_addr.codebook.grad is not None
|
| 159 |
+
if arm in ("tied", "keystone"):
|
| 160 |
+
assert sum(p.numel() for p in m.head.parameters()) == 0
|
| 161 |
+
assert m.emb.weight.grad is not None
|
| 162 |
+
m.zero_grad()
|
| 163 |
+
print("exp018 (reran 12) smoke passed (farmed_init needs the specimen on disk)")
|
| 164 |
+
|
| 165 |
+
|
| 166 |
+
def _in_notebook():
|
| 167 |
+
try:
|
| 168 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 169 |
+
return True
|
| 170 |
+
except NameError:
|
| 171 |
+
return False
|
| 172 |
+
|
| 173 |
+
|
| 174 |
+
if __name__ == "__main__":
|
| 175 |
+
smoke() if not _in_notebook() else (smoke(),
|
| 176 |
+
print("Notebook: run_rerun() on GPU."))
|
exp018_r12/geolip_vitals.py
ADDED
|
@@ -0,0 +1,219 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""geolip_vitals.py — shared diagnostic harness for the GeoLIP aleph experiments.
|
| 2 |
+
ALL functions are READOUTS: no gradients, no losses. CV is a readout, never a
|
| 3 |
+
force. Addressing is judged by drift->0.29154 and CV->0.20, never by recon cosine
|
| 4 |
+
(judgment criteria per the aleph-void article: https://huggingface.co/blog/AbstractPhil/geometric-vocabulary-patchwork-aleph-void).
|
| 5 |
+
|
| 6 |
+
Vitals provided:
|
| 7 |
+
anchor_drift — geodesic drift of anchors from init; binding fraction @0.29154
|
| 8 |
+
pentachoron_cv — CM 4-volume CV over random 5-row subsets (geovocab2 import)
|
| 9 |
+
axis_aliveness — oriented-address usage: axes alive, hppl, collapse flag
|
| 10 |
+
gate_stats — gate means vs the 0.012-0.03 band
|
| 11 |
+
path_diversity — unique-path counting, FIXED high-bits hash (low-16 bug is the
|
| 12 |
+
retracted artifact — never use the low bits)
|
| 13 |
+
grad_norm_spread — gradient democracy monitor (orders-of-magnitude spread)
|
| 14 |
+
CVScreen — CV@1000-batch early band screen (<0.30 LOW / .35-.50 MID / >.80 HIGH)
|
| 15 |
+
|
| 16 |
+
Smoke on a torch-capable env: python geolip_vitals.py
|
| 17 |
+
"""
|
| 18 |
+
from __future__ import annotations
|
| 19 |
+
import math
|
| 20 |
+
import torch
|
| 21 |
+
|
| 22 |
+
BINDING = 0.29154 # radians; the binding/separation constant
|
| 23 |
+
CV_BAND = (0.13, 0.30) # CM CV band (discovery_catalog #4)
|
| 24 |
+
GATE_BAND = (0.012, 0.03) # live invariant candidate (acd_campaign)
|
| 25 |
+
KNUTH32 = 2654435761
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
# ----------------------------------------------------------------------------- drift
|
| 29 |
+
@torch.no_grad()
|
| 30 |
+
def anchor_drift(current: torch.Tensor, init: torch.Tensor, tol: float = 0.05) -> dict:
|
| 31 |
+
"""Geodesic drift (radians) of each row of `current` from its row in `init`,
|
| 32 |
+
both row-normalized. Returns mean/std/per-row drift and the fraction of rows
|
| 33 |
+
within +/-tol of BINDING (the GLFM '46%' readout)."""
|
| 34 |
+
a = torch.nn.functional.normalize(current.float(), dim=-1)
|
| 35 |
+
b = torch.nn.functional.normalize(init.float(), dim=-1)
|
| 36 |
+
cos = (a * b).sum(-1).clamp(-1.0, 1.0)
|
| 37 |
+
drift = torch.arccos(cos)
|
| 38 |
+
frac = ((drift - BINDING).abs() <= tol).float().mean()
|
| 39 |
+
return {"mean": drift.mean().item(), "std": drift.std().item(),
|
| 40 |
+
"per_row": drift, "binding_fraction": frac.item()}
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
# -------------------------------------------------------------------------------- cv
|
| 44 |
+
@torch.no_grad()
|
| 45 |
+
def _pentachoron_volumes(pts: torch.Tensor) -> torch.Tensor:
|
| 46 |
+
"""Batched Cayley-Menger 4-simplex volumes. pts: (B, 5, D) -> (B,) volumes.
|
| 47 |
+
One float64 det over all samples (vol^2 = -det(CM)/9216 for n=4). Built-in
|
| 48 |
+
for speed (the per-sample reference path is ~260x slower in a vitals loop);
|
| 49 |
+
geovocab2 remains the formula's reference implementation, parity-checked
|
| 50 |
+
via cv_reference_check()."""
|
| 51 |
+
B = pts.shape[0]
|
| 52 |
+
d2 = torch.cdist(pts.double(), pts.double()).pow(2) # (B,5,5)
|
| 53 |
+
cm = torch.ones(B, 6, 6, dtype=torch.float64, device=pts.device)
|
| 54 |
+
cm[:, 0, 0] = 0.0
|
| 55 |
+
cm[:, 1:, 1:] = d2
|
| 56 |
+
det = torch.linalg.det(cm)
|
| 57 |
+
return (-det / 9216.0).clamp_min(0.0).sqrt().float()
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
@torch.no_grad()
|
| 61 |
+
def pentachoron_cv(rows: torch.Tensor, n_samples: int = 200,
|
| 62 |
+
generator: torch.Generator | None = None) -> float:
|
| 63 |
+
"""CV (std/mean) of Cayley-Menger 4-simplex volumes over n_samples random
|
| 64 |
+
5-row subsets. Rows are row-normalized before measurement. Uses the built-in
|
| 65 |
+
batched CM (float64 det); validate against geovocab2 with
|
| 66 |
+
cv_reference_check() after any change to the volume math."""
|
| 67 |
+
x = torch.nn.functional.normalize(rows.float(), dim=-1)
|
| 68 |
+
n = x.shape[0]
|
| 69 |
+
if n < 5:
|
| 70 |
+
raise ValueError(f"pentachoron_cv needs >=5 rows, got {n}")
|
| 71 |
+
g = generator or torch.Generator(device="cpu").manual_seed(0)
|
| 72 |
+
idx = torch.stack([torch.randperm(n, generator=g)[:5]
|
| 73 |
+
for _ in range(n_samples)]) # (B,5)
|
| 74 |
+
v = _pentachoron_volumes(x[idx].cpu())
|
| 75 |
+
return (v.std() / v.mean().clamp_min(1e-12)).item()
|
| 76 |
+
|
| 77 |
+
|
| 78 |
+
@torch.no_grad()
|
| 79 |
+
def cv_reference_check(n_trials: int = 50, tol: float = 1e-5) -> float:
|
| 80 |
+
"""Parity check of the built-in batched CM against geovocab2's reference
|
| 81 |
+
implementation (the formula's source of truth). Returns max |rel diff|;
|
| 82 |
+
raises if geovocab2 is absent or parity fails. Run after touching
|
| 83 |
+
_pentachoron_volumes."""
|
| 84 |
+
try:
|
| 85 |
+
from geovocab2.shapes.formula.symbolic.cayley_menger import (
|
| 86 |
+
CayleyMengerFromSimplex)
|
| 87 |
+
except Exception as e: # pragma: no cover
|
| 88 |
+
raise ImportError(
|
| 89 |
+
"cv_reference_check requires geovocab2 (install via the geolip-svae "
|
| 90 |
+
"umbrella: pip install git+https://github.com/AbstractEyes/"
|
| 91 |
+
"geolip-svae).") from e
|
| 92 |
+
ref = CayleyMengerFromSimplex()
|
| 93 |
+
g = torch.Generator().manual_seed(0)
|
| 94 |
+
pts = torch.nn.functional.normalize(
|
| 95 |
+
torch.randn(n_trials, 5, 4, generator=g), dim=-1)
|
| 96 |
+
mine = _pentachoron_volumes(pts)
|
| 97 |
+
# compare at float64: the reference computes in the INPUT dtype, and fp32
|
| 98 |
+
# dets lose up to ~4% on near-degenerate pentachora (measured 2026-07-11)
|
| 99 |
+
theirs = torch.stack([ref.forward(p.double())["volume"].float() for p in pts])
|
| 100 |
+
rel = ((mine - theirs).abs() / theirs.abs().clamp_min(1e-12)).max().item()
|
| 101 |
+
if rel > tol:
|
| 102 |
+
raise AssertionError(f"CM parity vs geovocab2 failed: max rel {rel}")
|
| 103 |
+
return rel
|
| 104 |
+
|
| 105 |
+
|
| 106 |
+
# ------------------------------------------------------------------------- aliveness
|
| 107 |
+
@torch.no_grad()
|
| 108 |
+
def axis_aliveness(oriented_weights: torch.Tensor, alive_thresh: float = 1e-3) -> dict:
|
| 109 |
+
"""`oriented_weights`: (..., 2K) nonnegative oriented-softmax address rows
|
| 110 |
+
(sum to 1 on the last dim). Returns axes-alive count, mean-usage perplexity
|
| 111 |
+
(hppl analogue; healthy hosted reference 125-126/128), and a collapse flag.
|
| 112 |
+
Reference behavior: near-uniform aliveness at div_weight=0 (discovery #22)."""
|
| 113 |
+
w = oriented_weights.reshape(-1, oriented_weights.shape[-1]).float()
|
| 114 |
+
usage = w.mean(0)
|
| 115 |
+
usage = usage / usage.sum().clamp_min(1e-12)
|
| 116 |
+
# an axis is alive if its mean usage exceeds alive_thresh x the uniform share
|
| 117 |
+
alive = int((usage > alive_thresh * (1.0 / usage.numel())).sum())
|
| 118 |
+
ent = -(usage.clamp_min(1e-12) * usage.clamp_min(1e-12).log()).sum()
|
| 119 |
+
ppl = float(ent.exp())
|
| 120 |
+
return {"axes_total": usage.numel(), "axes_alive": alive, "usage_ppl": ppl,
|
| 121 |
+
"collapsed": ppl < 0.05 * usage.numel()}
|
| 122 |
+
|
| 123 |
+
|
| 124 |
+
# ------------------------------------------------------------------------------ gates
|
| 125 |
+
@torch.no_grad()
|
| 126 |
+
def gate_stats(gates: torch.Tensor) -> dict:
|
| 127 |
+
"""Gate values (post-sigmoid/clamp). Reports mean and whether it sits in the
|
| 128 |
+
0.012-0.03 band (read-only — the band is a candidate invariant, never a target)."""
|
| 129 |
+
g = gates.float().flatten()
|
| 130 |
+
m = g.mean().item()
|
| 131 |
+
return {"mean": m, "std": g.std().item(),
|
| 132 |
+
"in_band": GATE_BAND[0] <= m <= GATE_BAND[1]}
|
| 133 |
+
|
| 134 |
+
|
| 135 |
+
# ------------------------------------------------------------------------------ paths
|
| 136 |
+
@torch.no_grad()
|
| 137 |
+
def path_diversity(ids: torch.Tensor) -> dict:
|
| 138 |
+
"""Unique-path counting with the FIXED multiplicative hash:
|
| 139 |
+
((ids * 2654435761) % 2^32) >> 16 — Knuth needs the HIGH bits; the low-16
|
| 140 |
+
variant produced a retracted ~1,500 path ceiling in a prior campaign.
|
| 141 |
+
`ids`: integer tensor, one composed path id per row (any shape)."""
|
| 142 |
+
x = ids.reshape(-1).to(torch.int64)
|
| 143 |
+
hashed = ((x * KNUTH32) % (1 << 32)) >> 16
|
| 144 |
+
return {"n": int(x.numel()),
|
| 145 |
+
"unique_raw": int(torch.unique(x).numel()),
|
| 146 |
+
"unique_hashed": int(torch.unique(hashed).numel())}
|
| 147 |
+
|
| 148 |
+
|
| 149 |
+
@torch.no_grad()
|
| 150 |
+
def compose_path_ids(stage_indices: list[torch.Tensor], radix: int) -> torch.Tensor:
|
| 151 |
+
"""Compose per-stage discrete indices (each (...,) int in [0, radix)) into a
|
| 152 |
+
single path id, positional base-`radix` — construction, not hashing."""
|
| 153 |
+
out = torch.zeros_like(stage_indices[0], dtype=torch.int64)
|
| 154 |
+
for s in stage_indices:
|
| 155 |
+
out = out * radix + s.to(torch.int64)
|
| 156 |
+
return out
|
| 157 |
+
|
| 158 |
+
|
| 159 |
+
# --------------------------------------------------------------------- grad democracy
|
| 160 |
+
@torch.no_grad()
|
| 161 |
+
def grad_norm_spread(groups: dict[str, list[torch.nn.Parameter]]) -> dict:
|
| 162 |
+
"""Gradient-democracy monitor. `groups`: name -> params of one parallel member
|
| 163 |
+
(tower/expert). Reports per-group grad norms and the orders-of-magnitude spread.
|
| 164 |
+
Reference: unequalized heterogeneous towers spread ~20 orders (fibonacci dead at
|
| 165 |
+
2.25e-21 under helix); equalized ~0.0 (geofractal gradient-democracy result)."""
|
| 166 |
+
norms = {}
|
| 167 |
+
for name, params in groups.items():
|
| 168 |
+
gs = [p.grad for p in params if p.grad is not None]
|
| 169 |
+
norms[name] = float(torch.sqrt(sum((g.float() ** 2).sum() for g in gs)).item()) \
|
| 170 |
+
if gs else 0.0
|
| 171 |
+
vals = [v for v in norms.values() if v > 0]
|
| 172 |
+
spread = (math.log10(max(vals)) - math.log10(min(vals))) if len(vals) >= 2 else 0.0
|
| 173 |
+
return {"norms": norms, "spread_orders": spread, "dead": [k for k, v in norms.items() if v == 0.0]}
|
| 174 |
+
|
| 175 |
+
|
| 176 |
+
# ----------------------------------------------------------------------------- screen
|
| 177 |
+
class CVScreen:
|
| 178 |
+
"""CV@N early band screen (tri-band ft1): record pentachoron CV at `step_mark`
|
| 179 |
+
batches; classify <0.30 LOW / 0.35-0.50 MID / >0.80 HIGH. Turns ~2h/config
|
| 180 |
+
into ~7min. Readout only."""
|
| 181 |
+
def __init__(self, step_mark: int = 1000):
|
| 182 |
+
self.step_mark = step_mark
|
| 183 |
+
self.recorded: float | None = None
|
| 184 |
+
|
| 185 |
+
def maybe_record(self, step: int, rows: torch.Tensor) -> float | None:
|
| 186 |
+
if self.recorded is None and step >= self.step_mark:
|
| 187 |
+
self.recorded = pentachoron_cv(rows)
|
| 188 |
+
return self.recorded
|
| 189 |
+
|
| 190 |
+
@property
|
| 191 |
+
def band(self) -> str | None:
|
| 192 |
+
c = self.recorded
|
| 193 |
+
if c is None:
|
| 194 |
+
return None
|
| 195 |
+
if c < 0.30:
|
| 196 |
+
return "LOW"
|
| 197 |
+
if 0.35 <= c <= 0.50:
|
| 198 |
+
return "MID"
|
| 199 |
+
if c > 0.80:
|
| 200 |
+
return "HIGH"
|
| 201 |
+
return "BETWEEN"
|
| 202 |
+
|
| 203 |
+
|
| 204 |
+
# ------------------------------------------------------------------------------ smoke
|
| 205 |
+
if __name__ == "__main__": # shapes/parse smoke ONLY — no training, ever.
|
| 206 |
+
g = torch.Generator().manual_seed(0)
|
| 207 |
+
K, D = 64, 4
|
| 208 |
+
init = torch.nn.functional.normalize(torch.randn(K, D, generator=g), dim=-1)
|
| 209 |
+
cur = torch.nn.functional.normalize(init + 0.29 * torch.randn(K, D, generator=g), dim=-1)
|
| 210 |
+
print("drift:", {k: v for k, v in anchor_drift(cur, init).items() if k != "per_row"})
|
| 211 |
+
w = torch.softmax(torch.randn(32, 2 * K, generator=g), dim=-1)
|
| 212 |
+
print("aliveness:", axis_aliveness(w))
|
| 213 |
+
print("gates:", gate_stats(torch.full((8,), 0.024)))
|
| 214 |
+
ids = compose_path_ids([torch.randint(0, 16, (4096,), generator=g) for _ in range(4)], 16)
|
| 215 |
+
print("paths:", path_diversity(ids))
|
| 216 |
+
lin = torch.nn.Linear(8, 8)
|
| 217 |
+
lin(torch.randn(4, 8)).sum().backward()
|
| 218 |
+
print("democracy:", grad_norm_spread({"a": list(lin.parameters())}))
|
| 219 |
+
print("OK — vitals smoke passed (pentachoron_cv needs geovocab2; run on GPU env)")
|
exp018_r12/read_codebook.py
ADDED
|
@@ -0,0 +1,159 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""read_codebook.py — projective reading of cultivated aleph codebooks.
|
| 2 |
+
Antipodal-collapse extraction on trained codebooks + projective statistics on
|
| 3 |
+
RP^(D-1) — applied to the exp012 AR-bed specimens.
|
| 4 |
+
|
| 5 |
+
Recipe per the Polygonal Omega article (geometric-tri-band-ft2): collapse = (row_i - row_j)/2 normalized for each MUTUAL-STRONGEST
|
| 6 |
+
pair with cos < -0.9 — "a deterministic tensor operation," not clustering.
|
| 7 |
+
Projective metric ALWAYS arccos|<a,b>| (metric-alignment rule, reading-voids-ft1).
|
| 8 |
+
D=4 scope is the validated regime (D=5 walked back; axis count grows with D).
|
| 9 |
+
|
| 10 |
+
Readouts per specimen:
|
| 11 |
+
pairs / n_axes / unpaired — antipodal structure
|
| 12 |
+
proj_angle mean vs uniform baseline, deviation — near-uniform RP^(D-1)?
|
| 13 |
+
drift from home + binding fraction @0.29154 — cultivation record
|
| 14 |
+
erank of the axis set — spectral occupancy
|
| 15 |
+
verdict: PROJECTIVE-CLEAN (|dev|<0.05, util>0.95, secondary pairs<=3) /
|
| 16 |
+
-MOSTLY / STRUCTURED / DEGENERATE (per Polygonal Omega thresholds)
|
| 17 |
+
|
| 18 |
+
Usage (terminal): python read_codebook.py <ckpt_or_dir> [more paths...]
|
| 19 |
+
Colab: paste geolip_vitals.py cell first (optional), then this file, then
|
| 20 |
+
read_all(r"/content/data/ar_ckpts").
|
| 21 |
+
"""
|
| 22 |
+
from __future__ import annotations
|
| 23 |
+
import math
|
| 24 |
+
import sys
|
| 25 |
+
import torch
|
| 26 |
+
import torch.nn.functional as F
|
| 27 |
+
|
| 28 |
+
BINDING = 0.29154
|
| 29 |
+
|
| 30 |
+
|
| 31 |
+
@torch.no_grad()
|
| 32 |
+
def antipodal_collapse(codebook: torch.Tensor, thresh: float = -0.9) -> dict:
|
| 33 |
+
"""Mutual-strongest antipodal pairing + collapse to axes on RP^(D-1)."""
|
| 34 |
+
A = F.normalize(codebook.float(), dim=-1)
|
| 35 |
+
K = A.shape[0]
|
| 36 |
+
cos = A @ A.T
|
| 37 |
+
cos.fill_diagonal_(2.0) # exclude self from minima
|
| 38 |
+
nearest_neg = cos.argmin(dim=-1) # most-antipodal partner
|
| 39 |
+
pairs = []
|
| 40 |
+
used = set()
|
| 41 |
+
for i in range(K):
|
| 42 |
+
j = int(nearest_neg[i])
|
| 43 |
+
if i < j and int(nearest_neg[j]) == i and cos[i, j] < thresh:
|
| 44 |
+
pairs.append((i, j))
|
| 45 |
+
used.update((i, j))
|
| 46 |
+
axes = [F.normalize((A[i] - A[j]) / 2.0, dim=-1) for i, j in pairs]
|
| 47 |
+
axes += [A[i] for i in range(K) if i not in used] # unpaired rows as axes
|
| 48 |
+
axes = torch.stack(axes) if axes else A[:0]
|
| 49 |
+
# sign-canon onto RP: first nonzero coordinate positive
|
| 50 |
+
for r in range(axes.shape[0]):
|
| 51 |
+
nz = torch.nonzero(axes[r].abs() > 1e-8)
|
| 52 |
+
if nz.numel() and axes[r, nz[0, 0]] < 0:
|
| 53 |
+
axes[r] = -axes[r]
|
| 54 |
+
return {"pairs": len(pairs), "n_axes": axes.shape[0],
|
| 55 |
+
"unpaired": K - 2 * len(pairs), "axes": axes}
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
@torch.no_grad()
|
| 59 |
+
def projective_stats(axes: torch.Tensor, n_baseline: int = 20000,
|
| 60 |
+
seed: int = 0) -> dict:
|
| 61 |
+
"""Mean projective angle arccos|<a,b>| vs a uniform-RP baseline at same (n, D)."""
|
| 62 |
+
n, D = axes.shape
|
| 63 |
+
if n < 2:
|
| 64 |
+
return {"proj_angle_mean": None, "uniform_baseline": None,
|
| 65 |
+
"deviation": None, "erank": None}
|
| 66 |
+
def mean_angle(rows):
|
| 67 |
+
c = (rows @ rows.T).abs().clamp(max=1.0)
|
| 68 |
+
iu = torch.triu_indices(rows.shape[0], rows.shape[0], offset=1)
|
| 69 |
+
return torch.arccos(c[iu[0], iu[1]]).mean().item()
|
| 70 |
+
obs = mean_angle(axes)
|
| 71 |
+
g = torch.Generator().manual_seed(seed)
|
| 72 |
+
base_angles = []
|
| 73 |
+
m = max(2, n)
|
| 74 |
+
for _ in range(max(1, n_baseline // max(1, m * (m - 1) // 2))):
|
| 75 |
+
r = F.normalize(torch.randn(m, D, generator=g), dim=-1)
|
| 76 |
+
base_angles.append(mean_angle(r))
|
| 77 |
+
base = sum(base_angles) / len(base_angles)
|
| 78 |
+
s = torch.linalg.svdvals(axes)
|
| 79 |
+
p = (s / s.sum().clamp_min(1e-12))
|
| 80 |
+
erank = float(torch.exp(-(p.clamp_min(1e-12) * p.clamp_min(1e-12).log()).sum()))
|
| 81 |
+
return {"proj_angle_mean": round(obs, 4), "uniform_baseline": round(base, 4),
|
| 82 |
+
"deviation": round(obs - base, 4), "erank": round(erank, 3)}
|
| 83 |
+
|
| 84 |
+
|
| 85 |
+
@torch.no_grad()
|
| 86 |
+
def read_specimen(path: str) -> dict:
|
| 87 |
+
ck = torch.load(path, map_location="cpu", weights_only=True)
|
| 88 |
+
out = {"file": path.split("\\")[-1].split("/")[-1],
|
| 89 |
+
"arm": ck.get("arm"), "seed": ck.get("seed"),
|
| 90 |
+
"steps": ck.get("steps"), "val_bpb": round(ck.get("val_bpb", -1), 4)}
|
| 91 |
+
if "state_dict" in ck: # full specimen checkpoint
|
| 92 |
+
sd = ck["state_dict"]
|
| 93 |
+
books = {k[:-len(".codebook")]: sd[k] for k in sd
|
| 94 |
+
if k.endswith("addr.codebook") or k.endswith("head_addr.codebook")}
|
| 95 |
+
homes = {k[:-len(".home")]: sd[k] for k in sd if k.endswith(".home")}
|
| 96 |
+
else: # bare genome dict (exp014+ champion files):
|
| 97 |
+
# books under flat/root/branch* keys; *_proj entries are projections
|
| 98 |
+
books = {k: v for k, v in ck.items()
|
| 99 |
+
if torch.is_tensor(v) and v.ndim == 2
|
| 100 |
+
and (k in ("flat", "root") or k.startswith("branch"))}
|
| 101 |
+
homes = {}
|
| 102 |
+
reads = {}
|
| 103 |
+
for name, cb in books.items():
|
| 104 |
+
col = antipodal_collapse(cb)
|
| 105 |
+
stats = projective_stats(col["axes"])
|
| 106 |
+
home = homes.get(name)
|
| 107 |
+
drift = None
|
| 108 |
+
binding = None
|
| 109 |
+
if home is not None and home.shape == cb.shape:
|
| 110 |
+
a = F.normalize(cb.float(), dim=-1)
|
| 111 |
+
b = F.normalize(home.float(), dim=-1)
|
| 112 |
+
dr = torch.arccos((a * b).sum(-1).clamp(-1, 1))
|
| 113 |
+
drift = round(dr.mean().item(), 4)
|
| 114 |
+
binding = round(((dr - BINDING).abs() <= 0.05).float().mean().item(), 4)
|
| 115 |
+
util = col["n_axes"] / cb.shape[0]
|
| 116 |
+
dev = stats["deviation"]
|
| 117 |
+
if dev is not None and abs(dev) < 0.05 and util > 0.95 and col["pairs"] <= 3:
|
| 118 |
+
verdict = "PROJECTIVE-CLEAN"
|
| 119 |
+
elif dev is not None and abs(dev) < 0.05:
|
| 120 |
+
verdict = "PROJECTIVE-MOSTLY"
|
| 121 |
+
elif dev is not None and dev > 0.05:
|
| 122 |
+
verdict = "STRUCTURED(repulsive)"
|
| 123 |
+
else:
|
| 124 |
+
verdict = "DEGENERATE/CLUMPED" if dev is not None else "TOO-FEW-AXES"
|
| 125 |
+
reads[name] = {
|
| 126 |
+
"pairs": col["pairs"], "n_axes": col["n_axes"], **stats,
|
| 127 |
+
"drift": drift, "binding_frac": binding, "verdict": verdict}
|
| 128 |
+
out["codebooks"] = reads
|
| 129 |
+
return out
|
| 130 |
+
|
| 131 |
+
|
| 132 |
+
def read_all(root: str) -> list:
|
| 133 |
+
import glob, os
|
| 134 |
+
results = []
|
| 135 |
+
for p in sorted(glob.glob(os.path.join(root, "*.pt"))):
|
| 136 |
+
r = read_specimen(p)
|
| 137 |
+
print(r, flush=True)
|
| 138 |
+
results.append(r)
|
| 139 |
+
return results
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
def _in_notebook() -> bool:
|
| 143 |
+
try:
|
| 144 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 145 |
+
return True
|
| 146 |
+
except NameError:
|
| 147 |
+
return False
|
| 148 |
+
|
| 149 |
+
|
| 150 |
+
if __name__ == "__main__":
|
| 151 |
+
if _in_notebook():
|
| 152 |
+
print("Notebook mode: call read_all(r'<data_root>/ar_ckpts') in the next cell.")
|
| 153 |
+
else:
|
| 154 |
+
args = [a for a in sys.argv[1:] if not a.startswith("-")]
|
| 155 |
+
if not args:
|
| 156 |
+
print("usage: python read_codebook.py <ckpt_or_dir> [...]")
|
| 157 |
+
for a in args:
|
| 158 |
+
import os
|
| 159 |
+
read_all(a) if os.path.isdir(a) else print(read_specimen(a))
|
exp018_r12/repro.py
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""repro.py — standalone loader/runner for exp018_r12. Code dependencies live
|
| 2 |
+
in THIS folder (geolip_vitals.py, ar_differentiation_bed.py,
|
| 3 |
+
exp014_genetic_distillation.py, exp018_reran12.py, read_codebook.py).
|
| 4 |
+
|
| 5 |
+
python repro.py # CPU smoke: pentachoron regularity + arm shapes
|
| 6 |
+
python repro.py --run # all 4 arms x 2 seeds (GPU, ~40 min)
|
| 7 |
+
|
| 8 |
+
Data lands in ./data (override with GEOLIP_DATA). The farmed_init arm needs
|
| 9 |
+
the exp012 donor specimen: ../exp012_ar/specimens/addr_msl64_s0_t2000.pt
|
| 10 |
+
(present in this repo) or set GEOLIP_FARMED_SPECIMEN to a path.
|
| 11 |
+
"""
|
| 12 |
+
import os
|
| 13 |
+
import sys
|
| 14 |
+
|
| 15 |
+
HERE = os.path.dirname(os.path.abspath(__file__))
|
| 16 |
+
sys.path.insert(0, HERE)
|
| 17 |
+
|
| 18 |
+
if __name__ == "__main__":
|
| 19 |
+
import geolip_vitals # noqa: F401 (paste order)
|
| 20 |
+
import ar_differentiation_bed # noqa: F401
|
| 21 |
+
import exp014_genetic_distillation # noqa: F401
|
| 22 |
+
import exp018_reran12 as r12
|
| 23 |
+
if "--run" in sys.argv[1:]:
|
| 24 |
+
r12.run_rerun()
|
| 25 |
+
else:
|
| 26 |
+
r12.smoke()
|
| 27 |
+
print("repro smoke passed — run with --run for the arms (GPU)")
|
exp018_r12/results/ledger.jsonl
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"exp": "18", "arm": "penta_init", "seed": 0, "steps": 2000, "bpb": 2.4957, "drift": 0.2417, "binding_frac": 0.25}
|
| 2 |
+
{"exp": "18", "arm": "farmed_init", "seed": 0, "steps": 2000, "bpb": 2.4934, "drift": 0.1787, "binding_frac": 0.2031}
|
| 3 |
+
{"exp": "18", "arm": "tied", "seed": 0, "steps": 2000, "bpb": 3.5131, "drift": 0.0188, "binding_frac": 0.0}
|
| 4 |
+
{"exp": "18", "arm": "keystone", "seed": 0, "steps": 2000, "bpb": 3.5169, "drift": 0.0203, "binding_frac": 0.0}
|
| 5 |
+
{"exp": "18", "arm": "penta_init", "seed": 1, "steps": 2000, "bpb": 2.4389, "drift": 0.2114, "binding_frac": 0.2969}
|
| 6 |
+
{"exp": "18", "arm": "farmed_init", "seed": 1, "steps": 2000, "bpb": 2.4614, "drift": 0.1986, "binding_frac": 0.2812}
|
| 7 |
+
{"exp": "18", "arm": "tied", "seed": 1, "steps": 2000, "bpb": 3.4997, "drift": 0.0203, "binding_frac": 0.0}
|
| 8 |
+
{"exp": "18", "arm": "keystone", "seed": 1, "steps": 2000, "bpb": 3.5006, "drift": 0.0208, "binding_frac": 0.0}
|
exp018_r12/results/results.json
ADDED
|
@@ -0,0 +1,34 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"farmed_init_s0": {
|
| 3 |
+
"bpb": 2.4934,
|
| 4 |
+
"drift": 0.1787
|
| 5 |
+
},
|
| 6 |
+
"farmed_init_s1": {
|
| 7 |
+
"bpb": 2.4614,
|
| 8 |
+
"drift": 0.1986
|
| 9 |
+
},
|
| 10 |
+
"keystone_s0": {
|
| 11 |
+
"bpb": 3.5169,
|
| 12 |
+
"drift": 0.0203
|
| 13 |
+
},
|
| 14 |
+
"keystone_s1": {
|
| 15 |
+
"bpb": 3.5006,
|
| 16 |
+
"drift": 0.0208
|
| 17 |
+
},
|
| 18 |
+
"penta_init_s0": {
|
| 19 |
+
"bpb": 2.4957,
|
| 20 |
+
"drift": 0.2417
|
| 21 |
+
},
|
| 22 |
+
"penta_init_s1": {
|
| 23 |
+
"bpb": 2.4389,
|
| 24 |
+
"drift": 0.2114
|
| 25 |
+
},
|
| 26 |
+
"tied_s0": {
|
| 27 |
+
"bpb": 3.5131,
|
| 28 |
+
"drift": 0.0188
|
| 29 |
+
},
|
| 30 |
+
"tied_s1": {
|
| 31 |
+
"bpb": 3.4997,
|
| 32 |
+
"drift": 0.0203
|
| 33 |
+
}
|
| 34 |
+
}
|
exp018_r12/specimens/r12_farmed_init_s0.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:863c95c2ada021129ed01776dd0505df338f1c3ba0134f7a2ff036f7bd535890
|
| 3 |
+
size 7977645
|
exp018_r12/specimens/r12_farmed_init_s1.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0ef0870f9a5ec4067e72dcf1950e183dd65e526d0e5a21316865681853798a16
|
| 3 |
+
size 7977645
|
exp018_r12/specimens/r12_keystone_s0.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:5b4700cd47a40d15928ea8e3f5cc2fbe6a86aa42368e6d9cb7090c8d972a9477
|
| 3 |
+
size 7713726
|
exp018_r12/specimens/r12_keystone_s1.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:1183c1dc78f0c1a2f20b277922103fb7192134d09939265575ecade8b87f524a
|
| 3 |
+
size 7713726
|
exp018_r12/specimens/r12_penta_init_s0.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:6e73f0554426e45fffcabeea895eff4e7d7a47076ca7a57f552c1962d6d11df7
|
| 3 |
+
size 7977590
|
exp018_r12/specimens/r12_penta_init_s1.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:f8b615f8a3ce5358b1286b5184f20c91aa055eb7b0a5b417f0c3b81d78f554f5
|
| 3 |
+
size 7977590
|
exp018_r12/specimens/r12_tied_s0.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:f12ee10ac833b18a3821ca9273ebc575713e334baa2fad06e93d732bace80222
|
| 3 |
+
size 7713514
|
exp018_r12/specimens/r12_tied_s1.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c1d8da581864f4fb18458b86d51148357c91b5cb7af0f958aae72a190912ae6f
|
| 3 |
+
size 7713514
|
exp019_cr/README.md
ADDED
|
@@ -0,0 +1,82 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# exp019_cr — the capacity for distilled content retention (+ generalization)
|
| 2 |
+
|
| 3 |
+
When content is **distilled** rather than directly learned, how much transfers,
|
| 4 |
+
how much is retained, and for how long? Sequel to [exp016](../exp016_ct/) and
|
| 5 |
+
the [exp014](../exp014_gd/) memory-substrate studies; the weak-to-strong
|
| 6 |
+
distillation cell made real (the teacher holds content the student lacks —
|
| 7 |
+
true headroom on the content axis).
|
| 8 |
+
|
| 9 |
+
**Instrument**: N fact records `\n@<key6>=<value12>\n` (random alphanumeric —
|
| 10 |
+
uncompletable from language statistics, so exact-match greedy completion IS
|
| 11 |
+
retention), mixed into the wikitext byte stream at rate 0.5. Teacher =
|
| 12 |
+
addr_msl64 bed model, 4000 steps, gated on its own recall. Channels (fresh
|
| 13 |
+
student, 2000 steps): `direct` (ground truth; ceiling) | `kd_facts` (fact rows
|
| 14 |
+
supervised ONLY by teacher logits — pure distilled content, α=1.0) |
|
| 15 |
+
`kd_general` (KD on clean text only — leakage probe) | `book_implant`
|
| 16 |
+
(teacher's trained codebook into a fresh student) | `none` (floor). Axes:
|
| 17 |
+
N ∈ {64, 256, 1024}; retention re-measured after 1000 further clean-stream
|
| 18 |
+
steps (interference). The **19b block** swaps rote facts for **rule-bearing
|
| 19 |
+
content** (value = fixed substitution cipher of the key; 256 train keys, 128
|
| 20 |
+
held out) plus prompt-format variants.
|
| 21 |
+
|
| 22 |
+
`build_results.py` re-asserts every claim below from `results/ledger.jsonl`
|
| 23 |
+
(36 rows).
|
| 24 |
+
|
| 25 |
+
## Findings — retention sweep
|
| 26 |
+
|
| 27 |
+
1. **Capacity curve**: teacher recall at 4k steps = 1.00 (N=64), 0.99 (N=256),
|
| 28 |
+
0.01–0.20 (N=1024) — the ~1.9M model holds ~256 facts near-perfectly and
|
| 29 |
+
hits a cliff before 1024 (cliff edge seed-chaotic).
|
| 30 |
+
2. **The distillation tax is zero-to-negative.** Where the teacher knows the
|
| 31 |
+
content, teacher logits alone transfer it at parity with ground truth
|
| 32 |
+
(N=64: exact 1.0 = 1.0, both seeds) or better (N=256: 0.953 vs 0.871 s0;
|
| 33 |
+
0.859 vs 0.856 s1). Soft targets from a teacher with headroom are at least
|
| 34 |
+
as content-efficient as the data itself. The cost is general modeling:
|
| 35 |
+
kd students' clean bpb 4.0–4.3 vs direct 2.8–3.1 vs floor 2.47.
|
| 36 |
+
3. **No logit leakage**: KD on clean text transfers zero facts (0.000, 6/6
|
| 37 |
+
cells) — content does not cross without exposure at this scale.
|
| 38 |
+
4. **Anchors carry no bytes**: the trained codebook of a teacher that knows
|
| 39 |
+
256 facts, implanted into a fresh student, transfers **none** of them
|
| 40 |
+
(0.000, both seeds). The codebook organizes addressing; content lives in
|
| 41 |
+
the trunk.
|
| 42 |
+
5. **Universal catastrophic forgetting**: after 1000 clean-stream steps, every
|
| 43 |
+
channel's recall is exactly 0.000 — including perfectly-learned direct
|
| 44 |
+
students. At this scale nothing persists without pressure; **persistence,
|
| 45 |
+
not transfer, is the unsolved axis of the memory substrate.**
|
| 46 |
+
|
| 47 |
+
## Findings — 19b generalization block (rule content)
|
| 48 |
+
|
| 49 |
+
6. **Rules are learned fragmentarily**: teachers memorize the rule-pairs
|
| 50 |
+
near-perfectly (train exact 0.98–1.00) but reach only ~0.25–0.27 held-out
|
| 51 |
+
byte accuracy (~10× chance) with **zero** exact completions — character
|
| 52 |
+
mappings absorbed, the full cipher never.
|
| 53 |
+
7. **The logit channel beats direct learning on structured content, 2/2
|
| 54 |
+
seeds** (train exact 0.758/0.844 vs 0.652/0.773) and transfers the partial
|
| 55 |
+
rule at full fidelity — the dark-knowledge advantage, certified on rules.
|
| 56 |
+
8. **Content is format-locked** — variant-format prompts collapse recall to
|
| 57 |
+
~0 even on trained keys — but **KD students are consistently less
|
| 58 |
+
format-locked** than direct students (variant byte 0.077–0.091 vs
|
| 59 |
+
0.000–0.024): soft targets bind content less rigidly to surface form.
|
| 60 |
+
9. Interference erases the rule too — structure forgets like instances.
|
| 61 |
+
|
| 62 |
+
## Files
|
| 63 |
+
- `exp019_content_retention.py` — fact/rule corpora, streams, exact-match
|
| 64 |
+
recall (with format variants), channel training, both runners, smoke.
|
| 65 |
+
- `geolip_vitals.py` / `ar_differentiation_bed.py` /
|
| 66 |
+
`exp014_genetic_distillation.py` / `read_codebook.py` — this package's own
|
| 67 |
+
harness copies. Standalone.
|
| 68 |
+
- `repro.py`, `build_results.py`, `results/ledger.jsonl` (36 rows),
|
| 69 |
+
`specimens/` (6 fact teachers + 2 rule teachers + 6 rule students).
|
| 70 |
+
|
| 71 |
+
## Reproduce (from inside this folder)
|
| 72 |
+
```bash
|
| 73 |
+
pip install torch --index-url https://download.pytorch.org/whl/cu128
|
| 74 |
+
pip install pyarrow huggingface_hub
|
| 75 |
+
python repro.py # CPU smoke
|
| 76 |
+
python repro.py --run # retention sweep (GPU, ~3h)
|
| 77 |
+
python repro.py --run19b # rule/generalization block (GPU, ~1h)
|
| 78 |
+
python build_results.py # re-assert every claim from the ledger
|
| 79 |
+
```
|
| 80 |
+
Data lands in `./data` (override with `GEOLIP_DATA`).
|
| 81 |
+
|
| 82 |
+
License: MIT · AbstractPhil + Claude Fable 5 · July 11, 2026
|
exp019_cr/ar_differentiation_bed.py
ADDED
|
@@ -0,0 +1,487 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""ar_differentiation_bed.py — exp012: autoregressive differentiation of the aleph.
|
| 2 |
+
|
| 3 |
+
Differentiation is cultivated by PREDICTIVE pressure along the sequence — the
|
| 4 |
+
address parameterizing the next-byte distribution (Law 2: chain-rule advantage pays
|
| 5 |
+
ONLY where the composed address directly parameterizes the predictive distribution).
|
| 6 |
+
This bed puts the aleph in the autoregressive gradient path and measures what
|
| 7 |
+
differentiates. The head arms enforce the employment law at its maximum: the
|
| 8 |
+
ENTIRE next-byte distribution is parameterized by the address.
|
| 9 |
+
|
| 10 |
+
Byte-level causal LM on wikitext-2-raw (HF parquet, CDN-fast), block 256. ARMS:
|
| 11 |
+
sdpa — standard causal transformer control (matched trunk).
|
| 12 |
+
hub — attention replaced by CAUSAL HUB: linear attention whose feature map
|
| 13 |
+
is the 2K-oriented aleph address, prefix-sum memories (no selection
|
| 14 |
+
event; O(n*K*d)). Differentiation cultivated INSIDE attention.
|
| 15 |
+
addr_head — sdpa trunk, but the OUTPUT HEAD reads ONLY the signed aleph
|
| 16 |
+
coefficient vector w_k = sinh(u_k)/sum_j cosh(u_j) of the final
|
| 17 |
+
hidden state (K -> 256 logits). The address MUST carry every bit of
|
| 18 |
+
next-byte information — the hardest Law-2 bottleneck.
|
| 19 |
+
|
| 20 |
+
JUDGED BY: val bits-per-byte per arm (task) + CULTIVATION VITALS on every aleph
|
| 21 |
+
codebook (readouts, never losses): axis aliveness/hppl, drift-from-init +
|
| 22 |
+
binding fraction @0.29154, winner-|cos| saturation (sign-code emergence), shadow
|
| 23 |
+
path diversity (fixed high-bits hash). Never by recon.
|
| 24 |
+
|
| 25 |
+
Riders: pure Adam wd=0; no BN/Dropout/GAP on geometric paths; orthogonal init;
|
| 26 |
+
Colab-cell-safe (paste-ahead imports, no bare argparse, no __file__ reliance);
|
| 27 |
+
GPU-only for verdict runs.
|
| 28 |
+
|
| 29 |
+
Terminal: python ar_differentiation_bed.py # shapes/parse smoke
|
| 30 |
+
python ar_differentiation_bed.py --train # verdict run
|
| 31 |
+
Colab: paste geolip_vitals.py cell, then this file (smoke auto-runs),
|
| 32 |
+
then train(steps=2000, data_root="/content/data") in the next cell.
|
| 33 |
+
"""
|
| 34 |
+
from __future__ import annotations
|
| 35 |
+
import math
|
| 36 |
+
import torch
|
| 37 |
+
import torch.nn as nn
|
| 38 |
+
import torch.nn.functional as F
|
| 39 |
+
|
| 40 |
+
if "anchor_drift" not in globals():
|
| 41 |
+
try:
|
| 42 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 43 |
+
except ImportError:
|
| 44 |
+
_here = globals().get("__file__")
|
| 45 |
+
if _here is not None:
|
| 46 |
+
import sys, pathlib
|
| 47 |
+
sys.path.insert(0, str(pathlib.Path(_here).parent))
|
| 48 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 49 |
+
else:
|
| 50 |
+
raise ImportError(
|
| 51 |
+
"geolip_vitals not found — paste/run its cell first, or "
|
| 52 |
+
"hf_hub_download exp012_ar/geolip_vitals.py from "
|
| 53 |
+
"AbstractPhil/geolip-aleph-differentiation.")
|
| 54 |
+
|
| 55 |
+
VOCAB = 256 # bytes
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
# ------------------------------------------------------------------ aleph address
|
| 59 |
+
def _super_fibonacci_s3(n: int) -> torch.Tensor:
|
| 60 |
+
"""Near-uniform unit quaternions (Alexa CVPR'22) —
|
| 61 |
+
starts the codebook INSIDE the RP^3 attractor basin. D=4 only."""
|
| 62 |
+
PHI, PSI = math.sqrt(2.0), 1.533751168755204288118041
|
| 63 |
+
i = torch.arange(n, dtype=torch.float64)
|
| 64 |
+
s = (i + 0.5) / n
|
| 65 |
+
r, R = torch.sqrt(s), torch.sqrt(1.0 - s)
|
| 66 |
+
a, b = 2 * math.pi * i / PHI, 2 * math.pi * i / PSI
|
| 67 |
+
q = torch.stack([r * torch.sin(a), r * torch.cos(a),
|
| 68 |
+
R * torch.sin(b), R * torch.cos(b)], dim=-1)
|
| 69 |
+
return F.normalize(q, dim=-1).float()
|
| 70 |
+
|
| 71 |
+
|
| 72 |
+
class AlephAddress(nn.Module):
|
| 73 |
+
"""Closed-form aleph over 2K oriented half-axes (aleph-void article).
|
| 74 |
+
signed(x): (..., K) w_k = sinh(u_k)/sum_j cosh(u_j) — the Law-2 head feature.
|
| 75 |
+
oriented(x): ((..., K), (..., K)) positive halves of the 2K softmax — HUB map."""
|
| 76 |
+
|
| 77 |
+
def __init__(self, K: int, D: int, tau: float = 0.1, init: str = "random"):
|
| 78 |
+
super().__init__()
|
| 79 |
+
self.K, self.D, self.tau = K, D, tau
|
| 80 |
+
if init == "fibonacci":
|
| 81 |
+
assert D == 4, "fibonacci init lives on S^3 (D=4)"
|
| 82 |
+
A = _super_fibonacci_s3(K)
|
| 83 |
+
else:
|
| 84 |
+
A = F.normalize(torch.randn(K, D), dim=-1)
|
| 85 |
+
self.codebook = nn.Parameter(A)
|
| 86 |
+
self.register_buffer("home", self.codebook.detach().clone())
|
| 87 |
+
|
| 88 |
+
def _u(self, x):
|
| 89 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 90 |
+
return (F.normalize(x, dim=-1) @ A.transpose(-1, -2)) / self.tau
|
| 91 |
+
|
| 92 |
+
def oriented(self, x):
|
| 93 |
+
u = self._u(x)
|
| 94 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 95 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 96 |
+
Z = (ep + en).sum(dim=-1, keepdim=True)
|
| 97 |
+
return ep / Z, en / Z
|
| 98 |
+
|
| 99 |
+
def signed(self, x):
|
| 100 |
+
u = self._u(x)
|
| 101 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 102 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 103 |
+
return (ep - en) / (ep + en).sum(dim=-1, keepdim=True)
|
| 104 |
+
|
| 105 |
+
def signed_at(self, x, taus):
|
| 106 |
+
"""Multi-tau stroboscope (rule of 3): signed coefficients at several
|
| 107 |
+
temperatures, concatenated — softer taus keep the vector dense while a
|
| 108 |
+
hard tau supplies the sign-code sharpness. v2 refinement (b)."""
|
| 109 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 110 |
+
cos = F.normalize(x, dim=-1) @ A.transpose(-1, -2)
|
| 111 |
+
outs = []
|
| 112 |
+
for t in taus:
|
| 113 |
+
u = cos / t
|
| 114 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 115 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 116 |
+
outs.append((ep - en) / (ep + en).sum(dim=-1, keepdim=True))
|
| 117 |
+
return torch.cat(outs, dim=-1)
|
| 118 |
+
|
| 119 |
+
def m_hat(self, x):
|
| 120 |
+
"""Closed-form soft read (decoders read M_hat, never M). v2 control (c)."""
|
| 121 |
+
u = self._u(x)
|
| 122 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 123 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 124 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 125 |
+
return ((ep - en) @ A) / (ep + en).sum(dim=-1, keepdim=True)
|
| 126 |
+
|
| 127 |
+
def m_hard_ste(self, x):
|
| 128 |
+
"""Hard mode (aleph-void article): M_hard = sign(cos_win) * A[win], straight-through to
|
| 129 |
+
the soft read — forward fully discrete SIGN CODE, backward soft gradient.
|
| 130 |
+
Legal per theme A (reconstructive sign code, not a one-hot roster pick)."""
|
| 131 |
+
u = self._u(x)
|
| 132 |
+
soft = self.m_hat(x)
|
| 133 |
+
win = u.abs().argmax(dim=-1)
|
| 134 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 135 |
+
sign = torch.sign(torch.gather(u, -1, win.unsqueeze(-1))).squeeze(-1)
|
| 136 |
+
hard = sign.unsqueeze(-1) * A[win]
|
| 137 |
+
return hard + soft - soft.detach()
|
| 138 |
+
|
| 139 |
+
@torch.no_grad()
|
| 140 |
+
def vitals(self, x_sample) -> dict:
|
| 141 |
+
u = self._u(x_sample.reshape(-1, x_sample.shape[-1]))
|
| 142 |
+
p, n = self.oriented(x_sample.reshape(-1, x_sample.shape[-1]))
|
| 143 |
+
two_k = torch.cat([p, n], dim=-1)
|
| 144 |
+
win = two_k.argmax(dim=-1)
|
| 145 |
+
cos_win = (u.abs().amax(dim=-1) * self.tau) # winner |cos| — sign-code sat.
|
| 146 |
+
d = anchor_drift(self.codebook, self.home)
|
| 147 |
+
return {"drift": round(d["mean"], 4),
|
| 148 |
+
"binding_frac": round(d["binding_fraction"], 4),
|
| 149 |
+
"aliveness": axis_aliveness(two_k),
|
| 150 |
+
"win_cos_mean": round(cos_win.mean().item(), 4),
|
| 151 |
+
"paths": path_diversity(win)}
|
| 152 |
+
|
| 153 |
+
|
| 154 |
+
# ------------------------------------------------------------------------- blocks
|
| 155 |
+
class CausalSDPA(nn.Module):
|
| 156 |
+
def __init__(self, d: int, heads: int = 4):
|
| 157 |
+
super().__init__()
|
| 158 |
+
self.h = heads
|
| 159 |
+
self.qkv = nn.Linear(d, 3 * d, bias=False)
|
| 160 |
+
self.o = nn.Linear(d, d, bias=False)
|
| 161 |
+
nn.init.orthogonal_(self.qkv.weight); nn.init.orthogonal_(self.o.weight)
|
| 162 |
+
|
| 163 |
+
def forward(self, x):
|
| 164 |
+
B, n, d = x.shape
|
| 165 |
+
q, k, v = self.qkv(x).chunk(3, dim=-1)
|
| 166 |
+
q, k, v = (t.view(B, n, self.h, d // self.h).transpose(1, 2) for t in (q, k, v))
|
| 167 |
+
y = F.scaled_dot_product_attention(q, k, v, is_causal=True)
|
| 168 |
+
return self.o(y.transpose(1, 2).reshape(B, n, d))
|
| 169 |
+
|
| 170 |
+
|
| 171 |
+
class CausalHUB(nn.Module):
|
| 172 |
+
"""Causal aleph linear attention: prefix-sum memories over the two K-wide
|
| 173 |
+
halves of the oriented address; 2K never materialized; no selection event."""
|
| 174 |
+
|
| 175 |
+
def __init__(self, d: int, K: int = 32, D: int = 4, tau: float = 0.1):
|
| 176 |
+
super().__init__()
|
| 177 |
+
self.addr = AlephAddress(K, D, tau)
|
| 178 |
+
self.q = nn.Linear(d, D, bias=False)
|
| 179 |
+
self.k = nn.Linear(d, D, bias=False)
|
| 180 |
+
self.v = nn.Linear(d, d, bias=False)
|
| 181 |
+
self.o = nn.Linear(d, d, bias=False)
|
| 182 |
+
for m in (self.q, self.k, self.v, self.o):
|
| 183 |
+
nn.init.orthogonal_(m.weight)
|
| 184 |
+
|
| 185 |
+
def forward(self, x):
|
| 186 |
+
qp, qn = self.addr.oriented(self.q(x)) # (B, n, K)
|
| 187 |
+
kp, kn = self.addr.oriented(self.k(x))
|
| 188 |
+
v = self.v(x) # (B, n, d)
|
| 189 |
+
Sp = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kp, v), dim=1)
|
| 190 |
+
Sn = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kn, v), dim=1)
|
| 191 |
+
zp = torch.cumsum(kp, dim=1)
|
| 192 |
+
zn = torch.cumsum(kn, dim=1)
|
| 193 |
+
num = torch.einsum("bnk,bnkd->bnd", qp, Sp) + torch.einsum("bnk,bnkd->bnd", qn, Sn)
|
| 194 |
+
den = (qp * zp).sum(-1, keepdim=True) + (qn * zn).sum(-1, keepdim=True)
|
| 195 |
+
return self.o(num / den.clamp_min(1e-12))
|
| 196 |
+
|
| 197 |
+
|
| 198 |
+
class MslRelay(nn.Module):
|
| 199 |
+
"""Depth-composition unit (chain-rule probe): multi-slot M_hat read entering
|
| 200 |
+
the trunk as a NEAR-ZERO gated residual (gate init -3.0, sigma~0.047 — theme D:
|
| 201 |
+
geometry enters as a nudge and grows only if it earns gradient)."""
|
| 202 |
+
|
| 203 |
+
def __init__(self, d: int, n_slots: int = 16, K: int = 64):
|
| 204 |
+
super().__init__()
|
| 205 |
+
self.n_slots = n_slots
|
| 206 |
+
self.proj = nn.Linear(d, n_slots * 4, bias=False)
|
| 207 |
+
self.out = nn.Linear(n_slots * 4, d, bias=False)
|
| 208 |
+
nn.init.orthogonal_(self.proj.weight)
|
| 209 |
+
nn.init.orthogonal_(self.out.weight)
|
| 210 |
+
self.addr = AlephAddress(K, 4)
|
| 211 |
+
self.gate = nn.Parameter(torch.tensor(-3.0))
|
| 212 |
+
|
| 213 |
+
def forward(self, x):
|
| 214 |
+
B, n, _ = x.shape
|
| 215 |
+
slots = self.proj(x).view(B, n, self.n_slots, 4)
|
| 216 |
+
m = self.addr.m_hat(slots).reshape(B, n, -1)
|
| 217 |
+
return x + self.gate.sigmoid() * self.out(m)
|
| 218 |
+
|
| 219 |
+
|
| 220 |
+
class Block(nn.Module):
|
| 221 |
+
def __init__(self, d: int, attn: nn.Module):
|
| 222 |
+
super().__init__()
|
| 223 |
+
self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)
|
| 224 |
+
self.attn = attn
|
| 225 |
+
self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
|
| 226 |
+
|
| 227 |
+
def forward(self, x):
|
| 228 |
+
x = x + self.attn(self.n1(x))
|
| 229 |
+
return x + self.mlp(self.n2(x))
|
| 230 |
+
|
| 231 |
+
|
| 232 |
+
class ByteLM(nn.Module):
|
| 233 |
+
def __init__(self, arm: str, d: int = 192, layers: int = 4, block: int = 256,
|
| 234 |
+
K: int = 32, D: int = 4):
|
| 235 |
+
super().__init__()
|
| 236 |
+
# "<arm>_tri" suffix = trigram byte embedding (AlephLM byte_emb x3 lineage):
|
| 237 |
+
# token embedding is the sum of embeddings of bytes t, t-1, t-2.
|
| 238 |
+
self.trigram = arm.endswith("_tri")
|
| 239 |
+
if self.trigram:
|
| 240 |
+
arm = arm[:-4]
|
| 241 |
+
# "_fib" = super-Fibonacci S^3 codebook init (basin test: starts INSIDE
|
| 242 |
+
# the RP^3 attractor; primary observable is init->final geodesic drift).
|
| 243 |
+
self.fib = arm.endswith("_fib")
|
| 244 |
+
if self.fib:
|
| 245 |
+
arm = arm[:-4]
|
| 246 |
+
# "relay*" = stacked addresses in depth: MslRelay after every block.
|
| 247 |
+
# relay -> sdpa trunk + standard head; relay_msl64 -> + addressed head.
|
| 248 |
+
self.use_relay = arm.startswith("relay")
|
| 249 |
+
if arm == "relay":
|
| 250 |
+
arm = "sdpa"
|
| 251 |
+
elif arm == "relay_msl64":
|
| 252 |
+
arm = "addr_msl64"
|
| 253 |
+
self.arm, self.block = arm, block
|
| 254 |
+
self.emb = nn.Embedding(VOCAB, d)
|
| 255 |
+
if self.trigram:
|
| 256 |
+
self.emb1 = nn.Embedding(VOCAB, d)
|
| 257 |
+
self.emb2 = nn.Embedding(VOCAB, d)
|
| 258 |
+
self.pos = nn.Parameter(torch.zeros(1, block, d) + 0.01 * torch.randn(1, block, d))
|
| 259 |
+
mk_attn = (lambda: CausalHUB(d, K, D)) if arm == "hub" else (lambda: CausalSDPA(d))
|
| 260 |
+
self.blocks = nn.ModuleList([Block(d, mk_attn()) for _ in range(layers)])
|
| 261 |
+
if self.use_relay:
|
| 262 |
+
self.relays = nn.ModuleList([MslRelay(d) for _ in range(layers)])
|
| 263 |
+
self.nf = nn.LayerNorm(d)
|
| 264 |
+
if arm == "addr_head":
|
| 265 |
+
self.head_addr = AlephAddress(K, d) # v1: codebook in model dim — COLLAPSED
|
| 266 |
+
self.head = nn.Linear(K, VOCAB, bias=True)
|
| 267 |
+
elif arm in ("addr_d4", "addr_3tau", "addr_mhat"):
|
| 268 |
+
# v2 refinements: LOW-D HOME — learned projection to the native D=4 home
|
| 269 |
+
# before addressing (mirrors the healthy HUB arms), K=64.
|
| 270 |
+
self.head_proj = nn.Linear(d, 4, bias=False)
|
| 271 |
+
nn.init.orthogonal_(self.head_proj.weight)
|
| 272 |
+
self.head_addr = AlephAddress(64, 4)
|
| 273 |
+
if arm == "addr_d4":
|
| 274 |
+
self.head = nn.Linear(64, VOCAB, bias=True) # w alone, D=4 home
|
| 275 |
+
elif arm == "addr_3tau":
|
| 276 |
+
self.taus = (0.05, 0.1, 0.3) # rule-of-3 strobe
|
| 277 |
+
self.head = nn.Linear(64 * 3, VOCAB, bias=True)
|
| 278 |
+
else: # addr_mhat
|
| 279 |
+
self.head = nn.Linear(4, VOCAB, bias=True) # tightest: M_hat
|
| 280 |
+
elif arm.startswith("addr_msl"):
|
| 281 |
+
# v3: MULTI-SLOT heads — the 16s funnel widening: P parallel D=4 slots
|
| 282 |
+
# over a SHARED codebook. addr_msl consumes the reconstructive M_hat per
|
| 283 |
+
# slot (Px4 dims); addr_msl_w consumes signed w per slot (Px64) — tests
|
| 284 |
+
# whether slot-parallel consumption alone rescues the coefficient path.
|
| 285 |
+
# addr_msl<P> = slot-count dose-response. addr_mslh<P> = HARD sign-code
|
| 286 |
+
# consumption (straight-through M_hard per slot).
|
| 287 |
+
self.hard = arm.startswith("addr_mslh")
|
| 288 |
+
if arm in ("addr_msl", "addr_msl_w"):
|
| 289 |
+
self.n_slots = 16
|
| 290 |
+
else:
|
| 291 |
+
self.n_slots = int(arm[len("addr_mslh" if self.hard else "addr_msl"):])
|
| 292 |
+
self.head_proj = nn.Linear(d, self.n_slots * 4, bias=False)
|
| 293 |
+
nn.init.orthogonal_(self.head_proj.weight)
|
| 294 |
+
self.head_addr = AlephAddress(
|
| 295 |
+
64, 4, init="fibonacci" if self.fib else "random")
|
| 296 |
+
width = self.n_slots * (64 if arm == "addr_msl_w" else 4)
|
| 297 |
+
self.head = nn.Linear(width, VOCAB, bias=True)
|
| 298 |
+
elif arm == "addr_3tau_mhat":
|
| 299 |
+
# v3: combine the two v2 winners — 3-tau stroboscope + reconstructive read.
|
| 300 |
+
self.head_proj = nn.Linear(d, 4, bias=False)
|
| 301 |
+
nn.init.orthogonal_(self.head_proj.weight)
|
| 302 |
+
self.head_addr = AlephAddress(64, 4)
|
| 303 |
+
self.taus = (0.05, 0.1, 0.3)
|
| 304 |
+
self.head = nn.Linear(64 * 3 + 4, VOCAB, bias=True)
|
| 305 |
+
else:
|
| 306 |
+
self.head = nn.Linear(d, VOCAB, bias=True)
|
| 307 |
+
self._last_h = None
|
| 308 |
+
|
| 309 |
+
def forward(self, idx):
|
| 310 |
+
x = self.emb(idx)
|
| 311 |
+
if self.trigram: # past-only shifts — causality preserved
|
| 312 |
+
x = x + self.emb1(F.pad(idx, (1, 0), value=0)[:, :-1]) \
|
| 313 |
+
+ self.emb2(F.pad(idx, (2, 0), value=0)[:, :-2])
|
| 314 |
+
x = x + self.pos[:, : idx.shape[1]]
|
| 315 |
+
if self.use_relay:
|
| 316 |
+
for b, r in zip(self.blocks, self.relays):
|
| 317 |
+
x = r(b(x))
|
| 318 |
+
else:
|
| 319 |
+
for b in self.blocks:
|
| 320 |
+
x = b(x)
|
| 321 |
+
h = self.nf(x)
|
| 322 |
+
self._last_h = h.detach()
|
| 323 |
+
if self.arm == "addr_head":
|
| 324 |
+
return self.head(self.head_addr.signed(h))
|
| 325 |
+
if self.arm == "addr_d4":
|
| 326 |
+
return self.head(self.head_addr.signed(self.head_proj(h)))
|
| 327 |
+
if self.arm == "addr_3tau":
|
| 328 |
+
return self.head(self.head_addr.signed_at(self.head_proj(h), self.taus))
|
| 329 |
+
if self.arm == "addr_mhat":
|
| 330 |
+
return self.head(self.head_addr.m_hat(self.head_proj(h)))
|
| 331 |
+
if self.arm.startswith("addr_msl"):
|
| 332 |
+
B, n, _ = h.shape
|
| 333 |
+
slots = self.head_proj(h).view(B, n, self.n_slots, 4)
|
| 334 |
+
if self.arm == "addr_msl_w":
|
| 335 |
+
feats = self.head_addr.signed(slots).reshape(B, n, -1)
|
| 336 |
+
elif getattr(self, "hard", False):
|
| 337 |
+
feats = self.head_addr.m_hard_ste(slots).reshape(B, n, -1)
|
| 338 |
+
else:
|
| 339 |
+
feats = self.head_addr.m_hat(slots).reshape(B, n, -1)
|
| 340 |
+
return self.head(feats)
|
| 341 |
+
if self.arm == "addr_3tau_mhat":
|
| 342 |
+
p = self.head_proj(h)
|
| 343 |
+
feats = torch.cat([self.head_addr.signed_at(p, self.taus),
|
| 344 |
+
self.head_addr.m_hat(p)], dim=-1)
|
| 345 |
+
return self.head(feats)
|
| 346 |
+
return self.head(h)
|
| 347 |
+
|
| 348 |
+
@torch.no_grad()
|
| 349 |
+
def vitals(self) -> dict:
|
| 350 |
+
out = {}
|
| 351 |
+
if self.arm == "hub":
|
| 352 |
+
for i, b in enumerate(self.blocks):
|
| 353 |
+
if self._last_h is not None:
|
| 354 |
+
out[f"L{i}"] = b.attn.addr.vitals(b.attn.q(self._last_h[:2]))
|
| 355 |
+
elif self.arm == "addr_head" and self._last_h is not None:
|
| 356 |
+
out["head"] = self.head_addr.vitals(self._last_h[:2])
|
| 357 |
+
elif self.arm in ("addr_d4", "addr_3tau", "addr_mhat",
|
| 358 |
+
"addr_3tau_mhat") and self._last_h is not None:
|
| 359 |
+
out["head"] = self.head_addr.vitals(self.head_proj(self._last_h[:2]))
|
| 360 |
+
elif self.arm.startswith("addr_msl") and self._last_h is not None:
|
| 361 |
+
slots = self.head_proj(self._last_h[:2])
|
| 362 |
+
out["head"] = self.head_addr.vitals(
|
| 363 |
+
slots.reshape(*slots.shape[:-1], self.n_slots, 4))
|
| 364 |
+
if self.use_relay and self._last_h is not None:
|
| 365 |
+
for i, r in enumerate(self.relays):
|
| 366 |
+
s = r.proj(self._last_h[:2])
|
| 367 |
+
v = r.addr.vitals(s.reshape(*s.shape[:-1], r.n_slots, 4))
|
| 368 |
+
out[f"relay{i}"] = {"gate": round(r.gate.sigmoid().item(), 4),
|
| 369 |
+
"drift": v["drift"],
|
| 370 |
+
"binding_frac": v["binding_frac"],
|
| 371 |
+
"ppl": round(v["aliveness"]["usage_ppl"], 1)}
|
| 372 |
+
return out
|
| 373 |
+
|
| 374 |
+
|
| 375 |
+
# --------------------------------------------------------------------------- data
|
| 376 |
+
def _wikitext_bytes(data_root: str):
|
| 377 |
+
"""wikitext-2-raw as flat uint8 tensors via the HF parquet CDN."""
|
| 378 |
+
from huggingface_hub import hf_hub_download
|
| 379 |
+
import pyarrow.parquet as pq
|
| 380 |
+
|
| 381 |
+
def load(split):
|
| 382 |
+
p = hf_hub_download("Salesforce/wikitext",
|
| 383 |
+
f"wikitext-2-raw-v1/{split}-00000-of-00001.parquet",
|
| 384 |
+
repo_type="dataset", local_dir=data_root)
|
| 385 |
+
text = "".join(pq.read_table(p).column("text").to_pylist())
|
| 386 |
+
return torch.frombuffer(bytearray(text.encode("utf-8")), dtype=torch.uint8).clone()
|
| 387 |
+
|
| 388 |
+
return load("train"), load("validation")
|
| 389 |
+
|
| 390 |
+
|
| 391 |
+
def _batch(data: torch.Tensor, batch: int, block: int, device, g: torch.Generator):
|
| 392 |
+
ix = torch.randint(0, data.numel() - block - 1, (batch,), generator=g)
|
| 393 |
+
x = torch.stack([data[i:i + block] for i in ix]).long().to(device)
|
| 394 |
+
y = torch.stack([data[i + 1:i + block + 1] for i in ix]).long().to(device)
|
| 395 |
+
return x, y
|
| 396 |
+
|
| 397 |
+
|
| 398 |
+
# -------------------------------------------------------------------- train/smoke
|
| 399 |
+
def train(arms=("sdpa", "hub", "addr_head"), steps: int = 2000, batch: int = 32,
|
| 400 |
+
block: int = 256, device: str = "cuda", data_root: str = "./data",
|
| 401 |
+
seed: int = 0, eval_every: int = 500, save: bool = True):
|
| 402 |
+
"""Verdict run — GPU only. Pure Adam wd=0. Reports val bits-per-byte + vitals.
|
| 403 |
+
save=True writes {data_root}/ar_ckpts/{arm}_s{seed}_t{steps}.pt per arm —
|
| 404 |
+
the cultivated codebooks are SPECIMENS for the projective reading instruments."""
|
| 405 |
+
import os
|
| 406 |
+
if device == "cuda" and not torch.cuda.is_available():
|
| 407 |
+
raise RuntimeError("Verdict runs are GPU-only (never CPU-train for accuracy).")
|
| 408 |
+
ckpt_dir = os.path.join(data_root, "ar_ckpts")
|
| 409 |
+
os.makedirs(ckpt_dir, exist_ok=True)
|
| 410 |
+
tr, va = _wikitext_bytes(data_root)
|
| 411 |
+
print(f"data ready: train {tr.numel():,} bytes, val {va.numel():,} bytes", flush=True)
|
| 412 |
+
results = {}
|
| 413 |
+
for arm in arms:
|
| 414 |
+
torch.manual_seed(seed)
|
| 415 |
+
g = torch.Generator().manual_seed(seed)
|
| 416 |
+
model = ByteLM(arm, block=block).to(device)
|
| 417 |
+
n_params = sum(p.numel() for p in model.parameters())
|
| 418 |
+
opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
|
| 419 |
+
for step in range(1, steps + 1):
|
| 420 |
+
x, y = _batch(tr, batch, block, device, g)
|
| 421 |
+
logits = model(x)
|
| 422 |
+
loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
|
| 423 |
+
opt.zero_grad(set_to_none=True)
|
| 424 |
+
loss.backward()
|
| 425 |
+
opt.step()
|
| 426 |
+
if step % eval_every == 0 or step == steps:
|
| 427 |
+
model.eval()
|
| 428 |
+
with torch.no_grad():
|
| 429 |
+
losses = []
|
| 430 |
+
for _ in range(20):
|
| 431 |
+
xv, yv = _batch(va, batch, block, device, g)
|
| 432 |
+
lv = F.cross_entropy(model(xv).reshape(-1, VOCAB),
|
| 433 |
+
yv.reshape(-1))
|
| 434 |
+
losses.append(lv.item())
|
| 435 |
+
bpb = sum(losses) / len(losses) / math.log(2)
|
| 436 |
+
print(f"[{arm}] step {step} val_bpb={bpb:.4f} vitals={model.vitals()}",
|
| 437 |
+
flush=True)
|
| 438 |
+
model.train()
|
| 439 |
+
results[arm] = {"val_bpb": bpb, "params": n_params, "vitals": model.vitals()}
|
| 440 |
+
if save:
|
| 441 |
+
path = os.path.join(ckpt_dir, f"{arm}_s{seed}_t{steps}.pt")
|
| 442 |
+
torch.save({"arm": arm, "seed": seed, "steps": steps, "val_bpb": bpb,
|
| 443 |
+
"state_dict": {k: v.cpu() for k, v in
|
| 444 |
+
model.state_dict().items()}}, path)
|
| 445 |
+
print(f"saved specimen: {path}", flush=True)
|
| 446 |
+
print(results, flush=True)
|
| 447 |
+
return results
|
| 448 |
+
|
| 449 |
+
|
| 450 |
+
def smoke():
|
| 451 |
+
"""Shapes/parse only — no accuracy claims."""
|
| 452 |
+
x = torch.randint(0, VOCAB, (2, 64))
|
| 453 |
+
for arm in ("sdpa", "hub", "addr_head"):
|
| 454 |
+
m = ByteLM(arm, d=96, layers=2, block=64, K=16)
|
| 455 |
+
logits = m(x)
|
| 456 |
+
assert logits.shape == (2, 64, VOCAB)
|
| 457 |
+
logits.sum().backward()
|
| 458 |
+
# causality check: future byte must not affect past logits
|
| 459 |
+
with torch.no_grad():
|
| 460 |
+
a = m(x)[0, 10]
|
| 461 |
+
x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
|
| 462 |
+
b = m(x2)[0, 10]
|
| 463 |
+
assert torch.allclose(a, b, atol=1e-4), f"{arm} leaks future context"
|
| 464 |
+
print(f"{arm}: OK params={sum(p.numel() for p in m.parameters()):,} "
|
| 465 |
+
f"vitals={m.vitals()}", flush=True)
|
| 466 |
+
print("OK — AR bed smoke passed (verdict run: train() on GPU)", flush=True)
|
| 467 |
+
|
| 468 |
+
|
| 469 |
+
def _in_notebook() -> bool:
|
| 470 |
+
try:
|
| 471 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 472 |
+
return True
|
| 473 |
+
except NameError:
|
| 474 |
+
return False
|
| 475 |
+
|
| 476 |
+
|
| 477 |
+
if __name__ == "__main__":
|
| 478 |
+
if _in_notebook():
|
| 479 |
+
smoke()
|
| 480 |
+
print("Notebook mode: call train(steps=2000) in the next cell (GPU).")
|
| 481 |
+
else:
|
| 482 |
+
import argparse
|
| 483 |
+
ap = argparse.ArgumentParser()
|
| 484 |
+
ap.add_argument("--train", action="store_true")
|
| 485 |
+
ap.add_argument("--steps", type=int, default=2000)
|
| 486 |
+
a, _ = ap.parse_known_args()
|
| 487 |
+
train(steps=a.steps) if a.train else smoke()
|
exp019_cr/build_results.py
ADDED
|
@@ -0,0 +1,81 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""build_results.py — exp019_cr: read results/ledger.jsonl and RE-ASSERT every
|
| 2 |
+
claim in the README (retention sweep + 19b generalization block).
|
| 3 |
+
Run from inside this folder: python build_results.py
|
| 4 |
+
"""
|
| 5 |
+
import json
|
| 6 |
+
import os
|
| 7 |
+
|
| 8 |
+
HERE = os.path.dirname(os.path.abspath(__file__))
|
| 9 |
+
rows = [json.loads(l) for l in
|
| 10 |
+
open(os.path.join(HERE, "results", "ledger.jsonl"), encoding="utf-8")]
|
| 11 |
+
a = [r for r in rows if r["exp"] == "19"]
|
| 12 |
+
b = [r for r in rows if r["exp"] == "19b"]
|
| 13 |
+
assert len(a) == 28 and len(b) == 8
|
| 14 |
+
|
| 15 |
+
def cellA(ch, n, s):
|
| 16 |
+
return next(r for r in a if r["channel"] == ch and r["n"] == n
|
| 17 |
+
and r["seed"] == s)
|
| 18 |
+
|
| 19 |
+
def cellB(ch, s):
|
| 20 |
+
return next(r for r in b if r["channel"] == ch and r["seed"] == s)
|
| 21 |
+
|
| 22 |
+
# retention claim 1: capacity curve — teachers near-perfect through N=256,
|
| 23 |
+
# cliff before 1024
|
| 24 |
+
for s in (0, 1):
|
| 25 |
+
assert cellA("teacher", 64, s)["recall"]["exact"] == 1.0
|
| 26 |
+
assert cellA("teacher", 256, s)["recall"]["exact"] > 0.98
|
| 27 |
+
assert cellA("teacher", 1024, s)["recall"]["exact"] < 0.20
|
| 28 |
+
|
| 29 |
+
# retention claim 2: distillation tax zero-to-negative where the teacher knows
|
| 30 |
+
# the content (kd_facts >= direct - 0.01 at N=64/256, both seeds)
|
| 31 |
+
for s in (0, 1):
|
| 32 |
+
for n in (64, 256):
|
| 33 |
+
kd = cellA("kd_facts", n, s)["recall"]["exact"]
|
| 34 |
+
di = cellA("direct", n, s)["recall"]["exact"]
|
| 35 |
+
assert kd >= di - 0.01, (n, s, kd, di)
|
| 36 |
+
|
| 37 |
+
# retention claim 3: no logit leakage on clean text (kd_general exact 0.0 all)
|
| 38 |
+
for r in a:
|
| 39 |
+
if r["channel"] == "kd_general":
|
| 40 |
+
assert r["recall"]["exact"] == 0.0
|
| 41 |
+
|
| 42 |
+
# retention claim 4: anchors carry no bytes (book_implant exact 0.0, both seeds)
|
| 43 |
+
for s in (0, 1):
|
| 44 |
+
assert cellA("book_implant", 256, s)["recall"]["exact"] == 0.0
|
| 45 |
+
|
| 46 |
+
# retention claim 5: universal catastrophic forgetting (every student channel
|
| 47 |
+
# -> 0.0 exact after 1k clean steps)
|
| 48 |
+
for r in a:
|
| 49 |
+
if "recall_interf" in r and r["recall_interf"] is not None:
|
| 50 |
+
assert r["recall_interf"]["exact"] == 0.0, r
|
| 51 |
+
|
| 52 |
+
# 19b claim 1: partial rule learning — teachers memorize, held-out byte acc
|
| 53 |
+
# ~10x chance but exact 0
|
| 54 |
+
for s in (0, 1):
|
| 55 |
+
t = cellB("teacher", s)["gauges"]
|
| 56 |
+
assert t["train"]["exact"] > 0.98
|
| 57 |
+
assert 0.15 < t["heldout"]["byte_acc"] < 0.35
|
| 58 |
+
assert t["heldout"]["exact"] == 0.0
|
| 59 |
+
# format lock: variant-format recall collapses even on train keys
|
| 60 |
+
assert t["train_varfmt"]["byte_acc"] < 0.05
|
| 61 |
+
|
| 62 |
+
# 19b claim 2: the logit channel beats direct on rule content, both seeds,
|
| 63 |
+
# and transfers the partial rule at full fidelity
|
| 64 |
+
for s in (0, 1):
|
| 65 |
+
kd, di = cellB("kd_facts", s)["gauges"], cellB("direct", s)["gauges"]
|
| 66 |
+
assert kd["train"]["exact"] > di["train"]["exact"], s
|
| 67 |
+
assert kd["heldout"]["byte_acc"] > 0.24
|
| 68 |
+
# kd students are less format-locked than direct students
|
| 69 |
+
assert kd["train_varfmt"]["byte_acc"] > di["train_varfmt"]["byte_acc"], s
|
| 70 |
+
|
| 71 |
+
out = {"capacity_curve": {f"N{n}_s{s}": cellA("teacher", n, s)["recall"]["exact"]
|
| 72 |
+
for n in (64, 256, 1024) for s in (0, 1)},
|
| 73 |
+
"kd_vs_direct_rule_train_exact": {
|
| 74 |
+
f"s{s}": [cellB("kd_facts", s)["gauges"]["train"]["exact"],
|
| 75 |
+
cellB("direct", s)["gauges"]["train"]["exact"]]
|
| 76 |
+
for s in (0, 1)},
|
| 77 |
+
"n_rows": len(rows)}
|
| 78 |
+
json.dump(out, open(os.path.join(HERE, "results", "results.json"), "w",
|
| 79 |
+
encoding="utf-8"), indent=1)
|
| 80 |
+
print(f"{len(rows)} rows -> results/results.json")
|
| 81 |
+
print("all README claims asserted OK")
|
exp019_cr/exp014_genetic_distillation.py
ADDED
|
@@ -0,0 +1,515 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""exp014_genetic_distillation.py — genetic distillation + memory substrate.
|
| 2 |
+
14_a: multi-generational tournament (GM3 paradigm) where the aleph codebook is the
|
| 3 |
+
explicit heritable genome. Lineages: ALEPH-FLAT (consensus book + KD) |
|
| 4 |
+
ALEPH-TREE (structured genome: root book + branch books) | MLP-KD
|
| 5 |
+
(traditional: best-parent weights + KD) | NO-INHERIT (evolution floor).
|
| 6 |
+
Both sides intentionally inherit logits (KD); only ours inherits geometry.
|
| 7 |
+
Consensus = Procrustes/GPA alignment of parents' books to mean shape
|
| 8 |
+
(placement by construction — replaces GM3's k-means-on-consensus init).
|
| 9 |
+
14_b: memory substrate — the D=4 home makes books size-agnostic. Implant books
|
| 10 |
+
cultivated in a small organism into a larger one (frozen / trainable), and
|
| 11 |
+
into GPT-2 relay adapters (cross-architecture frozen distillation).
|
| 12 |
+
|
| 13 |
+
Riders: pure Adam wd=0; KD = KL to detached teacher probs (predictive pressure, no
|
| 14 |
+
contrastive); tree routing is DENSE SOFT (oriented weights; collapse monitor on the
|
| 15 |
+
root); drift-check precedes every freeze claim; GPU-only verdict runs; Colab-safe.
|
| 16 |
+
Founders share a COMMON-ANCESTOR book so GPA row correspondence is inherited.
|
| 17 |
+
|
| 18 |
+
Colab paste order: geolip_vitals.py -> ar_differentiation_bed.py ->
|
| 19 |
+
exp013_augmentation_bed.py (only for run_b2) -> this file.
|
| 20 |
+
"""
|
| 21 |
+
from __future__ import annotations
|
| 22 |
+
import copy
|
| 23 |
+
import json
|
| 24 |
+
import math
|
| 25 |
+
import os
|
| 26 |
+
import torch
|
| 27 |
+
import torch.nn as nn
|
| 28 |
+
import torch.nn.functional as F
|
| 29 |
+
|
| 30 |
+
if "anchor_drift" not in globals():
|
| 31 |
+
try:
|
| 32 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 33 |
+
except ImportError:
|
| 34 |
+
_here = globals().get("__file__")
|
| 35 |
+
if _here is None:
|
| 36 |
+
raise ImportError("paste/run geolip_vitals.py first")
|
| 37 |
+
import sys, pathlib
|
| 38 |
+
sys.path.insert(0, str(pathlib.Path(_here).parent))
|
| 39 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 40 |
+
if "ByteLM" not in globals():
|
| 41 |
+
try:
|
| 42 |
+
from ar_differentiation_bed import (ByteLM, AlephAddress, _wikitext_bytes,
|
| 43 |
+
_batch, VOCAB)
|
| 44 |
+
except ImportError:
|
| 45 |
+
raise ImportError("paste/run ar_differentiation_bed.py first")
|
| 46 |
+
|
| 47 |
+
DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
|
| 48 |
+
EXP_DIR = os.path.join(DATA_ROOT, "exp014")
|
| 49 |
+
|
| 50 |
+
|
| 51 |
+
class SquaredReLU(nn.Module):
|
| 52 |
+
def forward(self, x):
|
| 53 |
+
return F.relu(x) ** 2
|
| 54 |
+
|
| 55 |
+
|
| 56 |
+
# ===================================================== consensus (the germline) ===
|
| 57 |
+
@torch.no_grad()
|
| 58 |
+
def procrustes_rotation(A: torch.Tensor, M: torch.Tensor) -> torch.Tensor:
|
| 59 |
+
"""Orthogonal R minimizing ||A R - M||_F (rows correspond)."""
|
| 60 |
+
U, _, Vt = torch.linalg.svd(A.T.double() @ M.double())
|
| 61 |
+
return (U @ Vt).float()
|
| 62 |
+
|
| 63 |
+
|
| 64 |
+
@torch.no_grad()
|
| 65 |
+
def align_to(A: torch.Tensor, ref: torch.Tensor, iters: int = 20) -> torch.Tensor:
|
| 66 |
+
"""Projective Procrustes (rows correspond, signs free): alternate the
|
| 67 |
+
orthogonal rotation and per-row sign flips (books live on RP^(D-1))."""
|
| 68 |
+
s = torch.ones(A.shape[0], 1)
|
| 69 |
+
for _ in range(iters):
|
| 70 |
+
R = procrustes_rotation(s * A, ref)
|
| 71 |
+
AR = (s * A) @ R
|
| 72 |
+
s_upd = torch.where((AR * ref).sum(-1, keepdim=True) < 0, -s, s)
|
| 73 |
+
if torch.equal(s_upd, s):
|
| 74 |
+
return AR
|
| 75 |
+
s = s_upd
|
| 76 |
+
return (s * A) @ procrustes_rotation(s * A, ref)
|
| 77 |
+
|
| 78 |
+
|
| 79 |
+
@torch.no_grad()
|
| 80 |
+
def consensus_codebook(books: list, iters: int = 50, tol: float = 1e-8):
|
| 81 |
+
"""GPA to mean shape (GM3 machinery, applied to aleph books), anchored to the
|
| 82 |
+
FIRST parent's frame. Rows must correspond (common-ancestor convention); signs
|
| 83 |
+
are projective. Returns (consensus, n_iters, delta)."""
|
| 84 |
+
# device-pin to CPU: parent models may live on CUDA after KD teacher moves
|
| 85 |
+
Bs = [F.normalize(b.detach().float().cpu(), dim=-1).clone() for b in books]
|
| 86 |
+
# pairwise projective alignment to parent-0's frame, THEN GPA refinement
|
| 87 |
+
aligned = [Bs[0]] + [align_to(b, Bs[0]) for b in Bs[1:]]
|
| 88 |
+
M = F.normalize(torch.stack(aligned).mean(0), dim=-1)
|
| 89 |
+
delta, it = 0.0, 0
|
| 90 |
+
for it in range(1, iters + 1):
|
| 91 |
+
aligned = [align_to(b, M) for b in Bs]
|
| 92 |
+
M_new = F.normalize(torch.stack(aligned).mean(0), dim=-1)
|
| 93 |
+
delta = (M_new - M).norm().item()
|
| 94 |
+
M = M_new
|
| 95 |
+
if delta < tol:
|
| 96 |
+
break
|
| 97 |
+
# re-anchor to parent-0 (GPA drift of the global frame stays measurable)
|
| 98 |
+
M = align_to(M, Bs[0])
|
| 99 |
+
return M, it, delta
|
| 100 |
+
|
| 101 |
+
|
| 102 |
+
@torch.no_grad()
|
| 103 |
+
def implant_book(addr: "AlephAddress", book: torch.Tensor, trainable: bool = True):
|
| 104 |
+
"""Load a book into an AlephAddress: codebook + home (drift measured from the
|
| 105 |
+
implant). Freeze only via trainable=False AFTER a drift-check justifies it."""
|
| 106 |
+
b = F.normalize(book.float(), dim=-1).to(addr.codebook.device)
|
| 107 |
+
assert b.shape == addr.codebook.shape, (b.shape, addr.codebook.shape)
|
| 108 |
+
addr.codebook.data.copy_(b)
|
| 109 |
+
addr.home.copy_(b)
|
| 110 |
+
addr.codebook.requires_grad_(trainable)
|
| 111 |
+
|
| 112 |
+
|
| 113 |
+
# ============================================================= tree head ==========
|
| 114 |
+
class TreeHead(nn.Module):
|
| 115 |
+
"""Autoregressive tree (the constellation-anchor analogue, exp011 TREE operator
|
| 116 |
+
in the healthy consumption regime): a ROOT aleph (K=2, D=4) yields the 4
|
| 117 |
+
oriented weights (2K half-axes = the 4 branches, dense soft, sums to 1);
|
| 118 |
+
each BRANCH is a 64-slot... shared slot projection read by a branch-specific
|
| 119 |
+
book (K=64, D=4); output = branch-weighted mixture of branch reads -> vocab.
|
| 120 |
+
Heritable genome: root book (2,4) + 4 branch books (64,4)."""
|
| 121 |
+
|
| 122 |
+
ROOT_SLOTS = 4 # slot-parallel root consumption (the collapse cure)
|
| 123 |
+
ROOT_TAU = 0.3 # softer root temperature (wave-1 fix: single hard-tau
|
| 124 |
+
# root partially collapsed, usage [.85,.12,.01,.02])
|
| 125 |
+
|
| 126 |
+
def __init__(self, d: int, vocab: int = 256, n_slots: int = 16):
|
| 127 |
+
super().__init__()
|
| 128 |
+
self.n_slots = n_slots
|
| 129 |
+
self.root_proj = nn.Linear(d, self.ROOT_SLOTS * 4, bias=False)
|
| 130 |
+
self.slot_proj = nn.Linear(d, n_slots * 4, bias=False)
|
| 131 |
+
nn.init.orthogonal_(self.root_proj.weight)
|
| 132 |
+
nn.init.orthogonal_(self.slot_proj.weight)
|
| 133 |
+
self.root = AlephAddress(2, 4, tau=self.ROOT_TAU)
|
| 134 |
+
self.branches = nn.ModuleList([AlephAddress(64, 4) for _ in range(4)])
|
| 135 |
+
self.out = nn.Linear(n_slots * 4, vocab, bias=True)
|
| 136 |
+
self._last_root = None
|
| 137 |
+
|
| 138 |
+
def forward(self, h):
|
| 139 |
+
B, n, _ = h.shape
|
| 140 |
+
rs = self.root_proj(h).view(B, n, self.ROOT_SLOTS, 4)
|
| 141 |
+
p, m = self.root.oriented(rs) # (B,n,S,2) x2
|
| 142 |
+
w = torch.cat([p, m], dim=-1).mean(dim=-2) # slot-avg -> (B,n,4)
|
| 143 |
+
self._last_root = w.detach()
|
| 144 |
+
slots = self.slot_proj(h).view(B, n, self.n_slots, 4)
|
| 145 |
+
mix = 0
|
| 146 |
+
for b, br in enumerate(self.branches):
|
| 147 |
+
mix = mix + w[..., b:b + 1] * br.m_hat(slots).reshape(B, n, -1)
|
| 148 |
+
return self.out(mix)
|
| 149 |
+
|
| 150 |
+
def genome(self):
|
| 151 |
+
return {"root": self.root.codebook.detach().clone(),
|
| 152 |
+
**{f"branch{i}": br.codebook.detach().clone()
|
| 153 |
+
for i, br in enumerate(self.branches)}}
|
| 154 |
+
|
| 155 |
+
@torch.no_grad()
|
| 156 |
+
def inherit(self, genomes: list):
|
| 157 |
+
c, it, dl = consensus_codebook([g["root"] for g in genomes])
|
| 158 |
+
implant_book(self.root, c)
|
| 159 |
+
for i, br in enumerate(self.branches):
|
| 160 |
+
c, _, _ = consensus_codebook([g[f"branch{i}"] for g in genomes])
|
| 161 |
+
implant_book(br, c)
|
| 162 |
+
|
| 163 |
+
@torch.no_grad()
|
| 164 |
+
def vitals(self):
|
| 165 |
+
out = {"root_drift": round(anchor_drift(self.root.codebook,
|
| 166 |
+
self.root.home)["mean"], 4)}
|
| 167 |
+
if self._last_root is not None:
|
| 168 |
+
w = self._last_root.reshape(-1, 4)
|
| 169 |
+
usage = w.mean(0)
|
| 170 |
+
usage = usage / usage.sum()
|
| 171 |
+
out["root_usage"] = [round(float(u), 3) for u in usage]
|
| 172 |
+
ent = -(usage.clamp_min(1e-9) * usage.clamp_min(1e-9).log()).sum()
|
| 173 |
+
out["root_ppl4"] = round(float(ent.exp()), 3)
|
| 174 |
+
d = [anchor_drift(br.codebook, br.home)["mean"] for br in self.branches]
|
| 175 |
+
out["branch_drift"] = [round(x, 3) for x in d]
|
| 176 |
+
return out
|
| 177 |
+
|
| 178 |
+
|
| 179 |
+
# ============================================================ organisms ===========
|
| 180 |
+
def make_organism(lineage: str, d: int = 192, layers: int = 4, block: int = 256,
|
| 181 |
+
seed: int = 0):
|
| 182 |
+
"""lineage in {aleph_flat, aleph_tree, mlp_kd, no_inherit}. no_inherit uses the
|
| 183 |
+
aleph_flat architecture (the control isolates INHERITANCE, not architecture)."""
|
| 184 |
+
torch.manual_seed(seed)
|
| 185 |
+
if lineage in ("aleph_flat", "no_inherit", "aleph_full", "aleph_weights"):
|
| 186 |
+
return ByteLM("addr_msl64", d=d, layers=layers, block=block)
|
| 187 |
+
if lineage == "aleph_tree":
|
| 188 |
+
m = ByteLM("sdpa", d=d, layers=layers, block=block)
|
| 189 |
+
m.head = TreeHead(d)
|
| 190 |
+
return m
|
| 191 |
+
if lineage == "mlp_kd":
|
| 192 |
+
m = ByteLM("sdpa", d=d, layers=layers, block=block)
|
| 193 |
+
m.head = nn.Sequential(nn.Linear(d, 224), SquaredReLU(),
|
| 194 |
+
nn.LayerNorm(224), nn.Linear(224, VOCAB))
|
| 195 |
+
return m
|
| 196 |
+
raise ValueError(lineage)
|
| 197 |
+
|
| 198 |
+
|
| 199 |
+
def genome_of(model):
|
| 200 |
+
"""The heritable organ = co-adapted (projection, book) pair(s). Books inherit
|
| 201 |
+
by CONSENSUS (the geometric germline); projections inherit from the BEST
|
| 202 |
+
parent (weight copy) — implanting a book against a random projection puts the
|
| 203 |
+
child below random init (campaign-v2 lesson)."""
|
| 204 |
+
if isinstance(model.head, TreeHead):
|
| 205 |
+
g = model.head.genome()
|
| 206 |
+
g["root_proj"] = model.head.root_proj.weight.detach().cpu().clone()
|
| 207 |
+
g["slot_proj"] = model.head.slot_proj.weight.detach().cpu().clone()
|
| 208 |
+
return g
|
| 209 |
+
if hasattr(model, "head_addr"):
|
| 210 |
+
return {"flat": model.head_addr.codebook.detach().cpu().clone(),
|
| 211 |
+
"proj": model.head_proj.weight.detach().cpu().clone()}
|
| 212 |
+
return None
|
| 213 |
+
|
| 214 |
+
|
| 215 |
+
@torch.no_grad()
|
| 216 |
+
def inherit_genome(model, genomes: list):
|
| 217 |
+
"""genomes[0] = the BEST parent (selection order matters)."""
|
| 218 |
+
if isinstance(model.head, TreeHead):
|
| 219 |
+
model.head.inherit(genomes)
|
| 220 |
+
model.head.root_proj.weight.copy_(genomes[0]["root_proj"].to(
|
| 221 |
+
model.head.root_proj.weight.device))
|
| 222 |
+
model.head.slot_proj.weight.copy_(genomes[0]["slot_proj"].to(
|
| 223 |
+
model.head.slot_proj.weight.device))
|
| 224 |
+
elif hasattr(model, "head_addr"):
|
| 225 |
+
c, it, dl = consensus_codebook([g["flat"] for g in genomes])
|
| 226 |
+
implant_book(model.head_addr, c)
|
| 227 |
+
model.head_proj.weight.copy_(genomes[0]["proj"].to(
|
| 228 |
+
model.head_proj.weight.device))
|
| 229 |
+
|
| 230 |
+
|
| 231 |
+
def organism_vitals(model):
|
| 232 |
+
if isinstance(model.head, TreeHead):
|
| 233 |
+
return model.head.vitals()
|
| 234 |
+
return model.vitals() if hasattr(model, "vitals") else {}
|
| 235 |
+
|
| 236 |
+
|
| 237 |
+
# ========================================================= train one member ======
|
| 238 |
+
def train_member(model, tr, va, steps=2000, batch=32, block=256, device="cuda",
|
| 239 |
+
seed=0, teachers=None, kd_alpha=1.0):
|
| 240 |
+
"""CE (+ KL to detached mean teacher probs when teachers given). Pure Adam."""
|
| 241 |
+
g = torch.Generator().manual_seed(seed)
|
| 242 |
+
model = model.to(device)
|
| 243 |
+
opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
|
| 244 |
+
if teachers:
|
| 245 |
+
teachers = [t.to(device).eval() for t in teachers]
|
| 246 |
+
for step in range(1, steps + 1):
|
| 247 |
+
x, y = _batch(tr, batch, block, device, g)
|
| 248 |
+
logits = model(x)
|
| 249 |
+
loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
|
| 250 |
+
if teachers:
|
| 251 |
+
with torch.no_grad():
|
| 252 |
+
tp = torch.stack([F.softmax(t(x), -1) for t in teachers]).mean(0)
|
| 253 |
+
loss = loss + kd_alpha * F.kl_div(
|
| 254 |
+
F.log_softmax(logits, -1), tp, reduction="batchmean")
|
| 255 |
+
opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
|
| 256 |
+
model.eval()
|
| 257 |
+
with torch.no_grad():
|
| 258 |
+
ls = []
|
| 259 |
+
for _ in range(20):
|
| 260 |
+
xv, yv = _batch(va, batch, block, device, g)
|
| 261 |
+
ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
|
| 262 |
+
yv.reshape(-1)).item())
|
| 263 |
+
return sum(ls) / len(ls) / math.log(2) # bpb
|
| 264 |
+
|
| 265 |
+
|
| 266 |
+
# ============================================================ the tournament ======
|
| 267 |
+
def run_tournament(lineage: str, gens: int = 4, pop: int = 4, steps: int = 2000,
|
| 268 |
+
seed: int = 0, device: str = "cuda",
|
| 269 |
+
catastrophic_at: int | None = None):
|
| 270 |
+
"""One lineage, one tournament seed. Logs per-gen to the ledger; saves the
|
| 271 |
+
champion genome per generation. catastrophic_at=G injects a 0-step random
|
| 272 |
+
parent into the consensus at generation G (the GM3 robustness probe)."""
|
| 273 |
+
if not torch.cuda.is_available():
|
| 274 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 275 |
+
os.makedirs(EXP_DIR, exist_ok=True)
|
| 276 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 277 |
+
ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 278 |
+
# common ancestor: every founder book starts identical within a tournament
|
| 279 |
+
torch.manual_seed(9000 + seed)
|
| 280 |
+
ancestor = make_organism(lineage, seed=9000 + seed)
|
| 281 |
+
anc_genome = genome_of(ancestor)
|
| 282 |
+
parents, parent_models, champion_genomes = [], [], []
|
| 283 |
+
for gen in range(gens):
|
| 284 |
+
members = []
|
| 285 |
+
for i in range(pop):
|
| 286 |
+
mseed = seed * 1000 + gen * 100 + i
|
| 287 |
+
m = make_organism(lineage, seed=mseed)
|
| 288 |
+
if anc_genome and gen == 0:
|
| 289 |
+
inherit_genome(m, [anc_genome]) # common ancestor
|
| 290 |
+
if gen > 0:
|
| 291 |
+
is_fresh = (i == pop - 1) # gene flow founder
|
| 292 |
+
if not is_fresh:
|
| 293 |
+
if lineage in ("aleph_flat", "aleph_tree"):
|
| 294 |
+
gs = [genome_of(pm) for pm in parent_models]
|
| 295 |
+
if catastrophic_at == gen:
|
| 296 |
+
bad = make_organism(lineage, seed=666 + i)
|
| 297 |
+
gs = gs + [genome_of(bad)]
|
| 298 |
+
inherit_genome(m, gs)
|
| 299 |
+
elif lineage == "aleph_full":
|
| 300 |
+
# v4 arm (v3 lesson: continuity is what pays) — inherit the
|
| 301 |
+
# WHOLE best parent, then overwrite the book with the
|
| 302 |
+
# two-parent consensus: germline ON TOP of continuity.
|
| 303 |
+
m.load_state_dict(copy.deepcopy(
|
| 304 |
+
parent_models[0].state_dict()))
|
| 305 |
+
gs = [genome_of(pm) for pm in parent_models]
|
| 306 |
+
if catastrophic_at == gen:
|
| 307 |
+
bad = make_organism(lineage, seed=666 + i)
|
| 308 |
+
gs = gs + [genome_of(bad)]
|
| 309 |
+
c, _, _ = consensus_codebook([g["flat"] for g in gs])
|
| 310 |
+
implant_book(m.head_addr, c)
|
| 311 |
+
elif lineage in ("mlp_kd", "aleph_weights"):
|
| 312 |
+
# pure continuity (no germline op) — aleph_weights is the
|
| 313 |
+
# within-architecture control for aleph_full
|
| 314 |
+
m.load_state_dict(copy.deepcopy(
|
| 315 |
+
parent_models[0].state_dict()))
|
| 316 |
+
# no_inherit: nothing
|
| 317 |
+
# KD: alpha 0.25 (campaign-v2 lesson: alpha=1.0 from near-parity
|
| 318 |
+
# teachers COMPOUNDS DOWNWARD — inverse evolution; fresh-founder
|
| 319 |
+
# control isolated it). Fresh founders get NO KD (clean gene flow).
|
| 320 |
+
is_fresh_now = (gen > 0 and i == pop - 1)
|
| 321 |
+
teachers = parent_models if (gen > 0 and not is_fresh_now
|
| 322 |
+
and lineage != "no_inherit") else None
|
| 323 |
+
bpb = train_member(m, tr, va, steps=steps, device=device,
|
| 324 |
+
seed=mseed, teachers=teachers, kd_alpha=0.25)
|
| 325 |
+
vit = organism_vitals(m)
|
| 326 |
+
members.append((bpb, m))
|
| 327 |
+
rec = {"exp": "14a", "lineage": lineage, "tseed": seed, "gen": gen,
|
| 328 |
+
"member": i, "fresh": gen > 0 and i == pop - 1,
|
| 329 |
+
"catastrophic": catastrophic_at == gen, "bpb": round(bpb, 4),
|
| 330 |
+
"vitals": vit}
|
| 331 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 332 |
+
print(f"[14a {lineage} t{seed} g{gen} m{i}] bpb={bpb:.4f} {vit}",
|
| 333 |
+
flush=True)
|
| 334 |
+
members.sort(key=lambda t: t[0])
|
| 335 |
+
parent_models = [members[0][1].cpu(), members[1][1].cpu()]
|
| 336 |
+
best = members[0][0]
|
| 337 |
+
gene = genome_of(members[0][1])
|
| 338 |
+
if gene:
|
| 339 |
+
torch.save(gene, os.path.join(
|
| 340 |
+
EXP_DIR, f"champion_{lineage}_t{seed}_g{gen}.pt"))
|
| 341 |
+
champion_genomes.append(gene)
|
| 342 |
+
print(f"[14a {lineage} t{seed} g{gen}] BEST={best:.4f} "
|
| 343 |
+
f"mean={sum(b for b, _ in members)/pop:.4f}", flush=True)
|
| 344 |
+
for _, mm in members[2:]:
|
| 345 |
+
del mm
|
| 346 |
+
torch.cuda.empty_cache()
|
| 347 |
+
ledger.close()
|
| 348 |
+
return best
|
| 349 |
+
|
| 350 |
+
|
| 351 |
+
# ============================================================ 14_b implants ======
|
| 352 |
+
def run_b1(steps: int = 2000, seed: int = 0, device: str = "cuda",
|
| 353 |
+
donor_book: torch.Tensor | None = None, tag: str = "small_cultivated"):
|
| 354 |
+
"""Cross-size: donor book (default: cultivate in a small organism) implanted
|
| 355 |
+
into a LARGE organism. Arms: fresh | implant-trainable | implant-frozen | mlp."""
|
| 356 |
+
if not torch.cuda.is_available():
|
| 357 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 358 |
+
os.makedirs(EXP_DIR, exist_ok=True)
|
| 359 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 360 |
+
ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 361 |
+
if donor_book is None:
|
| 362 |
+
small = make_organism("aleph_flat", d=128, layers=2, seed=seed)
|
| 363 |
+
bpb_small = train_member(small, tr, va, steps=steps, device=device, seed=seed)
|
| 364 |
+
donor_book = genome_of(small)["flat"]
|
| 365 |
+
print(f"[14b donor small] bpb={bpb_small:.4f}", flush=True)
|
| 366 |
+
results = {}
|
| 367 |
+
for arm in ("fresh", "implant_train", "implant_frozen", "mlp"):
|
| 368 |
+
lineage = "mlp_kd" if arm == "mlp" else "aleph_flat"
|
| 369 |
+
m = make_organism(lineage, d=384, layers=6, seed=seed + 10)
|
| 370 |
+
if arm.startswith("implant"):
|
| 371 |
+
implant_book(m.head_addr, donor_book, trainable=(arm == "implant_train"))
|
| 372 |
+
bpb = train_member(m, tr, va, steps=steps, device=device, seed=seed + 10)
|
| 373 |
+
vit = organism_vitals(m)
|
| 374 |
+
results[arm] = {"bpb": round(bpb, 4), "vitals": vit}
|
| 375 |
+
rec = {"exp": "14b1", "arm": arm, "donor": tag, "seed": seed,
|
| 376 |
+
"bpb": round(bpb, 4), "vitals": vit}
|
| 377 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 378 |
+
print(f"[14b1 {arm} donor={tag}] bpb={bpb:.4f} {vit}", flush=True)
|
| 379 |
+
del m; torch.cuda.empty_cache()
|
| 380 |
+
ledger.close()
|
| 381 |
+
return results, donor_book
|
| 382 |
+
|
| 383 |
+
|
| 384 |
+
def run_b2(donor_book: torch.Tensor, steps: int = 1500, seed: int = 0,
|
| 385 |
+
device: str = "cuda", tag: str = "small_cultivated"):
|
| 386 |
+
"""Cross-architecture: implant the donor book into every GPT-2 relay adapter
|
| 387 |
+
(exp013 Track C bed) vs random-init relays. Books are (64,4) — size-agnostic."""
|
| 388 |
+
from exp013_augmentation_bed import _wikitext_lines
|
| 389 |
+
from transformers import GPT2LMHeadModel, GPT2TokenizerFast
|
| 390 |
+
if "MslRelay" not in globals():
|
| 391 |
+
from ar_differentiation_bed import MslRelay
|
| 392 |
+
from exp013_augmentation_bed import _BlockWithAdapter
|
| 393 |
+
tok = GPT2TokenizerFast.from_pretrained("gpt2")
|
| 394 |
+
tr_lines, va_lines = _wikitext_lines(DATA_ROOT)
|
| 395 |
+
stream_tr = tok("\n\n".join(tr_lines[:8000]), return_tensors="pt").input_ids[0]
|
| 396 |
+
stream_va = tok("\n\n".join(va_lines[:1000]), return_tensors="pt").input_ids[0]
|
| 397 |
+
ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 398 |
+
out = {}
|
| 399 |
+
for arm in ("random_relays", "implanted_relays"):
|
| 400 |
+
torch.manual_seed(seed)
|
| 401 |
+
g = torch.Generator().manual_seed(seed)
|
| 402 |
+
model = GPT2LMHeadModel.from_pretrained("gpt2").to(device)
|
| 403 |
+
for p in model.parameters():
|
| 404 |
+
p.requires_grad_(False)
|
| 405 |
+
adapters = []
|
| 406 |
+
for i, blk in enumerate(model.transformer.h):
|
| 407 |
+
ad = MslRelay(model.config.n_embd).to(device)
|
| 408 |
+
if arm == "implanted_relays":
|
| 409 |
+
implant_book(ad.addr, donor_book, trainable=True)
|
| 410 |
+
model.transformer.h[i] = _BlockWithAdapter(blk, ad)
|
| 411 |
+
adapters.append(ad)
|
| 412 |
+
params = [p for ad in adapters for p in ad.parameters()
|
| 413 |
+
if p.requires_grad]
|
| 414 |
+
opt = torch.optim.Adam(params, lr=1e-3, weight_decay=0.0)
|
| 415 |
+
block = 256
|
| 416 |
+
for step in range(1, steps + 1):
|
| 417 |
+
ix = torch.randint(0, stream_tr.numel() - block - 1, (8,), generator=g)
|
| 418 |
+
x = torch.stack([stream_tr[i:i + block] for i in ix]).to(device)
|
| 419 |
+
loss = model(x, labels=x).loss
|
| 420 |
+
opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
|
| 421 |
+
model.eval()
|
| 422 |
+
with torch.no_grad():
|
| 423 |
+
ls = []
|
| 424 |
+
for j in range(0, stream_va.numel() - block - 1, block * 4):
|
| 425 |
+
x = stream_va[j:j + block].unsqueeze(0).to(device)
|
| 426 |
+
ls.append(model(x, labels=x).loss.item())
|
| 427 |
+
ppl = math.exp(sum(ls) / len(ls))
|
| 428 |
+
gates = [round(ad.gate.sigmoid().item(), 4) for ad in adapters]
|
| 429 |
+
drifts = [round(anchor_drift(ad.addr.codebook, ad.addr.home)["mean"], 3)
|
| 430 |
+
for ad in adapters]
|
| 431 |
+
out[arm] = {"ppl": round(ppl, 3), "gates": gates, "drift": drifts}
|
| 432 |
+
rec = {"exp": "14b2", "arm": arm, "donor": tag, "seed": seed,
|
| 433 |
+
"ppl": round(ppl, 3), "gates": gates, "drift": drifts}
|
| 434 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 435 |
+
print(f"[14b2 {arm} donor={tag}] ppl={ppl:.3f} gates={gates[:3]}.. "
|
| 436 |
+
f"drift={drifts[:3]}..", flush=True)
|
| 437 |
+
del model; torch.cuda.empty_cache()
|
| 438 |
+
ledger.close()
|
| 439 |
+
return out
|
| 440 |
+
|
| 441 |
+
|
| 442 |
+
# ================================================================ smoke ===========
|
| 443 |
+
def smoke():
|
| 444 |
+
"""CPU shapes/parse only: GPA ground truth, tree causality, implant, KD path."""
|
| 445 |
+
g = torch.Generator().manual_seed(0)
|
| 446 |
+
# GPA: two rotated (+row-sign-flipped) copies of one book must align back to it
|
| 447 |
+
A = F.normalize(torch.randn(64, 4, generator=g), dim=-1)
|
| 448 |
+
q, _ = torch.linalg.qr(torch.randn(4, 4, generator=g))
|
| 449 |
+
B = A @ q
|
| 450 |
+
B[::3] = -B[::3]
|
| 451 |
+
C, it, dl = consensus_codebook([A, B])
|
| 452 |
+
cos = (F.normalize(C, dim=-1) * A).sum(-1).abs().mean()
|
| 453 |
+
assert cos > 0.999, cos
|
| 454 |
+
print(f"GPA OK (iters={it}, delta={dl:.2e}, |cos to truth|={cos:.5f})")
|
| 455 |
+
x = torch.randint(0, 256, (2, 64))
|
| 456 |
+
for lineage in ("aleph_flat", "aleph_tree", "mlp_kd", "no_inherit"):
|
| 457 |
+
m = make_organism(lineage, d=96, layers=2, block=64, seed=0)
|
| 458 |
+
lg = m(x); assert lg.shape == (2, 64, 256); lg.sum().backward()
|
| 459 |
+
with torch.no_grad():
|
| 460 |
+
a = m(x)[0, 10]; x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
|
| 461 |
+
b = m(x2)[0, 10]
|
| 462 |
+
assert torch.allclose(a, b, atol=1e-4), lineage + " leaks"
|
| 463 |
+
gnm = genome_of(m)
|
| 464 |
+
if gnm:
|
| 465 |
+
inherit_genome(m, [gnm, gnm]) # self-consensus = identity-ish
|
| 466 |
+
print(lineage, "OK params",
|
| 467 |
+
f"{sum(p.numel() for p in m.parameters()):,}",
|
| 468 |
+
organism_vitals(m) if lineage != "mlp_kd" else {})
|
| 469 |
+
# KD path: teacher forward + KL backward
|
| 470 |
+
t = make_organism("mlp_kd", d=96, layers=2, block=64, seed=1)
|
| 471 |
+
s = make_organism("aleph_flat", d=96, layers=2, block=64, seed=2)
|
| 472 |
+
tp = F.softmax(t(x), -1).detach()
|
| 473 |
+
loss = F.kl_div(F.log_softmax(s(x), -1), tp, reduction="batchmean")
|
| 474 |
+
loss.backward()
|
| 475 |
+
print("KD OK — exp014 smoke passed (tournament on GPU: run_tournament(...))")
|
| 476 |
+
|
| 477 |
+
|
| 478 |
+
def _in_notebook():
|
| 479 |
+
try:
|
| 480 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 481 |
+
return True
|
| 482 |
+
except NameError:
|
| 483 |
+
return False
|
| 484 |
+
|
| 485 |
+
|
| 486 |
+
if __name__ == "__main__":
|
| 487 |
+
if _in_notebook():
|
| 488 |
+
smoke()
|
| 489 |
+
print("Notebook: run_tournament('aleph_flat'), run_b1(), run_b2(book).")
|
| 490 |
+
else:
|
| 491 |
+
import argparse
|
| 492 |
+
ap = argparse.ArgumentParser()
|
| 493 |
+
ap.add_argument("--mode", default="smoke",
|
| 494 |
+
choices=["smoke", "tournament", "b1", "b2"])
|
| 495 |
+
ap.add_argument("--lineage", default="aleph_full",
|
| 496 |
+
help="tournament lineage: aleph_flat|aleph_full|"
|
| 497 |
+
"aleph_weights|aleph_tree|mlp_kd|no_inherit")
|
| 498 |
+
ap.add_argument("--seed", type=int, default=0)
|
| 499 |
+
ap.add_argument("--steps", type=int, default=2000)
|
| 500 |
+
ap.add_argument("--genome", default="genomes/champion_aleph_full_t0_g3.pt",
|
| 501 |
+
help="donor genome .pt for --mode b1/b2 (uses its 'flat' book)")
|
| 502 |
+
a, _ = ap.parse_known_args()
|
| 503 |
+
if a.mode == "smoke":
|
| 504 |
+
smoke()
|
| 505 |
+
elif a.mode == "tournament":
|
| 506 |
+
run_tournament(a.lineage, steps=a.steps, seed=a.seed)
|
| 507 |
+
elif a.mode == "b1":
|
| 508 |
+
donor = (torch.load(a.genome, map_location="cpu")["flat"]
|
| 509 |
+
if os.path.exists(a.genome) else None)
|
| 510 |
+
run_b1(steps=a.steps, seed=a.seed, donor_book=donor,
|
| 511 |
+
tag=os.path.basename(a.genome) if donor is not None
|
| 512 |
+
else "small_cultivated")
|
| 513 |
+
elif a.mode == "b2":
|
| 514 |
+
donor = torch.load(a.genome, map_location="cpu")["flat"]
|
| 515 |
+
run_b2(donor, seed=a.seed, tag=os.path.basename(a.genome))
|
exp019_cr/exp019_content_retention.py
ADDED
|
@@ -0,0 +1,402 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""exp019_content_retention.py — CAPACITY FOR DISTILLED CONTENT RETENTION.
|
| 2 |
+
When content is DISTILLED rather than directly learned, how much is retained,
|
| 3 |
+
through which channel, and for how long? The memory-substrate compass made a
|
| 4 |
+
measurable instrument — and the weak-to-strong distillation cell made real: the
|
| 5 |
+
teacher holds content the student lacks (true headroom on the content axis).
|
| 6 |
+
|
| 7 |
+
THE CONTENT: N fact records "\n@<key6>=<value12>\n" with random alphanumeric
|
| 8 |
+
keys/values — uncompletable from language statistics, so exact-match completion
|
| 9 |
+
IS retention. Facts are mixed into the byte stream (fact-packed blocks at
|
| 10 |
+
FACT_RATE against wikitext blocks).
|
| 11 |
+
|
| 12 |
+
THE TEACHER: the certified bed model (addr_msl64 aleph substrate) trained
|
| 13 |
+
TEACHER_STEPS on the mix; gated on its own recall (the gate doubles as the
|
| 14 |
+
substrate's DIRECT capacity datum at each N).
|
| 15 |
+
|
| 16 |
+
THE CHANNELS (fresh student each, STUDENT_STEPS budget):
|
| 17 |
+
direct — ground-truth CE on the mix (ceiling: learning, not distillation)
|
| 18 |
+
kd_facts — CE on wikitext blocks; on fact blocks the ONLY signal is the
|
| 19 |
+
teacher's logits (KL) — pure distilled content
|
| 20 |
+
kd_general — CE + KL to teacher on CLEAN wikitext only; facts never shown —
|
| 21 |
+
does content leak through logits without exposure?
|
| 22 |
+
book_implant — teacher's farmed codebook implanted (trainable) into a fresh
|
| 23 |
+
student, clean-stream training — do the anchors carry
|
| 24 |
+
byte-content? (the open question from the exp014 implant studies, asked directly)
|
| 25 |
+
none — clean-stream only (floor)
|
| 26 |
+
THE AXES: capacity N in {64, 256, 1024} (main channels); retention = recall
|
| 27 |
+
right after training AND after INTERFERE_STEPS further clean-stream steps
|
| 28 |
+
(the forgetting measurement). KD alpha 1.0 here is LEGAL: the teacher has real
|
| 29 |
+
headroom on the content axis (the inverse-evolution failure was alpha 1.0 at
|
| 30 |
+
NEAR-PARITY — regime, not constant).
|
| 31 |
+
|
| 32 |
+
Preregistered forks:
|
| 33 |
+
F1 capacity curve: direct recall vs N = the substrate's raw content capacity.
|
| 34 |
+
F2 distillation tax: kd_facts vs direct at each N (what survives the logit
|
| 35 |
+
channel).
|
| 36 |
+
F3 leakage: kd_general recall > floor => content crosses on clean text alone.
|
| 37 |
+
F4 anchors: book_implant recall ~ floor => codebooks do not carry byte
|
| 38 |
+
content (mean-shape/content question closed in the direct sense).
|
| 39 |
+
F5 half-life: post-interference retention per channel (does distilled content
|
| 40 |
+
decay faster than learned content?).
|
| 41 |
+
Riders: pure Adam wd=0; GPU-only verdicts; >=2 seeds; Colab-safe. Paste order:
|
| 42 |
+
geolip_vitals -> ar_differentiation_bed -> exp014_genetic_distillation ->
|
| 43 |
+
this file.
|
| 44 |
+
"""
|
| 45 |
+
from __future__ import annotations
|
| 46 |
+
import json
|
| 47 |
+
import math
|
| 48 |
+
import os
|
| 49 |
+
import string
|
| 50 |
+
import torch
|
| 51 |
+
import torch.nn.functional as F
|
| 52 |
+
|
| 53 |
+
if "ByteLM" not in globals():
|
| 54 |
+
try:
|
| 55 |
+
from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
|
| 56 |
+
from exp014_genetic_distillation import implant_book
|
| 57 |
+
from geolip_vitals import anchor_drift
|
| 58 |
+
except ImportError:
|
| 59 |
+
_here = globals().get("__file__")
|
| 60 |
+
if _here is None:
|
| 61 |
+
raise ImportError("paste geolip_vitals + ar_differentiation_bed + "
|
| 62 |
+
"exp014_genetic_distillation first")
|
| 63 |
+
import sys, pathlib
|
| 64 |
+
sys.path.insert(0, str(pathlib.Path(_here).parent))
|
| 65 |
+
from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
|
| 66 |
+
from exp014_genetic_distillation import implant_book
|
| 67 |
+
from geolip_vitals import anchor_drift
|
| 68 |
+
|
| 69 |
+
DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
|
| 70 |
+
EXP19_DIR = os.path.join(DATA_ROOT, "exp019")
|
| 71 |
+
|
| 72 |
+
KEY_LEN, VAL_LEN = 6, 12
|
| 73 |
+
FACT_RATE = 0.5 # fraction of training blocks drawn from fact stream
|
| 74 |
+
TEACHER_STEPS = 4000
|
| 75 |
+
STUDENT_STEPS = 2000
|
| 76 |
+
INTERFERE_STEPS = 1000
|
| 77 |
+
ALNUM = (string.ascii_lowercase + string.digits).encode()
|
| 78 |
+
|
| 79 |
+
|
| 80 |
+
def make_facts(n: int, seed: int = 0):
|
| 81 |
+
"""N records '\\n@<key>=<value>\\n'; returns (records list, fact byte stream)."""
|
| 82 |
+
g = torch.Generator().manual_seed(4000 + seed)
|
| 83 |
+
recs = []
|
| 84 |
+
seen = set()
|
| 85 |
+
while len(recs) < n:
|
| 86 |
+
k = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (KEY_LEN,),
|
| 87 |
+
generator=g))
|
| 88 |
+
if k in seen:
|
| 89 |
+
continue
|
| 90 |
+
seen.add(k)
|
| 91 |
+
v = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (VAL_LEN,),
|
| 92 |
+
generator=g))
|
| 93 |
+
recs.append((k, v))
|
| 94 |
+
return recs
|
| 95 |
+
|
| 96 |
+
|
| 97 |
+
def fact_stream(recs, copies: int = 50, seed: int = 0) -> torch.Tensor:
|
| 98 |
+
"""Byte stream of shuffled fact records (each record appears `copies` times)."""
|
| 99 |
+
g = torch.Generator().manual_seed(5000 + seed)
|
| 100 |
+
order = torch.cat([torch.randperm(len(recs), generator=g)
|
| 101 |
+
for _ in range(copies)])
|
| 102 |
+
blob = b"".join(b"\n@" + recs[i][0] + b"=" + recs[i][1] + b"\n"
|
| 103 |
+
for i in order.tolist())
|
| 104 |
+
return torch.frombuffer(bytearray(blob), dtype=torch.uint8).clone()
|
| 105 |
+
|
| 106 |
+
|
| 107 |
+
def _mix_batch(tr, fs, batch, block, device, g):
|
| 108 |
+
"""Blocks drawn from the fact stream with prob FACT_RATE, else wikitext.
|
| 109 |
+
Returns (x, y, fact_mask (B,)) — mask marks fact-sourced rows."""
|
| 110 |
+
xw, yw = _batch(tr, batch, block, device, g)
|
| 111 |
+
xf, yf = _batch(fs, batch, block, device, g)
|
| 112 |
+
m = (torch.rand(batch, generator=g) < FACT_RATE).to(device)
|
| 113 |
+
x = torch.where(m[:, None], xf, xw)
|
| 114 |
+
y = torch.where(m[:, None], yf, yw)
|
| 115 |
+
return x, y, m
|
| 116 |
+
|
| 117 |
+
|
| 118 |
+
@torch.no_grad()
|
| 119 |
+
def recall(model, recs, device="cuda", max_eval: int = 256,
|
| 120 |
+
batch: int = 64) -> dict:
|
| 121 |
+
"""Exact-match greedy completion: prompt '\\n@<key>=' -> VAL_LEN bytes."""
|
| 122 |
+
model = model.to(device).eval()
|
| 123 |
+
recs = recs[:max_eval]
|
| 124 |
+
prompts = torch.stack([torch.frombuffer(
|
| 125 |
+
bytearray(b"\n@" + k + b"="), dtype=torch.uint8).long()
|
| 126 |
+
for k, _ in recs]).to(device)
|
| 127 |
+
outs = []
|
| 128 |
+
for i in range(0, len(recs), batch):
|
| 129 |
+
x = prompts[i:i + batch]
|
| 130 |
+
for _ in range(VAL_LEN):
|
| 131 |
+
nxt = model(x)[:, -1].argmax(-1, keepdim=True)
|
| 132 |
+
x = torch.cat([x, nxt], dim=1)
|
| 133 |
+
outs.append(x[:, -VAL_LEN:].cpu())
|
| 134 |
+
got = torch.cat(outs)
|
| 135 |
+
tgt = torch.stack([torch.frombuffer(bytearray(v), dtype=torch.uint8).long()
|
| 136 |
+
for _, v in recs])
|
| 137 |
+
byte_acc = (got == tgt).float().mean().item()
|
| 138 |
+
exact = (got == tgt).all(dim=1).float().mean().item()
|
| 139 |
+
return {"exact": round(exact, 4), "byte_acc": round(byte_acc, 4)}
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
def train_stream(model, tr, va, fs=None, teacher=None, channel="direct",
|
| 143 |
+
steps=2000, batch=32, block=256, device="cuda", seed=0):
|
| 144 |
+
"""One training run under a channel's signal routing (docstring above)."""
|
| 145 |
+
g = torch.Generator().manual_seed(seed)
|
| 146 |
+
model = model.to(device)
|
| 147 |
+
if teacher is not None:
|
| 148 |
+
teacher = teacher.to(device).eval()
|
| 149 |
+
opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
|
| 150 |
+
for step in range(1, steps + 1):
|
| 151 |
+
if channel in ("direct", "kd_facts") and fs is not None:
|
| 152 |
+
x, y, m = _mix_batch(tr, fs, batch, block, device, g)
|
| 153 |
+
else: # clean wikitext stream
|
| 154 |
+
x, y = _batch(tr, batch, block, device, g)
|
| 155 |
+
m = torch.zeros(batch, dtype=torch.bool, device=device)
|
| 156 |
+
logits = model(x)
|
| 157 |
+
if channel == "kd_facts":
|
| 158 |
+
# ground truth on wiki rows only; teacher logits are the ONLY
|
| 159 |
+
# signal on fact rows (pure distilled content)
|
| 160 |
+
ce_rows = ~m
|
| 161 |
+
loss = torch.tensor(0.0, device=device)
|
| 162 |
+
if ce_rows.any():
|
| 163 |
+
loss = F.cross_entropy(logits[ce_rows].reshape(-1, VOCAB),
|
| 164 |
+
y[ce_rows].reshape(-1))
|
| 165 |
+
if m.any():
|
| 166 |
+
with torch.no_grad():
|
| 167 |
+
tp = F.softmax(teacher(x[m]), -1)
|
| 168 |
+
loss = loss + F.kl_div(F.log_softmax(logits[m], -1), tp,
|
| 169 |
+
reduction="batchmean")
|
| 170 |
+
elif channel == "kd_general":
|
| 171 |
+
loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
|
| 172 |
+
with torch.no_grad():
|
| 173 |
+
tp = F.softmax(teacher(x), -1)
|
| 174 |
+
loss = loss + F.kl_div(F.log_softmax(logits, -1), tp,
|
| 175 |
+
reduction="batchmean")
|
| 176 |
+
else: # direct / book_implant / none
|
| 177 |
+
loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
|
| 178 |
+
opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
|
| 179 |
+
model.eval()
|
| 180 |
+
with torch.no_grad():
|
| 181 |
+
ls = []
|
| 182 |
+
for _ in range(10):
|
| 183 |
+
xv, yv = _batch(va, batch, block, device, g)
|
| 184 |
+
ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
|
| 185 |
+
yv.reshape(-1)).item())
|
| 186 |
+
return sum(ls) / len(ls) / math.log(2)
|
| 187 |
+
|
| 188 |
+
|
| 189 |
+
def run_retention(n_facts=(64, 256, 1024), seeds=(0, 1), device="cuda"):
|
| 190 |
+
if not torch.cuda.is_available():
|
| 191 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 192 |
+
os.makedirs(EXP19_DIR, exist_ok=True)
|
| 193 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 194 |
+
ledger = open(os.path.join(EXP19_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 195 |
+
|
| 196 |
+
def log(rec):
|
| 197 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 198 |
+
print(f"[19 {rec['channel']} N={rec['n']} s{rec['seed']}] "
|
| 199 |
+
f"recall={rec['recall']} after_interf={rec.get('recall_interf')} "
|
| 200 |
+
f"bpb={rec['bpb']}", flush=True)
|
| 201 |
+
|
| 202 |
+
for seed in seeds:
|
| 203 |
+
for n in n_facts:
|
| 204 |
+
recs = make_facts(n, seed=seed)
|
| 205 |
+
fs = fact_stream(recs, seed=seed)
|
| 206 |
+
# ---- teacher (also the DIRECT capacity datum at TEACHER_STEPS)
|
| 207 |
+
torch.manual_seed(seed)
|
| 208 |
+
teacher = ByteLM("addr_msl64")
|
| 209 |
+
t_bpb = train_stream(teacher, tr, va, fs=fs, channel="direct",
|
| 210 |
+
steps=TEACHER_STEPS, device=device, seed=seed)
|
| 211 |
+
t_rec = recall(teacher, recs, device=device)
|
| 212 |
+
log({"exp": "19", "channel": "teacher", "n": n, "seed": seed,
|
| 213 |
+
"steps": TEACHER_STEPS, "recall": t_rec, "bpb": round(t_bpb, 4)})
|
| 214 |
+
torch.save({"n": n, "seed": seed,
|
| 215 |
+
"state_dict": {k: v.cpu() for k, v in
|
| 216 |
+
teacher.state_dict().items()}},
|
| 217 |
+
os.path.join(EXP19_DIR, f"teacher_N{n}_s{seed}.pt"))
|
| 218 |
+
# ---- channels
|
| 219 |
+
chans = ["direct", "kd_facts", "kd_general"]
|
| 220 |
+
if n == 256:
|
| 221 |
+
chans += ["book_implant", "none"]
|
| 222 |
+
for ch in chans:
|
| 223 |
+
torch.manual_seed(1000 + seed)
|
| 224 |
+
m = ByteLM("addr_msl64")
|
| 225 |
+
if ch == "book_implant":
|
| 226 |
+
implant_book(m.head_addr,
|
| 227 |
+
teacher.head_addr.codebook.detach().cpu())
|
| 228 |
+
bpb = train_stream(
|
| 229 |
+
m, tr, va, fs=fs if ch in ("direct", "kd_facts") else None,
|
| 230 |
+
teacher=teacher if ch.startswith("kd") else None,
|
| 231 |
+
channel=ch, steps=STUDENT_STEPS, device=device,
|
| 232 |
+
seed=1000 + seed)
|
| 233 |
+
r0 = recall(m, recs, device=device)
|
| 234 |
+
# retention under interference: further CLEAN-stream training
|
| 235 |
+
bpb2 = train_stream(m, tr, va, channel="none",
|
| 236 |
+
steps=INTERFERE_STEPS, device=device,
|
| 237 |
+
seed=2000 + seed)
|
| 238 |
+
r1 = recall(m, recs, device=device)
|
| 239 |
+
log({"exp": "19", "channel": ch, "n": n, "seed": seed,
|
| 240 |
+
"steps": STUDENT_STEPS, "recall": r0, "recall_interf": r1,
|
| 241 |
+
"bpb": round(bpb, 4), "bpb_after_interf": round(bpb2, 4)})
|
| 242 |
+
del m
|
| 243 |
+
torch.cuda.empty_cache()
|
| 244 |
+
del teacher
|
| 245 |
+
torch.cuda.empty_cache()
|
| 246 |
+
ledger.close()
|
| 247 |
+
|
| 248 |
+
|
| 249 |
+
# ==================== exp019b — GENERALIZATION block =========================
|
| 250 |
+
# Rule-bearing content: value = fixed random substitution cipher applied to the
|
| 251 |
+
# key, extended to VAL_LEN (v[i] = subst(k[i % KEY_LEN])). Teacher sees
|
| 252 |
+
# N_TRAIN rule-keys; N_TEST keys are HELD OUT. Held-out recall = the RULE
|
| 253 |
+
# generalizing, not the list. Sharp question: does the logit channel transfer
|
| 254 |
+
# the rule better than it transfers the rote list? Plus prompt-format variants
|
| 255 |
+
# (content vs surface form disentangled).
|
| 256 |
+
|
| 257 |
+
def make_rule_facts(n_train: int = 256, n_test: int = 128, seed: int = 0):
|
| 258 |
+
g = torch.Generator().manual_seed(6000 + seed)
|
| 259 |
+
subst = {ALNUM[i]: ALNUM[j] for i, j in
|
| 260 |
+
enumerate(torch.randperm(len(ALNUM), generator=g).tolist())}
|
| 261 |
+
keys, seen = [], set()
|
| 262 |
+
while len(keys) < n_train + n_test:
|
| 263 |
+
k = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (KEY_LEN,),
|
| 264 |
+
generator=g))
|
| 265 |
+
if k not in seen:
|
| 266 |
+
seen.add(k)
|
| 267 |
+
keys.append(k)
|
| 268 |
+
def val(k):
|
| 269 |
+
return bytes(subst[k[i % KEY_LEN]] for i in range(VAL_LEN))
|
| 270 |
+
train = [(k, val(k)) for k in keys[:n_train]]
|
| 271 |
+
test = [(k, val(k)) for k in keys[n_train:]]
|
| 272 |
+
return train, test
|
| 273 |
+
|
| 274 |
+
|
| 275 |
+
@torch.no_grad()
|
| 276 |
+
def recall_fmt(model, recs, fmt: bytes = b"\n@%s=", device="cuda",
|
| 277 |
+
max_eval: int = 256, batch: int = 64) -> dict:
|
| 278 |
+
"""recall() under an arbitrary prompt format (b'\\n@%s=' = the training
|
| 279 |
+
format; variants probe surface-form generalization)."""
|
| 280 |
+
model = model.to(device).eval()
|
| 281 |
+
recs = recs[:max_eval]
|
| 282 |
+
proms = [torch.frombuffer(bytearray(fmt.replace(b"%s", k)),
|
| 283 |
+
dtype=torch.uint8).long() for k, _ in recs]
|
| 284 |
+
L = max(p.numel() for p in proms)
|
| 285 |
+
# left-pad with newlines to equal length (causal — padding is prefix noise)
|
| 286 |
+
prompts = torch.stack([torch.cat([torch.full((L - p.numel(),), 10,
|
| 287 |
+
dtype=torch.long), p])
|
| 288 |
+
for p in proms]).to(device)
|
| 289 |
+
outs = []
|
| 290 |
+
for i in range(0, len(recs), batch):
|
| 291 |
+
x = prompts[i:i + batch]
|
| 292 |
+
for _ in range(VAL_LEN):
|
| 293 |
+
nxt = model(x)[:, -1].argmax(-1, keepdim=True)
|
| 294 |
+
x = torch.cat([x, nxt], dim=1)
|
| 295 |
+
outs.append(x[:, -VAL_LEN:].cpu())
|
| 296 |
+
got = torch.cat(outs)
|
| 297 |
+
tgt = torch.stack([torch.frombuffer(bytearray(v), dtype=torch.uint8).long()
|
| 298 |
+
for _, v in recs])
|
| 299 |
+
return {"exact": round((got == tgt).all(dim=1).float().mean().item(), 4),
|
| 300 |
+
"byte_acc": round((got == tgt).float().mean().item(), 4)}
|
| 301 |
+
|
| 302 |
+
|
| 303 |
+
FMT_TRAIN = b"\n@%s="
|
| 304 |
+
FMT_VARIANT = b" @%s= " # never seen in training: pure format shift
|
| 305 |
+
|
| 306 |
+
|
| 307 |
+
def run_generalization(n_train: int = 256, n_test: int = 128, seeds=(0, 1),
|
| 308 |
+
device="cuda"):
|
| 309 |
+
if not torch.cuda.is_available():
|
| 310 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 311 |
+
os.makedirs(EXP19_DIR, exist_ok=True)
|
| 312 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 313 |
+
ledger = open(os.path.join(EXP19_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 314 |
+
|
| 315 |
+
def gauges(model, train_recs, test_recs):
|
| 316 |
+
return {"train": recall_fmt(model, train_recs, FMT_TRAIN, device=device),
|
| 317 |
+
"heldout": recall_fmt(model, test_recs, FMT_TRAIN, device=device),
|
| 318 |
+
"train_varfmt": recall_fmt(model, train_recs, FMT_VARIANT,
|
| 319 |
+
device=device)}
|
| 320 |
+
|
| 321 |
+
for seed in seeds:
|
| 322 |
+
train_recs, test_recs = make_rule_facts(n_train, n_test, seed=seed)
|
| 323 |
+
fs = fact_stream(train_recs, seed=seed) # held-out NEVER streamed
|
| 324 |
+
torch.manual_seed(seed)
|
| 325 |
+
teacher = ByteLM("addr_msl64")
|
| 326 |
+
t_bpb = train_stream(teacher, tr, va, fs=fs, channel="direct",
|
| 327 |
+
steps=TEACHER_STEPS, device=device, seed=seed)
|
| 328 |
+
gt = gauges(teacher, train_recs, test_recs)
|
| 329 |
+
rec = {"exp": "19b", "channel": "teacher", "n": n_train, "seed": seed,
|
| 330 |
+
"steps": TEACHER_STEPS, "gauges": gt, "bpb": round(t_bpb, 4)}
|
| 331 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 332 |
+
print(f"[19b teacher s{seed}] {gt} bpb={t_bpb:.4f}", flush=True)
|
| 333 |
+
torch.save({"seed": seed, "state_dict": {k: v.cpu() for k, v in
|
| 334 |
+
teacher.state_dict().items()}},
|
| 335 |
+
os.path.join(EXP19_DIR, f"rule_teacher_s{seed}.pt"))
|
| 336 |
+
for ch in ("direct", "kd_facts", "kd_general"):
|
| 337 |
+
torch.manual_seed(1000 + seed)
|
| 338 |
+
m = ByteLM("addr_msl64")
|
| 339 |
+
bpb = train_stream(
|
| 340 |
+
m, tr, va, fs=fs if ch in ("direct", "kd_facts") else None,
|
| 341 |
+
teacher=teacher if ch.startswith("kd") else None,
|
| 342 |
+
channel=ch, steps=STUDENT_STEPS, device=device, seed=1000 + seed)
|
| 343 |
+
g0 = gauges(m, train_recs, test_recs)
|
| 344 |
+
bpb2 = train_stream(m, tr, va, channel="none",
|
| 345 |
+
steps=INTERFERE_STEPS, device=device,
|
| 346 |
+
seed=2000 + seed)
|
| 347 |
+
g1 = gauges(m, train_recs, test_recs)
|
| 348 |
+
rec = {"exp": "19b", "channel": ch, "n": n_train, "seed": seed,
|
| 349 |
+
"steps": STUDENT_STEPS, "gauges": g0, "gauges_interf": g1,
|
| 350 |
+
"bpb": round(bpb, 4), "bpb_after_interf": round(bpb2, 4)}
|
| 351 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 352 |
+
print(f"[19b {ch} s{seed}] {g0} interf_heldout="
|
| 353 |
+
f"{g1['heldout']} bpb={bpb:.4f}", flush=True)
|
| 354 |
+
torch.save({"channel": ch, "seed": seed,
|
| 355 |
+
"state_dict": {k: v.cpu() for k, v in
|
| 356 |
+
m.state_dict().items()}},
|
| 357 |
+
os.path.join(EXP19_DIR, f"rule_{ch}_s{seed}.pt"))
|
| 358 |
+
del m
|
| 359 |
+
torch.cuda.empty_cache()
|
| 360 |
+
del teacher
|
| 361 |
+
torch.cuda.empty_cache()
|
| 362 |
+
ledger.close()
|
| 363 |
+
|
| 364 |
+
|
| 365 |
+
def smoke():
|
| 366 |
+
recs = make_facts(8, seed=0)
|
| 367 |
+
assert len(recs) == 8 and all(len(k) == KEY_LEN and len(v) == VAL_LEN
|
| 368 |
+
for k, v in recs)
|
| 369 |
+
fs = fact_stream(recs, copies=3, seed=0)
|
| 370 |
+
assert fs.dtype == torch.uint8 and fs.numel() == 3 * 8 * (KEY_LEN + VAL_LEN + 4)
|
| 371 |
+
m = ByteLM("addr_msl64", d=96, layers=2, block=64)
|
| 372 |
+
r = recall(m, recs, device="cpu", max_eval=8, batch=4)
|
| 373 |
+
assert 0.0 <= r["exact"] <= 1.0 and 0.0 <= r["byte_acc"] <= 1.0
|
| 374 |
+
g = torch.Generator().manual_seed(0)
|
| 375 |
+
x, y, mask = _mix_batch(torch.randint(0, 256, (50000,),
|
| 376 |
+
dtype=torch.uint8, generator=g),
|
| 377 |
+
fs, 8, 64, "cpu", g)
|
| 378 |
+
assert x.shape == (8, 64) and mask.shape == (8,)
|
| 379 |
+
assert (x[:, 1:] == y[:, :-1]).all() # stream alignment
|
| 380 |
+
# 19b: rule facts are rule-consistent + disjoint; variant recall runs
|
| 381 |
+
tr8, te4 = make_rule_facts(8, 4, seed=0)
|
| 382 |
+
assert len(tr8) == 8 and len(te4) == 4
|
| 383 |
+
assert not set(k for k, _ in tr8) & set(k for k, _ in te4)
|
| 384 |
+
k0, v0 = tr8[0]
|
| 385 |
+
assert len(v0) == VAL_LEN and v0[:KEY_LEN] == v0[KEY_LEN:2 * KEY_LEN]
|
| 386 |
+
rv = recall_fmt(m, tr8, FMT_VARIANT, device="cpu", max_eval=8, batch=4)
|
| 387 |
+
assert 0.0 <= rv["exact"] <= 1.0
|
| 388 |
+
print(f"exp019 smoke passed (untrained recall exact={r['exact']} "
|
| 389 |
+
f"byte={r['byte_acc']} ~ chance; 19b rule+variant OK)")
|
| 390 |
+
|
| 391 |
+
|
| 392 |
+
def _in_notebook():
|
| 393 |
+
try:
|
| 394 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 395 |
+
return True
|
| 396 |
+
except NameError:
|
| 397 |
+
return False
|
| 398 |
+
|
| 399 |
+
|
| 400 |
+
if __name__ == "__main__":
|
| 401 |
+
smoke() if not _in_notebook() else (smoke(),
|
| 402 |
+
print("Notebook: run_retention() on GPU."))
|
exp019_cr/geolip_vitals.py
ADDED
|
@@ -0,0 +1,219 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""geolip_vitals.py — shared diagnostic harness for the GeoLIP aleph experiments.
|
| 2 |
+
ALL functions are READOUTS: no gradients, no losses. CV is a readout, never a
|
| 3 |
+
force. Addressing is judged by drift->0.29154 and CV->0.20, never by recon cosine
|
| 4 |
+
(judgment criteria per the aleph-void article: https://huggingface.co/blog/AbstractPhil/geometric-vocabulary-patchwork-aleph-void).
|
| 5 |
+
|
| 6 |
+
Vitals provided:
|
| 7 |
+
anchor_drift — geodesic drift of anchors from init; binding fraction @0.29154
|
| 8 |
+
pentachoron_cv — CM 4-volume CV over random 5-row subsets (geovocab2 import)
|
| 9 |
+
axis_aliveness — oriented-address usage: axes alive, hppl, collapse flag
|
| 10 |
+
gate_stats — gate means vs the 0.012-0.03 band
|
| 11 |
+
path_diversity — unique-path counting, FIXED high-bits hash (low-16 bug is the
|
| 12 |
+
retracted artifact — never use the low bits)
|
| 13 |
+
grad_norm_spread — gradient democracy monitor (orders-of-magnitude spread)
|
| 14 |
+
CVScreen — CV@1000-batch early band screen (<0.30 LOW / .35-.50 MID / >.80 HIGH)
|
| 15 |
+
|
| 16 |
+
Smoke on a torch-capable env: python geolip_vitals.py
|
| 17 |
+
"""
|
| 18 |
+
from __future__ import annotations
|
| 19 |
+
import math
|
| 20 |
+
import torch
|
| 21 |
+
|
| 22 |
+
BINDING = 0.29154 # radians; the binding/separation constant
|
| 23 |
+
CV_BAND = (0.13, 0.30) # CM CV band (discovery_catalog #4)
|
| 24 |
+
GATE_BAND = (0.012, 0.03) # live invariant candidate (acd_campaign)
|
| 25 |
+
KNUTH32 = 2654435761
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
# ----------------------------------------------------------------------------- drift
|
| 29 |
+
@torch.no_grad()
|
| 30 |
+
def anchor_drift(current: torch.Tensor, init: torch.Tensor, tol: float = 0.05) -> dict:
|
| 31 |
+
"""Geodesic drift (radians) of each row of `current` from its row in `init`,
|
| 32 |
+
both row-normalized. Returns mean/std/per-row drift and the fraction of rows
|
| 33 |
+
within +/-tol of BINDING (the GLFM '46%' readout)."""
|
| 34 |
+
a = torch.nn.functional.normalize(current.float(), dim=-1)
|
| 35 |
+
b = torch.nn.functional.normalize(init.float(), dim=-1)
|
| 36 |
+
cos = (a * b).sum(-1).clamp(-1.0, 1.0)
|
| 37 |
+
drift = torch.arccos(cos)
|
| 38 |
+
frac = ((drift - BINDING).abs() <= tol).float().mean()
|
| 39 |
+
return {"mean": drift.mean().item(), "std": drift.std().item(),
|
| 40 |
+
"per_row": drift, "binding_fraction": frac.item()}
|
| 41 |
+
|
| 42 |
+
|
| 43 |
+
# -------------------------------------------------------------------------------- cv
|
| 44 |
+
@torch.no_grad()
|
| 45 |
+
def _pentachoron_volumes(pts: torch.Tensor) -> torch.Tensor:
|
| 46 |
+
"""Batched Cayley-Menger 4-simplex volumes. pts: (B, 5, D) -> (B,) volumes.
|
| 47 |
+
One float64 det over all samples (vol^2 = -det(CM)/9216 for n=4). Built-in
|
| 48 |
+
for speed (the per-sample reference path is ~260x slower in a vitals loop);
|
| 49 |
+
geovocab2 remains the formula's reference implementation, parity-checked
|
| 50 |
+
via cv_reference_check()."""
|
| 51 |
+
B = pts.shape[0]
|
| 52 |
+
d2 = torch.cdist(pts.double(), pts.double()).pow(2) # (B,5,5)
|
| 53 |
+
cm = torch.ones(B, 6, 6, dtype=torch.float64, device=pts.device)
|
| 54 |
+
cm[:, 0, 0] = 0.0
|
| 55 |
+
cm[:, 1:, 1:] = d2
|
| 56 |
+
det = torch.linalg.det(cm)
|
| 57 |
+
return (-det / 9216.0).clamp_min(0.0).sqrt().float()
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
@torch.no_grad()
|
| 61 |
+
def pentachoron_cv(rows: torch.Tensor, n_samples: int = 200,
|
| 62 |
+
generator: torch.Generator | None = None) -> float:
|
| 63 |
+
"""CV (std/mean) of Cayley-Menger 4-simplex volumes over n_samples random
|
| 64 |
+
5-row subsets. Rows are row-normalized before measurement. Uses the built-in
|
| 65 |
+
batched CM (float64 det); validate against geovocab2 with
|
| 66 |
+
cv_reference_check() after any change to the volume math."""
|
| 67 |
+
x = torch.nn.functional.normalize(rows.float(), dim=-1)
|
| 68 |
+
n = x.shape[0]
|
| 69 |
+
if n < 5:
|
| 70 |
+
raise ValueError(f"pentachoron_cv needs >=5 rows, got {n}")
|
| 71 |
+
g = generator or torch.Generator(device="cpu").manual_seed(0)
|
| 72 |
+
idx = torch.stack([torch.randperm(n, generator=g)[:5]
|
| 73 |
+
for _ in range(n_samples)]) # (B,5)
|
| 74 |
+
v = _pentachoron_volumes(x[idx].cpu())
|
| 75 |
+
return (v.std() / v.mean().clamp_min(1e-12)).item()
|
| 76 |
+
|
| 77 |
+
|
| 78 |
+
@torch.no_grad()
|
| 79 |
+
def cv_reference_check(n_trials: int = 50, tol: float = 1e-5) -> float:
|
| 80 |
+
"""Parity check of the built-in batched CM against geovocab2's reference
|
| 81 |
+
implementation (the formula's source of truth). Returns max |rel diff|;
|
| 82 |
+
raises if geovocab2 is absent or parity fails. Run after touching
|
| 83 |
+
_pentachoron_volumes."""
|
| 84 |
+
try:
|
| 85 |
+
from geovocab2.shapes.formula.symbolic.cayley_menger import (
|
| 86 |
+
CayleyMengerFromSimplex)
|
| 87 |
+
except Exception as e: # pragma: no cover
|
| 88 |
+
raise ImportError(
|
| 89 |
+
"cv_reference_check requires geovocab2 (install via the geolip-svae "
|
| 90 |
+
"umbrella: pip install git+https://github.com/AbstractEyes/"
|
| 91 |
+
"geolip-svae).") from e
|
| 92 |
+
ref = CayleyMengerFromSimplex()
|
| 93 |
+
g = torch.Generator().manual_seed(0)
|
| 94 |
+
pts = torch.nn.functional.normalize(
|
| 95 |
+
torch.randn(n_trials, 5, 4, generator=g), dim=-1)
|
| 96 |
+
mine = _pentachoron_volumes(pts)
|
| 97 |
+
# compare at float64: the reference computes in the INPUT dtype, and fp32
|
| 98 |
+
# dets lose up to ~4% on near-degenerate pentachora (measured 2026-07-11)
|
| 99 |
+
theirs = torch.stack([ref.forward(p.double())["volume"].float() for p in pts])
|
| 100 |
+
rel = ((mine - theirs).abs() / theirs.abs().clamp_min(1e-12)).max().item()
|
| 101 |
+
if rel > tol:
|
| 102 |
+
raise AssertionError(f"CM parity vs geovocab2 failed: max rel {rel}")
|
| 103 |
+
return rel
|
| 104 |
+
|
| 105 |
+
|
| 106 |
+
# ------------------------------------------------------------------------- aliveness
|
| 107 |
+
@torch.no_grad()
|
| 108 |
+
def axis_aliveness(oriented_weights: torch.Tensor, alive_thresh: float = 1e-3) -> dict:
|
| 109 |
+
"""`oriented_weights`: (..., 2K) nonnegative oriented-softmax address rows
|
| 110 |
+
(sum to 1 on the last dim). Returns axes-alive count, mean-usage perplexity
|
| 111 |
+
(hppl analogue; healthy hosted reference 125-126/128), and a collapse flag.
|
| 112 |
+
Reference behavior: near-uniform aliveness at div_weight=0 (discovery #22)."""
|
| 113 |
+
w = oriented_weights.reshape(-1, oriented_weights.shape[-1]).float()
|
| 114 |
+
usage = w.mean(0)
|
| 115 |
+
usage = usage / usage.sum().clamp_min(1e-12)
|
| 116 |
+
# an axis is alive if its mean usage exceeds alive_thresh x the uniform share
|
| 117 |
+
alive = int((usage > alive_thresh * (1.0 / usage.numel())).sum())
|
| 118 |
+
ent = -(usage.clamp_min(1e-12) * usage.clamp_min(1e-12).log()).sum()
|
| 119 |
+
ppl = float(ent.exp())
|
| 120 |
+
return {"axes_total": usage.numel(), "axes_alive": alive, "usage_ppl": ppl,
|
| 121 |
+
"collapsed": ppl < 0.05 * usage.numel()}
|
| 122 |
+
|
| 123 |
+
|
| 124 |
+
# ------------------------------------------------------------------------------ gates
|
| 125 |
+
@torch.no_grad()
|
| 126 |
+
def gate_stats(gates: torch.Tensor) -> dict:
|
| 127 |
+
"""Gate values (post-sigmoid/clamp). Reports mean and whether it sits in the
|
| 128 |
+
0.012-0.03 band (read-only — the band is a candidate invariant, never a target)."""
|
| 129 |
+
g = gates.float().flatten()
|
| 130 |
+
m = g.mean().item()
|
| 131 |
+
return {"mean": m, "std": g.std().item(),
|
| 132 |
+
"in_band": GATE_BAND[0] <= m <= GATE_BAND[1]}
|
| 133 |
+
|
| 134 |
+
|
| 135 |
+
# ------------------------------------------------------------------------------ paths
|
| 136 |
+
@torch.no_grad()
|
| 137 |
+
def path_diversity(ids: torch.Tensor) -> dict:
|
| 138 |
+
"""Unique-path counting with the FIXED multiplicative hash:
|
| 139 |
+
((ids * 2654435761) % 2^32) >> 16 — Knuth needs the HIGH bits; the low-16
|
| 140 |
+
variant produced a retracted ~1,500 path ceiling in a prior campaign.
|
| 141 |
+
`ids`: integer tensor, one composed path id per row (any shape)."""
|
| 142 |
+
x = ids.reshape(-1).to(torch.int64)
|
| 143 |
+
hashed = ((x * KNUTH32) % (1 << 32)) >> 16
|
| 144 |
+
return {"n": int(x.numel()),
|
| 145 |
+
"unique_raw": int(torch.unique(x).numel()),
|
| 146 |
+
"unique_hashed": int(torch.unique(hashed).numel())}
|
| 147 |
+
|
| 148 |
+
|
| 149 |
+
@torch.no_grad()
|
| 150 |
+
def compose_path_ids(stage_indices: list[torch.Tensor], radix: int) -> torch.Tensor:
|
| 151 |
+
"""Compose per-stage discrete indices (each (...,) int in [0, radix)) into a
|
| 152 |
+
single path id, positional base-`radix` — construction, not hashing."""
|
| 153 |
+
out = torch.zeros_like(stage_indices[0], dtype=torch.int64)
|
| 154 |
+
for s in stage_indices:
|
| 155 |
+
out = out * radix + s.to(torch.int64)
|
| 156 |
+
return out
|
| 157 |
+
|
| 158 |
+
|
| 159 |
+
# --------------------------------------------------------------------- grad democracy
|
| 160 |
+
@torch.no_grad()
|
| 161 |
+
def grad_norm_spread(groups: dict[str, list[torch.nn.Parameter]]) -> dict:
|
| 162 |
+
"""Gradient-democracy monitor. `groups`: name -> params of one parallel member
|
| 163 |
+
(tower/expert). Reports per-group grad norms and the orders-of-magnitude spread.
|
| 164 |
+
Reference: unequalized heterogeneous towers spread ~20 orders (fibonacci dead at
|
| 165 |
+
2.25e-21 under helix); equalized ~0.0 (geofractal gradient-democracy result)."""
|
| 166 |
+
norms = {}
|
| 167 |
+
for name, params in groups.items():
|
| 168 |
+
gs = [p.grad for p in params if p.grad is not None]
|
| 169 |
+
norms[name] = float(torch.sqrt(sum((g.float() ** 2).sum() for g in gs)).item()) \
|
| 170 |
+
if gs else 0.0
|
| 171 |
+
vals = [v for v in norms.values() if v > 0]
|
| 172 |
+
spread = (math.log10(max(vals)) - math.log10(min(vals))) if len(vals) >= 2 else 0.0
|
| 173 |
+
return {"norms": norms, "spread_orders": spread, "dead": [k for k, v in norms.items() if v == 0.0]}
|
| 174 |
+
|
| 175 |
+
|
| 176 |
+
# ----------------------------------------------------------------------------- screen
|
| 177 |
+
class CVScreen:
|
| 178 |
+
"""CV@N early band screen (tri-band ft1): record pentachoron CV at `step_mark`
|
| 179 |
+
batches; classify <0.30 LOW / 0.35-0.50 MID / >0.80 HIGH. Turns ~2h/config
|
| 180 |
+
into ~7min. Readout only."""
|
| 181 |
+
def __init__(self, step_mark: int = 1000):
|
| 182 |
+
self.step_mark = step_mark
|
| 183 |
+
self.recorded: float | None = None
|
| 184 |
+
|
| 185 |
+
def maybe_record(self, step: int, rows: torch.Tensor) -> float | None:
|
| 186 |
+
if self.recorded is None and step >= self.step_mark:
|
| 187 |
+
self.recorded = pentachoron_cv(rows)
|
| 188 |
+
return self.recorded
|
| 189 |
+
|
| 190 |
+
@property
|
| 191 |
+
def band(self) -> str | None:
|
| 192 |
+
c = self.recorded
|
| 193 |
+
if c is None:
|
| 194 |
+
return None
|
| 195 |
+
if c < 0.30:
|
| 196 |
+
return "LOW"
|
| 197 |
+
if 0.35 <= c <= 0.50:
|
| 198 |
+
return "MID"
|
| 199 |
+
if c > 0.80:
|
| 200 |
+
return "HIGH"
|
| 201 |
+
return "BETWEEN"
|
| 202 |
+
|
| 203 |
+
|
| 204 |
+
# ------------------------------------------------------------------------------ smoke
|
| 205 |
+
if __name__ == "__main__": # shapes/parse smoke ONLY — no training, ever.
|
| 206 |
+
g = torch.Generator().manual_seed(0)
|
| 207 |
+
K, D = 64, 4
|
| 208 |
+
init = torch.nn.functional.normalize(torch.randn(K, D, generator=g), dim=-1)
|
| 209 |
+
cur = torch.nn.functional.normalize(init + 0.29 * torch.randn(K, D, generator=g), dim=-1)
|
| 210 |
+
print("drift:", {k: v for k, v in anchor_drift(cur, init).items() if k != "per_row"})
|
| 211 |
+
w = torch.softmax(torch.randn(32, 2 * K, generator=g), dim=-1)
|
| 212 |
+
print("aliveness:", axis_aliveness(w))
|
| 213 |
+
print("gates:", gate_stats(torch.full((8,), 0.024)))
|
| 214 |
+
ids = compose_path_ids([torch.randint(0, 16, (4096,), generator=g) for _ in range(4)], 16)
|
| 215 |
+
print("paths:", path_diversity(ids))
|
| 216 |
+
lin = torch.nn.Linear(8, 8)
|
| 217 |
+
lin(torch.randn(4, 8)).sum().backward()
|
| 218 |
+
print("democracy:", grad_norm_spread({"a": list(lin.parameters())}))
|
| 219 |
+
print("OK — vitals smoke passed (pentachoron_cv needs geovocab2; run on GPU env)")
|
exp019_cr/read_codebook.py
ADDED
|
@@ -0,0 +1,159 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""read_codebook.py — projective reading of cultivated aleph codebooks.
|
| 2 |
+
Antipodal-collapse extraction on trained codebooks + projective statistics on
|
| 3 |
+
RP^(D-1) — applied to the exp012 AR-bed specimens.
|
| 4 |
+
|
| 5 |
+
Recipe per the Polygonal Omega article (geometric-tri-band-ft2): collapse = (row_i - row_j)/2 normalized for each MUTUAL-STRONGEST
|
| 6 |
+
pair with cos < -0.9 — "a deterministic tensor operation," not clustering.
|
| 7 |
+
Projective metric ALWAYS arccos|<a,b>| (metric-alignment rule, reading-voids-ft1).
|
| 8 |
+
D=4 scope is the validated regime (D=5 walked back; axis count grows with D).
|
| 9 |
+
|
| 10 |
+
Readouts per specimen:
|
| 11 |
+
pairs / n_axes / unpaired — antipodal structure
|
| 12 |
+
proj_angle mean vs uniform baseline, deviation — near-uniform RP^(D-1)?
|
| 13 |
+
drift from home + binding fraction @0.29154 — cultivation record
|
| 14 |
+
erank of the axis set — spectral occupancy
|
| 15 |
+
verdict: PROJECTIVE-CLEAN (|dev|<0.05, util>0.95, secondary pairs<=3) /
|
| 16 |
+
-MOSTLY / STRUCTURED / DEGENERATE (per Polygonal Omega thresholds)
|
| 17 |
+
|
| 18 |
+
Usage (terminal): python read_codebook.py <ckpt_or_dir> [more paths...]
|
| 19 |
+
Colab: paste geolip_vitals.py cell first (optional), then this file, then
|
| 20 |
+
read_all(r"/content/data/ar_ckpts").
|
| 21 |
+
"""
|
| 22 |
+
from __future__ import annotations
|
| 23 |
+
import math
|
| 24 |
+
import sys
|
| 25 |
+
import torch
|
| 26 |
+
import torch.nn.functional as F
|
| 27 |
+
|
| 28 |
+
BINDING = 0.29154
|
| 29 |
+
|
| 30 |
+
|
| 31 |
+
@torch.no_grad()
|
| 32 |
+
def antipodal_collapse(codebook: torch.Tensor, thresh: float = -0.9) -> dict:
|
| 33 |
+
"""Mutual-strongest antipodal pairing + collapse to axes on RP^(D-1)."""
|
| 34 |
+
A = F.normalize(codebook.float(), dim=-1)
|
| 35 |
+
K = A.shape[0]
|
| 36 |
+
cos = A @ A.T
|
| 37 |
+
cos.fill_diagonal_(2.0) # exclude self from minima
|
| 38 |
+
nearest_neg = cos.argmin(dim=-1) # most-antipodal partner
|
| 39 |
+
pairs = []
|
| 40 |
+
used = set()
|
| 41 |
+
for i in range(K):
|
| 42 |
+
j = int(nearest_neg[i])
|
| 43 |
+
if i < j and int(nearest_neg[j]) == i and cos[i, j] < thresh:
|
| 44 |
+
pairs.append((i, j))
|
| 45 |
+
used.update((i, j))
|
| 46 |
+
axes = [F.normalize((A[i] - A[j]) / 2.0, dim=-1) for i, j in pairs]
|
| 47 |
+
axes += [A[i] for i in range(K) if i not in used] # unpaired rows as axes
|
| 48 |
+
axes = torch.stack(axes) if axes else A[:0]
|
| 49 |
+
# sign-canon onto RP: first nonzero coordinate positive
|
| 50 |
+
for r in range(axes.shape[0]):
|
| 51 |
+
nz = torch.nonzero(axes[r].abs() > 1e-8)
|
| 52 |
+
if nz.numel() and axes[r, nz[0, 0]] < 0:
|
| 53 |
+
axes[r] = -axes[r]
|
| 54 |
+
return {"pairs": len(pairs), "n_axes": axes.shape[0],
|
| 55 |
+
"unpaired": K - 2 * len(pairs), "axes": axes}
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
@torch.no_grad()
|
| 59 |
+
def projective_stats(axes: torch.Tensor, n_baseline: int = 20000,
|
| 60 |
+
seed: int = 0) -> dict:
|
| 61 |
+
"""Mean projective angle arccos|<a,b>| vs a uniform-RP baseline at same (n, D)."""
|
| 62 |
+
n, D = axes.shape
|
| 63 |
+
if n < 2:
|
| 64 |
+
return {"proj_angle_mean": None, "uniform_baseline": None,
|
| 65 |
+
"deviation": None, "erank": None}
|
| 66 |
+
def mean_angle(rows):
|
| 67 |
+
c = (rows @ rows.T).abs().clamp(max=1.0)
|
| 68 |
+
iu = torch.triu_indices(rows.shape[0], rows.shape[0], offset=1)
|
| 69 |
+
return torch.arccos(c[iu[0], iu[1]]).mean().item()
|
| 70 |
+
obs = mean_angle(axes)
|
| 71 |
+
g = torch.Generator().manual_seed(seed)
|
| 72 |
+
base_angles = []
|
| 73 |
+
m = max(2, n)
|
| 74 |
+
for _ in range(max(1, n_baseline // max(1, m * (m - 1) // 2))):
|
| 75 |
+
r = F.normalize(torch.randn(m, D, generator=g), dim=-1)
|
| 76 |
+
base_angles.append(mean_angle(r))
|
| 77 |
+
base = sum(base_angles) / len(base_angles)
|
| 78 |
+
s = torch.linalg.svdvals(axes)
|
| 79 |
+
p = (s / s.sum().clamp_min(1e-12))
|
| 80 |
+
erank = float(torch.exp(-(p.clamp_min(1e-12) * p.clamp_min(1e-12).log()).sum()))
|
| 81 |
+
return {"proj_angle_mean": round(obs, 4), "uniform_baseline": round(base, 4),
|
| 82 |
+
"deviation": round(obs - base, 4), "erank": round(erank, 3)}
|
| 83 |
+
|
| 84 |
+
|
| 85 |
+
@torch.no_grad()
|
| 86 |
+
def read_specimen(path: str) -> dict:
|
| 87 |
+
ck = torch.load(path, map_location="cpu", weights_only=True)
|
| 88 |
+
out = {"file": path.split("\\")[-1].split("/")[-1],
|
| 89 |
+
"arm": ck.get("arm"), "seed": ck.get("seed"),
|
| 90 |
+
"steps": ck.get("steps"), "val_bpb": round(ck.get("val_bpb", -1), 4)}
|
| 91 |
+
if "state_dict" in ck: # full specimen checkpoint
|
| 92 |
+
sd = ck["state_dict"]
|
| 93 |
+
books = {k[:-len(".codebook")]: sd[k] for k in sd
|
| 94 |
+
if k.endswith("addr.codebook") or k.endswith("head_addr.codebook")}
|
| 95 |
+
homes = {k[:-len(".home")]: sd[k] for k in sd if k.endswith(".home")}
|
| 96 |
+
else: # bare genome dict (exp014+ champion files):
|
| 97 |
+
# books under flat/root/branch* keys; *_proj entries are projections
|
| 98 |
+
books = {k: v for k, v in ck.items()
|
| 99 |
+
if torch.is_tensor(v) and v.ndim == 2
|
| 100 |
+
and (k in ("flat", "root") or k.startswith("branch"))}
|
| 101 |
+
homes = {}
|
| 102 |
+
reads = {}
|
| 103 |
+
for name, cb in books.items():
|
| 104 |
+
col = antipodal_collapse(cb)
|
| 105 |
+
stats = projective_stats(col["axes"])
|
| 106 |
+
home = homes.get(name)
|
| 107 |
+
drift = None
|
| 108 |
+
binding = None
|
| 109 |
+
if home is not None and home.shape == cb.shape:
|
| 110 |
+
a = F.normalize(cb.float(), dim=-1)
|
| 111 |
+
b = F.normalize(home.float(), dim=-1)
|
| 112 |
+
dr = torch.arccos((a * b).sum(-1).clamp(-1, 1))
|
| 113 |
+
drift = round(dr.mean().item(), 4)
|
| 114 |
+
binding = round(((dr - BINDING).abs() <= 0.05).float().mean().item(), 4)
|
| 115 |
+
util = col["n_axes"] / cb.shape[0]
|
| 116 |
+
dev = stats["deviation"]
|
| 117 |
+
if dev is not None and abs(dev) < 0.05 and util > 0.95 and col["pairs"] <= 3:
|
| 118 |
+
verdict = "PROJECTIVE-CLEAN"
|
| 119 |
+
elif dev is not None and abs(dev) < 0.05:
|
| 120 |
+
verdict = "PROJECTIVE-MOSTLY"
|
| 121 |
+
elif dev is not None and dev > 0.05:
|
| 122 |
+
verdict = "STRUCTURED(repulsive)"
|
| 123 |
+
else:
|
| 124 |
+
verdict = "DEGENERATE/CLUMPED" if dev is not None else "TOO-FEW-AXES"
|
| 125 |
+
reads[name] = {
|
| 126 |
+
"pairs": col["pairs"], "n_axes": col["n_axes"], **stats,
|
| 127 |
+
"drift": drift, "binding_frac": binding, "verdict": verdict}
|
| 128 |
+
out["codebooks"] = reads
|
| 129 |
+
return out
|
| 130 |
+
|
| 131 |
+
|
| 132 |
+
def read_all(root: str) -> list:
|
| 133 |
+
import glob, os
|
| 134 |
+
results = []
|
| 135 |
+
for p in sorted(glob.glob(os.path.join(root, "*.pt"))):
|
| 136 |
+
r = read_specimen(p)
|
| 137 |
+
print(r, flush=True)
|
| 138 |
+
results.append(r)
|
| 139 |
+
return results
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
def _in_notebook() -> bool:
|
| 143 |
+
try:
|
| 144 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 145 |
+
return True
|
| 146 |
+
except NameError:
|
| 147 |
+
return False
|
| 148 |
+
|
| 149 |
+
|
| 150 |
+
if __name__ == "__main__":
|
| 151 |
+
if _in_notebook():
|
| 152 |
+
print("Notebook mode: call read_all(r'<data_root>/ar_ckpts') in the next cell.")
|
| 153 |
+
else:
|
| 154 |
+
args = [a for a in sys.argv[1:] if not a.startswith("-")]
|
| 155 |
+
if not args:
|
| 156 |
+
print("usage: python read_codebook.py <ckpt_or_dir> [...]")
|
| 157 |
+
for a in args:
|
| 158 |
+
import os
|
| 159 |
+
read_all(a) if os.path.isdir(a) else print(read_specimen(a))
|
exp019_cr/repro.py
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""repro.py — standalone loader/runner for exp019_cr. Code dependencies live
|
| 2 |
+
in THIS folder (geolip_vitals.py, ar_differentiation_bed.py,
|
| 3 |
+
exp014_genetic_distillation.py, exp019_content_retention.py, read_codebook.py).
|
| 4 |
+
|
| 5 |
+
python repro.py # CPU smoke: facts, streams, recall, rule block
|
| 6 |
+
python repro.py --run # retention sweep (2 seeds x 3 N x channels, ~3h GPU)
|
| 7 |
+
python repro.py --run19b # generalization block (rule content, ~1h GPU)
|
| 8 |
+
|
| 9 |
+
Data lands in ./data (override with GEOLIP_DATA); ledger + checkpoints in
|
| 10 |
+
./data/exp019.
|
| 11 |
+
"""
|
| 12 |
+
import os
|
| 13 |
+
import sys
|
| 14 |
+
|
| 15 |
+
HERE = os.path.dirname(os.path.abspath(__file__))
|
| 16 |
+
sys.path.insert(0, HERE)
|
| 17 |
+
|
| 18 |
+
if __name__ == "__main__":
|
| 19 |
+
import geolip_vitals # noqa: F401 (paste order)
|
| 20 |
+
import ar_differentiation_bed # noqa: F401
|
| 21 |
+
import exp014_genetic_distillation # noqa: F401
|
| 22 |
+
import exp019_content_retention as cr
|
| 23 |
+
if "--run" in sys.argv[1:]:
|
| 24 |
+
cr.run_retention()
|
| 25 |
+
elif "--run19b" in sys.argv[1:]:
|
| 26 |
+
cr.run_generalization()
|
| 27 |
+
else:
|
| 28 |
+
cr.smoke()
|
| 29 |
+
print("repro smoke passed — --run (retention) / --run19b (rule block)")
|
exp019_cr/results/ledger.jsonl
ADDED
|
@@ -0,0 +1,36 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"exp": "19", "channel": "teacher", "n": 64, "seed": 0, "steps": 4000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "bpb": 2.3987}
|
| 2 |
+
{"exp": "19", "channel": "direct", "n": 64, "seed": 0, "steps": 2000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0221}, "bpb": 2.7867, "bpb_after_interf": 2.3879}
|
| 3 |
+
{"exp": "19", "channel": "kd_facts", "n": 64, "seed": 0, "steps": 2000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0013}, "bpb": 4.0938, "bpb_after_interf": 2.5318}
|
| 4 |
+
{"exp": "19", "channel": "kd_general", "n": 64, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0065}, "bpb": 2.5001, "bpb_after_interf": 2.2824}
|
| 5 |
+
{"exp": "19", "channel": "teacher", "n": 256, "seed": 0, "steps": 4000, "recall": {"exact": 0.9883, "byte_acc": 0.9938}, "bpb": 2.4127}
|
| 6 |
+
{"exp": "19", "channel": "direct", "n": 256, "seed": 0, "steps": 2000, "recall": {"exact": 0.8711, "byte_acc": 0.9095}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0072}, "bpb": 2.8685, "bpb_after_interf": 2.3991}
|
| 7 |
+
{"exp": "19", "channel": "kd_facts", "n": 256, "seed": 0, "steps": 2000, "recall": {"exact": 0.9531, "byte_acc": 0.9694}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0186}, "bpb": 4.1136, "bpb_after_interf": 2.5174}
|
| 8 |
+
{"exp": "19", "channel": "kd_general", "n": 256, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0091}, "bpb": 2.5322, "bpb_after_interf": 2.3352}
|
| 9 |
+
{"exp": "19", "channel": "book_implant", "n": 256, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0085}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0075}, "bpb": 2.4608, "bpb_after_interf": 2.2463}
|
| 10 |
+
{"exp": "19", "channel": "none", "n": 256, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0101}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0078}, "bpb": 2.4685, "bpb_after_interf": 2.2399}
|
| 11 |
+
{"exp": "19", "channel": "teacher", "n": 1024, "seed": 0, "steps": 4000, "recall": {"exact": 0.0117, "byte_acc": 0.0693}, "bpb": 2.5179}
|
| 12 |
+
{"exp": "19", "channel": "direct", "n": 1024, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0319}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0065}, "bpb": 3.0269, "bpb_after_interf": 2.489}
|
| 13 |
+
{"exp": "19", "channel": "kd_facts", "n": 1024, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0352}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0199}, "bpb": 4.1922, "bpb_after_interf": 2.6121}
|
| 14 |
+
{"exp": "19", "channel": "kd_general", "n": 1024, "seed": 0, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0013}, "bpb": 2.5957, "bpb_after_interf": 2.3458}
|
| 15 |
+
{"exp": "19", "channel": "teacher", "n": 64, "seed": 1, "steps": 4000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "bpb": 2.3943}
|
| 16 |
+
{"exp": "19", "channel": "direct", "n": 64, "seed": 1, "steps": 2000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.013}, "bpb": 2.752, "bpb_after_interf": 2.3639}
|
| 17 |
+
{"exp": "19", "channel": "kd_facts", "n": 64, "seed": 1, "steps": 2000, "recall": {"exact": 1.0, "byte_acc": 1.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0052}, "bpb": 3.9946, "bpb_after_interf": 2.5456}
|
| 18 |
+
{"exp": "19", "channel": "kd_general", "n": 64, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0}, "bpb": 2.5233, "bpb_after_interf": 2.3492}
|
| 19 |
+
{"exp": "19", "channel": "teacher", "n": 256, "seed": 1, "steps": 4000, "recall": {"exact": 0.9922, "byte_acc": 0.9961}, "bpb": 2.3974}
|
| 20 |
+
{"exp": "19", "channel": "direct", "n": 256, "seed": 1, "steps": 2000, "recall": {"exact": 0.8555, "byte_acc": 0.9108}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0179}, "bpb": 2.9598, "bpb_after_interf": 2.4577}
|
| 21 |
+
{"exp": "19", "channel": "kd_facts", "n": 256, "seed": 1, "steps": 2000, "recall": {"exact": 0.8594, "byte_acc": 0.891}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0104}, "bpb": 4.149, "bpb_after_interf": 2.5129}
|
| 22 |
+
{"exp": "19", "channel": "kd_general", "n": 256, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0}, "bpb": 2.5208, "bpb_after_interf": 2.3414}
|
| 23 |
+
{"exp": "19", "channel": "book_implant", "n": 256, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0052}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0065}, "bpb": 2.4782, "bpb_after_interf": 2.2825}
|
| 24 |
+
{"exp": "19", "channel": "none", "n": 256, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0016}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0}, "bpb": 2.4772, "bpb_after_interf": 2.274}
|
| 25 |
+
{"exp": "19", "channel": "teacher", "n": 1024, "seed": 1, "steps": 4000, "recall": {"exact": 0.1992, "byte_acc": 0.3747}, "bpb": 2.4916}
|
| 26 |
+
{"exp": "19", "channel": "direct", "n": 1024, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0286}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0026}, "bpb": 3.0763, "bpb_after_interf": 2.5039}
|
| 27 |
+
{"exp": "19", "channel": "kd_facts", "n": 1024, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0319}, "recall_interf": {"exact": 0.0, "byte_acc": 0.0179}, "bpb": 4.2629, "bpb_after_interf": 2.6531}
|
| 28 |
+
{"exp": "19", "channel": "kd_general", "n": 1024, "seed": 1, "steps": 2000, "recall": {"exact": 0.0, "byte_acc": 0.0003}, "recall_interf": {"exact": 0.0, "byte_acc": 0.001}, "bpb": 2.5645, "bpb_after_interf": 2.3509}
|
| 29 |
+
{"exp": "19b", "channel": "teacher", "n": 256, "seed": 0, "steps": 4000, "gauges": {"train": {"exact": 1.0, "byte_acc": 1.0}, "heldout": {"exact": 0.0, "byte_acc": 0.2702}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.016}}, "bpb": 2.3906}
|
| 30 |
+
{"exp": "19b", "channel": "direct", "n": 256, "seed": 0, "steps": 2000, "gauges": {"train": {"exact": 0.6523, "byte_acc": 0.8337}, "heldout": {"exact": 0.0, "byte_acc": 0.2695}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0241}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0391}, "heldout": {"exact": 0.0, "byte_acc": 0.0228}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0033}}, "bpb": 2.8785, "bpb_after_interf": 2.4644}
|
| 31 |
+
{"exp": "19b", "channel": "kd_facts", "n": 256, "seed": 0, "steps": 2000, "gauges": {"train": {"exact": 0.7578, "byte_acc": 0.8981}, "heldout": {"exact": 0.0, "byte_acc": 0.2637}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0908}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0247}, "heldout": {"exact": 0.0, "byte_acc": 0.0176}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0225}}, "bpb": 4.1461, "bpb_after_interf": 2.5325}
|
| 32 |
+
{"exp": "19b", "channel": "kd_general", "n": 256, "seed": 0, "steps": 2000, "gauges": {"train": {"exact": 0.0, "byte_acc": 0.0}, "heldout": {"exact": 0.0, "byte_acc": 0.0}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0029}, "heldout": {"exact": 0.0, "byte_acc": 0.0039}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0094}}, "bpb": 2.5091, "bpb_after_interf": 2.3152}
|
| 33 |
+
{"exp": "19b", "channel": "teacher", "n": 256, "seed": 1, "steps": 4000, "gauges": {"train": {"exact": 0.9844, "byte_acc": 0.9948}, "heldout": {"exact": 0.0, "byte_acc": 0.2448}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.002}}, "bpb": 2.4099}
|
| 34 |
+
{"exp": "19b", "channel": "direct", "n": 256, "seed": 1, "steps": 2000, "gauges": {"train": {"exact": 0.7734, "byte_acc": 0.9027}, "heldout": {"exact": 0.0, "byte_acc": 0.2559}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0238}, "heldout": {"exact": 0.0, "byte_acc": 0.0221}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0029}}, "bpb": 2.928, "bpb_after_interf": 2.4109}
|
| 35 |
+
{"exp": "19b", "channel": "kd_facts", "n": 256, "seed": 1, "steps": 2000, "gauges": {"train": {"exact": 0.8438, "byte_acc": 0.932}, "heldout": {"exact": 0.0, "byte_acc": 0.2786}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0771}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0075}, "heldout": {"exact": 0.0, "byte_acc": 0.0085}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0163}}, "bpb": 4.0264, "bpb_after_interf": 2.5571}
|
| 36 |
+
{"exp": "19b", "channel": "kd_general", "n": 256, "seed": 1, "steps": 2000, "gauges": {"train": {"exact": 0.0, "byte_acc": 0.0}, "heldout": {"exact": 0.0, "byte_acc": 0.0}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0}}, "gauges_interf": {"train": {"exact": 0.0, "byte_acc": 0.0}, "heldout": {"exact": 0.0, "byte_acc": 0.0}, "train_varfmt": {"exact": 0.0, "byte_acc": 0.0013}}, "bpb": 2.5298, "bpb_after_interf": 2.3408}
|
exp019_cr/results/results.json
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"capacity_curve": {
|
| 3 |
+
"N64_s0": 1.0,
|
| 4 |
+
"N64_s1": 1.0,
|
| 5 |
+
"N256_s0": 0.9883,
|
| 6 |
+
"N256_s1": 0.9922,
|
| 7 |
+
"N1024_s0": 0.0117,
|
| 8 |
+
"N1024_s1": 0.1992
|
| 9 |
+
},
|
| 10 |
+
"kd_vs_direct_rule_train_exact": {
|
| 11 |
+
"s0": [
|
| 12 |
+
0.7578,
|
| 13 |
+
0.6523
|
| 14 |
+
],
|
| 15 |
+
"s1": [
|
| 16 |
+
0.8438,
|
| 17 |
+
0.7734
|
| 18 |
+
]
|
| 19 |
+
},
|
| 20 |
+
"n_rows": 36
|
| 21 |
+
}
|
exp019_cr/specimens/rule_direct_s0.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7cb0de713dc1277787f40717c4de271815ed5e8480500c9fc5aab8ed12b35e91
|
| 3 |
+
size 7977425
|
exp019_cr/specimens/rule_direct_s1.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b63fcab4464246727a8dff2fa2208c4e304e50557397dce45ee5962ef3acc7e6
|
| 3 |
+
size 7977425
|
exp019_cr/specimens/rule_kd_facts_s0.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:637f1a4de3270b34ed59b589fa083516faf1bcf93c7a96d6118ebcc14cdea7c0
|
| 3 |
+
size 7977535
|
exp019_cr/specimens/rule_kd_facts_s1.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:771d084ad26ddb878371e385ef91f7d433ac31cfb14412aef4bc171529bf423f
|
| 3 |
+
size 7977535
|
exp019_cr/specimens/rule_kd_general_s0.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c9ad649e1a8fb6fcacd67866757eb7f26c5e561dfd94f42b3e470f84b374b3dc
|
| 3 |
+
size 7977645
|
exp019_cr/specimens/rule_kd_general_s1.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:56acd7c7ef2291bfceae6106902b90a3376a7f4fbbd00ab1e81b8fb2731dd9fb
|
| 3 |
+
size 7977645
|
exp019_cr/specimens/rule_teacher_s0.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9e2fd2b00e87f8e380e45169830225adb15b7c7883b6e6fb027a4e5c973c2154
|
| 3 |
+
size 7977480
|
exp019_cr/specimens/rule_teacher_s1.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:9394e030f940e21b8f4cb9ffc4637bc2db1961c90040d01dca94d4fa66543b1b
|
| 3 |
+
size 7977480
|
exp019_cr/specimens/teacher_N1024_s0.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:af3c50e397e9443ebbaa1f2e7b9a231ee48e05eb1a9cadd8576f2c7ae162ba51
|
| 3 |
+
size 7977535
|
exp019_cr/specimens/teacher_N1024_s1.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:41211d812d569c081998188e7ba209bf6e411592dc300919a943ba32a42356c4
|
| 3 |
+
size 7977535
|
exp019_cr/specimens/teacher_N256_s0.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:f33ce09f77188f0275b71d35abb510e58e6855242c8a0faf75e459b4d2420ff0
|
| 3 |
+
size 7977480
|
exp019_cr/specimens/teacher_N256_s1.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7e6d6917c2c1a78172b5ccd7f73b8f5ba34674224f0dae0c548191b18a95ba6f
|
| 3 |
+
size 7977480
|
exp019_cr/specimens/teacher_N64_s0.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c171d410e06adb6eb5e9a4d078095d92ba54f91598f5b24e3e5e406dc72c165f
|
| 3 |
+
size 7977425
|
exp019_cr/specimens/teacher_N64_s1.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:ead81b62b73c759c75af1a84b02c477380cec45b835062fad06b64f427524f6c
|
| 3 |
+
size 7977425
|
exp020_gen/README.md
ADDED
|
@@ -0,0 +1,73 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# exp020_gen — generalization: structure vs capacity on a hidden rule
|
| 2 |
+
|
| 3 |
+
Sequel to [exp019](../exp019_cr/), whose 19b block found rules learned only
|
| 4 |
+
fragmentarily (held-out byte ~0.26, exact 0.000) and content locked to its
|
| 5 |
+
surface format. exp020 races the program's structural methodologies —
|
| 6 |
+
codebooks and constellations — on that same rule task, against the honest
|
| 7 |
+
capacity control, asking what (if anything) unlocks generalization.
|
| 8 |
+
|
| 9 |
+
**Task** (identical to 19b): value = fixed random substitution cipher of the
|
| 10 |
+
key; 256 train keys in-stream, 128 held out; 4000 steps direct training.
|
| 11 |
+
**Judge**: held-out recall (rule induction), variant-format recall (the
|
| 12 |
+
lock), clean bpb.
|
| 13 |
+
|
| 14 |
+
## The bakeoff (held-out byte accuracy; exact = 0.000 in all 12 cells)
|
| 15 |
+
|
| 16 |
+
| arm | head params | s0 | s1 | clean bpb |
|
| 17 |
+
|---|---|---|---|---|
|
| 18 |
+
| **aleph** (addr_msl64 bottleneck) | 115,200 | **.2611** | **.2493** | 2.40 / 2.41 |
|
| 19 |
+
| relay (addresses in depth) | 214,532 | .2305 | .2493 | 2.36 / 2.41 |
|
| 20 |
+
| aleph_fmtdiv (3 fact formats in-stream) | 115,200 | .2148 | .2363 | 2.35 / 2.40 |
|
| 21 |
+
| mlp (param-matched free head) | 115,261 | .1992 | .2044 | 2.39 / 2.42 |
|
| 22 |
+
| aleph_tri (trigram byte embeddings) | 115,200 | .1634 | .1582 | 2.37 / 2.30 |
|
| 23 |
+
| const (constellation 768 address) | 1,978,432 | .1133 | .1309 | **2.26 / 2.23** |
|
| 24 |
+
|
| 25 |
+
`build_results.py` re-asserts every claim below from `results/ledger.jsonl`.
|
| 26 |
+
|
| 27 |
+
## Findings
|
| 28 |
+
|
| 29 |
+
1. **The plateau is not a structure problem.** No methodology in the toolkit
|
| 30 |
+
lifts held-out accuracy off the ~0.26 byte plateau, and no cell produces a
|
| 31 |
+
single exact held-out completion. Cracking rule induction here needs
|
| 32 |
+
scale, budget, or curriculum — not a different head.
|
| 33 |
+
2. **The aleph bottleneck is the best generalizer — certified both seeds** —
|
| 34 |
+
beating its param-matched free head by ~25% relative at a head budget
|
| 35 |
+
matched to within 0.2%. Structure beats capacity for generalization even
|
| 36 |
+
while losing on bpb.
|
| 37 |
+
3. **The inverse law**: generalization ordering roughly REVERSES modeling
|
| 38 |
+
strength, both seeds. The constellation head (best clean bpb, 17× larger)
|
| 39 |
+
generalizes worst; the trigram memorizer lineage second-worst; the
|
| 40 |
+
tightest bottleneck best. **Memorization ease substitutes for rule
|
| 41 |
+
induction** — surplus representational capacity soaks the instances and
|
| 42 |
+
removes the pressure to induce. (This measured memorization↔generalization
|
| 43 |
+
axis is the empirical foundation for slider/registry-style composites
|
| 44 |
+
where blocks are graded by this property and gated accordingly.)
|
| 45 |
+
4. **Format diversity generalizes across trained surface forms, not to novel
|
| 46 |
+
ones**: the diverse-format arm completes held-out keys at full rule level
|
| 47 |
+
in its *seen* alternate format (byte .205/.260) while the *unseen* format
|
| 48 |
+
stays locked (.017/.041).
|
| 49 |
+
5. **Depth relays ≈ the head bottleneck** (no lift, no cost) — distributing
|
| 50 |
+
the address through depth neither helps nor hurts rule induction.
|
| 51 |
+
|
| 52 |
+
## Files
|
| 53 |
+
- `exp020_generalization.py` — the six arms, format-diverse fact streams,
|
| 54 |
+
matched-head construction, bakeoff runner, smoke.
|
| 55 |
+
- `geolip_vitals.py` / `ar_differentiation_bed.py` /
|
| 56 |
+
`exp014_genetic_distillation.py` / `exp017_aleph_constellation.py` /
|
| 57 |
+
`exp019_content_retention.py` / `read_codebook.py` — this package's own
|
| 58 |
+
harness copies (the bed, the constellation head, the rule-task machinery).
|
| 59 |
+
Standalone.
|
| 60 |
+
- `repro.py`, `build_results.py`, `results/ledger.jsonl` (12 rows),
|
| 61 |
+
`specimens/` (all 12 checkpoints).
|
| 62 |
+
|
| 63 |
+
## Reproduce (from inside this folder)
|
| 64 |
+
```bash
|
| 65 |
+
pip install torch --index-url https://download.pytorch.org/whl/cu128
|
| 66 |
+
pip install pyarrow huggingface_hub
|
| 67 |
+
python repro.py # CPU smoke
|
| 68 |
+
python repro.py --run # the bakeoff (GPU, ~2h)
|
| 69 |
+
python build_results.py # re-assert every claim from the ledger
|
| 70 |
+
```
|
| 71 |
+
Data lands in `./data` (override with `GEOLIP_DATA`).
|
| 72 |
+
|
| 73 |
+
License: MIT · AbstractPhil + Claude Fable 5 · July 11, 2026
|
exp020_gen/ar_differentiation_bed.py
ADDED
|
@@ -0,0 +1,487 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""ar_differentiation_bed.py — exp012: autoregressive differentiation of the aleph.
|
| 2 |
+
|
| 3 |
+
Differentiation is cultivated by PREDICTIVE pressure along the sequence — the
|
| 4 |
+
address parameterizing the next-byte distribution (Law 2: chain-rule advantage pays
|
| 5 |
+
ONLY where the composed address directly parameterizes the predictive distribution).
|
| 6 |
+
This bed puts the aleph in the autoregressive gradient path and measures what
|
| 7 |
+
differentiates. The head arms enforce the employment law at its maximum: the
|
| 8 |
+
ENTIRE next-byte distribution is parameterized by the address.
|
| 9 |
+
|
| 10 |
+
Byte-level causal LM on wikitext-2-raw (HF parquet, CDN-fast), block 256. ARMS:
|
| 11 |
+
sdpa — standard causal transformer control (matched trunk).
|
| 12 |
+
hub — attention replaced by CAUSAL HUB: linear attention whose feature map
|
| 13 |
+
is the 2K-oriented aleph address, prefix-sum memories (no selection
|
| 14 |
+
event; O(n*K*d)). Differentiation cultivated INSIDE attention.
|
| 15 |
+
addr_head — sdpa trunk, but the OUTPUT HEAD reads ONLY the signed aleph
|
| 16 |
+
coefficient vector w_k = sinh(u_k)/sum_j cosh(u_j) of the final
|
| 17 |
+
hidden state (K -> 256 logits). The address MUST carry every bit of
|
| 18 |
+
next-byte information — the hardest Law-2 bottleneck.
|
| 19 |
+
|
| 20 |
+
JUDGED BY: val bits-per-byte per arm (task) + CULTIVATION VITALS on every aleph
|
| 21 |
+
codebook (readouts, never losses): axis aliveness/hppl, drift-from-init +
|
| 22 |
+
binding fraction @0.29154, winner-|cos| saturation (sign-code emergence), shadow
|
| 23 |
+
path diversity (fixed high-bits hash). Never by recon.
|
| 24 |
+
|
| 25 |
+
Riders: pure Adam wd=0; no BN/Dropout/GAP on geometric paths; orthogonal init;
|
| 26 |
+
Colab-cell-safe (paste-ahead imports, no bare argparse, no __file__ reliance);
|
| 27 |
+
GPU-only for verdict runs.
|
| 28 |
+
|
| 29 |
+
Terminal: python ar_differentiation_bed.py # shapes/parse smoke
|
| 30 |
+
python ar_differentiation_bed.py --train # verdict run
|
| 31 |
+
Colab: paste geolip_vitals.py cell, then this file (smoke auto-runs),
|
| 32 |
+
then train(steps=2000, data_root="/content/data") in the next cell.
|
| 33 |
+
"""
|
| 34 |
+
from __future__ import annotations
|
| 35 |
+
import math
|
| 36 |
+
import torch
|
| 37 |
+
import torch.nn as nn
|
| 38 |
+
import torch.nn.functional as F
|
| 39 |
+
|
| 40 |
+
if "anchor_drift" not in globals():
|
| 41 |
+
try:
|
| 42 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 43 |
+
except ImportError:
|
| 44 |
+
_here = globals().get("__file__")
|
| 45 |
+
if _here is not None:
|
| 46 |
+
import sys, pathlib
|
| 47 |
+
sys.path.insert(0, str(pathlib.Path(_here).parent))
|
| 48 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 49 |
+
else:
|
| 50 |
+
raise ImportError(
|
| 51 |
+
"geolip_vitals not found — paste/run its cell first, or "
|
| 52 |
+
"hf_hub_download exp012_ar/geolip_vitals.py from "
|
| 53 |
+
"AbstractPhil/geolip-aleph-differentiation.")
|
| 54 |
+
|
| 55 |
+
VOCAB = 256 # bytes
|
| 56 |
+
|
| 57 |
+
|
| 58 |
+
# ------------------------------------------------------------------ aleph address
|
| 59 |
+
def _super_fibonacci_s3(n: int) -> torch.Tensor:
|
| 60 |
+
"""Near-uniform unit quaternions (Alexa CVPR'22) —
|
| 61 |
+
starts the codebook INSIDE the RP^3 attractor basin. D=4 only."""
|
| 62 |
+
PHI, PSI = math.sqrt(2.0), 1.533751168755204288118041
|
| 63 |
+
i = torch.arange(n, dtype=torch.float64)
|
| 64 |
+
s = (i + 0.5) / n
|
| 65 |
+
r, R = torch.sqrt(s), torch.sqrt(1.0 - s)
|
| 66 |
+
a, b = 2 * math.pi * i / PHI, 2 * math.pi * i / PSI
|
| 67 |
+
q = torch.stack([r * torch.sin(a), r * torch.cos(a),
|
| 68 |
+
R * torch.sin(b), R * torch.cos(b)], dim=-1)
|
| 69 |
+
return F.normalize(q, dim=-1).float()
|
| 70 |
+
|
| 71 |
+
|
| 72 |
+
class AlephAddress(nn.Module):
|
| 73 |
+
"""Closed-form aleph over 2K oriented half-axes (aleph-void article).
|
| 74 |
+
signed(x): (..., K) w_k = sinh(u_k)/sum_j cosh(u_j) — the Law-2 head feature.
|
| 75 |
+
oriented(x): ((..., K), (..., K)) positive halves of the 2K softmax — HUB map."""
|
| 76 |
+
|
| 77 |
+
def __init__(self, K: int, D: int, tau: float = 0.1, init: str = "random"):
|
| 78 |
+
super().__init__()
|
| 79 |
+
self.K, self.D, self.tau = K, D, tau
|
| 80 |
+
if init == "fibonacci":
|
| 81 |
+
assert D == 4, "fibonacci init lives on S^3 (D=4)"
|
| 82 |
+
A = _super_fibonacci_s3(K)
|
| 83 |
+
else:
|
| 84 |
+
A = F.normalize(torch.randn(K, D), dim=-1)
|
| 85 |
+
self.codebook = nn.Parameter(A)
|
| 86 |
+
self.register_buffer("home", self.codebook.detach().clone())
|
| 87 |
+
|
| 88 |
+
def _u(self, x):
|
| 89 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 90 |
+
return (F.normalize(x, dim=-1) @ A.transpose(-1, -2)) / self.tau
|
| 91 |
+
|
| 92 |
+
def oriented(self, x):
|
| 93 |
+
u = self._u(x)
|
| 94 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 95 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 96 |
+
Z = (ep + en).sum(dim=-1, keepdim=True)
|
| 97 |
+
return ep / Z, en / Z
|
| 98 |
+
|
| 99 |
+
def signed(self, x):
|
| 100 |
+
u = self._u(x)
|
| 101 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 102 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 103 |
+
return (ep - en) / (ep + en).sum(dim=-1, keepdim=True)
|
| 104 |
+
|
| 105 |
+
def signed_at(self, x, taus):
|
| 106 |
+
"""Multi-tau stroboscope (rule of 3): signed coefficients at several
|
| 107 |
+
temperatures, concatenated — softer taus keep the vector dense while a
|
| 108 |
+
hard tau supplies the sign-code sharpness. v2 refinement (b)."""
|
| 109 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 110 |
+
cos = F.normalize(x, dim=-1) @ A.transpose(-1, -2)
|
| 111 |
+
outs = []
|
| 112 |
+
for t in taus:
|
| 113 |
+
u = cos / t
|
| 114 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 115 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 116 |
+
outs.append((ep - en) / (ep + en).sum(dim=-1, keepdim=True))
|
| 117 |
+
return torch.cat(outs, dim=-1)
|
| 118 |
+
|
| 119 |
+
def m_hat(self, x):
|
| 120 |
+
"""Closed-form soft read (decoders read M_hat, never M). v2 control (c)."""
|
| 121 |
+
u = self._u(x)
|
| 122 |
+
m = u.abs().amax(dim=-1, keepdim=True)
|
| 123 |
+
ep, en = torch.exp(u - m), torch.exp(-u - m)
|
| 124 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 125 |
+
return ((ep - en) @ A) / (ep + en).sum(dim=-1, keepdim=True)
|
| 126 |
+
|
| 127 |
+
def m_hard_ste(self, x):
|
| 128 |
+
"""Hard mode (aleph-void article): M_hard = sign(cos_win) * A[win], straight-through to
|
| 129 |
+
the soft read — forward fully discrete SIGN CODE, backward soft gradient.
|
| 130 |
+
Legal per theme A (reconstructive sign code, not a one-hot roster pick)."""
|
| 131 |
+
u = self._u(x)
|
| 132 |
+
soft = self.m_hat(x)
|
| 133 |
+
win = u.abs().argmax(dim=-1)
|
| 134 |
+
A = F.normalize(self.codebook, dim=-1)
|
| 135 |
+
sign = torch.sign(torch.gather(u, -1, win.unsqueeze(-1))).squeeze(-1)
|
| 136 |
+
hard = sign.unsqueeze(-1) * A[win]
|
| 137 |
+
return hard + soft - soft.detach()
|
| 138 |
+
|
| 139 |
+
@torch.no_grad()
|
| 140 |
+
def vitals(self, x_sample) -> dict:
|
| 141 |
+
u = self._u(x_sample.reshape(-1, x_sample.shape[-1]))
|
| 142 |
+
p, n = self.oriented(x_sample.reshape(-1, x_sample.shape[-1]))
|
| 143 |
+
two_k = torch.cat([p, n], dim=-1)
|
| 144 |
+
win = two_k.argmax(dim=-1)
|
| 145 |
+
cos_win = (u.abs().amax(dim=-1) * self.tau) # winner |cos| — sign-code sat.
|
| 146 |
+
d = anchor_drift(self.codebook, self.home)
|
| 147 |
+
return {"drift": round(d["mean"], 4),
|
| 148 |
+
"binding_frac": round(d["binding_fraction"], 4),
|
| 149 |
+
"aliveness": axis_aliveness(two_k),
|
| 150 |
+
"win_cos_mean": round(cos_win.mean().item(), 4),
|
| 151 |
+
"paths": path_diversity(win)}
|
| 152 |
+
|
| 153 |
+
|
| 154 |
+
# ------------------------------------------------------------------------- blocks
|
| 155 |
+
class CausalSDPA(nn.Module):
|
| 156 |
+
def __init__(self, d: int, heads: int = 4):
|
| 157 |
+
super().__init__()
|
| 158 |
+
self.h = heads
|
| 159 |
+
self.qkv = nn.Linear(d, 3 * d, bias=False)
|
| 160 |
+
self.o = nn.Linear(d, d, bias=False)
|
| 161 |
+
nn.init.orthogonal_(self.qkv.weight); nn.init.orthogonal_(self.o.weight)
|
| 162 |
+
|
| 163 |
+
def forward(self, x):
|
| 164 |
+
B, n, d = x.shape
|
| 165 |
+
q, k, v = self.qkv(x).chunk(3, dim=-1)
|
| 166 |
+
q, k, v = (t.view(B, n, self.h, d // self.h).transpose(1, 2) for t in (q, k, v))
|
| 167 |
+
y = F.scaled_dot_product_attention(q, k, v, is_causal=True)
|
| 168 |
+
return self.o(y.transpose(1, 2).reshape(B, n, d))
|
| 169 |
+
|
| 170 |
+
|
| 171 |
+
class CausalHUB(nn.Module):
|
| 172 |
+
"""Causal aleph linear attention: prefix-sum memories over the two K-wide
|
| 173 |
+
halves of the oriented address; 2K never materialized; no selection event."""
|
| 174 |
+
|
| 175 |
+
def __init__(self, d: int, K: int = 32, D: int = 4, tau: float = 0.1):
|
| 176 |
+
super().__init__()
|
| 177 |
+
self.addr = AlephAddress(K, D, tau)
|
| 178 |
+
self.q = nn.Linear(d, D, bias=False)
|
| 179 |
+
self.k = nn.Linear(d, D, bias=False)
|
| 180 |
+
self.v = nn.Linear(d, d, bias=False)
|
| 181 |
+
self.o = nn.Linear(d, d, bias=False)
|
| 182 |
+
for m in (self.q, self.k, self.v, self.o):
|
| 183 |
+
nn.init.orthogonal_(m.weight)
|
| 184 |
+
|
| 185 |
+
def forward(self, x):
|
| 186 |
+
qp, qn = self.addr.oriented(self.q(x)) # (B, n, K)
|
| 187 |
+
kp, kn = self.addr.oriented(self.k(x))
|
| 188 |
+
v = self.v(x) # (B, n, d)
|
| 189 |
+
Sp = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kp, v), dim=1)
|
| 190 |
+
Sn = torch.cumsum(torch.einsum("bnk,bnd->bnkd", kn, v), dim=1)
|
| 191 |
+
zp = torch.cumsum(kp, dim=1)
|
| 192 |
+
zn = torch.cumsum(kn, dim=1)
|
| 193 |
+
num = torch.einsum("bnk,bnkd->bnd", qp, Sp) + torch.einsum("bnk,bnkd->bnd", qn, Sn)
|
| 194 |
+
den = (qp * zp).sum(-1, keepdim=True) + (qn * zn).sum(-1, keepdim=True)
|
| 195 |
+
return self.o(num / den.clamp_min(1e-12))
|
| 196 |
+
|
| 197 |
+
|
| 198 |
+
class MslRelay(nn.Module):
|
| 199 |
+
"""Depth-composition unit (chain-rule probe): multi-slot M_hat read entering
|
| 200 |
+
the trunk as a NEAR-ZERO gated residual (gate init -3.0, sigma~0.047 — theme D:
|
| 201 |
+
geometry enters as a nudge and grows only if it earns gradient)."""
|
| 202 |
+
|
| 203 |
+
def __init__(self, d: int, n_slots: int = 16, K: int = 64):
|
| 204 |
+
super().__init__()
|
| 205 |
+
self.n_slots = n_slots
|
| 206 |
+
self.proj = nn.Linear(d, n_slots * 4, bias=False)
|
| 207 |
+
self.out = nn.Linear(n_slots * 4, d, bias=False)
|
| 208 |
+
nn.init.orthogonal_(self.proj.weight)
|
| 209 |
+
nn.init.orthogonal_(self.out.weight)
|
| 210 |
+
self.addr = AlephAddress(K, 4)
|
| 211 |
+
self.gate = nn.Parameter(torch.tensor(-3.0))
|
| 212 |
+
|
| 213 |
+
def forward(self, x):
|
| 214 |
+
B, n, _ = x.shape
|
| 215 |
+
slots = self.proj(x).view(B, n, self.n_slots, 4)
|
| 216 |
+
m = self.addr.m_hat(slots).reshape(B, n, -1)
|
| 217 |
+
return x + self.gate.sigmoid() * self.out(m)
|
| 218 |
+
|
| 219 |
+
|
| 220 |
+
class Block(nn.Module):
|
| 221 |
+
def __init__(self, d: int, attn: nn.Module):
|
| 222 |
+
super().__init__()
|
| 223 |
+
self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)
|
| 224 |
+
self.attn = attn
|
| 225 |
+
self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
|
| 226 |
+
|
| 227 |
+
def forward(self, x):
|
| 228 |
+
x = x + self.attn(self.n1(x))
|
| 229 |
+
return x + self.mlp(self.n2(x))
|
| 230 |
+
|
| 231 |
+
|
| 232 |
+
class ByteLM(nn.Module):
|
| 233 |
+
def __init__(self, arm: str, d: int = 192, layers: int = 4, block: int = 256,
|
| 234 |
+
K: int = 32, D: int = 4):
|
| 235 |
+
super().__init__()
|
| 236 |
+
# "<arm>_tri" suffix = trigram byte embedding (AlephLM byte_emb x3 lineage):
|
| 237 |
+
# token embedding is the sum of embeddings of bytes t, t-1, t-2.
|
| 238 |
+
self.trigram = arm.endswith("_tri")
|
| 239 |
+
if self.trigram:
|
| 240 |
+
arm = arm[:-4]
|
| 241 |
+
# "_fib" = super-Fibonacci S^3 codebook init (basin test: starts INSIDE
|
| 242 |
+
# the RP^3 attractor; primary observable is init->final geodesic drift).
|
| 243 |
+
self.fib = arm.endswith("_fib")
|
| 244 |
+
if self.fib:
|
| 245 |
+
arm = arm[:-4]
|
| 246 |
+
# "relay*" = stacked addresses in depth: MslRelay after every block.
|
| 247 |
+
# relay -> sdpa trunk + standard head; relay_msl64 -> + addressed head.
|
| 248 |
+
self.use_relay = arm.startswith("relay")
|
| 249 |
+
if arm == "relay":
|
| 250 |
+
arm = "sdpa"
|
| 251 |
+
elif arm == "relay_msl64":
|
| 252 |
+
arm = "addr_msl64"
|
| 253 |
+
self.arm, self.block = arm, block
|
| 254 |
+
self.emb = nn.Embedding(VOCAB, d)
|
| 255 |
+
if self.trigram:
|
| 256 |
+
self.emb1 = nn.Embedding(VOCAB, d)
|
| 257 |
+
self.emb2 = nn.Embedding(VOCAB, d)
|
| 258 |
+
self.pos = nn.Parameter(torch.zeros(1, block, d) + 0.01 * torch.randn(1, block, d))
|
| 259 |
+
mk_attn = (lambda: CausalHUB(d, K, D)) if arm == "hub" else (lambda: CausalSDPA(d))
|
| 260 |
+
self.blocks = nn.ModuleList([Block(d, mk_attn()) for _ in range(layers)])
|
| 261 |
+
if self.use_relay:
|
| 262 |
+
self.relays = nn.ModuleList([MslRelay(d) for _ in range(layers)])
|
| 263 |
+
self.nf = nn.LayerNorm(d)
|
| 264 |
+
if arm == "addr_head":
|
| 265 |
+
self.head_addr = AlephAddress(K, d) # v1: codebook in model dim — COLLAPSED
|
| 266 |
+
self.head = nn.Linear(K, VOCAB, bias=True)
|
| 267 |
+
elif arm in ("addr_d4", "addr_3tau", "addr_mhat"):
|
| 268 |
+
# v2 refinements: LOW-D HOME — learned projection to the native D=4 home
|
| 269 |
+
# before addressing (mirrors the healthy HUB arms), K=64.
|
| 270 |
+
self.head_proj = nn.Linear(d, 4, bias=False)
|
| 271 |
+
nn.init.orthogonal_(self.head_proj.weight)
|
| 272 |
+
self.head_addr = AlephAddress(64, 4)
|
| 273 |
+
if arm == "addr_d4":
|
| 274 |
+
self.head = nn.Linear(64, VOCAB, bias=True) # w alone, D=4 home
|
| 275 |
+
elif arm == "addr_3tau":
|
| 276 |
+
self.taus = (0.05, 0.1, 0.3) # rule-of-3 strobe
|
| 277 |
+
self.head = nn.Linear(64 * 3, VOCAB, bias=True)
|
| 278 |
+
else: # addr_mhat
|
| 279 |
+
self.head = nn.Linear(4, VOCAB, bias=True) # tightest: M_hat
|
| 280 |
+
elif arm.startswith("addr_msl"):
|
| 281 |
+
# v3: MULTI-SLOT heads — the 16s funnel widening: P parallel D=4 slots
|
| 282 |
+
# over a SHARED codebook. addr_msl consumes the reconstructive M_hat per
|
| 283 |
+
# slot (Px4 dims); addr_msl_w consumes signed w per slot (Px64) — tests
|
| 284 |
+
# whether slot-parallel consumption alone rescues the coefficient path.
|
| 285 |
+
# addr_msl<P> = slot-count dose-response. addr_mslh<P> = HARD sign-code
|
| 286 |
+
# consumption (straight-through M_hard per slot).
|
| 287 |
+
self.hard = arm.startswith("addr_mslh")
|
| 288 |
+
if arm in ("addr_msl", "addr_msl_w"):
|
| 289 |
+
self.n_slots = 16
|
| 290 |
+
else:
|
| 291 |
+
self.n_slots = int(arm[len("addr_mslh" if self.hard else "addr_msl"):])
|
| 292 |
+
self.head_proj = nn.Linear(d, self.n_slots * 4, bias=False)
|
| 293 |
+
nn.init.orthogonal_(self.head_proj.weight)
|
| 294 |
+
self.head_addr = AlephAddress(
|
| 295 |
+
64, 4, init="fibonacci" if self.fib else "random")
|
| 296 |
+
width = self.n_slots * (64 if arm == "addr_msl_w" else 4)
|
| 297 |
+
self.head = nn.Linear(width, VOCAB, bias=True)
|
| 298 |
+
elif arm == "addr_3tau_mhat":
|
| 299 |
+
# v3: combine the two v2 winners — 3-tau stroboscope + reconstructive read.
|
| 300 |
+
self.head_proj = nn.Linear(d, 4, bias=False)
|
| 301 |
+
nn.init.orthogonal_(self.head_proj.weight)
|
| 302 |
+
self.head_addr = AlephAddress(64, 4)
|
| 303 |
+
self.taus = (0.05, 0.1, 0.3)
|
| 304 |
+
self.head = nn.Linear(64 * 3 + 4, VOCAB, bias=True)
|
| 305 |
+
else:
|
| 306 |
+
self.head = nn.Linear(d, VOCAB, bias=True)
|
| 307 |
+
self._last_h = None
|
| 308 |
+
|
| 309 |
+
def forward(self, idx):
|
| 310 |
+
x = self.emb(idx)
|
| 311 |
+
if self.trigram: # past-only shifts — causality preserved
|
| 312 |
+
x = x + self.emb1(F.pad(idx, (1, 0), value=0)[:, :-1]) \
|
| 313 |
+
+ self.emb2(F.pad(idx, (2, 0), value=0)[:, :-2])
|
| 314 |
+
x = x + self.pos[:, : idx.shape[1]]
|
| 315 |
+
if self.use_relay:
|
| 316 |
+
for b, r in zip(self.blocks, self.relays):
|
| 317 |
+
x = r(b(x))
|
| 318 |
+
else:
|
| 319 |
+
for b in self.blocks:
|
| 320 |
+
x = b(x)
|
| 321 |
+
h = self.nf(x)
|
| 322 |
+
self._last_h = h.detach()
|
| 323 |
+
if self.arm == "addr_head":
|
| 324 |
+
return self.head(self.head_addr.signed(h))
|
| 325 |
+
if self.arm == "addr_d4":
|
| 326 |
+
return self.head(self.head_addr.signed(self.head_proj(h)))
|
| 327 |
+
if self.arm == "addr_3tau":
|
| 328 |
+
return self.head(self.head_addr.signed_at(self.head_proj(h), self.taus))
|
| 329 |
+
if self.arm == "addr_mhat":
|
| 330 |
+
return self.head(self.head_addr.m_hat(self.head_proj(h)))
|
| 331 |
+
if self.arm.startswith("addr_msl"):
|
| 332 |
+
B, n, _ = h.shape
|
| 333 |
+
slots = self.head_proj(h).view(B, n, self.n_slots, 4)
|
| 334 |
+
if self.arm == "addr_msl_w":
|
| 335 |
+
feats = self.head_addr.signed(slots).reshape(B, n, -1)
|
| 336 |
+
elif getattr(self, "hard", False):
|
| 337 |
+
feats = self.head_addr.m_hard_ste(slots).reshape(B, n, -1)
|
| 338 |
+
else:
|
| 339 |
+
feats = self.head_addr.m_hat(slots).reshape(B, n, -1)
|
| 340 |
+
return self.head(feats)
|
| 341 |
+
if self.arm == "addr_3tau_mhat":
|
| 342 |
+
p = self.head_proj(h)
|
| 343 |
+
feats = torch.cat([self.head_addr.signed_at(p, self.taus),
|
| 344 |
+
self.head_addr.m_hat(p)], dim=-1)
|
| 345 |
+
return self.head(feats)
|
| 346 |
+
return self.head(h)
|
| 347 |
+
|
| 348 |
+
@torch.no_grad()
|
| 349 |
+
def vitals(self) -> dict:
|
| 350 |
+
out = {}
|
| 351 |
+
if self.arm == "hub":
|
| 352 |
+
for i, b in enumerate(self.blocks):
|
| 353 |
+
if self._last_h is not None:
|
| 354 |
+
out[f"L{i}"] = b.attn.addr.vitals(b.attn.q(self._last_h[:2]))
|
| 355 |
+
elif self.arm == "addr_head" and self._last_h is not None:
|
| 356 |
+
out["head"] = self.head_addr.vitals(self._last_h[:2])
|
| 357 |
+
elif self.arm in ("addr_d4", "addr_3tau", "addr_mhat",
|
| 358 |
+
"addr_3tau_mhat") and self._last_h is not None:
|
| 359 |
+
out["head"] = self.head_addr.vitals(self.head_proj(self._last_h[:2]))
|
| 360 |
+
elif self.arm.startswith("addr_msl") and self._last_h is not None:
|
| 361 |
+
slots = self.head_proj(self._last_h[:2])
|
| 362 |
+
out["head"] = self.head_addr.vitals(
|
| 363 |
+
slots.reshape(*slots.shape[:-1], self.n_slots, 4))
|
| 364 |
+
if self.use_relay and self._last_h is not None:
|
| 365 |
+
for i, r in enumerate(self.relays):
|
| 366 |
+
s = r.proj(self._last_h[:2])
|
| 367 |
+
v = r.addr.vitals(s.reshape(*s.shape[:-1], r.n_slots, 4))
|
| 368 |
+
out[f"relay{i}"] = {"gate": round(r.gate.sigmoid().item(), 4),
|
| 369 |
+
"drift": v["drift"],
|
| 370 |
+
"binding_frac": v["binding_frac"],
|
| 371 |
+
"ppl": round(v["aliveness"]["usage_ppl"], 1)}
|
| 372 |
+
return out
|
| 373 |
+
|
| 374 |
+
|
| 375 |
+
# --------------------------------------------------------------------------- data
|
| 376 |
+
def _wikitext_bytes(data_root: str):
|
| 377 |
+
"""wikitext-2-raw as flat uint8 tensors via the HF parquet CDN."""
|
| 378 |
+
from huggingface_hub import hf_hub_download
|
| 379 |
+
import pyarrow.parquet as pq
|
| 380 |
+
|
| 381 |
+
def load(split):
|
| 382 |
+
p = hf_hub_download("Salesforce/wikitext",
|
| 383 |
+
f"wikitext-2-raw-v1/{split}-00000-of-00001.parquet",
|
| 384 |
+
repo_type="dataset", local_dir=data_root)
|
| 385 |
+
text = "".join(pq.read_table(p).column("text").to_pylist())
|
| 386 |
+
return torch.frombuffer(bytearray(text.encode("utf-8")), dtype=torch.uint8).clone()
|
| 387 |
+
|
| 388 |
+
return load("train"), load("validation")
|
| 389 |
+
|
| 390 |
+
|
| 391 |
+
def _batch(data: torch.Tensor, batch: int, block: int, device, g: torch.Generator):
|
| 392 |
+
ix = torch.randint(0, data.numel() - block - 1, (batch,), generator=g)
|
| 393 |
+
x = torch.stack([data[i:i + block] for i in ix]).long().to(device)
|
| 394 |
+
y = torch.stack([data[i + 1:i + block + 1] for i in ix]).long().to(device)
|
| 395 |
+
return x, y
|
| 396 |
+
|
| 397 |
+
|
| 398 |
+
# -------------------------------------------------------------------- train/smoke
|
| 399 |
+
def train(arms=("sdpa", "hub", "addr_head"), steps: int = 2000, batch: int = 32,
|
| 400 |
+
block: int = 256, device: str = "cuda", data_root: str = "./data",
|
| 401 |
+
seed: int = 0, eval_every: int = 500, save: bool = True):
|
| 402 |
+
"""Verdict run — GPU only. Pure Adam wd=0. Reports val bits-per-byte + vitals.
|
| 403 |
+
save=True writes {data_root}/ar_ckpts/{arm}_s{seed}_t{steps}.pt per arm —
|
| 404 |
+
the cultivated codebooks are SPECIMENS for the projective reading instruments."""
|
| 405 |
+
import os
|
| 406 |
+
if device == "cuda" and not torch.cuda.is_available():
|
| 407 |
+
raise RuntimeError("Verdict runs are GPU-only (never CPU-train for accuracy).")
|
| 408 |
+
ckpt_dir = os.path.join(data_root, "ar_ckpts")
|
| 409 |
+
os.makedirs(ckpt_dir, exist_ok=True)
|
| 410 |
+
tr, va = _wikitext_bytes(data_root)
|
| 411 |
+
print(f"data ready: train {tr.numel():,} bytes, val {va.numel():,} bytes", flush=True)
|
| 412 |
+
results = {}
|
| 413 |
+
for arm in arms:
|
| 414 |
+
torch.manual_seed(seed)
|
| 415 |
+
g = torch.Generator().manual_seed(seed)
|
| 416 |
+
model = ByteLM(arm, block=block).to(device)
|
| 417 |
+
n_params = sum(p.numel() for p in model.parameters())
|
| 418 |
+
opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
|
| 419 |
+
for step in range(1, steps + 1):
|
| 420 |
+
x, y = _batch(tr, batch, block, device, g)
|
| 421 |
+
logits = model(x)
|
| 422 |
+
loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
|
| 423 |
+
opt.zero_grad(set_to_none=True)
|
| 424 |
+
loss.backward()
|
| 425 |
+
opt.step()
|
| 426 |
+
if step % eval_every == 0 or step == steps:
|
| 427 |
+
model.eval()
|
| 428 |
+
with torch.no_grad():
|
| 429 |
+
losses = []
|
| 430 |
+
for _ in range(20):
|
| 431 |
+
xv, yv = _batch(va, batch, block, device, g)
|
| 432 |
+
lv = F.cross_entropy(model(xv).reshape(-1, VOCAB),
|
| 433 |
+
yv.reshape(-1))
|
| 434 |
+
losses.append(lv.item())
|
| 435 |
+
bpb = sum(losses) / len(losses) / math.log(2)
|
| 436 |
+
print(f"[{arm}] step {step} val_bpb={bpb:.4f} vitals={model.vitals()}",
|
| 437 |
+
flush=True)
|
| 438 |
+
model.train()
|
| 439 |
+
results[arm] = {"val_bpb": bpb, "params": n_params, "vitals": model.vitals()}
|
| 440 |
+
if save:
|
| 441 |
+
path = os.path.join(ckpt_dir, f"{arm}_s{seed}_t{steps}.pt")
|
| 442 |
+
torch.save({"arm": arm, "seed": seed, "steps": steps, "val_bpb": bpb,
|
| 443 |
+
"state_dict": {k: v.cpu() for k, v in
|
| 444 |
+
model.state_dict().items()}}, path)
|
| 445 |
+
print(f"saved specimen: {path}", flush=True)
|
| 446 |
+
print(results, flush=True)
|
| 447 |
+
return results
|
| 448 |
+
|
| 449 |
+
|
| 450 |
+
def smoke():
|
| 451 |
+
"""Shapes/parse only — no accuracy claims."""
|
| 452 |
+
x = torch.randint(0, VOCAB, (2, 64))
|
| 453 |
+
for arm in ("sdpa", "hub", "addr_head"):
|
| 454 |
+
m = ByteLM(arm, d=96, layers=2, block=64, K=16)
|
| 455 |
+
logits = m(x)
|
| 456 |
+
assert logits.shape == (2, 64, VOCAB)
|
| 457 |
+
logits.sum().backward()
|
| 458 |
+
# causality check: future byte must not affect past logits
|
| 459 |
+
with torch.no_grad():
|
| 460 |
+
a = m(x)[0, 10]
|
| 461 |
+
x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
|
| 462 |
+
b = m(x2)[0, 10]
|
| 463 |
+
assert torch.allclose(a, b, atol=1e-4), f"{arm} leaks future context"
|
| 464 |
+
print(f"{arm}: OK params={sum(p.numel() for p in m.parameters()):,} "
|
| 465 |
+
f"vitals={m.vitals()}", flush=True)
|
| 466 |
+
print("OK — AR bed smoke passed (verdict run: train() on GPU)", flush=True)
|
| 467 |
+
|
| 468 |
+
|
| 469 |
+
def _in_notebook() -> bool:
|
| 470 |
+
try:
|
| 471 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 472 |
+
return True
|
| 473 |
+
except NameError:
|
| 474 |
+
return False
|
| 475 |
+
|
| 476 |
+
|
| 477 |
+
if __name__ == "__main__":
|
| 478 |
+
if _in_notebook():
|
| 479 |
+
smoke()
|
| 480 |
+
print("Notebook mode: call train(steps=2000) in the next cell (GPU).")
|
| 481 |
+
else:
|
| 482 |
+
import argparse
|
| 483 |
+
ap = argparse.ArgumentParser()
|
| 484 |
+
ap.add_argument("--train", action="store_true")
|
| 485 |
+
ap.add_argument("--steps", type=int, default=2000)
|
| 486 |
+
a, _ = ap.parse_known_args()
|
| 487 |
+
train(steps=a.steps) if a.train else smoke()
|
exp020_gen/build_results.py
ADDED
|
@@ -0,0 +1,52 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""build_results.py — exp020_gen: read results/ledger.jsonl and RE-ASSERT every
|
| 2 |
+
claim in the README. Run from inside this folder: python build_results.py
|
| 3 |
+
"""
|
| 4 |
+
import json
|
| 5 |
+
import os
|
| 6 |
+
|
| 7 |
+
HERE = os.path.dirname(os.path.abspath(__file__))
|
| 8 |
+
rows = [json.loads(l) for l in
|
| 9 |
+
open(os.path.join(HERE, "results", "ledger.jsonl"), encoding="utf-8")]
|
| 10 |
+
assert all(r["exp"] == "20" for r in rows) and len(rows) == 12
|
| 11 |
+
cell = {(r["arm"], r["seed"]): r for r in rows}
|
| 12 |
+
|
| 13 |
+
def held(arm, s):
|
| 14 |
+
return cell[(arm, s)]["gauges"]["heldout"]["byte_acc"]
|
| 15 |
+
|
| 16 |
+
# claim 1: no arm lifts rule induction off the plateau — held-out EXACT is
|
| 17 |
+
# 0.000 in all 12 cells, byte acc <= 0.27 everywhere
|
| 18 |
+
for r in rows:
|
| 19 |
+
assert r["gauges"]["heldout"]["exact"] == 0.0
|
| 20 |
+
assert r["gauges"]["heldout"]["byte_acc"] <= 0.27
|
| 21 |
+
|
| 22 |
+
# claim 2: the aleph bottleneck beats its param-matched free head, both seeds,
|
| 23 |
+
# at matched head budget (within 0.2%)
|
| 24 |
+
for s in (0, 1):
|
| 25 |
+
assert held("aleph", s) > held("mlp", s), s
|
| 26 |
+
p_al, p_ml = cell[("aleph", 0)]["head_params"], cell[("mlp", 0)]["head_params"]
|
| 27 |
+
assert abs(p_al - p_ml) / p_al < 0.002, (p_al, p_ml)
|
| 28 |
+
|
| 29 |
+
# claim 3 (the inverse law): the best clean-bpb arm (const) generalizes WORST,
|
| 30 |
+
# and the memorizer lineage (trigram) is second-worst — both seeds
|
| 31 |
+
for s in (0, 1):
|
| 32 |
+
bpbs = {a: cell[(a, s)]["bpb"] for a, _ in cell if _ == s}
|
| 33 |
+
assert min(bpbs, key=bpbs.get) == "const", s
|
| 34 |
+
order = sorted((held(a, s), a) for a, ss in cell if ss == s)
|
| 35 |
+
assert order[0][1] == "const" and order[1][1] == "aleph_tri", (s, order)
|
| 36 |
+
assert order[-1][1] in ("aleph", "relay"), (s, order)
|
| 37 |
+
|
| 38 |
+
# claim 4: format diversity teaches the SEEN alternate format at rule level
|
| 39 |
+
# but does NOT unlock the unseen variant
|
| 40 |
+
for s in (0, 1):
|
| 41 |
+
g = cell[("aleph_fmtdiv", s)]["gauges"]
|
| 42 |
+
assert g["heldout_r1"]["byte_acc"] > 0.20, s # seen format: rule-level
|
| 43 |
+
assert g["train_varfmt"]["byte_acc"] < 0.06, s # unseen: still locked
|
| 44 |
+
|
| 45 |
+
out = {"heldout_byte": {f"{a}_s{s}": held(a, s) for (a, s) in sorted(cell)},
|
| 46 |
+
"head_params": {a: cell[(a, 0)]["head_params"]
|
| 47 |
+
for a in {k[0] for k in cell}},
|
| 48 |
+
"n_rows": len(rows)}
|
| 49 |
+
json.dump(out, open(os.path.join(HERE, "results", "results.json"), "w",
|
| 50 |
+
encoding="utf-8"), indent=1)
|
| 51 |
+
print(f"{len(rows)} rows -> results/results.json")
|
| 52 |
+
print("all README claims asserted OK")
|
exp020_gen/exp014_genetic_distillation.py
ADDED
|
@@ -0,0 +1,515 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""exp014_genetic_distillation.py — genetic distillation + memory substrate.
|
| 2 |
+
14_a: multi-generational tournament (GM3 paradigm) where the aleph codebook is the
|
| 3 |
+
explicit heritable genome. Lineages: ALEPH-FLAT (consensus book + KD) |
|
| 4 |
+
ALEPH-TREE (structured genome: root book + branch books) | MLP-KD
|
| 5 |
+
(traditional: best-parent weights + KD) | NO-INHERIT (evolution floor).
|
| 6 |
+
Both sides intentionally inherit logits (KD); only ours inherits geometry.
|
| 7 |
+
Consensus = Procrustes/GPA alignment of parents' books to mean shape
|
| 8 |
+
(placement by construction — replaces GM3's k-means-on-consensus init).
|
| 9 |
+
14_b: memory substrate — the D=4 home makes books size-agnostic. Implant books
|
| 10 |
+
cultivated in a small organism into a larger one (frozen / trainable), and
|
| 11 |
+
into GPT-2 relay adapters (cross-architecture frozen distillation).
|
| 12 |
+
|
| 13 |
+
Riders: pure Adam wd=0; KD = KL to detached teacher probs (predictive pressure, no
|
| 14 |
+
contrastive); tree routing is DENSE SOFT (oriented weights; collapse monitor on the
|
| 15 |
+
root); drift-check precedes every freeze claim; GPU-only verdict runs; Colab-safe.
|
| 16 |
+
Founders share a COMMON-ANCESTOR book so GPA row correspondence is inherited.
|
| 17 |
+
|
| 18 |
+
Colab paste order: geolip_vitals.py -> ar_differentiation_bed.py ->
|
| 19 |
+
exp013_augmentation_bed.py (only for run_b2) -> this file.
|
| 20 |
+
"""
|
| 21 |
+
from __future__ import annotations
|
| 22 |
+
import copy
|
| 23 |
+
import json
|
| 24 |
+
import math
|
| 25 |
+
import os
|
| 26 |
+
import torch
|
| 27 |
+
import torch.nn as nn
|
| 28 |
+
import torch.nn.functional as F
|
| 29 |
+
|
| 30 |
+
if "anchor_drift" not in globals():
|
| 31 |
+
try:
|
| 32 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 33 |
+
except ImportError:
|
| 34 |
+
_here = globals().get("__file__")
|
| 35 |
+
if _here is None:
|
| 36 |
+
raise ImportError("paste/run geolip_vitals.py first")
|
| 37 |
+
import sys, pathlib
|
| 38 |
+
sys.path.insert(0, str(pathlib.Path(_here).parent))
|
| 39 |
+
from geolip_vitals import anchor_drift, axis_aliveness, path_diversity
|
| 40 |
+
if "ByteLM" not in globals():
|
| 41 |
+
try:
|
| 42 |
+
from ar_differentiation_bed import (ByteLM, AlephAddress, _wikitext_bytes,
|
| 43 |
+
_batch, VOCAB)
|
| 44 |
+
except ImportError:
|
| 45 |
+
raise ImportError("paste/run ar_differentiation_bed.py first")
|
| 46 |
+
|
| 47 |
+
DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
|
| 48 |
+
EXP_DIR = os.path.join(DATA_ROOT, "exp014")
|
| 49 |
+
|
| 50 |
+
|
| 51 |
+
class SquaredReLU(nn.Module):
|
| 52 |
+
def forward(self, x):
|
| 53 |
+
return F.relu(x) ** 2
|
| 54 |
+
|
| 55 |
+
|
| 56 |
+
# ===================================================== consensus (the germline) ===
|
| 57 |
+
@torch.no_grad()
|
| 58 |
+
def procrustes_rotation(A: torch.Tensor, M: torch.Tensor) -> torch.Tensor:
|
| 59 |
+
"""Orthogonal R minimizing ||A R - M||_F (rows correspond)."""
|
| 60 |
+
U, _, Vt = torch.linalg.svd(A.T.double() @ M.double())
|
| 61 |
+
return (U @ Vt).float()
|
| 62 |
+
|
| 63 |
+
|
| 64 |
+
@torch.no_grad()
|
| 65 |
+
def align_to(A: torch.Tensor, ref: torch.Tensor, iters: int = 20) -> torch.Tensor:
|
| 66 |
+
"""Projective Procrustes (rows correspond, signs free): alternate the
|
| 67 |
+
orthogonal rotation and per-row sign flips (books live on RP^(D-1))."""
|
| 68 |
+
s = torch.ones(A.shape[0], 1)
|
| 69 |
+
for _ in range(iters):
|
| 70 |
+
R = procrustes_rotation(s * A, ref)
|
| 71 |
+
AR = (s * A) @ R
|
| 72 |
+
s_upd = torch.where((AR * ref).sum(-1, keepdim=True) < 0, -s, s)
|
| 73 |
+
if torch.equal(s_upd, s):
|
| 74 |
+
return AR
|
| 75 |
+
s = s_upd
|
| 76 |
+
return (s * A) @ procrustes_rotation(s * A, ref)
|
| 77 |
+
|
| 78 |
+
|
| 79 |
+
@torch.no_grad()
|
| 80 |
+
def consensus_codebook(books: list, iters: int = 50, tol: float = 1e-8):
|
| 81 |
+
"""GPA to mean shape (GM3 machinery, applied to aleph books), anchored to the
|
| 82 |
+
FIRST parent's frame. Rows must correspond (common-ancestor convention); signs
|
| 83 |
+
are projective. Returns (consensus, n_iters, delta)."""
|
| 84 |
+
# device-pin to CPU: parent models may live on CUDA after KD teacher moves
|
| 85 |
+
Bs = [F.normalize(b.detach().float().cpu(), dim=-1).clone() for b in books]
|
| 86 |
+
# pairwise projective alignment to parent-0's frame, THEN GPA refinement
|
| 87 |
+
aligned = [Bs[0]] + [align_to(b, Bs[0]) for b in Bs[1:]]
|
| 88 |
+
M = F.normalize(torch.stack(aligned).mean(0), dim=-1)
|
| 89 |
+
delta, it = 0.0, 0
|
| 90 |
+
for it in range(1, iters + 1):
|
| 91 |
+
aligned = [align_to(b, M) for b in Bs]
|
| 92 |
+
M_new = F.normalize(torch.stack(aligned).mean(0), dim=-1)
|
| 93 |
+
delta = (M_new - M).norm().item()
|
| 94 |
+
M = M_new
|
| 95 |
+
if delta < tol:
|
| 96 |
+
break
|
| 97 |
+
# re-anchor to parent-0 (GPA drift of the global frame stays measurable)
|
| 98 |
+
M = align_to(M, Bs[0])
|
| 99 |
+
return M, it, delta
|
| 100 |
+
|
| 101 |
+
|
| 102 |
+
@torch.no_grad()
|
| 103 |
+
def implant_book(addr: "AlephAddress", book: torch.Tensor, trainable: bool = True):
|
| 104 |
+
"""Load a book into an AlephAddress: codebook + home (drift measured from the
|
| 105 |
+
implant). Freeze only via trainable=False AFTER a drift-check justifies it."""
|
| 106 |
+
b = F.normalize(book.float(), dim=-1).to(addr.codebook.device)
|
| 107 |
+
assert b.shape == addr.codebook.shape, (b.shape, addr.codebook.shape)
|
| 108 |
+
addr.codebook.data.copy_(b)
|
| 109 |
+
addr.home.copy_(b)
|
| 110 |
+
addr.codebook.requires_grad_(trainable)
|
| 111 |
+
|
| 112 |
+
|
| 113 |
+
# ============================================================= tree head ==========
|
| 114 |
+
class TreeHead(nn.Module):
|
| 115 |
+
"""Autoregressive tree (the constellation-anchor analogue, exp011 TREE operator
|
| 116 |
+
in the healthy consumption regime): a ROOT aleph (K=2, D=4) yields the 4
|
| 117 |
+
oriented weights (2K half-axes = the 4 branches, dense soft, sums to 1);
|
| 118 |
+
each BRANCH is a 64-slot... shared slot projection read by a branch-specific
|
| 119 |
+
book (K=64, D=4); output = branch-weighted mixture of branch reads -> vocab.
|
| 120 |
+
Heritable genome: root book (2,4) + 4 branch books (64,4)."""
|
| 121 |
+
|
| 122 |
+
ROOT_SLOTS = 4 # slot-parallel root consumption (the collapse cure)
|
| 123 |
+
ROOT_TAU = 0.3 # softer root temperature (wave-1 fix: single hard-tau
|
| 124 |
+
# root partially collapsed, usage [.85,.12,.01,.02])
|
| 125 |
+
|
| 126 |
+
def __init__(self, d: int, vocab: int = 256, n_slots: int = 16):
|
| 127 |
+
super().__init__()
|
| 128 |
+
self.n_slots = n_slots
|
| 129 |
+
self.root_proj = nn.Linear(d, self.ROOT_SLOTS * 4, bias=False)
|
| 130 |
+
self.slot_proj = nn.Linear(d, n_slots * 4, bias=False)
|
| 131 |
+
nn.init.orthogonal_(self.root_proj.weight)
|
| 132 |
+
nn.init.orthogonal_(self.slot_proj.weight)
|
| 133 |
+
self.root = AlephAddress(2, 4, tau=self.ROOT_TAU)
|
| 134 |
+
self.branches = nn.ModuleList([AlephAddress(64, 4) for _ in range(4)])
|
| 135 |
+
self.out = nn.Linear(n_slots * 4, vocab, bias=True)
|
| 136 |
+
self._last_root = None
|
| 137 |
+
|
| 138 |
+
def forward(self, h):
|
| 139 |
+
B, n, _ = h.shape
|
| 140 |
+
rs = self.root_proj(h).view(B, n, self.ROOT_SLOTS, 4)
|
| 141 |
+
p, m = self.root.oriented(rs) # (B,n,S,2) x2
|
| 142 |
+
w = torch.cat([p, m], dim=-1).mean(dim=-2) # slot-avg -> (B,n,4)
|
| 143 |
+
self._last_root = w.detach()
|
| 144 |
+
slots = self.slot_proj(h).view(B, n, self.n_slots, 4)
|
| 145 |
+
mix = 0
|
| 146 |
+
for b, br in enumerate(self.branches):
|
| 147 |
+
mix = mix + w[..., b:b + 1] * br.m_hat(slots).reshape(B, n, -1)
|
| 148 |
+
return self.out(mix)
|
| 149 |
+
|
| 150 |
+
def genome(self):
|
| 151 |
+
return {"root": self.root.codebook.detach().clone(),
|
| 152 |
+
**{f"branch{i}": br.codebook.detach().clone()
|
| 153 |
+
for i, br in enumerate(self.branches)}}
|
| 154 |
+
|
| 155 |
+
@torch.no_grad()
|
| 156 |
+
def inherit(self, genomes: list):
|
| 157 |
+
c, it, dl = consensus_codebook([g["root"] for g in genomes])
|
| 158 |
+
implant_book(self.root, c)
|
| 159 |
+
for i, br in enumerate(self.branches):
|
| 160 |
+
c, _, _ = consensus_codebook([g[f"branch{i}"] for g in genomes])
|
| 161 |
+
implant_book(br, c)
|
| 162 |
+
|
| 163 |
+
@torch.no_grad()
|
| 164 |
+
def vitals(self):
|
| 165 |
+
out = {"root_drift": round(anchor_drift(self.root.codebook,
|
| 166 |
+
self.root.home)["mean"], 4)}
|
| 167 |
+
if self._last_root is not None:
|
| 168 |
+
w = self._last_root.reshape(-1, 4)
|
| 169 |
+
usage = w.mean(0)
|
| 170 |
+
usage = usage / usage.sum()
|
| 171 |
+
out["root_usage"] = [round(float(u), 3) for u in usage]
|
| 172 |
+
ent = -(usage.clamp_min(1e-9) * usage.clamp_min(1e-9).log()).sum()
|
| 173 |
+
out["root_ppl4"] = round(float(ent.exp()), 3)
|
| 174 |
+
d = [anchor_drift(br.codebook, br.home)["mean"] for br in self.branches]
|
| 175 |
+
out["branch_drift"] = [round(x, 3) for x in d]
|
| 176 |
+
return out
|
| 177 |
+
|
| 178 |
+
|
| 179 |
+
# ============================================================ organisms ===========
|
| 180 |
+
def make_organism(lineage: str, d: int = 192, layers: int = 4, block: int = 256,
|
| 181 |
+
seed: int = 0):
|
| 182 |
+
"""lineage in {aleph_flat, aleph_tree, mlp_kd, no_inherit}. no_inherit uses the
|
| 183 |
+
aleph_flat architecture (the control isolates INHERITANCE, not architecture)."""
|
| 184 |
+
torch.manual_seed(seed)
|
| 185 |
+
if lineage in ("aleph_flat", "no_inherit", "aleph_full", "aleph_weights"):
|
| 186 |
+
return ByteLM("addr_msl64", d=d, layers=layers, block=block)
|
| 187 |
+
if lineage == "aleph_tree":
|
| 188 |
+
m = ByteLM("sdpa", d=d, layers=layers, block=block)
|
| 189 |
+
m.head = TreeHead(d)
|
| 190 |
+
return m
|
| 191 |
+
if lineage == "mlp_kd":
|
| 192 |
+
m = ByteLM("sdpa", d=d, layers=layers, block=block)
|
| 193 |
+
m.head = nn.Sequential(nn.Linear(d, 224), SquaredReLU(),
|
| 194 |
+
nn.LayerNorm(224), nn.Linear(224, VOCAB))
|
| 195 |
+
return m
|
| 196 |
+
raise ValueError(lineage)
|
| 197 |
+
|
| 198 |
+
|
| 199 |
+
def genome_of(model):
|
| 200 |
+
"""The heritable organ = co-adapted (projection, book) pair(s). Books inherit
|
| 201 |
+
by CONSENSUS (the geometric germline); projections inherit from the BEST
|
| 202 |
+
parent (weight copy) — implanting a book against a random projection puts the
|
| 203 |
+
child below random init (campaign-v2 lesson)."""
|
| 204 |
+
if isinstance(model.head, TreeHead):
|
| 205 |
+
g = model.head.genome()
|
| 206 |
+
g["root_proj"] = model.head.root_proj.weight.detach().cpu().clone()
|
| 207 |
+
g["slot_proj"] = model.head.slot_proj.weight.detach().cpu().clone()
|
| 208 |
+
return g
|
| 209 |
+
if hasattr(model, "head_addr"):
|
| 210 |
+
return {"flat": model.head_addr.codebook.detach().cpu().clone(),
|
| 211 |
+
"proj": model.head_proj.weight.detach().cpu().clone()}
|
| 212 |
+
return None
|
| 213 |
+
|
| 214 |
+
|
| 215 |
+
@torch.no_grad()
|
| 216 |
+
def inherit_genome(model, genomes: list):
|
| 217 |
+
"""genomes[0] = the BEST parent (selection order matters)."""
|
| 218 |
+
if isinstance(model.head, TreeHead):
|
| 219 |
+
model.head.inherit(genomes)
|
| 220 |
+
model.head.root_proj.weight.copy_(genomes[0]["root_proj"].to(
|
| 221 |
+
model.head.root_proj.weight.device))
|
| 222 |
+
model.head.slot_proj.weight.copy_(genomes[0]["slot_proj"].to(
|
| 223 |
+
model.head.slot_proj.weight.device))
|
| 224 |
+
elif hasattr(model, "head_addr"):
|
| 225 |
+
c, it, dl = consensus_codebook([g["flat"] for g in genomes])
|
| 226 |
+
implant_book(model.head_addr, c)
|
| 227 |
+
model.head_proj.weight.copy_(genomes[0]["proj"].to(
|
| 228 |
+
model.head_proj.weight.device))
|
| 229 |
+
|
| 230 |
+
|
| 231 |
+
def organism_vitals(model):
|
| 232 |
+
if isinstance(model.head, TreeHead):
|
| 233 |
+
return model.head.vitals()
|
| 234 |
+
return model.vitals() if hasattr(model, "vitals") else {}
|
| 235 |
+
|
| 236 |
+
|
| 237 |
+
# ========================================================= train one member ======
|
| 238 |
+
def train_member(model, tr, va, steps=2000, batch=32, block=256, device="cuda",
|
| 239 |
+
seed=0, teachers=None, kd_alpha=1.0):
|
| 240 |
+
"""CE (+ KL to detached mean teacher probs when teachers given). Pure Adam."""
|
| 241 |
+
g = torch.Generator().manual_seed(seed)
|
| 242 |
+
model = model.to(device)
|
| 243 |
+
opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
|
| 244 |
+
if teachers:
|
| 245 |
+
teachers = [t.to(device).eval() for t in teachers]
|
| 246 |
+
for step in range(1, steps + 1):
|
| 247 |
+
x, y = _batch(tr, batch, block, device, g)
|
| 248 |
+
logits = model(x)
|
| 249 |
+
loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
|
| 250 |
+
if teachers:
|
| 251 |
+
with torch.no_grad():
|
| 252 |
+
tp = torch.stack([F.softmax(t(x), -1) for t in teachers]).mean(0)
|
| 253 |
+
loss = loss + kd_alpha * F.kl_div(
|
| 254 |
+
F.log_softmax(logits, -1), tp, reduction="batchmean")
|
| 255 |
+
opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
|
| 256 |
+
model.eval()
|
| 257 |
+
with torch.no_grad():
|
| 258 |
+
ls = []
|
| 259 |
+
for _ in range(20):
|
| 260 |
+
xv, yv = _batch(va, batch, block, device, g)
|
| 261 |
+
ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
|
| 262 |
+
yv.reshape(-1)).item())
|
| 263 |
+
return sum(ls) / len(ls) / math.log(2) # bpb
|
| 264 |
+
|
| 265 |
+
|
| 266 |
+
# ============================================================ the tournament ======
|
| 267 |
+
def run_tournament(lineage: str, gens: int = 4, pop: int = 4, steps: int = 2000,
|
| 268 |
+
seed: int = 0, device: str = "cuda",
|
| 269 |
+
catastrophic_at: int | None = None):
|
| 270 |
+
"""One lineage, one tournament seed. Logs per-gen to the ledger; saves the
|
| 271 |
+
champion genome per generation. catastrophic_at=G injects a 0-step random
|
| 272 |
+
parent into the consensus at generation G (the GM3 robustness probe)."""
|
| 273 |
+
if not torch.cuda.is_available():
|
| 274 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 275 |
+
os.makedirs(EXP_DIR, exist_ok=True)
|
| 276 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 277 |
+
ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 278 |
+
# common ancestor: every founder book starts identical within a tournament
|
| 279 |
+
torch.manual_seed(9000 + seed)
|
| 280 |
+
ancestor = make_organism(lineage, seed=9000 + seed)
|
| 281 |
+
anc_genome = genome_of(ancestor)
|
| 282 |
+
parents, parent_models, champion_genomes = [], [], []
|
| 283 |
+
for gen in range(gens):
|
| 284 |
+
members = []
|
| 285 |
+
for i in range(pop):
|
| 286 |
+
mseed = seed * 1000 + gen * 100 + i
|
| 287 |
+
m = make_organism(lineage, seed=mseed)
|
| 288 |
+
if anc_genome and gen == 0:
|
| 289 |
+
inherit_genome(m, [anc_genome]) # common ancestor
|
| 290 |
+
if gen > 0:
|
| 291 |
+
is_fresh = (i == pop - 1) # gene flow founder
|
| 292 |
+
if not is_fresh:
|
| 293 |
+
if lineage in ("aleph_flat", "aleph_tree"):
|
| 294 |
+
gs = [genome_of(pm) for pm in parent_models]
|
| 295 |
+
if catastrophic_at == gen:
|
| 296 |
+
bad = make_organism(lineage, seed=666 + i)
|
| 297 |
+
gs = gs + [genome_of(bad)]
|
| 298 |
+
inherit_genome(m, gs)
|
| 299 |
+
elif lineage == "aleph_full":
|
| 300 |
+
# v4 arm (v3 lesson: continuity is what pays) — inherit the
|
| 301 |
+
# WHOLE best parent, then overwrite the book with the
|
| 302 |
+
# two-parent consensus: germline ON TOP of continuity.
|
| 303 |
+
m.load_state_dict(copy.deepcopy(
|
| 304 |
+
parent_models[0].state_dict()))
|
| 305 |
+
gs = [genome_of(pm) for pm in parent_models]
|
| 306 |
+
if catastrophic_at == gen:
|
| 307 |
+
bad = make_organism(lineage, seed=666 + i)
|
| 308 |
+
gs = gs + [genome_of(bad)]
|
| 309 |
+
c, _, _ = consensus_codebook([g["flat"] for g in gs])
|
| 310 |
+
implant_book(m.head_addr, c)
|
| 311 |
+
elif lineage in ("mlp_kd", "aleph_weights"):
|
| 312 |
+
# pure continuity (no germline op) — aleph_weights is the
|
| 313 |
+
# within-architecture control for aleph_full
|
| 314 |
+
m.load_state_dict(copy.deepcopy(
|
| 315 |
+
parent_models[0].state_dict()))
|
| 316 |
+
# no_inherit: nothing
|
| 317 |
+
# KD: alpha 0.25 (campaign-v2 lesson: alpha=1.0 from near-parity
|
| 318 |
+
# teachers COMPOUNDS DOWNWARD — inverse evolution; fresh-founder
|
| 319 |
+
# control isolated it). Fresh founders get NO KD (clean gene flow).
|
| 320 |
+
is_fresh_now = (gen > 0 and i == pop - 1)
|
| 321 |
+
teachers = parent_models if (gen > 0 and not is_fresh_now
|
| 322 |
+
and lineage != "no_inherit") else None
|
| 323 |
+
bpb = train_member(m, tr, va, steps=steps, device=device,
|
| 324 |
+
seed=mseed, teachers=teachers, kd_alpha=0.25)
|
| 325 |
+
vit = organism_vitals(m)
|
| 326 |
+
members.append((bpb, m))
|
| 327 |
+
rec = {"exp": "14a", "lineage": lineage, "tseed": seed, "gen": gen,
|
| 328 |
+
"member": i, "fresh": gen > 0 and i == pop - 1,
|
| 329 |
+
"catastrophic": catastrophic_at == gen, "bpb": round(bpb, 4),
|
| 330 |
+
"vitals": vit}
|
| 331 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 332 |
+
print(f"[14a {lineage} t{seed} g{gen} m{i}] bpb={bpb:.4f} {vit}",
|
| 333 |
+
flush=True)
|
| 334 |
+
members.sort(key=lambda t: t[0])
|
| 335 |
+
parent_models = [members[0][1].cpu(), members[1][1].cpu()]
|
| 336 |
+
best = members[0][0]
|
| 337 |
+
gene = genome_of(members[0][1])
|
| 338 |
+
if gene:
|
| 339 |
+
torch.save(gene, os.path.join(
|
| 340 |
+
EXP_DIR, f"champion_{lineage}_t{seed}_g{gen}.pt"))
|
| 341 |
+
champion_genomes.append(gene)
|
| 342 |
+
print(f"[14a {lineage} t{seed} g{gen}] BEST={best:.4f} "
|
| 343 |
+
f"mean={sum(b for b, _ in members)/pop:.4f}", flush=True)
|
| 344 |
+
for _, mm in members[2:]:
|
| 345 |
+
del mm
|
| 346 |
+
torch.cuda.empty_cache()
|
| 347 |
+
ledger.close()
|
| 348 |
+
return best
|
| 349 |
+
|
| 350 |
+
|
| 351 |
+
# ============================================================ 14_b implants ======
|
| 352 |
+
def run_b1(steps: int = 2000, seed: int = 0, device: str = "cuda",
|
| 353 |
+
donor_book: torch.Tensor | None = None, tag: str = "small_cultivated"):
|
| 354 |
+
"""Cross-size: donor book (default: cultivate in a small organism) implanted
|
| 355 |
+
into a LARGE organism. Arms: fresh | implant-trainable | implant-frozen | mlp."""
|
| 356 |
+
if not torch.cuda.is_available():
|
| 357 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 358 |
+
os.makedirs(EXP_DIR, exist_ok=True)
|
| 359 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 360 |
+
ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 361 |
+
if donor_book is None:
|
| 362 |
+
small = make_organism("aleph_flat", d=128, layers=2, seed=seed)
|
| 363 |
+
bpb_small = train_member(small, tr, va, steps=steps, device=device, seed=seed)
|
| 364 |
+
donor_book = genome_of(small)["flat"]
|
| 365 |
+
print(f"[14b donor small] bpb={bpb_small:.4f}", flush=True)
|
| 366 |
+
results = {}
|
| 367 |
+
for arm in ("fresh", "implant_train", "implant_frozen", "mlp"):
|
| 368 |
+
lineage = "mlp_kd" if arm == "mlp" else "aleph_flat"
|
| 369 |
+
m = make_organism(lineage, d=384, layers=6, seed=seed + 10)
|
| 370 |
+
if arm.startswith("implant"):
|
| 371 |
+
implant_book(m.head_addr, donor_book, trainable=(arm == "implant_train"))
|
| 372 |
+
bpb = train_member(m, tr, va, steps=steps, device=device, seed=seed + 10)
|
| 373 |
+
vit = organism_vitals(m)
|
| 374 |
+
results[arm] = {"bpb": round(bpb, 4), "vitals": vit}
|
| 375 |
+
rec = {"exp": "14b1", "arm": arm, "donor": tag, "seed": seed,
|
| 376 |
+
"bpb": round(bpb, 4), "vitals": vit}
|
| 377 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 378 |
+
print(f"[14b1 {arm} donor={tag}] bpb={bpb:.4f} {vit}", flush=True)
|
| 379 |
+
del m; torch.cuda.empty_cache()
|
| 380 |
+
ledger.close()
|
| 381 |
+
return results, donor_book
|
| 382 |
+
|
| 383 |
+
|
| 384 |
+
def run_b2(donor_book: torch.Tensor, steps: int = 1500, seed: int = 0,
|
| 385 |
+
device: str = "cuda", tag: str = "small_cultivated"):
|
| 386 |
+
"""Cross-architecture: implant the donor book into every GPT-2 relay adapter
|
| 387 |
+
(exp013 Track C bed) vs random-init relays. Books are (64,4) — size-agnostic."""
|
| 388 |
+
from exp013_augmentation_bed import _wikitext_lines
|
| 389 |
+
from transformers import GPT2LMHeadModel, GPT2TokenizerFast
|
| 390 |
+
if "MslRelay" not in globals():
|
| 391 |
+
from ar_differentiation_bed import MslRelay
|
| 392 |
+
from exp013_augmentation_bed import _BlockWithAdapter
|
| 393 |
+
tok = GPT2TokenizerFast.from_pretrained("gpt2")
|
| 394 |
+
tr_lines, va_lines = _wikitext_lines(DATA_ROOT)
|
| 395 |
+
stream_tr = tok("\n\n".join(tr_lines[:8000]), return_tensors="pt").input_ids[0]
|
| 396 |
+
stream_va = tok("\n\n".join(va_lines[:1000]), return_tensors="pt").input_ids[0]
|
| 397 |
+
ledger = open(os.path.join(EXP_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 398 |
+
out = {}
|
| 399 |
+
for arm in ("random_relays", "implanted_relays"):
|
| 400 |
+
torch.manual_seed(seed)
|
| 401 |
+
g = torch.Generator().manual_seed(seed)
|
| 402 |
+
model = GPT2LMHeadModel.from_pretrained("gpt2").to(device)
|
| 403 |
+
for p in model.parameters():
|
| 404 |
+
p.requires_grad_(False)
|
| 405 |
+
adapters = []
|
| 406 |
+
for i, blk in enumerate(model.transformer.h):
|
| 407 |
+
ad = MslRelay(model.config.n_embd).to(device)
|
| 408 |
+
if arm == "implanted_relays":
|
| 409 |
+
implant_book(ad.addr, donor_book, trainable=True)
|
| 410 |
+
model.transformer.h[i] = _BlockWithAdapter(blk, ad)
|
| 411 |
+
adapters.append(ad)
|
| 412 |
+
params = [p for ad in adapters for p in ad.parameters()
|
| 413 |
+
if p.requires_grad]
|
| 414 |
+
opt = torch.optim.Adam(params, lr=1e-3, weight_decay=0.0)
|
| 415 |
+
block = 256
|
| 416 |
+
for step in range(1, steps + 1):
|
| 417 |
+
ix = torch.randint(0, stream_tr.numel() - block - 1, (8,), generator=g)
|
| 418 |
+
x = torch.stack([stream_tr[i:i + block] for i in ix]).to(device)
|
| 419 |
+
loss = model(x, labels=x).loss
|
| 420 |
+
opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
|
| 421 |
+
model.eval()
|
| 422 |
+
with torch.no_grad():
|
| 423 |
+
ls = []
|
| 424 |
+
for j in range(0, stream_va.numel() - block - 1, block * 4):
|
| 425 |
+
x = stream_va[j:j + block].unsqueeze(0).to(device)
|
| 426 |
+
ls.append(model(x, labels=x).loss.item())
|
| 427 |
+
ppl = math.exp(sum(ls) / len(ls))
|
| 428 |
+
gates = [round(ad.gate.sigmoid().item(), 4) for ad in adapters]
|
| 429 |
+
drifts = [round(anchor_drift(ad.addr.codebook, ad.addr.home)["mean"], 3)
|
| 430 |
+
for ad in adapters]
|
| 431 |
+
out[arm] = {"ppl": round(ppl, 3), "gates": gates, "drift": drifts}
|
| 432 |
+
rec = {"exp": "14b2", "arm": arm, "donor": tag, "seed": seed,
|
| 433 |
+
"ppl": round(ppl, 3), "gates": gates, "drift": drifts}
|
| 434 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 435 |
+
print(f"[14b2 {arm} donor={tag}] ppl={ppl:.3f} gates={gates[:3]}.. "
|
| 436 |
+
f"drift={drifts[:3]}..", flush=True)
|
| 437 |
+
del model; torch.cuda.empty_cache()
|
| 438 |
+
ledger.close()
|
| 439 |
+
return out
|
| 440 |
+
|
| 441 |
+
|
| 442 |
+
# ================================================================ smoke ===========
|
| 443 |
+
def smoke():
|
| 444 |
+
"""CPU shapes/parse only: GPA ground truth, tree causality, implant, KD path."""
|
| 445 |
+
g = torch.Generator().manual_seed(0)
|
| 446 |
+
# GPA: two rotated (+row-sign-flipped) copies of one book must align back to it
|
| 447 |
+
A = F.normalize(torch.randn(64, 4, generator=g), dim=-1)
|
| 448 |
+
q, _ = torch.linalg.qr(torch.randn(4, 4, generator=g))
|
| 449 |
+
B = A @ q
|
| 450 |
+
B[::3] = -B[::3]
|
| 451 |
+
C, it, dl = consensus_codebook([A, B])
|
| 452 |
+
cos = (F.normalize(C, dim=-1) * A).sum(-1).abs().mean()
|
| 453 |
+
assert cos > 0.999, cos
|
| 454 |
+
print(f"GPA OK (iters={it}, delta={dl:.2e}, |cos to truth|={cos:.5f})")
|
| 455 |
+
x = torch.randint(0, 256, (2, 64))
|
| 456 |
+
for lineage in ("aleph_flat", "aleph_tree", "mlp_kd", "no_inherit"):
|
| 457 |
+
m = make_organism(lineage, d=96, layers=2, block=64, seed=0)
|
| 458 |
+
lg = m(x); assert lg.shape == (2, 64, 256); lg.sum().backward()
|
| 459 |
+
with torch.no_grad():
|
| 460 |
+
a = m(x)[0, 10]; x2 = x.clone(); x2[0, 40] = (x2[0, 40] + 7) % 256
|
| 461 |
+
b = m(x2)[0, 10]
|
| 462 |
+
assert torch.allclose(a, b, atol=1e-4), lineage + " leaks"
|
| 463 |
+
gnm = genome_of(m)
|
| 464 |
+
if gnm:
|
| 465 |
+
inherit_genome(m, [gnm, gnm]) # self-consensus = identity-ish
|
| 466 |
+
print(lineage, "OK params",
|
| 467 |
+
f"{sum(p.numel() for p in m.parameters()):,}",
|
| 468 |
+
organism_vitals(m) if lineage != "mlp_kd" else {})
|
| 469 |
+
# KD path: teacher forward + KL backward
|
| 470 |
+
t = make_organism("mlp_kd", d=96, layers=2, block=64, seed=1)
|
| 471 |
+
s = make_organism("aleph_flat", d=96, layers=2, block=64, seed=2)
|
| 472 |
+
tp = F.softmax(t(x), -1).detach()
|
| 473 |
+
loss = F.kl_div(F.log_softmax(s(x), -1), tp, reduction="batchmean")
|
| 474 |
+
loss.backward()
|
| 475 |
+
print("KD OK — exp014 smoke passed (tournament on GPU: run_tournament(...))")
|
| 476 |
+
|
| 477 |
+
|
| 478 |
+
def _in_notebook():
|
| 479 |
+
try:
|
| 480 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 481 |
+
return True
|
| 482 |
+
except NameError:
|
| 483 |
+
return False
|
| 484 |
+
|
| 485 |
+
|
| 486 |
+
if __name__ == "__main__":
|
| 487 |
+
if _in_notebook():
|
| 488 |
+
smoke()
|
| 489 |
+
print("Notebook: run_tournament('aleph_flat'), run_b1(), run_b2(book).")
|
| 490 |
+
else:
|
| 491 |
+
import argparse
|
| 492 |
+
ap = argparse.ArgumentParser()
|
| 493 |
+
ap.add_argument("--mode", default="smoke",
|
| 494 |
+
choices=["smoke", "tournament", "b1", "b2"])
|
| 495 |
+
ap.add_argument("--lineage", default="aleph_full",
|
| 496 |
+
help="tournament lineage: aleph_flat|aleph_full|"
|
| 497 |
+
"aleph_weights|aleph_tree|mlp_kd|no_inherit")
|
| 498 |
+
ap.add_argument("--seed", type=int, default=0)
|
| 499 |
+
ap.add_argument("--steps", type=int, default=2000)
|
| 500 |
+
ap.add_argument("--genome", default="genomes/champion_aleph_full_t0_g3.pt",
|
| 501 |
+
help="donor genome .pt for --mode b1/b2 (uses its 'flat' book)")
|
| 502 |
+
a, _ = ap.parse_known_args()
|
| 503 |
+
if a.mode == "smoke":
|
| 504 |
+
smoke()
|
| 505 |
+
elif a.mode == "tournament":
|
| 506 |
+
run_tournament(a.lineage, steps=a.steps, seed=a.seed)
|
| 507 |
+
elif a.mode == "b1":
|
| 508 |
+
donor = (torch.load(a.genome, map_location="cpu")["flat"]
|
| 509 |
+
if os.path.exists(a.genome) else None)
|
| 510 |
+
run_b1(steps=a.steps, seed=a.seed, donor_book=donor,
|
| 511 |
+
tag=os.path.basename(a.genome) if donor is not None
|
| 512 |
+
else "small_cultivated")
|
| 513 |
+
elif a.mode == "b2":
|
| 514 |
+
donor = torch.load(a.genome, map_location="cpu")["flat"]
|
| 515 |
+
run_b2(donor, seed=a.seed, tag=os.path.basename(a.genome))
|
exp020_gen/exp017_aleph_constellation.py
ADDED
|
@@ -0,0 +1,266 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""exp017_aleph_constellation.py — PROTOTYPE ALEPH CONSTELLATION.
|
| 2 |
+
The constellation element for the aleph address: the 16s bridge (per the
|
| 3 |
+
constellation-diffusion-bottleneck article: 16 anchors x 16 patches x 3
|
| 4 |
+
stroboscope SLERP phases = the 768 address; aleph D=4 home lifted to the S^15
|
| 5 |
+
measurement sphere via 16 = 4^2) implemented as a byte-LM head — the first
|
| 6 |
+
configuration in the AR bed where the constellation gauges legitimately apply.
|
| 7 |
+
|
| 8 |
+
Funnel (per position, the 16s numerology intact on the aleph D=4 home):
|
| 9 |
+
h(d) -> proj -> 16 patches x 32 rows of D=4 (in_per_patch = 32*4 = 128)
|
| 10 |
+
-> AlephAddress M_hat read per row (farms the codebook)
|
| 11 |
+
-> crush MLP 128 -> 48 -> 16, SquaredReLU (8x squeeze, rule-of-3 hidden)
|
| 12 |
+
-> row-normalize = S^15 point per patch (the measurement sphere; 16 = 4^2)
|
| 13 |
+
-> triangulate 16 constellation anchors x 3 SLERP strobe phases
|
| 14 |
+
tri_k(t) = cos((1-t) * theta_k), t in {0, 1/3, 2/3} (acos in fp32)
|
| 15 |
+
-> address 16 patches x 16 anchors x 3 = 768
|
| 16 |
+
-> patchwork Linear(768,1536) -> SquaredReLU -> LN -> Linear(1536, VOCAB)
|
| 17 |
+
Constellation rules honored (per the Constellation Forms Catalogue,
|
| 18 |
+
AbstractPhil/geolip-constellation-activations): SquaredReLU in
|
| 19 |
+
all constellation paths (never GELU); patchwork spec; anchor dropout 30%;
|
| 20 |
+
Procrustes CALIBRATION of the anchor init (non-negotiable rule; ablated in
|
| 21 |
+
const_uncal); acos fp32; Adam wd=0. No gate: this is the whole output path, not
|
| 22 |
+
a residual entry.
|
| 23 |
+
|
| 24 |
+
JUDGED BY (16s law): constellation anchor drift -> 0.29154, crushed CV -> 0.20,
|
| 25 |
+
and task bpb — NEVER recon cosine. The aleph codebook underneath keeps its own
|
| 26 |
+
regime (CV~0.9 volatile home; recon/CE gradient the only pressure on it).
|
| 27 |
+
|
| 28 |
+
Arms:
|
| 29 |
+
const — CV strictly a readout (aleph-side discipline)
|
| 30 |
+
const_cv — + micro CV loss 1e-3 on the constellation bank (Form 3:
|
| 31 |
+
"CV ON THE BANK is load-bearing"; forward loss, never backward
|
| 32 |
+
injection) — tests whether the constellation rule transfers
|
| 33 |
+
const_uncal — calibration ablation (random anchors, no Procrustes calibration;
|
| 34 |
+
the "without: 1/256 anchors" rule test)
|
| 35 |
+
Baselines (exp012 certified ledger, cited not rerun): addr_msl64 2.4990,
|
| 36 |
+
sdpa 2.5182 @2k/seed0. Head is NOT param-matched to addr_msl64 (~2.0M vs ~114K)
|
| 37 |
+
— the prototype question is health + gauge placement, not a matched win.
|
| 38 |
+
Riders: pure Adam wd=0; GPU-only verdicts; >=2 seeds; Colab-safe. Paste order:
|
| 39 |
+
geolip_vitals -> ar_differentiation_bed -> this file.
|
| 40 |
+
"""
|
| 41 |
+
from __future__ import annotations
|
| 42 |
+
import json
|
| 43 |
+
import math
|
| 44 |
+
import os
|
| 45 |
+
import torch
|
| 46 |
+
import torch.nn as nn
|
| 47 |
+
import torch.nn.functional as F
|
| 48 |
+
|
| 49 |
+
if "ByteLM" not in globals():
|
| 50 |
+
try:
|
| 51 |
+
from ar_differentiation_bed import (ByteLM, AlephAddress, _wikitext_bytes,
|
| 52 |
+
_batch, VOCAB)
|
| 53 |
+
from geolip_vitals import anchor_drift, pentachoron_cv, BINDING
|
| 54 |
+
except ImportError:
|
| 55 |
+
_here = globals().get("__file__")
|
| 56 |
+
if _here is None:
|
| 57 |
+
raise ImportError("paste geolip_vitals.py + ar_differentiation_bed.py first")
|
| 58 |
+
import sys, pathlib
|
| 59 |
+
sys.path.insert(0, str(pathlib.Path(_here).parent))
|
| 60 |
+
from ar_differentiation_bed import (ByteLM, AlephAddress, _wikitext_bytes,
|
| 61 |
+
_batch, VOCAB)
|
| 62 |
+
from geolip_vitals import anchor_drift, pentachoron_cv, BINDING
|
| 63 |
+
|
| 64 |
+
DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
|
| 65 |
+
EXP17_DIR = os.path.join(DATA_ROOT, "exp017")
|
| 66 |
+
|
| 67 |
+
N_PATCHES, N_ANCH, V_ROWS, GEO = 16, 16, 32, 16 # 16 = 4^2 = d^2 (16s law)
|
| 68 |
+
PHASES = (0.0, 1.0 / 3.0, 2.0 / 3.0) # stroboscope SLERP, rule of 3
|
| 69 |
+
|
| 70 |
+
|
| 71 |
+
class SquaredReLU(nn.Module):
|
| 72 |
+
def forward(self, x):
|
| 73 |
+
return F.relu(x) ** 2
|
| 74 |
+
|
| 75 |
+
|
| 76 |
+
class ConstellationHead(nn.Module):
|
| 77 |
+
"""The 16s funnel over the aleph address (docstring above)."""
|
| 78 |
+
|
| 79 |
+
def __init__(self, d: int, vocab: int = 256, calibrated: bool = True,
|
| 80 |
+
anchor_dropout: float = 0.30):
|
| 81 |
+
super().__init__()
|
| 82 |
+
self.calibrated = calibrated
|
| 83 |
+
self.anchor_dropout = anchor_dropout
|
| 84 |
+
self.proj = nn.Linear(d, N_PATCHES * V_ROWS * 4, bias=False)
|
| 85 |
+
nn.init.orthogonal_(self.proj.weight)
|
| 86 |
+
self.addr = AlephAddress(64, 4) # the aleph layer beneath
|
| 87 |
+
self.crush = nn.Sequential( # 128 -> 48 -> 16 (8x squeeze)
|
| 88 |
+
nn.Linear(V_ROWS * 4, GEO * 3), SquaredReLU(),
|
| 89 |
+
nn.Linear(GEO * 3, GEO))
|
| 90 |
+
A = F.normalize(torch.randn(N_ANCH, GEO), dim=-1)
|
| 91 |
+
self.anchors = nn.Parameter(A) # S^15 constellation bank
|
| 92 |
+
self.register_buffer("anchors_home", A.clone())
|
| 93 |
+
tri = N_PATCHES * N_ANCH * len(PHASES) # 768, the address
|
| 94 |
+
self.patchwork = nn.Sequential(
|
| 95 |
+
nn.Linear(tri, tri * 2), SquaredReLU(),
|
| 96 |
+
nn.LayerNorm(tri * 2), nn.Linear(tri * 2, vocab))
|
| 97 |
+
self._calibrated_done = not calibrated
|
| 98 |
+
|
| 99 |
+
@torch.no_grad()
|
| 100 |
+
def calibrate(self, h_sample: torch.Tensor):
|
| 101 |
+
"""Procrustes calibration of the anchor init (NON-NEGOTIABLE constellation
|
| 102 |
+
rule): rotate the anchor template into the principal frame of the actual
|
| 103 |
+
crushed S^15 activations, then re-home. Runs once, before training."""
|
| 104 |
+
x = self._crushed(h_sample.reshape(-1, h_sample.shape[-1])) # (N, P, GEO)
|
| 105 |
+
x = x.reshape(-1, GEO).float()
|
| 106 |
+
# data frame: eigenvectors of the crushed covariance (fp64)
|
| 107 |
+
C = (x.T.double() @ x.double()) / x.shape[0]
|
| 108 |
+
_, Vd = torch.linalg.eigh(C)
|
| 109 |
+
A = self.anchors.double()
|
| 110 |
+
Ca = (A.T @ A) / A.shape[0]
|
| 111 |
+
_, Va = torch.linalg.eigh(Ca)
|
| 112 |
+
R = Va @ Vd.T # template frame -> data frame
|
| 113 |
+
A_cal = F.normalize((A @ R).float(), dim=-1)
|
| 114 |
+
self.anchors.data.copy_(A_cal)
|
| 115 |
+
self.anchors_home.copy_(A_cal)
|
| 116 |
+
self._calibrated_done = True
|
| 117 |
+
|
| 118 |
+
def _crushed(self, h):
|
| 119 |
+
rows = self.proj(h).view(*h.shape[:-1], N_PATCHES, V_ROWS, 4)
|
| 120 |
+
m_hat = self.addr.m_hat(rows) # aleph read per row
|
| 121 |
+
flat = m_hat.reshape(*h.shape[:-1], N_PATCHES, V_ROWS * 4)
|
| 122 |
+
return F.normalize(self.crush(flat), dim=-1) # (..., P, GEO) on S^15
|
| 123 |
+
|
| 124 |
+
def forward(self, h):
|
| 125 |
+
x = self._crushed(h) # (..., P, GEO)
|
| 126 |
+
A = F.normalize(self.anchors, dim=-1)
|
| 127 |
+
cos = (x @ A.T).clamp(-1 + 1e-6, 1 - 1e-6) # (..., P, K)
|
| 128 |
+
theta = torch.acos(cos.float()) # acos in fp32
|
| 129 |
+
tri = torch.cat([torch.cos((1.0 - t) * theta) for t in PHASES],
|
| 130 |
+
dim=-1).to(h.dtype) # (..., P, K*3)
|
| 131 |
+
if self.training and self.anchor_dropout > 0:
|
| 132 |
+
keep = (torch.rand(N_ANCH, device=h.device)
|
| 133 |
+
> self.anchor_dropout).float()
|
| 134 |
+
tri = tri * keep.repeat(len(PHASES))
|
| 135 |
+
return self.patchwork(tri.reshape(*h.shape[:-1],
|
| 136 |
+
N_PATCHES * N_ANCH * len(PHASES)))
|
| 137 |
+
|
| 138 |
+
@torch.no_grad()
|
| 139 |
+
def constellation_vitals(self) -> dict:
|
| 140 |
+
d = anchor_drift(self.anchors, self.anchors_home)
|
| 141 |
+
return {"anchor_drift": round(d["mean"], 4),
|
| 142 |
+
"binding_frac": round(d["binding_fraction"], 4),
|
| 143 |
+
"crushed_cv": round(pentachoron_cv(self.anchors.detach()), 4),
|
| 144 |
+
"aleph": {"drift": round(anchor_drift(
|
| 145 |
+
self.addr.codebook, self.addr.home)["mean"], 4)}}
|
| 146 |
+
|
| 147 |
+
|
| 148 |
+
def make_const_model(arm: str, d: int = 192, layers: int = 4,
|
| 149 |
+
block: int = 256, seed: int = 0):
|
| 150 |
+
torch.manual_seed(seed)
|
| 151 |
+
m = ByteLM("sdpa", d=d, layers=layers, block=block)
|
| 152 |
+
m.head = ConstellationHead(d, calibrated=(arm != "const_uncal"))
|
| 153 |
+
return m
|
| 154 |
+
|
| 155 |
+
|
| 156 |
+
def cv_bank_loss(head: ConstellationHead) -> torch.Tensor:
|
| 157 |
+
"""Micro CV loss on the CONSTELLATION bank only (const_cv arm; Form 3).
|
| 158 |
+
Forward loss (differentiable through the anchors); never touches the aleph
|
| 159 |
+
codebook, whose only pressure stays the CE gradient through M_hat."""
|
| 160 |
+
A = F.normalize(head.anchors, dim=-1)
|
| 161 |
+
n = A.shape[0]
|
| 162 |
+
g = torch.Generator(device="cpu").manual_seed(0)
|
| 163 |
+
idx = torch.stack([torch.randperm(n, generator=g)[:5] for _ in range(64)])
|
| 164 |
+
pts = A[idx] # (64, 5, GEO)
|
| 165 |
+
d2 = torch.cdist(pts.double(), pts.double()).pow(2)
|
| 166 |
+
cm = torch.ones(64, 6, 6, dtype=torch.float64, device=A.device)
|
| 167 |
+
cm[:, 0, 0] = 0.0
|
| 168 |
+
cm[:, 1:, 1:] = d2
|
| 169 |
+
v = (-torch.linalg.det(cm) / 9216.0).clamp_min(1e-24).sqrt()
|
| 170 |
+
return (v.std() / v.mean().clamp_min(1e-12)).float()
|
| 171 |
+
|
| 172 |
+
|
| 173 |
+
def train_const(model, tr, va, arm: str, steps=2000, batch=32, block=256,
|
| 174 |
+
device="cuda", seed=0, cv_weight=1e-3, log_every=500,
|
| 175 |
+
ledger=None, tag=None):
|
| 176 |
+
g = torch.Generator().manual_seed(seed)
|
| 177 |
+
model = model.to(device)
|
| 178 |
+
if not model.head._calibrated_done: # calibration pass (1 batch)
|
| 179 |
+
x, _ = _batch(tr, batch, block, device, g)
|
| 180 |
+
model(x)
|
| 181 |
+
model.head.calibrate(model._last_h)
|
| 182 |
+
opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
|
| 183 |
+
for step in range(1, steps + 1):
|
| 184 |
+
x, y = _batch(tr, batch, block, device, g)
|
| 185 |
+
logits = model(x)
|
| 186 |
+
loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
|
| 187 |
+
if arm == "const_cv":
|
| 188 |
+
loss = loss + cv_weight * cv_bank_loss(model.head)
|
| 189 |
+
opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
|
| 190 |
+
if step % log_every == 0:
|
| 191 |
+
cv = model.head.constellation_vitals()
|
| 192 |
+
print(f"[17 {tag or arm} step {step}] {cv}", flush=True)
|
| 193 |
+
model.eval()
|
| 194 |
+
with torch.no_grad():
|
| 195 |
+
ls = []
|
| 196 |
+
for _ in range(20):
|
| 197 |
+
xv, yv = _batch(va, batch, block, device, g)
|
| 198 |
+
ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
|
| 199 |
+
yv.reshape(-1)).item())
|
| 200 |
+
bpb = sum(ls) / len(ls) / math.log(2)
|
| 201 |
+
cv = model.head.constellation_vitals()
|
| 202 |
+
if ledger is not None:
|
| 203 |
+
rec = {"exp": "17", "arm": arm, "seed": seed, "steps": steps,
|
| 204 |
+
"bpb": round(bpb, 4), "constellation": cv}
|
| 205 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 206 |
+
print(f"[17 {tag or arm} s{seed}] FINAL bpb={bpb:.4f} {cv}", flush=True)
|
| 207 |
+
return bpb
|
| 208 |
+
|
| 209 |
+
|
| 210 |
+
def run_prototype(arms=("const", "const_cv", "const_uncal"), seeds=(0, 1),
|
| 211 |
+
steps=2000, device="cuda"):
|
| 212 |
+
if not torch.cuda.is_available():
|
| 213 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 214 |
+
os.makedirs(EXP17_DIR, exist_ok=True)
|
| 215 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 216 |
+
ledger = open(os.path.join(EXP17_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 217 |
+
for seed in seeds:
|
| 218 |
+
for arm in arms:
|
| 219 |
+
m = make_const_model(arm, seed=seed)
|
| 220 |
+
train_const(m, tr, va, arm, steps=steps, device=device, seed=seed,
|
| 221 |
+
ledger=ledger, tag=f"{arm}")
|
| 222 |
+
torch.save({"arm": arm, "seed": seed, "steps": steps,
|
| 223 |
+
"state_dict": {k: v.cpu() for k, v in
|
| 224 |
+
m.state_dict().items()}},
|
| 225 |
+
os.path.join(EXP17_DIR, f"const_{arm}_s{seed}.pt"))
|
| 226 |
+
del m
|
| 227 |
+
torch.cuda.empty_cache()
|
| 228 |
+
ledger.close()
|
| 229 |
+
|
| 230 |
+
|
| 231 |
+
def smoke():
|
| 232 |
+
x = torch.randint(0, 256, (2, 64))
|
| 233 |
+
m = make_const_model("const", d=96, layers=2, block=64, seed=0)
|
| 234 |
+
lg = m(x)
|
| 235 |
+
assert lg.shape == (2, 64, 256)
|
| 236 |
+
lg.sum().backward()
|
| 237 |
+
assert m.head.addr.codebook.grad is not None # aleph farms through M_hat
|
| 238 |
+
assert m.head.anchors.grad is not None # constellation trains
|
| 239 |
+
m.zero_grad()
|
| 240 |
+
m(x)
|
| 241 |
+
m.head.calibrate(m._last_h) # calibration runs
|
| 242 |
+
cv = m.head.constellation_vitals()
|
| 243 |
+
assert "crushed_cv" in cv and cv["anchor_drift"] == 0.0
|
| 244 |
+
l = cv_bank_loss(m.head)
|
| 245 |
+
assert l.requires_grad and l.item() > 0
|
| 246 |
+
# causality: future byte must not affect past logits
|
| 247 |
+
x2 = x.clone(); x2[:, -1] = (x2[:, -1] + 1) % 256
|
| 248 |
+
m.eval()
|
| 249 |
+
with torch.no_grad():
|
| 250 |
+
a, b = m(x), m(x2)
|
| 251 |
+
assert torch.allclose(a[:, :-1], b[:, :-1], atol=1e-5)
|
| 252 |
+
print("exp017 smoke passed —", {k: cv[k] for k in ("crushed_cv",
|
| 253 |
+
"binding_frac")})
|
| 254 |
+
|
| 255 |
+
|
| 256 |
+
def _in_notebook():
|
| 257 |
+
try:
|
| 258 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 259 |
+
return True
|
| 260 |
+
except NameError:
|
| 261 |
+
return False
|
| 262 |
+
|
| 263 |
+
|
| 264 |
+
if __name__ == "__main__":
|
| 265 |
+
smoke() if not _in_notebook() else (smoke(),
|
| 266 |
+
print("Notebook: run_prototype() on GPU."))
|
exp020_gen/exp019_content_retention.py
ADDED
|
@@ -0,0 +1,402 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""exp019_content_retention.py — CAPACITY FOR DISTILLED CONTENT RETENTION.
|
| 2 |
+
When content is DISTILLED rather than directly learned, how much is retained,
|
| 3 |
+
through which channel, and for how long? The memory-substrate compass made a
|
| 4 |
+
measurable instrument — and the weak-to-strong distillation cell made real: the
|
| 5 |
+
teacher holds content the student lacks (true headroom on the content axis).
|
| 6 |
+
|
| 7 |
+
THE CONTENT: N fact records "\n@<key6>=<value12>\n" with random alphanumeric
|
| 8 |
+
keys/values — uncompletable from language statistics, so exact-match completion
|
| 9 |
+
IS retention. Facts are mixed into the byte stream (fact-packed blocks at
|
| 10 |
+
FACT_RATE against wikitext blocks).
|
| 11 |
+
|
| 12 |
+
THE TEACHER: the certified bed model (addr_msl64 aleph substrate) trained
|
| 13 |
+
TEACHER_STEPS on the mix; gated on its own recall (the gate doubles as the
|
| 14 |
+
substrate's DIRECT capacity datum at each N).
|
| 15 |
+
|
| 16 |
+
THE CHANNELS (fresh student each, STUDENT_STEPS budget):
|
| 17 |
+
direct — ground-truth CE on the mix (ceiling: learning, not distillation)
|
| 18 |
+
kd_facts — CE on wikitext blocks; on fact blocks the ONLY signal is the
|
| 19 |
+
teacher's logits (KL) — pure distilled content
|
| 20 |
+
kd_general — CE + KL to teacher on CLEAN wikitext only; facts never shown —
|
| 21 |
+
does content leak through logits without exposure?
|
| 22 |
+
book_implant — teacher's farmed codebook implanted (trainable) into a fresh
|
| 23 |
+
student, clean-stream training — do the anchors carry
|
| 24 |
+
byte-content? (the open question from the exp014 implant studies, asked directly)
|
| 25 |
+
none — clean-stream only (floor)
|
| 26 |
+
THE AXES: capacity N in {64, 256, 1024} (main channels); retention = recall
|
| 27 |
+
right after training AND after INTERFERE_STEPS further clean-stream steps
|
| 28 |
+
(the forgetting measurement). KD alpha 1.0 here is LEGAL: the teacher has real
|
| 29 |
+
headroom on the content axis (the inverse-evolution failure was alpha 1.0 at
|
| 30 |
+
NEAR-PARITY — regime, not constant).
|
| 31 |
+
|
| 32 |
+
Preregistered forks:
|
| 33 |
+
F1 capacity curve: direct recall vs N = the substrate's raw content capacity.
|
| 34 |
+
F2 distillation tax: kd_facts vs direct at each N (what survives the logit
|
| 35 |
+
channel).
|
| 36 |
+
F3 leakage: kd_general recall > floor => content crosses on clean text alone.
|
| 37 |
+
F4 anchors: book_implant recall ~ floor => codebooks do not carry byte
|
| 38 |
+
content (mean-shape/content question closed in the direct sense).
|
| 39 |
+
F5 half-life: post-interference retention per channel (does distilled content
|
| 40 |
+
decay faster than learned content?).
|
| 41 |
+
Riders: pure Adam wd=0; GPU-only verdicts; >=2 seeds; Colab-safe. Paste order:
|
| 42 |
+
geolip_vitals -> ar_differentiation_bed -> exp014_genetic_distillation ->
|
| 43 |
+
this file.
|
| 44 |
+
"""
|
| 45 |
+
from __future__ import annotations
|
| 46 |
+
import json
|
| 47 |
+
import math
|
| 48 |
+
import os
|
| 49 |
+
import string
|
| 50 |
+
import torch
|
| 51 |
+
import torch.nn.functional as F
|
| 52 |
+
|
| 53 |
+
if "ByteLM" not in globals():
|
| 54 |
+
try:
|
| 55 |
+
from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
|
| 56 |
+
from exp014_genetic_distillation import implant_book
|
| 57 |
+
from geolip_vitals import anchor_drift
|
| 58 |
+
except ImportError:
|
| 59 |
+
_here = globals().get("__file__")
|
| 60 |
+
if _here is None:
|
| 61 |
+
raise ImportError("paste geolip_vitals + ar_differentiation_bed + "
|
| 62 |
+
"exp014_genetic_distillation first")
|
| 63 |
+
import sys, pathlib
|
| 64 |
+
sys.path.insert(0, str(pathlib.Path(_here).parent))
|
| 65 |
+
from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
|
| 66 |
+
from exp014_genetic_distillation import implant_book
|
| 67 |
+
from geolip_vitals import anchor_drift
|
| 68 |
+
|
| 69 |
+
DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
|
| 70 |
+
EXP19_DIR = os.path.join(DATA_ROOT, "exp019")
|
| 71 |
+
|
| 72 |
+
KEY_LEN, VAL_LEN = 6, 12
|
| 73 |
+
FACT_RATE = 0.5 # fraction of training blocks drawn from fact stream
|
| 74 |
+
TEACHER_STEPS = 4000
|
| 75 |
+
STUDENT_STEPS = 2000
|
| 76 |
+
INTERFERE_STEPS = 1000
|
| 77 |
+
ALNUM = (string.ascii_lowercase + string.digits).encode()
|
| 78 |
+
|
| 79 |
+
|
| 80 |
+
def make_facts(n: int, seed: int = 0):
|
| 81 |
+
"""N records '\\n@<key>=<value>\\n'; returns (records list, fact byte stream)."""
|
| 82 |
+
g = torch.Generator().manual_seed(4000 + seed)
|
| 83 |
+
recs = []
|
| 84 |
+
seen = set()
|
| 85 |
+
while len(recs) < n:
|
| 86 |
+
k = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (KEY_LEN,),
|
| 87 |
+
generator=g))
|
| 88 |
+
if k in seen:
|
| 89 |
+
continue
|
| 90 |
+
seen.add(k)
|
| 91 |
+
v = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (VAL_LEN,),
|
| 92 |
+
generator=g))
|
| 93 |
+
recs.append((k, v))
|
| 94 |
+
return recs
|
| 95 |
+
|
| 96 |
+
|
| 97 |
+
def fact_stream(recs, copies: int = 50, seed: int = 0) -> torch.Tensor:
|
| 98 |
+
"""Byte stream of shuffled fact records (each record appears `copies` times)."""
|
| 99 |
+
g = torch.Generator().manual_seed(5000 + seed)
|
| 100 |
+
order = torch.cat([torch.randperm(len(recs), generator=g)
|
| 101 |
+
for _ in range(copies)])
|
| 102 |
+
blob = b"".join(b"\n@" + recs[i][0] + b"=" + recs[i][1] + b"\n"
|
| 103 |
+
for i in order.tolist())
|
| 104 |
+
return torch.frombuffer(bytearray(blob), dtype=torch.uint8).clone()
|
| 105 |
+
|
| 106 |
+
|
| 107 |
+
def _mix_batch(tr, fs, batch, block, device, g):
|
| 108 |
+
"""Blocks drawn from the fact stream with prob FACT_RATE, else wikitext.
|
| 109 |
+
Returns (x, y, fact_mask (B,)) — mask marks fact-sourced rows."""
|
| 110 |
+
xw, yw = _batch(tr, batch, block, device, g)
|
| 111 |
+
xf, yf = _batch(fs, batch, block, device, g)
|
| 112 |
+
m = (torch.rand(batch, generator=g) < FACT_RATE).to(device)
|
| 113 |
+
x = torch.where(m[:, None], xf, xw)
|
| 114 |
+
y = torch.where(m[:, None], yf, yw)
|
| 115 |
+
return x, y, m
|
| 116 |
+
|
| 117 |
+
|
| 118 |
+
@torch.no_grad()
|
| 119 |
+
def recall(model, recs, device="cuda", max_eval: int = 256,
|
| 120 |
+
batch: int = 64) -> dict:
|
| 121 |
+
"""Exact-match greedy completion: prompt '\\n@<key>=' -> VAL_LEN bytes."""
|
| 122 |
+
model = model.to(device).eval()
|
| 123 |
+
recs = recs[:max_eval]
|
| 124 |
+
prompts = torch.stack([torch.frombuffer(
|
| 125 |
+
bytearray(b"\n@" + k + b"="), dtype=torch.uint8).long()
|
| 126 |
+
for k, _ in recs]).to(device)
|
| 127 |
+
outs = []
|
| 128 |
+
for i in range(0, len(recs), batch):
|
| 129 |
+
x = prompts[i:i + batch]
|
| 130 |
+
for _ in range(VAL_LEN):
|
| 131 |
+
nxt = model(x)[:, -1].argmax(-1, keepdim=True)
|
| 132 |
+
x = torch.cat([x, nxt], dim=1)
|
| 133 |
+
outs.append(x[:, -VAL_LEN:].cpu())
|
| 134 |
+
got = torch.cat(outs)
|
| 135 |
+
tgt = torch.stack([torch.frombuffer(bytearray(v), dtype=torch.uint8).long()
|
| 136 |
+
for _, v in recs])
|
| 137 |
+
byte_acc = (got == tgt).float().mean().item()
|
| 138 |
+
exact = (got == tgt).all(dim=1).float().mean().item()
|
| 139 |
+
return {"exact": round(exact, 4), "byte_acc": round(byte_acc, 4)}
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
def train_stream(model, tr, va, fs=None, teacher=None, channel="direct",
|
| 143 |
+
steps=2000, batch=32, block=256, device="cuda", seed=0):
|
| 144 |
+
"""One training run under a channel's signal routing (docstring above)."""
|
| 145 |
+
g = torch.Generator().manual_seed(seed)
|
| 146 |
+
model = model.to(device)
|
| 147 |
+
if teacher is not None:
|
| 148 |
+
teacher = teacher.to(device).eval()
|
| 149 |
+
opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
|
| 150 |
+
for step in range(1, steps + 1):
|
| 151 |
+
if channel in ("direct", "kd_facts") and fs is not None:
|
| 152 |
+
x, y, m = _mix_batch(tr, fs, batch, block, device, g)
|
| 153 |
+
else: # clean wikitext stream
|
| 154 |
+
x, y = _batch(tr, batch, block, device, g)
|
| 155 |
+
m = torch.zeros(batch, dtype=torch.bool, device=device)
|
| 156 |
+
logits = model(x)
|
| 157 |
+
if channel == "kd_facts":
|
| 158 |
+
# ground truth on wiki rows only; teacher logits are the ONLY
|
| 159 |
+
# signal on fact rows (pure distilled content)
|
| 160 |
+
ce_rows = ~m
|
| 161 |
+
loss = torch.tensor(0.0, device=device)
|
| 162 |
+
if ce_rows.any():
|
| 163 |
+
loss = F.cross_entropy(logits[ce_rows].reshape(-1, VOCAB),
|
| 164 |
+
y[ce_rows].reshape(-1))
|
| 165 |
+
if m.any():
|
| 166 |
+
with torch.no_grad():
|
| 167 |
+
tp = F.softmax(teacher(x[m]), -1)
|
| 168 |
+
loss = loss + F.kl_div(F.log_softmax(logits[m], -1), tp,
|
| 169 |
+
reduction="batchmean")
|
| 170 |
+
elif channel == "kd_general":
|
| 171 |
+
loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
|
| 172 |
+
with torch.no_grad():
|
| 173 |
+
tp = F.softmax(teacher(x), -1)
|
| 174 |
+
loss = loss + F.kl_div(F.log_softmax(logits, -1), tp,
|
| 175 |
+
reduction="batchmean")
|
| 176 |
+
else: # direct / book_implant / none
|
| 177 |
+
loss = F.cross_entropy(logits.reshape(-1, VOCAB), y.reshape(-1))
|
| 178 |
+
opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
|
| 179 |
+
model.eval()
|
| 180 |
+
with torch.no_grad():
|
| 181 |
+
ls = []
|
| 182 |
+
for _ in range(10):
|
| 183 |
+
xv, yv = _batch(va, batch, block, device, g)
|
| 184 |
+
ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
|
| 185 |
+
yv.reshape(-1)).item())
|
| 186 |
+
return sum(ls) / len(ls) / math.log(2)
|
| 187 |
+
|
| 188 |
+
|
| 189 |
+
def run_retention(n_facts=(64, 256, 1024), seeds=(0, 1), device="cuda"):
|
| 190 |
+
if not torch.cuda.is_available():
|
| 191 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 192 |
+
os.makedirs(EXP19_DIR, exist_ok=True)
|
| 193 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 194 |
+
ledger = open(os.path.join(EXP19_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 195 |
+
|
| 196 |
+
def log(rec):
|
| 197 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 198 |
+
print(f"[19 {rec['channel']} N={rec['n']} s{rec['seed']}] "
|
| 199 |
+
f"recall={rec['recall']} after_interf={rec.get('recall_interf')} "
|
| 200 |
+
f"bpb={rec['bpb']}", flush=True)
|
| 201 |
+
|
| 202 |
+
for seed in seeds:
|
| 203 |
+
for n in n_facts:
|
| 204 |
+
recs = make_facts(n, seed=seed)
|
| 205 |
+
fs = fact_stream(recs, seed=seed)
|
| 206 |
+
# ---- teacher (also the DIRECT capacity datum at TEACHER_STEPS)
|
| 207 |
+
torch.manual_seed(seed)
|
| 208 |
+
teacher = ByteLM("addr_msl64")
|
| 209 |
+
t_bpb = train_stream(teacher, tr, va, fs=fs, channel="direct",
|
| 210 |
+
steps=TEACHER_STEPS, device=device, seed=seed)
|
| 211 |
+
t_rec = recall(teacher, recs, device=device)
|
| 212 |
+
log({"exp": "19", "channel": "teacher", "n": n, "seed": seed,
|
| 213 |
+
"steps": TEACHER_STEPS, "recall": t_rec, "bpb": round(t_bpb, 4)})
|
| 214 |
+
torch.save({"n": n, "seed": seed,
|
| 215 |
+
"state_dict": {k: v.cpu() for k, v in
|
| 216 |
+
teacher.state_dict().items()}},
|
| 217 |
+
os.path.join(EXP19_DIR, f"teacher_N{n}_s{seed}.pt"))
|
| 218 |
+
# ---- channels
|
| 219 |
+
chans = ["direct", "kd_facts", "kd_general"]
|
| 220 |
+
if n == 256:
|
| 221 |
+
chans += ["book_implant", "none"]
|
| 222 |
+
for ch in chans:
|
| 223 |
+
torch.manual_seed(1000 + seed)
|
| 224 |
+
m = ByteLM("addr_msl64")
|
| 225 |
+
if ch == "book_implant":
|
| 226 |
+
implant_book(m.head_addr,
|
| 227 |
+
teacher.head_addr.codebook.detach().cpu())
|
| 228 |
+
bpb = train_stream(
|
| 229 |
+
m, tr, va, fs=fs if ch in ("direct", "kd_facts") else None,
|
| 230 |
+
teacher=teacher if ch.startswith("kd") else None,
|
| 231 |
+
channel=ch, steps=STUDENT_STEPS, device=device,
|
| 232 |
+
seed=1000 + seed)
|
| 233 |
+
r0 = recall(m, recs, device=device)
|
| 234 |
+
# retention under interference: further CLEAN-stream training
|
| 235 |
+
bpb2 = train_stream(m, tr, va, channel="none",
|
| 236 |
+
steps=INTERFERE_STEPS, device=device,
|
| 237 |
+
seed=2000 + seed)
|
| 238 |
+
r1 = recall(m, recs, device=device)
|
| 239 |
+
log({"exp": "19", "channel": ch, "n": n, "seed": seed,
|
| 240 |
+
"steps": STUDENT_STEPS, "recall": r0, "recall_interf": r1,
|
| 241 |
+
"bpb": round(bpb, 4), "bpb_after_interf": round(bpb2, 4)})
|
| 242 |
+
del m
|
| 243 |
+
torch.cuda.empty_cache()
|
| 244 |
+
del teacher
|
| 245 |
+
torch.cuda.empty_cache()
|
| 246 |
+
ledger.close()
|
| 247 |
+
|
| 248 |
+
|
| 249 |
+
# ==================== exp019b — GENERALIZATION block =========================
|
| 250 |
+
# Rule-bearing content: value = fixed random substitution cipher applied to the
|
| 251 |
+
# key, extended to VAL_LEN (v[i] = subst(k[i % KEY_LEN])). Teacher sees
|
| 252 |
+
# N_TRAIN rule-keys; N_TEST keys are HELD OUT. Held-out recall = the RULE
|
| 253 |
+
# generalizing, not the list. Sharp question: does the logit channel transfer
|
| 254 |
+
# the rule better than it transfers the rote list? Plus prompt-format variants
|
| 255 |
+
# (content vs surface form disentangled).
|
| 256 |
+
|
| 257 |
+
def make_rule_facts(n_train: int = 256, n_test: int = 128, seed: int = 0):
|
| 258 |
+
g = torch.Generator().manual_seed(6000 + seed)
|
| 259 |
+
subst = {ALNUM[i]: ALNUM[j] for i, j in
|
| 260 |
+
enumerate(torch.randperm(len(ALNUM), generator=g).tolist())}
|
| 261 |
+
keys, seen = [], set()
|
| 262 |
+
while len(keys) < n_train + n_test:
|
| 263 |
+
k = bytes(ALNUM[i] for i in torch.randint(len(ALNUM), (KEY_LEN,),
|
| 264 |
+
generator=g))
|
| 265 |
+
if k not in seen:
|
| 266 |
+
seen.add(k)
|
| 267 |
+
keys.append(k)
|
| 268 |
+
def val(k):
|
| 269 |
+
return bytes(subst[k[i % KEY_LEN]] for i in range(VAL_LEN))
|
| 270 |
+
train = [(k, val(k)) for k in keys[:n_train]]
|
| 271 |
+
test = [(k, val(k)) for k in keys[n_train:]]
|
| 272 |
+
return train, test
|
| 273 |
+
|
| 274 |
+
|
| 275 |
+
@torch.no_grad()
|
| 276 |
+
def recall_fmt(model, recs, fmt: bytes = b"\n@%s=", device="cuda",
|
| 277 |
+
max_eval: int = 256, batch: int = 64) -> dict:
|
| 278 |
+
"""recall() under an arbitrary prompt format (b'\\n@%s=' = the training
|
| 279 |
+
format; variants probe surface-form generalization)."""
|
| 280 |
+
model = model.to(device).eval()
|
| 281 |
+
recs = recs[:max_eval]
|
| 282 |
+
proms = [torch.frombuffer(bytearray(fmt.replace(b"%s", k)),
|
| 283 |
+
dtype=torch.uint8).long() for k, _ in recs]
|
| 284 |
+
L = max(p.numel() for p in proms)
|
| 285 |
+
# left-pad with newlines to equal length (causal — padding is prefix noise)
|
| 286 |
+
prompts = torch.stack([torch.cat([torch.full((L - p.numel(),), 10,
|
| 287 |
+
dtype=torch.long), p])
|
| 288 |
+
for p in proms]).to(device)
|
| 289 |
+
outs = []
|
| 290 |
+
for i in range(0, len(recs), batch):
|
| 291 |
+
x = prompts[i:i + batch]
|
| 292 |
+
for _ in range(VAL_LEN):
|
| 293 |
+
nxt = model(x)[:, -1].argmax(-1, keepdim=True)
|
| 294 |
+
x = torch.cat([x, nxt], dim=1)
|
| 295 |
+
outs.append(x[:, -VAL_LEN:].cpu())
|
| 296 |
+
got = torch.cat(outs)
|
| 297 |
+
tgt = torch.stack([torch.frombuffer(bytearray(v), dtype=torch.uint8).long()
|
| 298 |
+
for _, v in recs])
|
| 299 |
+
return {"exact": round((got == tgt).all(dim=1).float().mean().item(), 4),
|
| 300 |
+
"byte_acc": round((got == tgt).float().mean().item(), 4)}
|
| 301 |
+
|
| 302 |
+
|
| 303 |
+
FMT_TRAIN = b"\n@%s="
|
| 304 |
+
FMT_VARIANT = b" @%s= " # never seen in training: pure format shift
|
| 305 |
+
|
| 306 |
+
|
| 307 |
+
def run_generalization(n_train: int = 256, n_test: int = 128, seeds=(0, 1),
|
| 308 |
+
device="cuda"):
|
| 309 |
+
if not torch.cuda.is_available():
|
| 310 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 311 |
+
os.makedirs(EXP19_DIR, exist_ok=True)
|
| 312 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 313 |
+
ledger = open(os.path.join(EXP19_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 314 |
+
|
| 315 |
+
def gauges(model, train_recs, test_recs):
|
| 316 |
+
return {"train": recall_fmt(model, train_recs, FMT_TRAIN, device=device),
|
| 317 |
+
"heldout": recall_fmt(model, test_recs, FMT_TRAIN, device=device),
|
| 318 |
+
"train_varfmt": recall_fmt(model, train_recs, FMT_VARIANT,
|
| 319 |
+
device=device)}
|
| 320 |
+
|
| 321 |
+
for seed in seeds:
|
| 322 |
+
train_recs, test_recs = make_rule_facts(n_train, n_test, seed=seed)
|
| 323 |
+
fs = fact_stream(train_recs, seed=seed) # held-out NEVER streamed
|
| 324 |
+
torch.manual_seed(seed)
|
| 325 |
+
teacher = ByteLM("addr_msl64")
|
| 326 |
+
t_bpb = train_stream(teacher, tr, va, fs=fs, channel="direct",
|
| 327 |
+
steps=TEACHER_STEPS, device=device, seed=seed)
|
| 328 |
+
gt = gauges(teacher, train_recs, test_recs)
|
| 329 |
+
rec = {"exp": "19b", "channel": "teacher", "n": n_train, "seed": seed,
|
| 330 |
+
"steps": TEACHER_STEPS, "gauges": gt, "bpb": round(t_bpb, 4)}
|
| 331 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 332 |
+
print(f"[19b teacher s{seed}] {gt} bpb={t_bpb:.4f}", flush=True)
|
| 333 |
+
torch.save({"seed": seed, "state_dict": {k: v.cpu() for k, v in
|
| 334 |
+
teacher.state_dict().items()}},
|
| 335 |
+
os.path.join(EXP19_DIR, f"rule_teacher_s{seed}.pt"))
|
| 336 |
+
for ch in ("direct", "kd_facts", "kd_general"):
|
| 337 |
+
torch.manual_seed(1000 + seed)
|
| 338 |
+
m = ByteLM("addr_msl64")
|
| 339 |
+
bpb = train_stream(
|
| 340 |
+
m, tr, va, fs=fs if ch in ("direct", "kd_facts") else None,
|
| 341 |
+
teacher=teacher if ch.startswith("kd") else None,
|
| 342 |
+
channel=ch, steps=STUDENT_STEPS, device=device, seed=1000 + seed)
|
| 343 |
+
g0 = gauges(m, train_recs, test_recs)
|
| 344 |
+
bpb2 = train_stream(m, tr, va, channel="none",
|
| 345 |
+
steps=INTERFERE_STEPS, device=device,
|
| 346 |
+
seed=2000 + seed)
|
| 347 |
+
g1 = gauges(m, train_recs, test_recs)
|
| 348 |
+
rec = {"exp": "19b", "channel": ch, "n": n_train, "seed": seed,
|
| 349 |
+
"steps": STUDENT_STEPS, "gauges": g0, "gauges_interf": g1,
|
| 350 |
+
"bpb": round(bpb, 4), "bpb_after_interf": round(bpb2, 4)}
|
| 351 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 352 |
+
print(f"[19b {ch} s{seed}] {g0} interf_heldout="
|
| 353 |
+
f"{g1['heldout']} bpb={bpb:.4f}", flush=True)
|
| 354 |
+
torch.save({"channel": ch, "seed": seed,
|
| 355 |
+
"state_dict": {k: v.cpu() for k, v in
|
| 356 |
+
m.state_dict().items()}},
|
| 357 |
+
os.path.join(EXP19_DIR, f"rule_{ch}_s{seed}.pt"))
|
| 358 |
+
del m
|
| 359 |
+
torch.cuda.empty_cache()
|
| 360 |
+
del teacher
|
| 361 |
+
torch.cuda.empty_cache()
|
| 362 |
+
ledger.close()
|
| 363 |
+
|
| 364 |
+
|
| 365 |
+
def smoke():
|
| 366 |
+
recs = make_facts(8, seed=0)
|
| 367 |
+
assert len(recs) == 8 and all(len(k) == KEY_LEN and len(v) == VAL_LEN
|
| 368 |
+
for k, v in recs)
|
| 369 |
+
fs = fact_stream(recs, copies=3, seed=0)
|
| 370 |
+
assert fs.dtype == torch.uint8 and fs.numel() == 3 * 8 * (KEY_LEN + VAL_LEN + 4)
|
| 371 |
+
m = ByteLM("addr_msl64", d=96, layers=2, block=64)
|
| 372 |
+
r = recall(m, recs, device="cpu", max_eval=8, batch=4)
|
| 373 |
+
assert 0.0 <= r["exact"] <= 1.0 and 0.0 <= r["byte_acc"] <= 1.0
|
| 374 |
+
g = torch.Generator().manual_seed(0)
|
| 375 |
+
x, y, mask = _mix_batch(torch.randint(0, 256, (50000,),
|
| 376 |
+
dtype=torch.uint8, generator=g),
|
| 377 |
+
fs, 8, 64, "cpu", g)
|
| 378 |
+
assert x.shape == (8, 64) and mask.shape == (8,)
|
| 379 |
+
assert (x[:, 1:] == y[:, :-1]).all() # stream alignment
|
| 380 |
+
# 19b: rule facts are rule-consistent + disjoint; variant recall runs
|
| 381 |
+
tr8, te4 = make_rule_facts(8, 4, seed=0)
|
| 382 |
+
assert len(tr8) == 8 and len(te4) == 4
|
| 383 |
+
assert not set(k for k, _ in tr8) & set(k for k, _ in te4)
|
| 384 |
+
k0, v0 = tr8[0]
|
| 385 |
+
assert len(v0) == VAL_LEN and v0[:KEY_LEN] == v0[KEY_LEN:2 * KEY_LEN]
|
| 386 |
+
rv = recall_fmt(m, tr8, FMT_VARIANT, device="cpu", max_eval=8, batch=4)
|
| 387 |
+
assert 0.0 <= rv["exact"] <= 1.0
|
| 388 |
+
print(f"exp019 smoke passed (untrained recall exact={r['exact']} "
|
| 389 |
+
f"byte={r['byte_acc']} ~ chance; 19b rule+variant OK)")
|
| 390 |
+
|
| 391 |
+
|
| 392 |
+
def _in_notebook():
|
| 393 |
+
try:
|
| 394 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 395 |
+
return True
|
| 396 |
+
except NameError:
|
| 397 |
+
return False
|
| 398 |
+
|
| 399 |
+
|
| 400 |
+
if __name__ == "__main__":
|
| 401 |
+
smoke() if not _in_notebook() else (smoke(),
|
| 402 |
+
print("Notebook: run_retention() on GPU."))
|
exp020_gen/exp020_generalization.py
ADDED
|
@@ -0,0 +1,231 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""exp020_generalization.py — GENERALIZATION via the program's best structural
|
| 2 |
+
methodologies (codebooks + constellations), raced on the exp019b rule task.
|
| 3 |
+
exp019b established the failure modes: teachers learn a substitution cipher
|
| 4 |
+
only fragmentarily (held-out byte ~0.26, exact 0.000) and bind content to its
|
| 5 |
+
surface format. exp020 asks what UNLOCKS generalization: geometric structure
|
| 6 |
+
(aleph codebook reads, trigram character embeddings, addresses in depth, the
|
| 7 |
+
constellation 768 address) and data-side format diversity — against the honest
|
| 8 |
+
capacity control (the exp017 lesson: always run the matched free head).
|
| 9 |
+
|
| 10 |
+
TASK (identical to exp019b for comparability): value = fixed random
|
| 11 |
+
substitution cipher of the key (v[i] = subst(k[i % 6]), VAL_LEN 12); 256
|
| 12 |
+
train keys in the stream at FACT_RATE 0.5, 128 keys HELD OUT; 4000 steps
|
| 13 |
+
direct training. JUDGE: held-out exact/byte (rule induction), train recall,
|
| 14 |
+
variant-format recall (the format lock), clean bpb.
|
| 15 |
+
|
| 16 |
+
ARMS (2 seeds each):
|
| 17 |
+
aleph — ByteLM('addr_msl64') (the exp019b reference, in-harness)
|
| 18 |
+
aleph_tri — ByteLM('addr_msl64_tri') (trigram byte embeddings: certified
|
| 19 |
+
-10% bpb; character structure is
|
| 20 |
+
the cipher's own factorization)
|
| 21 |
+
relay — ByteLM('relay_msl64') (addresses in depth, GPT-2-certified
|
| 22 |
+
methodology on the co-trained bed)
|
| 23 |
+
const — ByteLM('sdpa') + ConstellationHead (exp017: the 768 address;
|
| 24 |
+
params NOT matched — recorded)
|
| 25 |
+
mlp — matched free head (params == aleph head budget; capacity control)
|
| 26 |
+
aleph_fmtdiv — addr_msl64 with FORMAT-DIVERSE fact rendering (3 formats in the
|
| 27 |
+
stream; the exp019b format-lock finding turned into a training
|
| 28 |
+
methodology) — judged on the base format + the UNSEEN variant
|
| 29 |
+
Preregistered forks:
|
| 30 |
+
F1 does ANY structural arm lift held-out byte acc off the flat ~0.26 plateau
|
| 31 |
+
(rule induction unlocked by geometry)?
|
| 32 |
+
F2 structure vs capacity: const/aleph vs mlp on held-out at recorded budgets.
|
| 33 |
+
F3 trigram: strongest prior for a per-character cipher.
|
| 34 |
+
F4 format diversity: does it break the format lock (variant recall >> 0) and
|
| 35 |
+
does breaking the lock ALSO lift rule induction?
|
| 36 |
+
F5 relays in depth: does distributed addressing help where the head alone
|
| 37 |
+
plateaus?
|
| 38 |
+
Riders: pure Adam wd=0; GPU-only verdicts; >=2 seeds; Colab-safe. Paste order:
|
| 39 |
+
geolip_vitals -> ar_differentiation_bed -> exp014_genetic_distillation ->
|
| 40 |
+
exp017_aleph_constellation -> exp019_content_retention -> this file.
|
| 41 |
+
"""
|
| 42 |
+
from __future__ import annotations
|
| 43 |
+
import json
|
| 44 |
+
import math
|
| 45 |
+
import os
|
| 46 |
+
import torch
|
| 47 |
+
import torch.nn as nn
|
| 48 |
+
import torch.nn.functional as F
|
| 49 |
+
|
| 50 |
+
if "make_rule_facts" not in globals():
|
| 51 |
+
try:
|
| 52 |
+
from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
|
| 53 |
+
from exp017_aleph_constellation import ConstellationHead, SquaredReLU
|
| 54 |
+
from exp019_content_retention import (
|
| 55 |
+
make_rule_facts, recall_fmt, FMT_TRAIN, FMT_VARIANT,
|
| 56 |
+
KEY_LEN, VAL_LEN, FACT_RATE, ALNUM)
|
| 57 |
+
except ImportError:
|
| 58 |
+
_here = globals().get("__file__")
|
| 59 |
+
if _here is None:
|
| 60 |
+
raise ImportError("paste the stack (vitals, bed, exp017, exp019) first")
|
| 61 |
+
import sys, pathlib
|
| 62 |
+
sys.path.insert(0, str(pathlib.Path(_here).parent))
|
| 63 |
+
from ar_differentiation_bed import ByteLM, _wikitext_bytes, _batch, VOCAB
|
| 64 |
+
from exp017_aleph_constellation import ConstellationHead, SquaredReLU
|
| 65 |
+
from exp019_content_retention import (
|
| 66 |
+
make_rule_facts, recall_fmt, FMT_TRAIN, FMT_VARIANT,
|
| 67 |
+
KEY_LEN, VAL_LEN, FACT_RATE, ALNUM)
|
| 68 |
+
|
| 69 |
+
DATA_ROOT = os.environ.get("GEOLIP_DATA", "./data")
|
| 70 |
+
EXP20_DIR = os.path.join(DATA_ROOT, "exp020")
|
| 71 |
+
STEPS = 4000
|
| 72 |
+
|
| 73 |
+
# fact renderers: index 0 is the base format (matches FMT_TRAIN prompts);
|
| 74 |
+
# fmtdiv streams all three. FMT_VARIANT (" @%s= ") stays UNSEEN by every arm.
|
| 75 |
+
RENDERERS = (lambda k, v: b"\n@" + k + b"=" + v + b"\n",
|
| 76 |
+
lambda k, v: b"\n" + k + b" -> " + v + b"\n",
|
| 77 |
+
lambda k, v: b"\n<" + k + b"|" + v + b">\n")
|
| 78 |
+
FMT_R1 = b"\n%s -> " # prompt form of renderer 1 (fmtdiv-seen)
|
| 79 |
+
|
| 80 |
+
|
| 81 |
+
def fact_stream_fmt(recs, renderers, copies: int = 50, seed: int = 0):
|
| 82 |
+
g = torch.Generator().manual_seed(5000 + seed)
|
| 83 |
+
order = torch.cat([torch.randperm(len(recs), generator=g)
|
| 84 |
+
for _ in range(copies)])
|
| 85 |
+
rsel = torch.randint(len(renderers), (order.numel(),), generator=g)
|
| 86 |
+
blob = b"".join(renderers[rsel[i].item()](*recs[order[i].item()])
|
| 87 |
+
for i in range(order.numel()))
|
| 88 |
+
return torch.frombuffer(bytearray(blob), dtype=torch.uint8).clone()
|
| 89 |
+
|
| 90 |
+
|
| 91 |
+
def make_arm(arm: str, seed: int, d: int = 192, layers: int = 4,
|
| 92 |
+
block: int = 256):
|
| 93 |
+
torch.manual_seed(seed)
|
| 94 |
+
if arm in ("aleph", "aleph_fmtdiv"):
|
| 95 |
+
return ByteLM("addr_msl64", d=d, layers=layers, block=block)
|
| 96 |
+
if arm == "aleph_tri":
|
| 97 |
+
return ByteLM("addr_msl64_tri", d=d, layers=layers, block=block)
|
| 98 |
+
if arm == "relay":
|
| 99 |
+
return ByteLM("relay_msl64", d=d, layers=layers, block=block)
|
| 100 |
+
if arm == "const":
|
| 101 |
+
m = ByteLM("sdpa", d=d, layers=layers, block=block)
|
| 102 |
+
m.head = ConstellationHead(d)
|
| 103 |
+
return m
|
| 104 |
+
if arm == "mlp":
|
| 105 |
+
m = ByteLM("sdpa", d=d, layers=layers, block=block)
|
| 106 |
+
ref = ByteLM("addr_msl64", d=d, layers=layers, block=block)
|
| 107 |
+
target = (sum(p.numel() for p in ref.head_proj.parameters())
|
| 108 |
+
+ sum(p.numel() for p in ref.head_addr.parameters())
|
| 109 |
+
+ sum(p.numel() for p in ref.head.parameters()))
|
| 110 |
+
h = max(8, round((target - VOCAB) / (d + 1 + 2 + VOCAB)))
|
| 111 |
+
m.head = nn.Sequential(nn.Linear(d, h), SquaredReLU(),
|
| 112 |
+
nn.LayerNorm(h), nn.Linear(h, VOCAB))
|
| 113 |
+
del ref
|
| 114 |
+
return m
|
| 115 |
+
raise ValueError(arm)
|
| 116 |
+
|
| 117 |
+
|
| 118 |
+
def head_params(m) -> int:
|
| 119 |
+
tot = 0
|
| 120 |
+
for name in ("head", "head_proj", "head_addr"):
|
| 121 |
+
mod = getattr(m, name, None)
|
| 122 |
+
if mod is not None and isinstance(mod, nn.Module):
|
| 123 |
+
tot += sum(p.numel() for p in mod.parameters())
|
| 124 |
+
if hasattr(m, "relays"):
|
| 125 |
+
tot += sum(p.numel() for p in m.relays.parameters())
|
| 126 |
+
return tot
|
| 127 |
+
|
| 128 |
+
|
| 129 |
+
def train_direct(model, tr, va, fs, steps=STEPS, batch=32, block=256,
|
| 130 |
+
device="cuda", seed=0, needs_calibration=False):
|
| 131 |
+
g = torch.Generator().manual_seed(seed)
|
| 132 |
+
model = model.to(device)
|
| 133 |
+
if needs_calibration:
|
| 134 |
+
xw, _ = _batch(tr, batch, block, device, g)
|
| 135 |
+
model(xw)
|
| 136 |
+
model.head.calibrate(model._last_h)
|
| 137 |
+
opt = torch.optim.Adam(model.parameters(), lr=3e-4, weight_decay=0.0)
|
| 138 |
+
for step in range(1, steps + 1):
|
| 139 |
+
xw, yw = _batch(tr, batch, block, device, g)
|
| 140 |
+
xf, yf = _batch(fs, batch, block, device, g)
|
| 141 |
+
mk = (torch.rand(batch, generator=g) < FACT_RATE).to(device)
|
| 142 |
+
x = torch.where(mk[:, None], xf, xw)
|
| 143 |
+
y = torch.where(mk[:, None], yf, yw)
|
| 144 |
+
loss = F.cross_entropy(model(x).reshape(-1, VOCAB), y.reshape(-1))
|
| 145 |
+
opt.zero_grad(set_to_none=True); loss.backward(); opt.step()
|
| 146 |
+
model.eval()
|
| 147 |
+
with torch.no_grad():
|
| 148 |
+
ls = []
|
| 149 |
+
for _ in range(10):
|
| 150 |
+
xv, yv = _batch(va, batch, block, device, g)
|
| 151 |
+
ls.append(F.cross_entropy(model(xv).reshape(-1, VOCAB),
|
| 152 |
+
yv.reshape(-1)).item())
|
| 153 |
+
return sum(ls) / len(ls) / math.log(2)
|
| 154 |
+
|
| 155 |
+
|
| 156 |
+
ARMS = ("aleph", "aleph_tri", "relay", "const", "mlp", "aleph_fmtdiv")
|
| 157 |
+
|
| 158 |
+
|
| 159 |
+
def run_bakeoff(arms=ARMS, seeds=(0, 1), device="cuda"):
|
| 160 |
+
if not torch.cuda.is_available():
|
| 161 |
+
raise RuntimeError("verdict runs are GPU-only")
|
| 162 |
+
os.makedirs(EXP20_DIR, exist_ok=True)
|
| 163 |
+
tr, va = _wikitext_bytes(DATA_ROOT)
|
| 164 |
+
ledger = open(os.path.join(EXP20_DIR, "ledger.jsonl"), "a", encoding="utf-8")
|
| 165 |
+
for seed in seeds:
|
| 166 |
+
train_recs, test_recs = make_rule_facts(256, 128, seed=seed)
|
| 167 |
+
fs_base = fact_stream_fmt(train_recs, RENDERERS[:1], seed=seed)
|
| 168 |
+
fs_div = fact_stream_fmt(train_recs, RENDERERS, seed=seed)
|
| 169 |
+
for arm in arms:
|
| 170 |
+
m = make_arm(arm, seed=seed)
|
| 171 |
+
hp = head_params(m)
|
| 172 |
+
bpb = train_direct(m, tr, va,
|
| 173 |
+
fs_div if arm == "aleph_fmtdiv" else fs_base,
|
| 174 |
+
device=device, seed=seed,
|
| 175 |
+
needs_calibration=(arm == "const"))
|
| 176 |
+
gz = {"train": recall_fmt(m, train_recs, FMT_TRAIN, device=device),
|
| 177 |
+
"heldout": recall_fmt(m, test_recs, FMT_TRAIN, device=device),
|
| 178 |
+
"train_varfmt": recall_fmt(m, train_recs, FMT_VARIANT,
|
| 179 |
+
device=device)}
|
| 180 |
+
if arm == "aleph_fmtdiv":
|
| 181 |
+
gz["heldout_r1"] = recall_fmt(m, test_recs, FMT_R1,
|
| 182 |
+
device=device)
|
| 183 |
+
rec = {"exp": "20", "arm": arm, "seed": seed, "steps": STEPS,
|
| 184 |
+
"head_params": hp, "gauges": gz, "bpb": round(bpb, 4)}
|
| 185 |
+
ledger.write(json.dumps(rec) + "\n"); ledger.flush()
|
| 186 |
+
print(f"[20 {arm} s{seed}] heldout={gz['heldout']} "
|
| 187 |
+
f"train={gz['train']['exact']} varfmt="
|
| 188 |
+
f"{gz['train_varfmt']['byte_acc']} bpb={bpb:.4f} "
|
| 189 |
+
f"hp={hp}", flush=True)
|
| 190 |
+
torch.save({"arm": arm, "seed": seed,
|
| 191 |
+
"state_dict": {k: v.cpu() for k, v in
|
| 192 |
+
m.state_dict().items()}},
|
| 193 |
+
os.path.join(EXP20_DIR, f"gen_{arm}_s{seed}.pt"))
|
| 194 |
+
del m
|
| 195 |
+
torch.cuda.empty_cache()
|
| 196 |
+
ledger.close()
|
| 197 |
+
|
| 198 |
+
|
| 199 |
+
def smoke():
|
| 200 |
+
recs, test = make_rule_facts(8, 4, seed=0)
|
| 201 |
+
fs = fact_stream_fmt(recs, RENDERERS, copies=3, seed=0)
|
| 202 |
+
assert fs.dtype == torch.uint8 and fs.numel() > 0
|
| 203 |
+
blob = bytes(fs.tolist())
|
| 204 |
+
assert b"->" in blob and b"<" in blob and b"@" in blob # all renderers hit
|
| 205 |
+
x = torch.randint(0, 256, (2, 64))
|
| 206 |
+
for arm in ARMS:
|
| 207 |
+
m = make_arm(arm, seed=0, d=96, layers=2, block=64)
|
| 208 |
+
lg = m(x)
|
| 209 |
+
assert lg.shape == (2, 64, 256), arm
|
| 210 |
+
lg.sum().backward()
|
| 211 |
+
assert head_params(m) > 0, arm
|
| 212 |
+
m.zero_grad()
|
| 213 |
+
del m
|
| 214 |
+
ref = make_arm("aleph", seed=0, d=96, layers=2, block=64)
|
| 215 |
+
mm = make_arm("mlp", seed=0, d=96, layers=2, block=64)
|
| 216 |
+
rp, mp = head_params(ref), head_params(mm)
|
| 217 |
+
assert abs(rp - mp) / rp < 0.05, (rp, mp) # matched within 5%
|
| 218 |
+
print(f"exp020 smoke passed (mlp head {mp} ~ aleph head {rp})")
|
| 219 |
+
|
| 220 |
+
|
| 221 |
+
def _in_notebook():
|
| 222 |
+
try:
|
| 223 |
+
get_ipython() # type: ignore[name-defined] # noqa: F821
|
| 224 |
+
return True
|
| 225 |
+
except NameError:
|
| 226 |
+
return False
|
| 227 |
+
|
| 228 |
+
|
| 229 |
+
if __name__ == "__main__":
|
| 230 |
+
smoke() if not _in_notebook() else (smoke(),
|
| 231 |
+
print("Notebook: run_bakeoff() on GPU."))
|