GoLLeM-v5 64M Muon v1 (earlier recipe)
First 64M flagship of the series (Qwen3-style decoder, Muon optimizer). Earlier recipe than the final models (13.1B instead of 24.9B tokens); the final 64M model supersedes it on eff by +0.26.
Author: Arkadiusz Słota / Fabryka AI. Part of the GoLLeM-v5 series of small English language models trained from scratch. All checkpoints of the series, including intermediate ones, are archived in SlayerLab/gollem-v5-ckpts.
Model
- 62.9M parameters; 14 layers, d_model 576, 9 heads (head dim 64); context 1024 tokens.
- Architecture: Qwen3-style decoder: RoPE (theta 100,000), SwiGLU (FFN multiplier 2.667), RMSNorm, QK-norm, value residual; tied input/output embeddings.
- Tokenizer: BPE, 12,288 tokens (
tokenizer.json, sha2563733307577230bb4802d2c774d8e2323f7f64e4139c13712736a57daf91bdda1), 3.8605 bytes per token on WikiText-2 test. - Weights:
model.safetensors(its sha256 is shown by the Hub for the LFS file), converted tensor by tensor fromv1_muon/ckpt_400k.pt(sha25659f982c1516c6e19df86d4ee1ac6359cd1987830a624080c70d07db0c2d0455f) without changing any weight. Only weights are published here: no pickled checkpoint, optimizer state or training logs.
Training
- Recipe: Muon (hidden 2-D weights, lr 0.02) + AdamW for non-matrix parameters, cosine schedule 6e-4 to 6e-5, batch 32 x 1024 tokens, 400,000 steps = 13.1B tokens (~1.4 epochs over ARC-MIX), seed 1337.
- Data: ARC-MIX (9.42B BPE tokens): the base mixture SlayerLab/minimal-en-corpus-5b (~5.4B tokens: FineWeb-Edu, DCLM, StackExchange, open-web-math, FineMath, scientific papers, books/Gutenberg, code, CC-News) plus FineWeb-Edu and OpenStax textbooks, with ARC-relevant science, reasoning and Q&A web content upweighted (gold ~3x, related ~2x). Filtered against benchmark test sets with 13-gram matching at build time.
- Result checkpoint: The result is the last step of the planned run (400,000).
- Training run (loss curves, metrics): https://track.fabryka.ai/run/16b892d3-2a72-4b54-a595-6f03a9ad4a8b
- An independent re-benchmark by the board maintainer on the same checkpoint (sha256 matched) reproduced ARC-Easy 47.94 exactly.
- In a clean 64M A/B (same data, architecture and seed) Muon was at least as good as AdamW.
Evaluation (our measurement, not official)
Measured by us on the result checkpoint with the Glint Tiny-ML board protocol (BLiMP: 67 configs, 67,000 pairs, sentences clipped to 256 tokens, raw log-prob preference; ARC-Easy: test split, zero-shot, LL(question + choice) - LL(question); WikiText-2: byte-normalized perplexity). These are not official scores.
| Benchmark | Score |
|---|---|
| ARC-Easy (acc, %) | 47.94 |
| BLiMP (acc, %) | 75.83 |
| WikiText-2 byte perplexity (lower is better) | 2.3718 |
| eff (board formula, includes a size multiplier) | 75.81 |
| MultiBLiMP-pl (acc, %, 3,272 pairs) | 62.04 |
- eff = mean of BLiMP, ARC-Easy and normalized WikiText-2, times a size multiplier that is larger for smaller models; compare eff only between models under the same board formula.
- This is an English model. Its Polish MultiBLiMP score is close to the length baseline (always choosing the shorter sentence gives 59.67 %); it is reported for completeness, not as Polish ability.
- Board: https://track.fabryka.ai/models (row "GoLLeM-v5-64M-Muon-v1").
- Source of the numbers:
liczby_v5.json(sha25635c9022ee9dbd96d588093c6fc69604d0eaabd18843458671205db41e234dde8), built from the anonymous board read-out and our PL measurement files.
Benchmark overlap disclosure
A scan of the finished training corpus found that 7 of 62 WikiText-2 test articles appear almost completely in it (as web copies) and 30 of 2,376 ARC-Easy test questions have a match (mostly short factual sentences). On the 16M, 32M and 64M Muon v1 models, removing them changed WikiText-2 byte perplexity by less than 0.01 and ARC-Easy by -0.07 to +0.19 points, within noise (details in the root card of the archive repository). BLiMP was covered only by the build-time 13-gram filter.
How to use
The weights are a plain PyTorch state dict saved as safetensors; they are not a transformers AutoModel. The model class is GPT in train_gpt_ref.py in SlayerLab/gollem-v5-ckpts; build it from config.json and load the state dict, then trim logits to 12,288.
Limitations
- Small base model trained from scratch for research: not instruction-tuned, limited factual knowledge and coherence, not for production use.
- English only.
- Single seed; scores come from one checkpoint and our own implementation of the board protocol.
Training data
Base mixture: SlayerLab/minimal-en-corpus-5b (research-mix-5b, 5.0B tokens of its own tokenizer). It is an aggregate corpus: no unified licence is applied and each source keeps its upstream terms. Sources as named in its manifests/mixture.json:
| source | tokens (M) | upstream licence |
|---|---|---|
| fineweb-edu | 1,101 | ODC-By 1.0 |
| dclm-baseline | 801 | CC BY 4.0 |
| starcoderdata | 469 | The Stack terms (per-file licences, opt-out) |
| stackexchange | 447 | CC BY-SA |
| loc-pd-books | 402 | CC0 1.0 |
| wikipedia | 362 | CC BY-SA 3.0 / GFDL |
| open-web-math | 228 | ODC-By |
| finemath | 221 | ODC-By |
| project-gutenberg | 210 | public domain in the US |
| scientific-papers | 206 | unknown |
| ultrachat | 200 | MIT; text generated by OpenAI models |
| wildchat | 151 | ODC-By; text from conversations with OpenAI models |
| cc-news | 150 | unknown (news articles under copyright) |
| tiny-textbooks | 51 | Apache 2.0; model-generated text |
| open-subtitles | 1 | film subtitles, under copyright |
- Upstream repositories are not recorded in the corpus manifest; licences are those of the Hub datasets matching these source names (not confirmed against the build).
- Added on top of the base mixture for this series: FineWeb-Edu (ODC-By 1.0) and OpenStax textbooks (CC BY 4.0).
- Aggregate corpus; each source keeps its upstream terms. Commercial use: review upstream terms, in particular CC-News, scientific papers, StarCoderData and model-generated chat data.
Licence and attribution
- Weights: no single licence is claimed for the weights yet; they were trained on data under the licences listed in the Training data section above, and any use must respect those terms (license id
other, namegollem-v5-see-training-data). - OpenStax textbooks: CC BY 4.0, © Rice University, https://openstax.org; titles and editions in
OPENSTAX_ATTRIBUTION.mdin this repository.
- Downloads last month
- 107