Klyra-64M-Base

Klyra-64M-Base is the foundation checkpoint of the Klyra-64M model family.

It was trained from random initialization on approximately 1B tokens before any instruction tuning or reasoning-specific post-training.

Training

  • Dataset: openbmb/Ultra-FineWeb-L1
  • Configuration: CC-MAIN-2025-30
  • Training tokens: ~1B
  • Sequence length: 512
  • Peak learning rate: 5e-4
  • Precision: BF16

Final validation:

  • Loss: 2.3649
  • Perplexity: 10.643

Intended Use

This checkpoint is intended for:

  • continued pretraining,
  • representation-learning experiments,
  • tokenizer/model research,
  • studying the Klyra training lineage.

It is not instruction tuned.

About Klyra

Klyra-64M is a compact language-model research project initiated and developed by a student of Informatics Engineering at Politeknik Negeri Jakarta (State Polytechnic of Jakarta).

The model uses a MiniMind-compatible decoder-only architecture and was trained through a staged pipeline from random initialization.

Main project: Jahirrrr/Klyra-64M

Architecture

Component Value
Unique runtime parameters ~63.9M
Transformer layers 8
Hidden size 768
Attention heads 8
KV heads 4
Vocabulary ~6.4K
Context length 1,024
MoE No

Parameter note: Klyra uses tied input/output embeddings at runtime. Some exported safetensors files may contain both model.embed_tokens.weight and lm_head.weight as separate serialized tensors. The intended runtime model has 63,912,192 unique parameters after weight tying.

Downloads last month
228
Safetensors
Model size
68.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including Jahirrrr/Klyra-64M-Base