A newer version of this model is available: MoreThought/AresMath1.0-IT-GGUF

Profile:

AresMath1.0-IT

AresMath1.0-IT (AresMath1.0-Iterative Tiny) is an 11,008 parameter TLM made to solve exact-match arithmetic problems even ones including ().

It uses a heavily looped architecture similar to that of nanbeige-4.2-3b to achieve superior mathematic accuracy on unseen problems.

It has 21 vocab, a 32 token context window, a 32 token output window, and 4 attention heads.

Benchmarks

AresMath1.0-IT achieves comparable/exceeding performance against models over 10,000x bigger than it.

It exceeds almost every tested SLM's scores on arithmark 3.0 (converted into a format it can understand) while generating exact-match answers instead of guessing.

AresMath1.0-IT's scores at different tempatures:

Temp Accuracy Total
1.00 42.30% 423/1000
0.90 44.70% 447/1000
0.80 47.10% 471/1000
0.70 48.90% 489/1000
0.60 50.20% 502/1000
0.50 52.40% 524/1000
0.40 52.80% 528/1000
0.30 53.80% 538/1000
0.20 54.70% 547/1000
0.10 54.80% 548/1000

Top AresMath1.0-IT score vs top SLM scores: Scores:

Model Accuracy Total
MobileLLM-R1 65.70% 657/1000
AresMath1.0-IT 54.80% 548/1000
GPT-X3-150M 50.90% 509/1000
OdinNext-138M 41.80% 418/1000
OdinNext-138M-Instruct 40.90% 409/1000
Quark-135M 39.70% 397/1000
SmolLM2-135M 39.20% 392/1000
SmolLM-135M 36.80% 368/1000

Training

It was trained on 13B tokens of pretraining and 13B tokens of RL, totalling 26B tokens or 2,361,918 tokens per parameter.

The data was all synthetically generated and completely random, the problems were not made to be similar to arithmark 3.0 problems so it can score better.

It slowly climbed from around 30% accuracy at 5B tokens to the current 50% accuracy, jumping from negative RL rewards to +8.2 at the end.

It seemed like it hit ceilings many times in the training run but it always managed to slowly push past all of them.

Training Hardware

This model was trained on multiple v5e TPUs.

Downloads last month
2
Safetensors
Model size
11k params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MoreThought/AresMath1.0-IT

Quantizations
1 model