TinyBrainBot-100M-v4-Thinking

Try it in your browser: it runs on your own device with WebGPU, no install or account needed.

A 100M-parameter model that thinks before it answers — when the question needs it. It writes its reasoning inside <think>…</think> and then gives a short answer. Greetings and simple facts get a direct reply with no reasoning block; decisions, tradeoffs, advice and "why" questions get deliberation first.

The reasoning is general-purpose rather than maths-only: weighing options, naming the deciding factor, and saying so when it doesn't have enough information to be sure.

A sampled output, unedited:

You: Should I get a dog or a cat if I live in a small apartment?

<think>
The purpose is to compare the cost and time of each option for a small apartment. A cat's
care needs are low-maintenance, but a dog's need could be more demanding like feeding and
grooming. I don't know the exact size of your apartment or how much space you have available.
</think>
If your apartment has limited space and you want to stay comfortable without spending too
much on a pet, a cat would likely suit that better. However, if you value quality and can
afford the upfront costs, a dog might be worth it.

Built from tinybrainbot-100m-v3-base (the v4 round of post-training on that base):

  1. SFT on a thinking mix — verified reasoning traces, open-ended deliberation (decisions, tradeoffs, "not enough information"), refusal data, plus ordinary chat so it learns when not to think.
  2. A LoRA pass, merged in, that tunes where it draws the line on refusals and how it handles everyday disagreements.
  • Architecture: Llama-compatible, 100.1M params (768/12L/12h·4kv, ctx 1024, vocab 32k).
  • Chat template: <|user|>\n{msg}\n<|end|>\n<|assistant|>\n (apply_chat_template also works).

What it does

Measured on held-out prompts that were not used in training:

Behaviour Result
Thinks on everyday decision / advice / "why" questions 100%
Thinks on greetings and simple facts 0%
Declines clearly harmful requests 87%
Declines harmless requests that only sound alarming ("how do I kill mould?") ~5%

Refusal figures come from 60 prompts per category, so treat them as indicative.

Recommended settings

temperature 0.4, top_p 0.95, repetition_penalty 1.15 — these are already the defaults in generation_config.json, so a plain model.generate(...) uses them.

Don't raise the temperature much above ~0.7: at 1.0 it declines harmful requests noticeably less often. Keep the repetition penalty on; without it the model can repeat itself.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "nkthebass/tinybrainbot-100m-v4-thinking"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)

prompt = "<|user|>\nShould I get a dog or a cat if I live in a small apartment?\n<|end|>\n<|assistant|>\n"
ids = tok(prompt, return_tensors="pt").input_ids
out = model.generate(ids)                      # sampling defaults come from generation_config.json
text = tok.decode(out[0][ids.shape[1]:], skip_special_tokens=False).split("<|end|>")[0]

thinking, _, answer = text.rpartition("</think>")
thinking = thinking.replace("<think>", "").strip()
print(answer.strip())                          # the reply, with the reasoning hidden

thinking holds the reasoning if you want to show or log it; for a direct reply it is empty.

GGUF

An F16 GGUF (tinybrainbot-100m-v4-thinking-f16.gguf, 200 MB) and a Q8_0 GGUF (tinybrainbot-100m-v4-thinking-q8_0.gguf, 107 MB) are included for llama.cpp, LM Studio and Ollama. The chat template is embedded, and the tokenizer is set up so llama.cpp tokenizes chat messages exactly as the model was trained.

llama-server -m tinybrainbot-100m-v4-thinking-f16.gguf -c 1024 --temp 0.4 --top-p 0.95 --repeat-penalty 1.15

Set the same three values in LM Studio / Ollama — their default temperature (0.8) is higher than this model should run at. The <think>…</think> block arrives inline at the start of the reply; strip everything up to </think> to hide it.

Benchmarks

EleutherAI lm-eval v0.4.13, 0-shot, acc_norm (WinoGrande and MMLU: acc):

Benchmark This model 100M v3 Instruct Supra2-100M-Instruct
ARC-Easy 52.5 54.5 44.4
ARC-Easy (raw acc) 56.4 58.3 51.4
ARC-Challenge 27.7 29.2 24.7
ARC-Challenge (raw acc) 26.3 26.7 22.7
OpenBookQA 33.2 32.8 30.4
OpenBookQA (raw acc) 20.4 21.6 17.2
PIQA 64.9 65.5 64.4
PIQA (raw acc) 65.0 65.3 64.3
WinoGrande 49.5 51.4 50.5
MMLU 25.3 26.1 25.8
HellaSwag 33.3 32.8 35.9
HellaSwag (raw acc) 30.0 30.1 31.5

Multiple-choice benchmarks score answer likelihoods directly, so the model never gets to think during them — they measure what it knows, not whether its reasoning helps.

Limitations

  • It's a 100M model. Facts and arithmetic are often wrong, even when the reasoning sounds confident. Don't rely on it for anything that matters.
  • Not medical advice. It can't reliably tell a warning sign from a mild symptom and may decline health questions altogether. In an emergency, contact emergency services.
  • Refusals are learned behaviour, not a guarantee.
Downloads last month
1,003
Safetensors
Model size
0.1B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nkthebass/tinybrainbot-100m-v4-thinking

Quantized
(1)
this model

Space using nkthebass/tinybrainbot-100m-v4-thinking 1

Collection including nkthebass/tinybrainbot-100m-v4-thinking