Instructions to use nkthebass/tinybrainbot-100m-v4-thinking with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nkthebass/tinybrainbot-100m-v4-thinking with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="nkthebass/tinybrainbot-100m-v4-thinking") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("nkthebass/tinybrainbot-100m-v4-thinking") model = AutoModelForCausalLM.from_pretrained("nkthebass/tinybrainbot-100m-v4-thinking", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use nkthebass/tinybrainbot-100m-v4-thinking with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf nkthebass/tinybrainbot-100m-v4-thinking:F16 # Run inference directly in the terminal: llama cli -hf nkthebass/tinybrainbot-100m-v4-thinking:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf nkthebass/tinybrainbot-100m-v4-thinking:F16 # Run inference directly in the terminal: llama cli -hf nkthebass/tinybrainbot-100m-v4-thinking:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf nkthebass/tinybrainbot-100m-v4-thinking:F16 # Run inference directly in the terminal: ./llama-cli -hf nkthebass/tinybrainbot-100m-v4-thinking:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf nkthebass/tinybrainbot-100m-v4-thinking:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf nkthebass/tinybrainbot-100m-v4-thinking:F16
Use Docker
docker model run hf.co/nkthebass/tinybrainbot-100m-v4-thinking:F16
- LM Studio
- Jan
- vLLM
How to use nkthebass/tinybrainbot-100m-v4-thinking with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nkthebass/tinybrainbot-100m-v4-thinking" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nkthebass/tinybrainbot-100m-v4-thinking", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nkthebass/tinybrainbot-100m-v4-thinking:F16
- SGLang
How to use nkthebass/tinybrainbot-100m-v4-thinking with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nkthebass/tinybrainbot-100m-v4-thinking" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nkthebass/tinybrainbot-100m-v4-thinking", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nkthebass/tinybrainbot-100m-v4-thinking" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nkthebass/tinybrainbot-100m-v4-thinking", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use nkthebass/tinybrainbot-100m-v4-thinking with Ollama:
ollama run hf.co/nkthebass/tinybrainbot-100m-v4-thinking:F16
- Unsloth Desktop
- Docker Model Runner
How to use nkthebass/tinybrainbot-100m-v4-thinking with Docker Model Runner:
docker model run hf.co/nkthebass/tinybrainbot-100m-v4-thinking:F16
- Lemonade
How to use nkthebass/tinybrainbot-100m-v4-thinking with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull nkthebass/tinybrainbot-100m-v4-thinking:F16
Run and chat with the model
lemonade run user.tinybrainbot-100m-v4-thinking-F16
List all available models
lemonade list
- Atomic Chat
TinyBrainBot-100M-v4-Thinking
Try it in your browser: it runs on your own device with WebGPU, no install or account needed.
A 100M-parameter model that thinks before it answers — when the question needs it.
It writes its reasoning inside <think>…</think> and then gives a short answer. Greetings and
simple facts get a direct reply with no reasoning block; decisions, tradeoffs, advice and
"why" questions get deliberation first.
The reasoning is general-purpose rather than maths-only: weighing options, naming the deciding factor, and saying so when it doesn't have enough information to be sure.
A sampled output, unedited:
You: Should I get a dog or a cat if I live in a small apartment?
<think>
The purpose is to compare the cost and time of each option for a small apartment. A cat's
care needs are low-maintenance, but a dog's need could be more demanding like feeding and
grooming. I don't know the exact size of your apartment or how much space you have available.
</think>
If your apartment has limited space and you want to stay comfortable without spending too
much on a pet, a cat would likely suit that better. However, if you value quality and can
afford the upfront costs, a dog might be worth it.
Built from tinybrainbot-100m-v3-base (the v4 round of post-training on that base):
- SFT on a thinking mix — verified reasoning traces, open-ended deliberation (decisions, tradeoffs, "not enough information"), refusal data, plus ordinary chat so it learns when not to think.
- A LoRA pass, merged in, that tunes where it draws the line on refusals and how it handles everyday disagreements.
- Architecture: Llama-compatible, 100.1M params (768/12L/12h·4kv, ctx 1024, vocab 32k).
- Chat template:
<|user|>\n{msg}\n<|end|>\n<|assistant|>\n(apply_chat_templatealso works).
What it does
Measured on held-out prompts that were not used in training:
| Behaviour | Result |
|---|---|
| Thinks on everyday decision / advice / "why" questions | 100% |
| Thinks on greetings and simple facts | 0% |
| Declines clearly harmful requests | 87% |
| Declines harmless requests that only sound alarming ("how do I kill mould?") | ~5% |
Refusal figures come from 60 prompts per category, so treat them as indicative.
Recommended settings
temperature 0.4, top_p 0.95, repetition_penalty 1.15 — these are already the defaults in
generation_config.json, so a plain model.generate(...) uses them.
Don't raise the temperature much above ~0.7: at 1.0 it declines harmful requests noticeably less often. Keep the repetition penalty on; without it the model can repeat itself.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "nkthebass/tinybrainbot-100m-v4-thinking"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)
prompt = "<|user|>\nShould I get a dog or a cat if I live in a small apartment?\n<|end|>\n<|assistant|>\n"
ids = tok(prompt, return_tensors="pt").input_ids
out = model.generate(ids) # sampling defaults come from generation_config.json
text = tok.decode(out[0][ids.shape[1]:], skip_special_tokens=False).split("<|end|>")[0]
thinking, _, answer = text.rpartition("</think>")
thinking = thinking.replace("<think>", "").strip()
print(answer.strip()) # the reply, with the reasoning hidden
thinking holds the reasoning if you want to show or log it; for a direct reply it is empty.
GGUF
An F16 GGUF (tinybrainbot-100m-v4-thinking-f16.gguf, 200 MB) and a Q8_0 GGUF (tinybrainbot-100m-v4-thinking-q8_0.gguf, 107 MB) are included for llama.cpp,
LM Studio and Ollama. The chat template is embedded, and the tokenizer is set up so llama.cpp
tokenizes chat messages exactly as the model was trained.
llama-server -m tinybrainbot-100m-v4-thinking-f16.gguf -c 1024 --temp 0.4 --top-p 0.95 --repeat-penalty 1.15
Set the same three values in LM Studio / Ollama — their default temperature (0.8) is higher
than this model should run at. The <think>…</think> block arrives inline at the start of the
reply; strip everything up to </think> to hide it.
Benchmarks
EleutherAI lm-eval v0.4.13, 0-shot, acc_norm (WinoGrande and MMLU: acc):
| Benchmark | This model | 100M v3 Instruct | Supra2-100M-Instruct |
|---|---|---|---|
| ARC-Easy | 52.5 | 54.5 | 44.4 |
| ARC-Easy (raw acc) | 56.4 | 58.3 | 51.4 |
| ARC-Challenge | 27.7 | 29.2 | 24.7 |
| ARC-Challenge (raw acc) | 26.3 | 26.7 | 22.7 |
| OpenBookQA | 33.2 | 32.8 | 30.4 |
| OpenBookQA (raw acc) | 20.4 | 21.6 | 17.2 |
| PIQA | 64.9 | 65.5 | 64.4 |
| PIQA (raw acc) | 65.0 | 65.3 | 64.3 |
| WinoGrande | 49.5 | 51.4 | 50.5 |
| MMLU | 25.3 | 26.1 | 25.8 |
| HellaSwag | 33.3 | 32.8 | 35.9 |
| HellaSwag (raw acc) | 30.0 | 30.1 | 31.5 |
Multiple-choice benchmarks score answer likelihoods directly, so the model never gets to think during them — they measure what it knows, not whether its reasoning helps.
Limitations
- It's a 100M model. Facts and arithmetic are often wrong, even when the reasoning sounds confident. Don't rely on it for anything that matters.
- Not medical advice. It can't reliably tell a warning sign from a mild symptom and may decline health questions altogether. In an emergency, contact emergency services.
- Refusals are learned behaviour, not a guarantee.
- Downloads last month
- 1,003
Model tree for nkthebass/tinybrainbot-100m-v4-thinking
Base model
nkthebass/tinybrainbot-100m-v3-base