Instructions to use cisimcik/cisimcik-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cisimcik/cisimcik-4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="cisimcik/cisimcik-4b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("cisimcik/cisimcik-4b") model = AutoModelForCausalLM.from_pretrained("cisimcik/cisimcik-4b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use cisimcik/cisimcik-4b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "cisimcik/cisimcik-4b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cisimcik/cisimcik-4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/cisimcik/cisimcik-4b
- SGLang
How to use cisimcik/cisimcik-4b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "cisimcik/cisimcik-4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cisimcik/cisimcik-4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "cisimcik/cisimcik-4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cisimcik/cisimcik-4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use cisimcik/cisimcik-4b with Docker Model Runner:
docker model run hf.co/cisimcik/cisimcik-4b
cisimcik-4b
cisimcik-4b, Türkçe öncelikli data ile cisimcik-base modelinin 5 aşamada post-train'lenmiş halidir. Hibrit düşünme ve tool kullanımı destekler ve sadece yazı input'u kabul eder.
- 4,2 milyar parametre. Qwen3.5 hibrit mimarisi (Gated DeltaNet + full attention).
- Hibrit düşünme. Düşünme modu açıkken model her turda düşünüp düşünmeyeceğine (reasoning) kendisi karar verir.
- Tool kullanımı, chat template üzerinden ve Qwen3.5'in XML biçiminde.
- Apache 2.0 lisansıyla açık kaynak olarak yayınlandı.
Hızlı başlangıç
pip install "transformers>=5.17" torch
# isteğe bağlı, NVIDIA GPU'larda daha hızlı Gated DeltaNet kernel'leri:
pip install flash-linear-attention
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("cisimcik/cisimcik-4b")
model = AutoModelForCausalLM.from_pretrained("cisimcik/cisimcik-4b", dtype=torch.bfloat16)
model = model.to("cuda" if torch.cuda.is_available() else "cpu").eval()
messages = [{"role": "user", "content": "Bir manav 3 kg elmayı 75 TL'ye satıyor. 8 kg elma alan biri 50 TL indirim alırsa kaç TL öder?"}]
prompt = tok.apply_chat_template(messages, add_generation_prompt=True, tokenize=False, enable_thinking=True)
inputs = tok(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
out = model.generate(**inputs, max_new_tokens=4096, do_sample=True, temperature=0.6, top_p=0.95, top_k=20)
text = tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
reasoning, _, answer = text.rpartition("</think>")
print(answer.strip())
Düşünme. enable_thinking=True ile yanıtta önce düşünce, sonra </think>, sonra cevap gelir. Model bir turda
düşünmeye gerek görmezse düşünce kısmı boş kalır. enable_thinking=False ile düşünmeyi atlayarak cevap verir.
Sampling. Qwen'in önerdiği ayarları kullanıyoruz: Düşünme açıkken temperature 0.6, top_p 0.95, top_k 20. Düşünme kapalıyken temperature 0.7, top_p 0.8, top_k 20.
Donanım. Ağırlıklar bf16'da yaklaşık 8,4 GB: 12 GB veya üzeri belleğe sahip herhangi bir GPU ya da 16 GB RAM'li
bir CPU yeterlidir. flash-linear-attention kurulu değilse Gated DeltaNet katmanları transformers'ın PyTorch
uygulamasını kullanır. Bu doğru sonuç verir ama uzun prompt'larda daha yavaştır.
Araç çağırma
Araçları chat template'e verin:
tools = [{
"type": "function",
"function": {
"name": "hava_durumu",
"description": "Bir şehrin güncel hava durumunu verir.",
"parameters": {"type": "object", "properties": {"sehir": {"type": "string"}}, "required": ["sehir"]},
},
}]
messages = [{"role": "user", "content": "İzmir'de şu an hava nasıl?"}]
prompt = tok.apply_chat_template(messages, tools=tools, add_generation_prompt=True, tokenize=False, enable_thinking=False)
Tool kullanımı şu şekilde görünür:
<tool_call>
<function=hava_durumu>
<parameter=sehir>
İzmir
</parameter>
</function>
</tool_call>
Benchmark sonuçları
Dosyalar
| Dosya | İçerik |
|---|---|
model-0000*-of-00002.safetensors |
ağırlıklar (bf16) |
config.json, generation_config.json |
Qwen3_5ForCausalLM config'i ve durdurma token'ları |
tokenizer*.json, vocab.json, merges.txt |
Qwen3.5 tokenizer'ı ve chat template (düşünme, araçlar) |
Lisans
Bu depodaki ağırlıklar Apache License 2.0 ile yayımlanmıştır (bkz. LICENSE ve NOTICE). Model, Qwen/Qwen3.5-4B-Base © 2026 Alibaba Cloud, Apache License 2.0 lisanslı modelin değiştirilmiş bir sürümüdür.
Atıf
@misc{cisimcik4b2026,
title = {cisimcik-4b: a Turkish-first chat model},
author = {cisimcik},
year = {2026},
url = {https://huggingface.co/cisimcik/cisimcik-4b}
}
- Downloads last month
- 823
Model tree for cisimcik/cisimcik-4b
Dataset used to train cisimcik/cisimcik-4b
Evaluation results
- Average on CETVEL (zero-shot)Self-reported, CETVEL protocol57.310
- Extractive QA (EM) on CETVEL (zero-shot)Self-reported, CETVEL protocol41.090
- Multiple choice (acc) on CETVEL (zero-shot)Self-reported, CETVEL protocol64.560
- Text classification (acc) on CETVEL (zero-shot)Self-reported, CETVEL protocol74.780
- NLI (acc) on CETVEL (zero-shot)Self-reported, CETVEL protocol78.140
- Summarization (ROUGE-2) on CETVEL (zero-shot)Self-reported, CETVEL protocol25.500
- Grammar correction (EM) on CETVEL (zero-shot)Self-reported, CETVEL protocol89.550
- Machine translation (BLEU) on CETVEL (zero-shot)Self-reported, CETVEL protocol27.570