Have you experimented with or tested different vocabulary sizes to see if the vocab size directly impacts the overtraining threshold and benchmark peak?
Mustafa Tahir KANAT
Rodeszones
AI & ML interests
None yet
Recent Activity
commentedon an article about 18 hours ago
Extreme Overtraining in Tiny Language Models commentedon an article 1 day ago
Sleeper Agents and How to Tame Them liked a dataset 3 days ago
openbmb/UltraFeedbackOrganizations
None yet
commented on Extreme Overtraining in Tiny Language Models about 18 hours ago
commented on Sleeper Agents and How to Tame Them 1 day ago
Execute Order 66
upvoted an article 4 days ago
Article
Distillation in 2026 (so far): which frontier models use it and how
sergiopaniego
• • 21upvoted an article 9 days ago
Article
Making Knowledge Distillation Cheap Enough to Run at Scale
MultiverseComputingCAI
• • 32reacted to danielhanchen's post with 👍 about 1 month ago
Post
6206
Gemma 4 is now faster and much more accurate! 🚀
Google made huge improvements to tool-calling and chat accuracy, reliability + speed.
To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!
Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4
Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4
Google made huge improvements to tool-calling and chat accuracy, reliability + speed.
To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!
Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4
Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4
upvoted an article about 2 months ago
Article
VLX-Seek: Improving VLM Fine-Grained Perception via Region Reference Instead of Coordinate Generation
omlab
• • 14upvoted a paper about 2 months ago