VisionGuardrail EVO-2, a multimodal image-classification content-safety model based on Qwen/Qwen3.8-27B, is now available on the Hub!
Stricter image classification than before, with a dense 27-billion-parameter multimodal model, more precise reasoning, and improved captions for classifying visual media.
Scribble-Board-Fast is a sketch-to-image workspace powered by Klein-9B, transforming doodles, brush strokes, stickers, and uploaded images into high-fidelity visuals with 4-step distilled sampling.
VisionGuardrail, a multimodal content-safety classifier based on Qwen3.5, is now available on Hugging Face in 4B and 9B variants. It is a direct upgrade to ImageShield-MMCF, providing improved parental controls through conservative visual content-safety filtering.
Pretrained on 4x more tokens than the previous releases (20b vs 5b). Instruct tuned versions are coming soon. Very interesting models are coming soon too (hint: super long context).
ImageShield-MMCF β Multimodal Content Filter is a multimodal content-safety classifier built on top of Qwen3.5 and is now available on Hugging Face!
This is the preview initial version (v1.0) of the model, designed to classify visual content as Safe or Unsafe, with a particular focus on detecting Not Safe for Work (NSFW) and other potentially sensitive visual content.
The demo is implemented in the prithivMLmods/opencaption-4b-vl-sft Space, which serves as an active content-safety layer for computer vision tasks. It helps block Not Safe for Work (NSFW) content generation and paves the way for more meaningful and responsible creativity.
Just hit #14 and #15 with out FIRST models on Open SLM Leaderboard. The models were trained on 5B tokens, while competing with similarly sized models trained on more than 6-20x the data.
A new base model Speck1.5-140M being trained right now on a higher quality corpus and will be released soon. SpeckChat3 is coming very soon with 1 million samples, specifically designed to post train small base models.
Also, just to clarify stuff, we will NOT release anything that is NOT MIT licensed EVER. Openness is needed in small language research.
Thanks to everyone supporting the project, and stay tuned for new releases!
SPECK UPDATES: 1 New instruct model tuned on top of Speck1-140M: specklabs/Speck1-140M-Instruct 2 Instruction tuning datasets 2 GGUFs
Much more coming soon: Speck1.1-140M-Instruct that is post trained on SpeckChat2 will be coming very soon New base model Speck1.5-140M is coming with a much higher quality corpus
Thanks to everyone who is already supporting the project, and stay tuned for new releases!
new models coming very soon (both instruct and much better models), with much much higher training scale as i am getting marenostrum5 access soon! we will be looking at 100b-2t token budgets :)
I've open-sourced the trainer I've been using to build tiny language models from scratch, together with the 95M base model I trained with it.
The trainer runs on Deno (https://deno.com, cross-platform), trains on WebGPU, and it writes GGUF directly. No Python/PyTorch. The weights live in a GGUF file from the first step to the last, so every checkpoint is already something llama.cpp can load.
So far I've been pointing it at Markdown Minimap, an Obsidian plugin that adds a scrollable IDE-style minimap to your notes. This week I've been clearing a backlog of user-reported issues on it, with Claude often handling them end to end.
Made a demo for Text/Image-to-3D Video and Image-to-3D Video asset generation using TRELLIS.2. It is paired with Z-Image-Turbo to accelerate the input image preprocessing pipeline, streamlining the Image-to-3D workflow. The generated GLB (GL Transmission Format) files are converted into MP4 (MPEG-4) videos, making them easy to preview and share. Try it now on Hugging Face Spaces.π€
Introducing Inflect-v2, two exceptionally small, open-weight English TTS models at just 3.9M and 9.3M parameters. Both generate speech multiple times faster than real-time on CPU. Despite their size, Inflect-v2 delivers quality that is competitive with much larger lightweight TTS systems, including KittenTTS, Piper, and Supertonic-3.
CPU, CUDA, PyTorch, and ONNX are supported. Apache 2.0.
This dataset provides over 210,000 distinct enterprise software architectures generated using two open source models: GPT-OSS-120B and Qwen3-Coder-Next-FP8.
These architectures model realistic enterprise systems complete with client layers, edge security, API gateways, service meshes, compliance boundaries, and multi-cloud infrastructure topologies.
Decades before the modern scaling laws, this paper showed that neural networks behavior under scale follows remarkably predictable laws.
In 1993, researchers at Bell Labs were grappling with a constraint that feels entirely familiar (and contemporary): datasets were outgrowing the available hardware, and training a model to the end was becoming too expensive. To evaluate an architectural tweak to a state-of-the-art model (at the time it was LeNet) on 60,000 samples meant burning up to three weeks of compute time.
To save compute, people would train candidate architectures on small subsets of the data, assuming that the top performer at small scale would remain the top performer at full scale. But with our future wisdom, we know this is not the case.
In "Learning Curves: Asymptotic Values and Rate of Convergence (NeurIPS 93)", using insights from statistical mechanics, they proposed a practical and principled method for predicting the performance of classifiers trained on large datasets (at the time, models were assumed to be large enough). The method was based on a simple power-law modeling of the expected training and test errors.
It is often noted that many of today's breakthroughs in AI and deep learning are actually decades-old concepts that simply lacked the computational power to be tested at the time. While there is some truth to that, it highlights a more valuable lesson: there is immense worth in revisiting early literature and reflecting on foundational ideas we may have prematurely left behind.
So, go explore and find your own inspiration. The current trend has enough champions already!
It has been more than a decade now since the knowledge distillation paper came out.
Knowledge Distillation (KD) is one of my favorite topics, but I have to confess that I'm not a huge fan of the term because I find it confusing (or at least, it has became so over time).
The idea behind KD is not novel; it was there almost a decade before the paper came out (and arguably even a decade before that, back to 1990-91). But this paper is the one that clicked, the one that made the topic much more popular and introduced it to a broader audience.
First, the timing and the authors played a big role: we have Geoffrey Hinton, Oriol Vinyals, and Jeff Dean here. And second, Geoffrey Hinton is really good at idea branding: Model compression?! No, no, no! Let's call it "Knowledge Distillation" and use evocative terms such as "Dark Knowledge" to describe what is being transferred.
It's a great name, but as time has passed, the term became a bit of a relic. KD is no longer solely about compression (KD used to be introduced as a method for model compression, but now model compression is just one application of KD). And the other thing is that the word "distillation" implies some sort of potency here, that the student is somehow more powerful than the teacher, which is not the case (but many counterarguments could be made, for example, more powerful compared to another model trained with no teacher)
Nevertheless, the paper is incredibly well-written, short, and fun to read. It's one of few papers that I read several times. Check it out, and maybe share your thoughts on the topic with us here!
If you had to choose another name for Knowledge Distillation, what would it be?
Shell-Code-Large is a large-scale corpus of Shell scripting source code comprising approximately 640,000 code samples stored in JSON Lines (.jsonl) format. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, DevOps automation, cloud infrastructure engineering, system administration, and software engineering automation.
By providing a high-volume, language-specific corpus focused exclusively on Shell scripting, Shell-Code-Large enables systematic experimentation in automation workflows, deployment pipelines, infrastructure management, and command-line tooling. These domains remain foundational to Linux systems, cloud-native platforms, CI/CD environments, and modern DevOps practices.
Shell-Code-Large addresses the need for a dedicated Shell-focused dataset at substantial scale, enabling targeted research into scripting patterns, command composition, workflow orchestration, infrastructure automation, and operational engineering practices
Shipped v0.1.2 of vtx β a minimalist coding agent for the terminal.
Most agentic CLIs ship 10k+ token system prompts. Vtx is ~2,200. Less prompt overhead means more room for your code in the model's context window.
Vtx is a from-scratch Python implementation of the design philosophy behind pi-mono β same principles, pure Python, no transpiled runtime.
What ships out of the box:
β Textual TUI + headless CLI (vtx -p "fix the failing test") β 49 LLM provider gateways, all declared in a single provider.yaml β 5 core tools (read / edit / write / bash / find) plus web search and fetch β Session tree with compaction, handoff, and resume β AGENTS.md / CLAUDE.md auto-discovery β Skills system β drop SKILL.md files in .agents/skills/ and they become slash commands β Two OAuth flows (GitHub Copilot device flow, OpenAI Codex PKCE) β Two-mode permissions: prompt (default) or auto, with a safe-command allowlist
This release adds a proper extension system. Register new LLM-callable tools, intercept tool calls, hook lifecycle events, and add slash commands from a single register(api) function in a Python file under ~/.vtx/agent/extensions/. Extensions can override built-in tools by name and chain handler logic across subscribers.
Apache 2.0. uv tool install vtx-coding-agent and you're running.