istanbul-offsite
community
AI & ML interests
None defined yet.
sergiopaniego
posted an update 3 days ago
sergiopaniego
posted an update 13 days ago
Post
364
ICYMI, Async GRPO in TRL now supports LoRA and we wrote a looong blog testing it
> the adapter is a few megabytes, so the weight sync is a file instead of an NCCL transfer
> 3 HF Jobs: 1 trainer and 2 vLLM replicas
> the adapter travels through an HF Storage Bucket mounted in all 3 at the same path
> a proxy in front of the replicas routes each rollout to the one already holding its KV prefix
https://huggingface.co/blog/asyncgrpo-lora-hfjobs
> the adapter is a few megabytes, so the weight sync is a file instead of an NCCL transfer
> 3 HF Jobs: 1 trainer and 2 vLLM replicas
> the adapter travels through an HF Storage Bucket mounted in all 3 at the same path
> a proxy in front of the replicas routes each rollout to the one already holding its KV prefix
https://huggingface.co/blog/asyncgrpo-lora-hfjobs
sergiopaniego
posted an update 24 days ago
Post
3314
while preparing the last class of the Training Agents live series during the summer, i spent some time reading the post-training sections of many frontier model reports, to learn how they use RL environments to improve their models, and wrote a blog about it
if you use any kind of coding harness, or you saw the Blender scenes that went viral recently, this might be interesting to you
Blog: https://huggingface.co/blog/sergiopaniego/rl-environments-2026
if you use any kind of coding harness, or you saw the Blender scenes that went viral recently, this might be interesting to you
Blog: https://huggingface.co/blog/sergiopaniego/rl-environments-2026
sergiopaniego
posted an update about 1 month ago
Post
2174
Can you do RL over taste?
I've spent some time reproducing, in the open, Surya N's idea of training a model to paint with code. It's a coding model that learns to paint watercolours by writing JS code, trained with GRPO. I used TRL and OpenEnv for this, with the whole pipeline running on Hugging Face.
The interesting part is that the reward has no correct answer, unlike a math problem. In this case it's based on the artistic preferences of the person who builds the dataset.
Everything is published: the environment, the reference pool, the trained adapters, every painting of every run with the code that made it, and a write-up with all the decisions, including the ones that went wrong.
Blog post: https://huggingface.co/blog/train-to-paint-with-code
I've spent some time reproducing, in the open, Surya N's idea of training a model to paint with code. It's a coding model that learns to paint watercolours by writing JS code, trained with GRPO. I used TRL and OpenEnv for this, with the whole pipeline running on Hugging Face.
The interesting part is that the reward has no correct answer, unlike a math problem. In this case it's based on the artistic preferences of the person who builds the dataset.
Everything is published: the environment, the reference pool, the trained adapters, every painting of every run with the code that made it, and a write-up with all the decisions, including the ones that went wrong.
Blog post: https://huggingface.co/blog/train-to-paint-with-code
sergiopaniego
posted an update about 1 month ago
Post
556
catching up on some bookmarked reads from the summer, reading Antidoom from @liquidai
small reasoning models get stuck more easily when the task involves a long thinking trace and a hard problem. It starts repeating the same word over and over again ("Wait", "Alternatively"…), each repetition makes the next one likelier, and the generation is spent before it reaches an answer
they measured it, 10.2% of completions for an early LFM2.5-2.6B checkpoint and 22.9% for Qwen3.5-4B at greedy. After training those drop to 1.4% and 1.0%
the fix is FTPO (final token preference optimization). What I like is how narrow it is, it only touches the single token where the loop starts
three ways it differs from DPO:
> trains one token position, mid-generation, instead of whole sequences
> spreads probability across ~20 plausible alternatives instead of swapping one overtrained token for another
> keeps the regularizer in logit space, no softmax, so the rest of the vocabulary stays put
the third one is what makes it usable. If you want to edit one position without disturbing the model, you can't have a loss that reshuffles the other 150k logits on the way
and their explanation abt the result: the training teaches the model nothing new about math or code, it clears the failure mode that was blocking answers the model could already produce
full blog > https://www.liquid.ai/blog/antidoom
FTPO itself comes from Antislop, where it was built to strip overused phrasing. LiquidAI retargeted it to doom loops
and under the hood it's a subclass of TRL's DPOTrainer with compute_loss overridden, around 90 lines of loss and no new trainer
we documented that pattern in TRL's docs
https://huggingface.co/docs/trl/main/en/customization#change-the-training-objective
small reasoning models get stuck more easily when the task involves a long thinking trace and a hard problem. It starts repeating the same word over and over again ("Wait", "Alternatively"…), each repetition makes the next one likelier, and the generation is spent before it reaches an answer
they measured it, 10.2% of completions for an early LFM2.5-2.6B checkpoint and 22.9% for Qwen3.5-4B at greedy. After training those drop to 1.4% and 1.0%
the fix is FTPO (final token preference optimization). What I like is how narrow it is, it only touches the single token where the loop starts
three ways it differs from DPO:
> trains one token position, mid-generation, instead of whole sequences
> spreads probability across ~20 plausible alternatives instead of swapping one overtrained token for another
> keeps the regularizer in logit space, no softmax, so the rest of the vocabulary stays put
the third one is what makes it usable. If you want to edit one position without disturbing the model, you can't have a loss that reshuffles the other 150k logits on the way
and their explanation abt the result: the training teaches the model nothing new about math or code, it clears the failure mode that was blocking answers the model could already produce
full blog > https://www.liquid.ai/blog/antidoom
FTPO itself comes from Antislop, where it was built to strip overused phrasing. LiquidAI retargeted it to doom loops
and under the hood it's a subclass of TRL's DPOTrainer with compute_loss overridden, around 90 lines of loss and no new trainer
we documented that pattern in TRL's docs
https://huggingface.co/docs/trl/main/en/customization#change-the-training-objective
sergiopaniego
posted an update about 2 months ago
Post
381
super interesting new paper from Microsoft "Agent Lightning v1.0: Towards Harnessed Agentic RL" by Zhiyuan He et al.
same idea we've seen already several times: you train the agent inside the real harness it ships with, instead of a reimplementation of it
now that recipe has a name → harnessed agentic RL
paper: huggingface.co/papers/2608.17528
the tricky bit they nail down: one rollout is not one training sample
the harness calls the model many times, so a single episode → a variable number of (prompt, response) rows
you don't even know the batch size until the episode finishes running
its real contribution is being first to systematically map the four problems that fall out of that:
> retokenization + sample merging
> advantage calculation over a variable sample count
> loss normalization at the rollout level, not per sample
> backend scheduling when the batch size is dynamic
and it actually works → plain RL inside the real harness, no reimplementation
Qwen3.5-9B on SWE-bench Verified 41.8 → 56.4 (+14.6), with only ~6k examples
the whole thing is ~3,500 lines, any harness, self-hosted k8s
from our side, we've shared some materials on the same line you may want to check out :)
> Agentic RL: Token-In, Token-Out Done Right: https://huggingface.co/blog/huggingface/tito
> a full worked example, opencode owning its loop trained with GRPO: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
> Harness, Scaffold, and the AI Agent Terms Worth Getting Right: https://huggingface.co/blog/agent-glossary
on a similar line:
https://x.com/SergioPaniego/status/2062911580564496576
same idea we've seen already several times: you train the agent inside the real harness it ships with, instead of a reimplementation of it
now that recipe has a name → harnessed agentic RL
paper: huggingface.co/papers/2608.17528
the tricky bit they nail down: one rollout is not one training sample
the harness calls the model many times, so a single episode → a variable number of (prompt, response) rows
you don't even know the batch size until the episode finishes running
its real contribution is being first to systematically map the four problems that fall out of that:
> retokenization + sample merging
> advantage calculation over a variable sample count
> loss normalization at the rollout level, not per sample
> backend scheduling when the batch size is dynamic
and it actually works → plain RL inside the real harness, no reimplementation
Qwen3.5-9B on SWE-bench Verified 41.8 → 56.4 (+14.6), with only ~6k examples
the whole thing is ~3,500 lines, any harness, self-hosted k8s
from our side, we've shared some materials on the same line you may want to check out :)
> Agentic RL: Token-In, Token-Out Done Right: https://huggingface.co/blog/huggingface/tito
> a full worked example, opencode owning its loop trained with GRPO: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
> Harness, Scaffold, and the AI Agent Terms Worth Getting Right: https://huggingface.co/blog/agent-glossary
on a similar line:
https://x.com/SergioPaniego/status/2062911580564496576
sergiopaniego
posted an update about 2 months ago
Post
794
Something I really like when I study a subject is understanding its history, how it reached the point where it is today
I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words
This is the companion piece to Class 3 of our Training Agents series with @burtenshaw . The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale
https://huggingface.co/blog/sergiopaniego/agentic-rl-2026
I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words
This is the companion piece to Class 3 of our Training Agents series with @burtenshaw . The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale
https://huggingface.co/blog/sergiopaniego/agentic-rl-2026
sergiopaniego
posted an update 2 months ago
Post
363
we just released a new blog "Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv"
you can take a real coding agent (OpenCode), let it run its own tool loop against real coding problems, and train it with RL on the exact tokens it produced
and every rollout runs in its own remote HF sandbox, so rollouts scale out beyond one machine
the loop:
- OpenCode owns its tool loop inside an OpenEnv sandbox
- an in-sandbox proxy records the real token ids + logprobs, per turn
- a hidden-test verifier scores the result, and that is the reward
- TRL trains with AsyncGRPO, weights sync back to vLLM over NCCL
blog + runnable example: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
you can take a real coding agent (OpenCode), let it run its own tool loop against real coding problems, and train it with RL on the exact tokens it produced
and every rollout runs in its own remote HF sandbox, so rollouts scale out beyond one machine
the loop:
- OpenCode owns its tool loop inside an OpenEnv sandbox
- an in-sandbox proxy records the real token ids + logprobs, per turn
- a hidden-test verifier scores the result, and that is the reward
- TRL trains with AsyncGRPO, weights sync back to vLLM over NCCL
blog + runnable example: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
sergiopaniego
posted an update 2 months ago
Post
2686
LFM2.5-2.6B just dropped!
and the @liquidai blog comes with some nice details about the training procedure, so let's analyze it.
basically, a full agent training pipeline but compressed into 2.6B
base model → SFT → specialized teachers per domain (SFT + RLVR) → on-policy distillation back into one student → agentic RL
the two most interesting stages
→ MOPD: the student generates, each prompt routes to its domain teacher for token-level feedback. teachers branch from the same SFT checkpoint, so their signal stays close to the student's distribution
→ agentic RL: multi-turn GRPO inside real harnesses (OpenClaw, Hermes Agent), one sandbox per rollout, a proxy captures token-level trajectories while the harness stays a black box
this makes a 2.6B that beats much larger models on instruction following and tool use
SFT, distillation, RL, RL envs: exactly what we're covering in our Training Agents livestream series (next one coming soon!)
→ model: LiquidAI/LFM2.5-2.6B
→ blog: https://www.liquid.ai/blog/lfm2-5-2-6b
→ live series: https://www.youtube.com/playlist?list=PLo2EIpI_JMQvQZm-kVlz4wY1vWF0LBcf5
and the @liquidai blog comes with some nice details about the training procedure, so let's analyze it.
basically, a full agent training pipeline but compressed into 2.6B
base model → SFT → specialized teachers per domain (SFT + RLVR) → on-policy distillation back into one student → agentic RL
the two most interesting stages
→ MOPD: the student generates, each prompt routes to its domain teacher for token-level feedback. teachers branch from the same SFT checkpoint, so their signal stays close to the student's distribution
→ agentic RL: multi-turn GRPO inside real harnesses (OpenClaw, Hermes Agent), one sandbox per rollout, a proxy captures token-level trajectories while the harness stays a black box
this makes a 2.6B that beats much larger models on instruction following and tool use
SFT, distillation, RL, RL envs: exactly what we're covering in our Training Agents livestream series (next one coming soon!)
→ model: LiquidAI/LFM2.5-2.6B
→ blog: https://www.liquid.ai/blog/lfm2-5-2-6b
→ live series: https://www.youtube.com/playlist?list=PLo2EIpI_JMQvQZm-kVlz4wY1vWF0LBcf5
sergiopaniego
posted an update 2 months ago
Post
2706
Simon Willison (@simonw ) has asked every new model to draw a pelican riding a bicycle for some time now
you look at the drawing and you know. but there is no number, so nothing can train against it, no?
I turned this idea into an rl env in OpenEnv. now, you can eval any model against it, and train against it with TRL
read the details!🤓
https://huggingface.co/blog/sergiopaniego/pelican-env-openenv
you look at the drawing and you know. but there is no number, so nothing can train against it, no?
I turned this idea into an rl env in OpenEnv. now, you can eval any model against it, and train against it with TRL
read the details!🤓
https://huggingface.co/blog/sergiopaniego/pelican-env-openenv
sergiopaniego
posted an update 2 months ago
Post
264
yesterday we had Class 3 of the Training Agents live series
we went through how GRPO works in depth and applied it to real experiments with TRL
sharing the resources in case you want to dig in, enjoy!
🎥 session recording: https://www.youtube.com/watch?v=ztdTed5egrM
📄 slides with links: https://docs.google.com/presentation/d/19v5_HR5B-1CPHuoZ-RXjhgBrGFXGnBNd6TfLhsXWT1c/edit?usp=sharing
we went through how GRPO works in depth and applied it to real experiments with TRL
sharing the resources in case you want to dig in, enjoy!
🎥 session recording: https://www.youtube.com/watch?v=ztdTed5egrM
📄 slides with links: https://docs.google.com/presentation/d/19v5_HR5B-1CPHuoZ-RXjhgBrGFXGnBNd6TfLhsXWT1c/edit?usp=sharing
sergiopaniego
posted an update 2 months ago
Post
2954
quick reminder! 🚨
tomorrow (Tuesday, July 28), we're back with Class 3 of the Training Agents live series
🧠 what: reinforcement learning for training agents (GRPO): how it works, how to implement it in TRL, and end-to-end examples
🗓️ when: Tuesday, July 28 - 🕔 5:00 PM CEST / 8:30 PM IST
📍 where: Live on @huggingface 's X, YouTube, and LinkedIn
live: https://www.youtube.com/watch?v=ztdTed5egrM
class 1: https://x.com/SergioPaniego/status/2069382207618379813
class 2: https://x.com/SergioPaniego/status/2075180665184686187
tomorrow (Tuesday, July 28), we're back with Class 3 of the Training Agents live series
🧠 what: reinforcement learning for training agents (GRPO): how it works, how to implement it in TRL, and end-to-end examples
🗓️ when: Tuesday, July 28 - 🕔 5:00 PM CEST / 8:30 PM IST
📍 where: Live on @huggingface 's X, YouTube, and LinkedIn
live: https://www.youtube.com/watch?v=ztdTed5egrM
class 1: https://x.com/SergioPaniego/status/2069382207618379813
class 2: https://x.com/SergioPaniego/status/2075180665184686187
sergiopaniego
posted an update 2 months ago
Post
273
you can now train your own coding agents with trl + openenv, starting with opencode
we just added end-to-end support for training agent harnesses:
> TRL: a loop-owning training path (AsyncGRPOTrainer + HarnessRolloutWorker) that launches the agent in an OpenEnv session, reads back its trace, reconstructs the training samples, and trains with AsyncGRPO
> OpenEnv: the OpenCode harness environment plus a transparent proxy that forwards the agent's model calls and records each turn's token ids and logprobs
you train the actual opencode agent as is, it runs its own loop and tools and the policy learns from the exact tokens it produced
we're shipping a self-contained example: local subprocess sandbox, DeepCoder problems, validated on Qwen3-8B.
> example: https://github.com/huggingface/trl/blob/main/examples/scripts/openenv/opencode.py
> docs: https://huggingface.co/docs/trl/main/openenv
and we're working actively on both sides so expect more 🤓
we just added end-to-end support for training agent harnesses:
> TRL: a loop-owning training path (AsyncGRPOTrainer + HarnessRolloutWorker) that launches the agent in an OpenEnv session, reads back its trace, reconstructs the training samples, and trains with AsyncGRPO
> OpenEnv: the OpenCode harness environment plus a transparent proxy that forwards the agent's model calls and records each turn's token ids and logprobs
you train the actual opencode agent as is, it runs its own loop and tools and the policy learns from the exact tokens it produced
we're shipping a self-contained example: local subprocess sandbox, DeepCoder problems, validated on Qwen3-8B.
> example: https://github.com/huggingface/trl/blob/main/examples/scripts/openenv/opencode.py
> docs: https://huggingface.co/docs/trl/main/openenv
and we're working actively on both sides so expect more 🤓
sergiopaniego
posted an update 2 months ago
Post
1569
you can train DiffusionGemma (a block-diffusion LLM) in TRL! and we're sharing an example for it
TRL trainers are made to be easily extended and adapted to different real-world use cases.
in this one, with a single method overridden in SFTTrainer (compute_loss), you can train this model
> example: https://github.com/huggingface/trl/blob/main/examples/scripts/sft_diffusion_gemma.py
TRL trainers are made to be easily extended and adapted to different real-world use cases.
in this one, with a single method overridden in SFTTrainer (compute_loss), you can train this model
> example: https://github.com/huggingface/trl/blob/main/examples/scripts/sft_diffusion_gemma.py
sergiopaniego
posted an update 2 months ago
Post
284
join us next Tuesday, July 28, for Class 3 of the Training Agents live series!
we'll dive into reinforcement learning for agent training, covering the intuition behind GRPO, how it works, and how to implement it in TRL with practical, e2e examples
see you there 🤠
live: https://www.youtube.com/live/ztdTed5egrM
> in case you missed class 1:
https://x.com/SergioPaniego/status/2069382207618379813
> and in case you missed class 2: https://x.com/SergioPaniego/status/2075180665184686187
we'll dive into reinforcement learning for agent training, covering the intuition behind GRPO, how it works, and how to implement it in TRL with practical, e2e examples
see you there 🤠
live: https://www.youtube.com/live/ztdTed5egrM
> in case you missed class 1:
https://x.com/SergioPaniego/status/2069382207618379813
> and in case you missed class 2: https://x.com/SergioPaniego/status/2075180665184686187
sergiopaniego
posted an update 3 months ago
Post
7831
Frontier models use distillation as a step of their post-training pipelines.
In 2026 it has three jobs: compress a big model into a small one, merge RL experts into a single model, and let a model teach itself.
I wrote up which frontier models use each one and how: https://huggingface.co/blog/sergiopaniego/distillation-2026
It pairs with Class 2 of the Training an Agent series Ben and I are doing, where we teach these techniques hands-on with TRL!
In 2026 it has three jobs: compress a big model into a small one, merge RL experts into a single model, and let a model teach itself.
I wrote up which frontier models use each one and how: https://huggingface.co/blog/sergiopaniego/distillation-2026
It pairs with Class 2 of the Training an Agent series Ben and I are doing, where we teach these techniques hands-on with TRL!
sergiopaniego
posted an update 3 months ago
Post
411
TRL v1.7.0 is out‼️
+ continuous batching makes GRPO and RLOO 1.25x faster at -16 GB
+ proper MoE post-training across GRPO/RLOO/AsyncGRPO
+ new GMPO trainer
+ AsyncGRPO weight sync + padding-free
+ more
https://github.com/huggingface/trl/releases/tag/v1.7.0
wrote a small article about the continuous batching for GRPO feature
https://huggingface.co/blog/sergiopaniego/cb-trl-grpo
+ continuous batching makes GRPO and RLOO 1.25x faster at -16 GB
+ proper MoE post-training across GRPO/RLOO/AsyncGRPO
+ new GMPO trainer
+ AsyncGRPO weight sync + padding-free
+ more
https://github.com/huggingface/trl/releases/tag/v1.7.0
wrote a small article about the continuous batching for GRPO feature
https://huggingface.co/blog/sergiopaniego/cb-trl-grpo
sergiopaniego
posted an update 4 months ago
Post
389
Continuous batching just landed in TRL for GRPO!
At 64 generations it runs faster and uses less VRAM than plain generate, no vLLM needed
How it works and when to reach for it, below
https://huggingface.co/blog/sergiopaniego/cb-trl-grpo
At 64 generations it runs faster and uses less VRAM than plain generate, no vLLM needed
How it works and when to reach for it, below
https://huggingface.co/blog/sergiopaniego/cb-trl-grpo
sergiopaniego
posted an update 4 months ago
Post
369
GLM-5.2 is open and comes with competitive performance against opus 4.8
day-0 in transformers + vllm + sglang, mit license 🤗
on the post-training side: critic-based ppo for variable-length agentic rollouts (ppo is back!) + an online anti-reward-hacking module that feeds the agent dummy info when it tries to cheat
day-0 in transformers + vllm + sglang, mit license 🤗
on the post-training side: critic-based ppo for variable-length agentic rollouts (ppo is back!) + an online anti-reward-hacking module that feeds the agent dummy info when it tries to cheat
sergiopaniego
posted an update 4 months ago
Post
4020
OpenEnv has a new home: github.com/huggingface/OpenEnv
Starting today, it's coordinated by a committee that includes Meta-PyTorch, Reflection, Unsloth, Modal, Prime Intellect, Nvidia, Mercor, Fleet AI, and Hugging Face
frontier labs train their models and their harnesses together. Claude knows Claude Code. GPT-5.5 knows Codex. that's not an accident, it's training. open-source models deserve the same magic, but pulling that off requires infrastructure that belongs to everyone, not one lab
OpenEnv is that layer. one api, any harness, any trainer, any environment
Rewards and training loops stay in TRL, Unsloth, wherever you already work. OpenEnv is the socket they all plug into
Get involved!
Full announcement: https://huggingface.co/blog/openenv-agentic-rl
Starting today, it's coordinated by a committee that includes Meta-PyTorch, Reflection, Unsloth, Modal, Prime Intellect, Nvidia, Mercor, Fleet AI, and Hugging Face
frontier labs train their models and their harnesses together. Claude knows Claude Code. GPT-5.5 knows Codex. that's not an accident, it's training. open-source models deserve the same magic, but pulling that off requires infrastructure that belongs to everyone, not one lab
OpenEnv is that layer. one api, any harness, any trainer, any environment
Rewards and training loops stay in TRL, Unsloth, wherever you already work. OpenEnv is the socket they all plug into
Get involved!
Full announcement: https://huggingface.co/blog/openenv-agentic-rl