ReadyArt/gemma-4-31B-it-scotoma-2-GGUF Text Generation ⢠31B ⢠Updated 3 days ago ⢠6.2k ⢠23
view post Post 3111 š The reasoning backbone quadruples from 8B to 32B , while the action expert remains at 2.3B! š We took a closer look at the architectural evolution from nvidia/Alpamayo-1.5-10B to nvidia/Alpamayo2-Super .Read the analysis here:https://huggingface.co/blog/JonnaMat/alpamayo2-superOur analysis explores some implications of this design choice, especially from a distillation perspective where keeping the expert compact could be key for efficient deployment. š§ See translation 3 replies Ā· š„ 7 7 + Reply
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper ⢠2608.01964 ⢠Published 7 days ago ⢠162
SequenceMatch: Imitation Learning for Autoregressive Sequence Modelling with Backtracking Paper ⢠2306.05426 ⢠Published Jun 8, 2023 ⢠1