Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 18 days ago • 98
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability? Paper • 2609.00787 • Published 20 days ago • 19
PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments Paper • 2608.14441 • Published Aug 14 • 29
Rethinking the Role of Efficient Attention in Hybrid Architectures Paper • 2606.15378 • Published Jun 13 • 21
ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning Paper • 2505.04881 • Published May 8, 2025 • 1
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models Paper • 2505.22617 • Published May 28, 2025 • 132