X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published 5 days ago • 43
TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model Paper • 2609.09158 • Published 7 days ago • 23
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets Paper • 2609.05663 • Published 11 days ago • 22
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published 12 days ago • 49
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published 19 days ago • 155
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published 27 days ago • 54
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models Paper • 2608.13049 • Published Aug 13 • 20
MASS: Multiplayer World Models with Authoritative Shared State Paper • 2608.06257 • Published Aug 10 • 17
ChronoVision: Temporal Reasoning via Latent State Reconstruction Paper • 2608.05631 • Published Aug 6 • 40
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published Aug 5 • 23
CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Paper • 2608.02589 • Published Aug 3 • 25