Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results Paper β’ 2606.14516 β’ Published Jun 12 β’ 6
view article Article Featuring Every Eval Ever Results on Hugging Face Model Pages +5 deepmage121, evijit, SaylorTwift, janbatzner, borgr, irenesolaiman, julien-c β’ 23 days ago β’ 47
Runtime error 10 INTIMA Companionship Benchmark Responses πΊ 10 Visualizing model responses to companionship prompts
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers Paper β’ 2504.20752 β’ Published Apr 29, 2025 β’ 97