AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
datasets 25
benchflow/frontierphysics-pr889-evidence
Updated • 20
benchflow/frontierphysics-pr887-evidence
Updated • 47
benchflow/frontierphysics-pr888-evidence
Updated • 28
benchflow/frontierphysics-pr885-evidence
Updated • 25
benchflow/frontierphysics-pr884-evidence
Updated • 24
benchflow/frontierphysics-pr883-evidence
Updated • 23
benchflow/frontierphysics-pr881-evidence
Updated • 50
benchflow/frontierphysics-pr882-evidence
Updated • 36
benchflow/frontierphysics-pr879-evidence
Updated • 49
benchflow/frontierphysics-pr880-evidence
Updated • 12