Program-as-Weights: A Programming Paradigm for Fuzzy Functions Paper • 2607.02512 • Published 18 days ago • 120
AgenticDataBench: A Comprehensive Benchmark for Data Agents Paper • 2607.01647 • Published 18 days ago • 37
AutoTrainess: Teaching Language Models to Improve Language Models Autonomously Paper • 2606.31551 • Published 20 days ago • 24
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published 25 days ago • 56
Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Paper • 2606.15932 • Published Jun 16 • 38
Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention Paper • 2606.20945 • Published Jun 18 • 80
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions Paper • 2606.23654 • Published 28 days ago • 79
GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents Paper • 2606.18829 • Published Jun 17 • 18
💻 Qwopus-Coder Collection Reasoning-distilled coding models optimized for specialized domains like agentic workflows. • 10 items • Updated 19 days ago • 41
HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing Paper • 2412.01778 • Published Dec 2, 2024 • 3
Environment-Grounded Multi-Agent Workflow for Autonomous Penetration Testing Paper • 2603.24221 • Published Mar 25 • 1
Getting pwn'd by AI: Penetration Testing with Large Language Models Paper • 2308.00121 • Published Jul 24, 2023 • 1
Knowledge-Informed Auto-Penetration Testing Based on Reinforcement Learning with Reward Machine Paper • 2405.15908 • Published May 24, 2024 • 1
Pentest-R1: Towards Autonomous Penetration Testing Reasoning Optimized via Two-Stage Reinforcement Learning Paper • 2508.07382 • Published Aug 10, 2025 • 2
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool Paper • 2308.06782 • Published Aug 13, 2023 • 1
CIPHER: Cybersecurity Intelligent Penetration-testing Helper for Ethical Researcher Paper • 2408.11650 • Published Aug 21, 2024 • 3