SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents Paper • 2609.08149 • Published 17 days ago • 28
Intern-S2-Preview: Scientific Agentic Foundation Model Paper • 2608.13505 • Published Aug 13 • 76
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Paper • 2607.13705 • Published Jul 15 • 40
Rethinking Verification for LLM Code Generation: From Generation to Testing Paper • 2507.06920 • Published Jul 9, 2025 • 29
Coding Triangle: How Does Large Language Model Understand Code? Paper • 2507.06138 • Published Jul 8, 2025 • 23
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference Paper • 2502.18411 • Published Feb 25, 2025 • 74