Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 7 days ago • 202
Mi-Ripple: Restoring Images Degraded by Iterative AI Editing Paper • 2609.11317 • Published 7 days ago • 53
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 9 days ago • 419
VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published 22 days ago • 179
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published about 1 month ago • 341
Demystifying Agent Skills: Why They Work-Until They Don't Paper • 2608.14036 • Published Aug 14 • 170
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published Aug 12 • 293
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say Paper • 2606.00152 • Published Aug 6 • 4
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Paper • 2608.01755 • Published Aug 3 • 142