RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 9 days ago • 277
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks Paper • 2610.00972 • Published 9 days ago • 58
Beyond Accuracy: A Cognitive Load Framework for Mapping the Capability Boundaries of Tool-use Agents Paper • 2601.20412 • Published Jan 28 • 2
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families Paper • 2412.06540 • Published Dec 9, 2024 • 3
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 14 days ago • 326
AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation Paper • 2609.35530 • Published 12 days ago • 22
Chinese-Jev: Bringing System One Model to Chinese-Language Tasks Paper • 2609.36965 • Published 11 days ago • 26
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding? Paper • 2609.38079 • Published 11 days ago • 56
EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments? Paper • 2609.37686 • Published 11 days ago • 72
Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations Paper • 2609.32522 • Published 14 days ago • 82