SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 7 days ago • 129
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 3 days ago • 175
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 6 days ago • 27
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 9 days ago • 70
LimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence Paper • 2609.17488 • Published 9 days ago • 661
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 17 days ago • 372
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up Paper • 2609.13443 • Published 13 days ago • 13
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 10 days ago • 244
One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation Paper • 2608.25936 • Published 29 days ago • 16
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 21 days ago • 101
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning Paper • 2609.00638 • Published 23 days ago • 66
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall Paper • 2609.01532 • Published 23 days ago • 9
Evaluating the Hidden Costs of Personalization in Large Language Models Paper • 2608.28833 • Published 27 days ago • 30
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 24 days ago • 97
The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published Aug 13 • 16
DarwinX: Evolving Agent Harnesses Through Natural Selection Paper • 2608.07545 • Published Jul 31 • 116