When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models Paper • 2609.19671 • Published 6 days ago • 45
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published 14 days ago • 44
A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM Paper • 2609.07821 • Published 15 days ago • 15
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published 27 days ago • 155
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents Paper • 2608.18423 • Published Aug 19 • 21
Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness Paper • 2608.09900 • Published Aug 10 • 13
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory Paper • 2608.07169 • Published Aug 7 • 51
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Paper • 2608.01837 • Published Aug 3 • 40
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Paper • 2607.28617 • Published Jul 30 • 37