When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 5 days ago • 99
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published Aug 17 • 50
Modular TTT: Rethinking Test-Time Training as Composable Modules Paper • 2608.07110 • Published Aug 7 • 10
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 178
A Quantized Native Runtime for On-Device Semantic Audio Generation Paper • 2607.08526 • Published Jul 9 • 3
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published Jul 4 • 42
BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery Paper • 2606.20997 • Published Jun 19 • 11