Kalman Delta Networks: Uncertainty-aware Associative Memory Paper • 2609.07816 • Published 5 days ago • 27
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation Paper • 2609.02998 • Published 10 days ago • 16
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation Paper • 2609.05295 • Published 8 days ago • 13
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published 8 days ago • 20
MDN: Parallelizing Stepwise Momentum for Delta Linear Attention Paper • 2605.05838 • Published May 7 • 6
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 9 days ago • 177
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 9 days ago • 80
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 9 days ago • 91
Post-Training Language Models for Gold-Medal Performance in Coding Competitions Paper • 2609.02849 • Published 10 days ago • 11
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model Paper • 2607.22083 • Published Jul 27 • 11
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 12 days ago • 58
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 12 days ago • 151
Self-Improving Pretraining: using post-trained models to pretrain better models Paper • 2601.21343 • Published Jan 29 • 21
A Programming Paradigm for Spatiotemporal Composability Paper • 2608.25512 • Published 17 days ago • 16