From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix Paper • 2609.01572 • Published 3 days ago • 33
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 3 days ago • 86
Verification-Aware Training for Speculative Decoding Paper • 2608.30135 • Published 4 days ago • 8
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability Paper • 2608.30320 • Published 4 days ago • 50
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 4 days ago • 135
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published 8 days ago • 36
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation Paper • 2608.19098 • Published 16 days ago • 22
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation Paper • 2608.15062 • Published 9 days ago • 10
D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation Paper • 2608.24987 • Published 10 days ago • 25
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO Paper • 2608.27351 • Published 8 days ago • 21
Meta^n: Recursive Self-Improvement through Emergent Depth Paper • 2608.24735 • Published 10 days ago • 15
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts Paper • 2604.19835 • Published Apr 21 • 22
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs? Paper • 2603.24472 • Published Mar 25 • 59
CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment Paper • 2608.21278 • Published 14 days ago • 5