NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 3 days ago • 386
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 8 days ago • 89
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 8 days ago • 80
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization Paper • 2608.20281 • Published 22 days ago • 13
Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning Paper • 2608.18746 • Published 23 days ago • 18
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 263
CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG Paper • 2608.07458 • Published Aug 7 • 10
On-Policy Delta Distillation for Multilingual Math Reasoning Paper • 2608.05802 • Published Aug 6 • 32
Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming Paper • 2608.05108 • Published Aug 5 • 9
HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Image-Text-to-Text • 35B • Updated Apr 17 • 1.34M • 3.63k
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Paper • 2607.23782 • Published Jul 26 • 78