The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction Paper • 2609.18063 • Published 5 days ago • 13
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published 13 days ago • 162
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 18 days ago • 83
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 19 days ago • 400
Metis Collection Metis persistent-memory model family based on Qwen3.5, including 4B, 9B, and 27B parameter scales. • 4 items • Updated Jul 31 • 8
CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation Paper • 2607.03803 • Published Jul 4 • 17
Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe Paper • 2606.20381 • Published Jun 18 • 10
The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models Paper • 2606.03645 • Published May 29 • 6
LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation Paper • 2606.02553 • Published Jun 1 • 19
LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws Paper • 2605.23901 • Published May 22 • 11
BitCPM-CANN Collection Full-pipeline ternary quantized model trained on CANN. • 12 items • Updated 6 days ago • 28
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs Paper • 2605.20315 • Published May 19 • 26
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos Paper • 2605.18233 • Published May 18 • 92