VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 4 days ago • 35
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published 10 days ago • 68
view article Article Profiling in PyTorch (Part 3): Attention is all you profile +2 ariG23498, sergiopaniego, sayakpaul, ror • Jul 10 • 50
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 8 days ago • 714
The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements Paper • 2509.01809 • Published 10 days ago • 4
view article Article Bringing Nunchaku 4-bit Diffusion Inference to Diffusers rootonchair, sayakpaul • Jul 23 • 69
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 15 days ago • 153
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 15 days ago • 83
H3-World: Turning Language Understanding into World Control Paper • 2609.01560 • Published 17 days ago • 51