Pretraining Transformers with Quantized Softmax in Attention Paper • 2609.33591 • Published 13 days ago • 7
Softmax Reparameterization for Output-Head Quantization Paper • 2609.31291 • Published 15 days ago • 6
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration Paper • 2609.33586 • Published 13 days ago • 4
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration Paper • 2609.33586 • Published 13 days ago • 4
Pretraining Transformers with Quantized Softmax in Attention Paper • 2609.33591 • Published 13 days ago • 7
Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration Paper • 2609.33586 • Published 13 days ago • 4
Pretraining Transformers with Quantized Softmax in Attention Paper • 2609.33591 • Published 13 days ago • 7