Post
60
quant_eval v7.22: what quantization does to agent behavior
Six paired runs, four models, Q4_K_M / Q5_K_M / Q8_0 vs the full-weight F16 lane on identical fixtures: 19,200 per-case results across 8 agent capabilities.
- Mistral-Nemo Q4_K_M: 5 of 8 families degraded (mixed_brief_json −28 pts, toolcall_only −21.5)
- Q5_K_M and Q8_0: no significant change in any family. The cliff is between Q4 and Q5.
- Ministral-3-14B vs BF16: Q4_K_M lost 1 family; Q5_K_M and Q8_0 lost none.
Paper: https://doi.org/10.5281/zenodo.22948597
Data (CC BY 4.0): https://doi.org/10.5281/zenodo.22009419
Arena: pbhappliedsystems/quant-eval-agent-arena
Which model should we measure next?
Six paired runs, four models, Q4_K_M / Q5_K_M / Q8_0 vs the full-weight F16 lane on identical fixtures: 19,200 per-case results across 8 agent capabilities.
- Mistral-Nemo Q4_K_M: 5 of 8 families degraded (mixed_brief_json −28 pts, toolcall_only −21.5)
- Q5_K_M and Q8_0: no significant change in any family. The cliff is between Q4 and Q5.
- Ministral-3-14B vs BF16: Q4_K_M lost 1 family; Q5_K_M and Q8_0 lost none.
Paper: https://doi.org/10.5281/zenodo.22948597
Data (CC BY 4.0): https://doi.org/10.5281/zenodo.22009419
Arena: pbhappliedsystems/quant-eval-agent-arena
Which model should we measure next?