🪨 ONYX

Adaptive Precision Engine

This repo contains ONYX Quants of ornith-ai/Ornith-1.5-35B-A3B with a custom chat template

⚙️ The ONYX Architecture

🧠 Dynamic Layer Sensitivity

Replaces hardcoded edge boundaries with real activation variance measurements. ONYX autonomously identifies critical layers (like mid-network attention blocks) and protects them dynamically.

🎯 Router-Weighted Imatrix

Captures MoE router probabilities and multiplies them into activation scales. This forces the quantizer to aggressively crush "cold" experts while fiercely protecting "hot" ones within the same tensor block.

🏗️ Architecture-Agnostic

Dynamically reads HuggingFace modules and the generated F16 GGUF to map tensors. No hardcoded regex. Works out-of-the-box on Llama, DeepSeek, and custom hybrid SSM/MoE architectures.

Tier Name Target Quality Target Size Middle Layer Strategy
🪨 qualityQ8 Match~21 GBIQ4_XS
⚖️ balancedQ6 Match~24 GBQ5_K
📦 compactQ4 Match~16 GBQ3_K
🚀 miniQ2 Match~12 GBIQ2_XXS

📚 Credits & Foundations

👉 APEX Quantization Method
Ettore Di Giacinto & Richard Palethorpe (LocalAI Team). ONYX evolves the layer-wise precision gradients and MoE-aware tensor classification outlined in the APEX technical paper into a fully dynamic, data-driven engine.

👉 Bartowski and Lamim
For the excellent semantic imatrix calibration dataset that powers ONYX's activation scaling.

👉 llama.cpp
Georgi Gerganov and contributors for the foundational inference and quantization engine.

👉 HuggingFace Accelerate
For the init_empty_weights() context manager that makes the 0-RAM "Ghost Model" possible on consumer hardware.

Support the Project

A coffee in Ethereum would be cool! Although I don't drink coffee—I think it tastes like burnt water—but a pink lemonade would be fire! 🔥

0xDEE7fa8C421BD038D32e4441ea1aDe72fE973706

recommended sampling parameters:

--temp 0.85 --top-p 0.95 --top-k 40 --min-p 0.05 --repeat-penalty 1.1 --presence-penalty 0.0
Downloads last month
-
GGUF
Model size
36B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for el4/Ornith-1.5-35B-A3B-ONYX-GGUF

Quantized
(132)
this model

Collection including el4/Ornith-1.5-35B-A3B-ONYX-GGUF