Post
83
We put Qwen2.5-3B and its Instruct fine-tune under an X-ray — layer by layer, two scans running side by side on our GPU fleet.
🩺 Structural comparison (base → instruct): the change starts at layer 4 and spreads across 33 of 37 measurement points (~89% of the scanned depth) — in this pair, instruction tuning is not a "last few layers" story.
🧠 Knowledge delta, on a 20-item probe set: every fact the base model knew survived (18/20 → 18/20, zero broken), and fabrication-avoidance improved from 13/20 to 18/20. The internal true/false separation signal got slightly weaker (AUROC 0.915 → 0.878) — an honest trade-off worth knowing about.
Full interactive reports (no login needed):
🔗 structure: https://www.tetracta.ai/llm_tomografi/r/a048cc12934d488aa415d0ee0c1abeca/eGWvi64CrsRQAwFirfqmVg
🔗 knowledge: https://www.tetracta.ai/llm_tomografi/r/d0a0a6ea76f04b59b35307313b7e1cf0/WHH9qVypf4j0Fd6XPjP7Yw
Model X-Ray is in free open beta — 20 scans per account, any public safetensors checkpoint up to 7B, running on our own small GPU fleet: https://www.tetracta.ai/xray.html
If you fine-tune or merge models: scan your checkpoint before and after, and tell us what you see — we answer every message.
🩺 Structural comparison (base → instruct): the change starts at layer 4 and spreads across 33 of 37 measurement points (~89% of the scanned depth) — in this pair, instruction tuning is not a "last few layers" story.
🧠 Knowledge delta, on a 20-item probe set: every fact the base model knew survived (18/20 → 18/20, zero broken), and fabrication-avoidance improved from 13/20 to 18/20. The internal true/false separation signal got slightly weaker (AUROC 0.915 → 0.878) — an honest trade-off worth knowing about.
Full interactive reports (no login needed):
🔗 structure: https://www.tetracta.ai/llm_tomografi/r/a048cc12934d488aa415d0ee0c1abeca/eGWvi64CrsRQAwFirfqmVg
🔗 knowledge: https://www.tetracta.ai/llm_tomografi/r/d0a0a6ea76f04b59b35307313b7e1cf0/WHH9qVypf4j0Fd6XPjP7Yw
Model X-Ray is in free open beta — 20 scans per account, any public safetensors checkpoint up to 7B, running on our own small GPU fleet: https://www.tetracta.ai/xray.html
If you fine-tune or merge models: scan your checkpoint before and after, and tell us what you see — we answer every message.