Qwen3 $8B Post Training Ternarisation
Covered by
arxiv.org
Timeline
- Qwen3-8B ternarised post-training retains 78.5% of FP16 accuracy, packs to 8.24 GiB
It gives a reproducible, matched 4B/8B baseline showing that larger models tolerate post-training ternarisation better, while confirming that packed execution is feasible but not yet faster than FP16.
1 source · excerpt-only
All sources (1)
- arXiv cs.LG2026-09-10
Related subjects
- Qwen3 8 $27B Task Aware QuantLocal AI Scene
- Embedding Model Migration Without ReembeddingLocal AI Scene
- Cosmos3 Int4 Quant Local ReleaseLocal AI Scene