Task-aware quant of Qwen3 27B hits 99% of BF16 reasoning at 15% size
A developer has released TAK, a task-aware quantization pipeline, and reports it preserves nearly all of a model's reasoning ability at a fraction of the size. The TAK quant of Qwen3 27B scored 82.81% on a reasoning benchmark, compared with 83.59% for the BF16 original and 77.34% for Unsloth's byte-matched UD IQ2_S, which the author included as an industry-standard reference. That puts TAK at roughly 99% of BF16 reasoning performance at about 15% of the size.
The author describes TAK, short for Task Aware Knapsack, as a blend of two earlier approaches, TASA and TAQ. The pipeline builds an imatrix from a task-specific corpus, finds the smallest model size before complete collapse, then promotes and demotes tensors within a byte-specific budget. The result is a purpose-built quantization rather than a general-purpose recovery.
The work grew out of Qlab, a broader measurement-heavy system the author used to decide what to test. Once a reliable pipeline emerged, the author retired Qlab and specialized it into TAK. Coding is the next target domain. The author did not release benchmark methodology details or model weights in the post.
If the results hold, task-aware quantization could let users run near-full-precision reasoning models on far smaller hardware footprints.