TUESDAY 8 SEPTEMBER 2026latent·wire25 PIECES ON FILE
← Local AI SceneLocal AI Scene

Strix Halo users say official llama.cpp wastes silicon, tout optimized forks

A Reddit user on r/LocalLLaMA argues that official llama.cpp is poorly optimized for AMD's Strix Halo APU (gfx1150), struggling to reach 50 percent of hardware theoretical throughput, and points to community forks that claim far higher performance. The post links three alternatives, each tuned for the Qwen 3.8 Flash Next (Q38FN) model that the author calls the best fit for the chip.

Halogen Flash Server, described as "Ninfer for Strix Halo," claims roughly 50 tokens per second decode and 1,200 t/s prefill, about 90 percent of theoretical peak. A llama.cpp experiment fork from myhacsint reports nearly 60 t/s decode and 600 t/s prefill, around 80 percent of theory. A third fork from halo-box, which the author calls the community's first Strix Halo fork and says has updated continuously, delivers almost 30 t/s decode and 800 t/s prefill, about 75 percent of theory.

The author estimates 90 percent of Strix Halo users run official llama.cpp, which they say wastes the device's silicon. The claims are user-reported and not independently verified.

Why it matters

If accurate, the performance gap suggests official llama.cpp leaves large headroom on AMD's Strix Halo APU, and community forks may be the practical path to full throughput for local model users.