Apple A20 Pro chip doubles Neural Engine, widens memory bus to 96-bit
Apple's A20 Pro chip, reportedly built on a 2nm process, debuts with a 7-core GPU, a 32-core Neural Engine, and roughly 115 GB/s of memory bandwidth, according to a post on the r/LocalLLaMA subreddit. The figures, which have not been confirmed by Apple, describe a chip that widens the memory interface from the 64-bit buses used in previous Apple silicon to a 96-bit LPDDR5X bus, a change that accounts for the roughly 50 percent jump in bandwidth over prior generations.
The memory bandwidth increase is the headline number for local AI workloads. At about 115 GB/s, the A20 Pro moves meaningfully closer to the throughput that makes running larger language models on-device practical, since memory bandwidth is the primary constraint on token generation speed for models that fit in unified memory. The post frames the wider bus as expensive silicon on a 2nm node, suggesting Apple absorbed a real cost to close the bandwidth gap.
The Neural Engine also doubles in size, from 16 to 32 cores total. That is a structural change rather than a clock-speed bump, and it points to Apple pushing more of its on-device AI inference through the dedicated accelerator rather than the GPU. The 7-core GPU count is unchanged from what the post describes as the prior configuration, meaning the added bandwidth and Neural Engine capacity are the substantive upgrades in this generation.
The post does not specify which device the A20 Pro ships in, nor does it cite a teardown, benchmark, or Apple documentation. The information appears to be derived from early reporting or analysis of the chip's specifications, and none of it has been independently verified. Apple has not commented on the A20 Pro, and the company typically does not disclose detailed memory bus or Neural Engine core counts in its marketing materials, leaving such figures to chip analysts and post-launch teardowns.
For the local LLM community that the subreddit serves, the practical question is what the bandwidth and Neural Engine changes mean for real inference performance. Doubling the Neural Engine cores could substantially accelerate Apple's Core ML and on-device transformer workloads, while the wider memory bus raises the ceiling on model size and context length that runs comfortably in unified memory. Neither effect has been benchmarked publicly yet.
What remains unknown is whether the A20 Pro's specifications as described hold up under scrutiny, which devices will carry the chip, and when Apple will announce it. The post offers no release timeline. Until Apple confirms the silicon or independent measurements surface, the 96-bit bus, 32-core Neural Engine, and 115 GB/s figure should be treated as unverified early specifications.
The A20 Pro's wider memory bus and doubled Neural Engine could meaningfully raise the ceiling for on-device LLM inference, but the specifications are unconfirmed.