A20 Pro Runs On-Device Language Models Up to 85% Faster
Summary
Ricky Takkar compares an iPhone 18 Pro with A20 Pro against an iPhone 17 Pro Max with A19 Pro using the same iOS version, model files, prompts, and MLX-based app build. Across Qwen3.5 0.8B and Gemma 4 E2B Instruct, the A20 Pro delivered the first response 42-53% sooner and generated visible output 53-85% faster; the answers matched word for word under greedy sampling. The comparison used four model and context conditions, with 12 post-warm-up measurements per phone and condition. In wall-clock terms, long-context Gemma responses fell from 7.58 to 4.13 seconds, while short-context Qwen responses fell from 3.27 to 2.14 seconds. Every MLX workload exceeded the 1.5x ceiling implied by the A20 Pro’s 50% theoretical memory-bandwidth increase, suggesting bandwidth alone does not explain the gains, although the test could not isolate GPU, power-management, cooling, or scheduling effects. Apple’s on-device AFM 3 Core Advanced model also improved, but by 35-42% in output speed and 23-33% in first-response latency, remaining below 1.5x. The author also observed much tighter run-to-run consistency on the A20 Pro. Results are limited to one phone per generation, short burst tests with pauses, and app-level measurements; they do not establish battery efficiency, sustained performance, Neural Engine usage, or general performance for all models.