Local AI Is Usable Now
Summary
A hands-on report examines whether local AI has become practical for continuous, private work. On a Mac Studio with an Apple M5 Ultra and 256 GB of unified memory, the author organizes models into four roles: a large text model, small text and vision models, and a decision model. Loading DeepSeek V4 Flash alongside three smaller models used about 220 GB, leaving too little working room, so two profiles were created. The everyday Work profile budgets 99 GB for Qwen3.8 Flash Next, 21 GB each for Gemma 4 26B and Qwen 3.6 35B-A3B, and 18.5 GB for Clef-Flash, leaving about 82 GB for context after macOS. A Smart profile gives DeepSeek the large slot and leaves about 64 GB. In benchmarks using uncached, streamed prompts, the small models generated roughly 91 to 115 tokens per second, while Qwen3.8 Flash Next produced 51 to 58 and DeepSeek V4 Flash 27 to 31. Flash Next used speculative decoding, accepting about 55% to 65% of draft tokens. These are speed measurements rather than quality evaluations: the author did not compare against cloud APIs, test task accuracy, or measure how much work moved local. The local model was unreliable as a plan manager, so frontier models still handle repository reading, ambiguity, planning, and task contracts, while local sessions execute bounded contracts and return results. For privacy-sensitive work, planning and execution can remain on the machine. A 9B Clef-Flash model also provides typed decisions and probabilities without free-form output; it was retained because it used 19 GB, compared with 52 GB for a 27B version, though calibration was not measured. The setup includes monitoring, redacted logs, memory-aware profile changes, and a planned router that will choose among local models and paid APIs. The author concludes that local inference is now useful for some always-running jobs, with a single small model requiring about 15 to 20 GB of resident memory.