Back to News
User submissionlws.io

My Local Model Setup on an M4 Pro Mac mini

Summary

The author runs Qwen3.6-35B-A3B and Gemma-4-E4B locally on a 48 GB M4 Pro Mac mini through oMLX. Hermes, Tailscale, and several client apps share the same private inference endpoint across a MacBook and iPhone. The setup is intended to cover routine workloads with predictable costs, low latency, offline access, and greater control over sensitive data, while cloud models remain available for more demanding tasks. The article also explains how MoE active parameters, quantization, context memory, and system overhead affect hardware requirements.