Back to News
RSS feedwww.itsthatlady.dev

A Beginner’s Guide to Running AI Models Locally

Summary

This beginner-focused guide explains local AI as running a downloadable model on your own laptop, desktop, Mac, or other hardware instead of sending prompts to a cloud service. It highlights privacy, offline access, no recurring subscription fees, control over model and tools, and the value of learning how models work, while noting that cloud systems remain powerful and convenient. The guide recommends at least 8GB of RAM and 5GB or more of free storage as a starting point, and explains that storage holds the model while RAM or VRAM loads it for inference; models that exceed available memory may run slowly, fail to load, or crash. It introduces parameter counts such as 3B, 7B, 13B, 30B, and 70B, linking larger models with greater hardware demands, and advises beginners to start with a smaller model. Quantization, including Q4, Q5, and Q8 variants, reduces memory needs by trading some quality for a smaller file. The article also defines tokens, context windows, embeddings, evaluations, benchmarks, RAM, VRAM, Apple Silicon unified memory, and GGUF files. Hugging Face is presented as a place to browse and download open-source models. For actually running them, the guide lists Ollama, LM Studio, llama.cpp, Jan, and Msty, then demonstrates LM Studio on a MacBook Pro with an M1 chip and 16GB of unified memory. LM Studio recommends Gemma4’s gemma-4-e4b model, which the author says needs about 7GB of RAM; after downloading and loading it, the user can begin a local inference session. The closing advice is to choose the smallest model that fits the task, compare models with real prompts, and observe how the computer responds.