Back to News
RSS feedgithub.com

AI Workload Placement Tests Local, Edge, and Cloud Trade-offs

Summary

This GitHub project explores how to choose an execution location and model for an AI workload while balancing quality, privacy, latency, connectivity, budget, energy, and hardware constraints. Its decision engine is available as both a command-line tool and a Streamlit web app, and returns a recommendation for local/device or cloud execution with reasoning based on benchmark data. The first benchmark, run in October 2026 on a task summarizing five short news-style paragraphs, compared llama3.2:1b through Ollama on a full-power laptop, the same model with one CPU thread as a rough edge simulation, and Claude Haiku 4.5 through AWS Bedrock. Local execution took about 2.9 to 5.3 seconds per request, while the simulated edge run took about 3.5 to 4.7 seconds; both had no reported per-request cost. Cloud execution took about 1.1 to 1.5 seconds and cost approximately $0.0001 to $0.00015 per request. In this test, cloud was consistently faster and its cost was very small, so the main practical trade-off was privacy and offline capability, which the project says favor local or edge execution. The author cautions that the edge result is not a real constrained-device test because it only limits laptop CPU threads, and plans Raspberry Pi testing and broader model and task coverage. Reproducing the benchmark requires Ollama, the local model, Python dependencies, and AWS Bedrock access.