Back to News
RSS feedarxiv.org

Carbon-Aware Routing Cuts Edge-Cloud LLM Emissions Fourfold

Summary

A new carbon-aware routing framework targets the energy and emissions costs of deploying function-calling large language models primarily in the cloud. It distributes queries across a three-tier edge-cloud architecture that combines models running on heterogeneous edge and cloud hardware. A lightweight k-nearest-neighbor predictor estimates, for each query and edge tier, expected accuracy, delay, and power consumption in a shared semantic-lexical embedding space. The system combines those estimates with real-time electricity-grid carbon intensity and routes each query to the lowest-emission tier judged capable of completing it successfully. Evaluations on function-calling benchmarks and multiple LLM families found that the framework matched cloud-level accuracy while reducing operational carbon emissions by an average of four times.