papersTODAY 04:00 UTC
Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems
A new arXiv preprint proposes a routing approach that decides whether function-calling LLM requests should be handled at the edge or sent to cloud models, with the aim of reducing energy consumption and carbon emissions. The work targets agentic systems where inference is currently concentrated in large cloud-hosted models. It argues that latency, compute placement, and grid carbon intensity can be balanced when choosing where an inference runs.