papersSEP 10 04:00 UTC
PELM paper proposes speculative decoding and DVFS for power-efficient on-device LLM inference
A new arXiv paper introduces PELM, a system aimed at making large language model inference more power-efficient on mobile devices. The approach combines speculative decoding with dynamic voltage and frequency scaling to reduce the energy demands of running LLMs at the edge, where privacy, personalization, and lower latency are key motivations.