papersTODAY 04:00 UTC
OpWeave: Operator-Level Disaggregation for Heterogeneous LLM Serving
A new arXiv paper introduces OpWeave, a system that breaks LLM inference into finer-grained operators rather than coarse stages, extending recent work that separates attention from FFN or MoE execution during decoding. The authors argue this operator-level disaggregation improves how workloads are matched to heterogeneous hardware during serving.