Abstract
Edge computing enables distributed inference for computation-intensive applications. However, the autoregressive nature and large model size of large language models (LLMs) pose challenges for their deployment in wireless edge networks. Existing edge inference methods mainly assume an equivalent delay for each token generation step, which fails to capture the dynamic computational and memory overhead incurred during the decoding process. This paper proposes a latency-sensitive wireless edge inference framework for LLMs, where tasks are grouped into multiple batches and processed in parallel by partitioning the LLM into multiple pipeline stages across heterogeneous edge GPUs. An accurate latency model is established, where the latency of each token generation step increases during the autoregressive generation process. Based on this model, the end-to-end inference latency is minimized, which is formulated as a joint optimization problem of bandwidth allocation, model partitioning, and batch scheduling subject to heterogeneous GPU memory constraints. To solve this NP-hard problem with coupled variables, we develop a polynomial-time alternating optimization algorithm that iteratively optimizes model partitioning and batch scheduling via dynamic programming. The closed-form solutions of wireless bandwidth allocation are derived. Extensive simulations show that our approach reduces latency by up to 42.1% versus state-of-the-art baselines across diverse edge scenarios.
| Original language | English |
|---|---|
| Pages (from-to) | 8390-8406 |
| Number of pages | 17 |
| Journal | IEEE Transactions on Communications |
| Volume | 74 |
| DOIs | |
| Publication status | Published - 2026 |
Bibliographical note
Publisher Copyright:© 1972-2012 IEEE.
Keywords
- Edge inference
- batching
- large language models (LLMs)
- pipeline parallelism
- resource optimization
Fingerprint
Dive into the research topics of 'Edge Inference for Large Language Models with Pipeline Parallelism and Batching'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver