Janus: Joint Prefill/Decode Disaggregation with KV-Cache-Aware Multi-Cloud Routing for Edge-Adjacent LLM Serving

Authors

DOI:

https://doi.org/10.4108/eetiot.13349

Keywords:

LLM inference, disaggregated serving, multi-cloud, edge serving, KV-cache, Lyapunov optimization, scheduling, submodular optimization

Abstract

INTRODUCTION: Disaggregated large language model (LLM) serving separates the compute-bound prefill phase from the memory-bound decode phase and is increasingly deployed across heterogeneous multi-cloud and edge-adjacent fleets serving geographically distributed (including IoT and edge) clients. The key-value (KV) cache that couples the two phases raises a stateful routing problem—migrate, recompute, or partially ship the cache across inter-cloud links of varying bandwidth, on hardware of varying capability, under spot prices that change every few minutes—that, to our knowledge, no published framework fully addresses.

OBJECTIVES: To jointly optimize prefill placement, decode placement, KV-cache transport policy, and slow-timescale pool sizing across clouds with heterogeneous link bandwidths, GPU capabilities, and volatile spot prices, with explicit provable guarantees.

METHODS: We present Janus, an online scheduler that formulates per-request scheduling as a constrained graph-routing problem with stateful edges and decomposes it into a monotone-submodular prefix-aware placement subproblem and a Lyapunov drift-plus-penalty control subproblem, with four KV-transport policies including a hybrid layer-pipelined policy admitting a closed-form layer-split optimum. A 17.1K-line prototype implements the scheduling logic; evaluation uses a trace-driven, discrete-event simulator whose timing and cost models are calibrated against measured single-pod microbenchmarks, configured to model a 96-pod (512-GPU) three-cloud, six-region fleet.

RESULTS: We prove a (1 1/e) approximation for prefix reuse under continuous greedy (with a 1/2 guarantee for the deployed combinatorial greedy under slack capacity, degrading to 1/3 when heterogeneous KV capacity binds), an O(1/V ) gap to the best policy in the decomposed class with O(V ) queue bound stated with its explicit additive constants, a sample-path robustness guarantee under adversarially time-varying prices and bandwidth, and a hybrid-transport optimality theorem. In simulation, versus the strongest multi-cloud baseline we construct, Janus attains 3.8 median and 4.6 P99 time-to-first-token reduction, 2.1 goodput, a 71% reuse-capture rate, and 38% cost reduction, with graceful degradation under spot-preemption, WAN-bandwidth-collapse, and region-failure scenarios.

CONCLUSION: Janus is, to our knowledge, the first scheduler to treat the KV-transport decision as a first-class scheduling variable jointly with prefill and decode routing across heterogeneous multi-cloud fleets, with provable guarantees; physical multi-cloud deployment and hardware validation of the simulated results are explicitly left as future work.

Downloads

Download data is not yet available.

References

[1] Agrawal A, Kedia N, Panwar A, Mohan J, Kwatra N,Gulavani B, et al. Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2024.

[2] Kwon W, Li Z, Zhuang S, Sheng Y, Zheng L, Yu CH, et al. Efficient memory management for large language model serving with PagedAttention. In: Proceedings of the ACM Symposium on Operating Systems Principles (SOSP); 2023.

[3] Red Hat, IBM, Google, NVIDIA, CoreWeave, and contributors. llm-d: A Kubernetes-native distributed inference framework; 2025. Available from: https:// github.com/llm-d/llm-d.

[4] Zheng L, Yin L, Xie Z, Huang J, Sun C, Yu CH, et al. SGLang: Efficient execution of structured language model programs. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS); 2024.

[5] Zhong Y, Liu S, Chen J, Hu J, Zhu Y, Liu X, et al. DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2024.

[6] Patel P, Choukse E, Zhang C, Shah A, Goiri I, Maleki S, et al. Splitwise: Efficient generative LLM inference using phase splitting. In: Proceedings of the International Symposium on Computer Architecture (ISCA); 2024.

[7] Qin R, Li Z, He W, Zhang M, Wu Y, Zheng W, et al. Mooncake: A KVCache-centric disaggregated architecture for LLM serving. In: Proceedings of the USENIX Conference on File and Storage Technologies (FAST); 2025.

[8] Dao T, Fu DY, Ermon S, Rudra A, Re C. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS); 2022.

[9] Hu J, Wu X, Liu D, Yang F, Yang Z, Zhou L. Bench-marking multi-cloud LLM inference under heteroge-neous accelerators and pricing. In: Proceedings of the Conference on Machine Learning and Systems (MLSys); 2025.

[10] Pope R, Douglas S, Chowdhery A, Devlin J, Bradbury J, Heek J, et al. Efficiently scaling transformer inference. In: Proceedings of the Conference on Machine Learning and Systems (MLSys); 2023.

[11] Kwon W, Li Z, et al. vLLM: Easy, fast, and cheap LLM serving with PagedAttention; 2023. Available from: https://github.com/vllm-project/vllm.

[12] Nemhauser GL, Wolsey LA, Fisher ML. An analysis of approximations for maximizing submodular set functions—I. Mathematical Programming. 1978;14(1):265–294.

[13] Neely MJ. Stochastic Network Optimization with Application to Communication and Queueing Systems. Morgan & Claypool; 2010.

[14] Dao T. FlashAttention-2: Faster attention with better parallelism and work partitioning. In: Proceedings of the International Conference on Learning Representations (ICLR); 2024.

[15] Sun B, Huang Z, Zhao H, Xiao W, Zhang X, Li Y, et al. Llumnix: Dynamic scheduling for large language model serving. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2024.

[16] Yu GI, Jeong JS, Kim GW, Kim S, Chun BG. Orca: A distributed serving system for transformer-based generative models. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2022.

[17] Kosaian J, Rashmi KV, Venkataraman S. Parity models: Erasure-coded resilience for prediction serving systems. In: Proceedings of the ACM Symposium on Operating Systems Principles (SOSP); 2019.

[18] Stoica I, Shenker S. From cloud computing to sky computing. In: Proceedings of the Workshop on Hot Topics in Operating Systems (HotOS); 2021.

[19] Yang Z, Wu Z, Luo M, Chiang WL, Bhardwaj R, Kwon W, et al. SkyPilot: An intercloud broker for sky computing. In: Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI); 2023.

[20] Dehghan M, Jiang B, Seetharam A, He T, Salonidis T, Kurose J, et al. On the complexity of optimal request routing and content caching in heterogeneous cache networks. IEEE/ACM Transactions on Networking. 2017;25(3).

[21] Mei Y, Zhuang Y, Miao X, Yang J, Jia Z, Vinayak R. Helix: Distributed serving of large language models via max-flow on heterogeneous GPUs. In: Proceedings of the ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS); 2025.

[22] Hsieh K, Harlap A, Vijaykumar N, Konomis D, Ganger GR, Gibbons PB, et al. Gaia: Geo-distributed machine learning approaching LAN speeds. In: Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI); 2017.

[23] Ren J, Rajbhandari S, Aminabadi RY, Ruwase O, Yang S, Zhang M, et al. ZeRO-Offload: Democratizing billion-scale model training. In: Proceedings of the USENIX Annual Technical Conference (ATC); 2021.

[24] Ongaro D, Ousterhout J. In search of an understandable consensus algorithm. In: Proceedings of the USENIX Annual Technical Conference (ATC); 2014.

[25] Verma A, Pedrosa L, Korupolu M, Oppenheimer D, Tune E, Wilkes J. Large-scale cluster management at Google with Borg. In: Proceedings of the European Conference on Computer Systems (EuroSys); 2015.

[26] Grandl R, Ananthanarayanan G, Kandula S, Rao S, Akella A. Multi-resource packing for cluster schedulers. In: Proceedings of the ACM SIGCOMM Conference; 2014.

[27] Delgado P, Dinu F, Kermarrec AM, Zwaenepoel W. Hawk: Hybrid datacenter scheduling. In: Proceedings of the USENIX Annual Technical Conference (ATC); 2015.

[28] Ousterhout K, Wendell P, Zaharia M, Stoica I. Sparrow: Distributed, low-latency scheduling. In: Proceedings of the ACM Symposium on Operating Systems Principles (SOSP); 2013.

[29] Nishihara R, Moritz P, Wang S, Tumanov A, Paul W, Schleier-Smith J, et al. Real-time machine learning: The missing pieces. In: Proceedings of the Workshop on Hot Topics in Operating Systems (HotOS); 2017.

[30] Nishtala R, Fugal H, Grimm S, Kwiatkowski M, Lee H, Li HC, et al. Scaling Memcache at Facebook. In: Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI); 2013.

[31] Shoeybi M, Patwary M, Puri R, LeGresley P, Casper J, Catanzaro B. Megatron-LM: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053. 2019.

[32] Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, et al. Language models are few-shot learners. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS); 2020.

[33] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS); 2017.

[34] Touvron H, Lavril T, Izacard G, Martinet X, Lachaux MA, Lacroix T, et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971. 2023.

[35] Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. 2023.

[36] Jiang AQ, Sablayrolles A, Roux A, Mensch A, Savary B, Bamford C, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088. 2024.

[37] Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A, Letman A, et al. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783. 2024.

[38] Leviathan Y, Kalman M, Matias Y. Fast inference from transformers via speculative decoding. In: Proceedings of the International Conference on Machine Learning (ICML); 2023.

[39] Liu Y, Li H, Cheng Y, Ray S, Huang Y, Zhang Q, et al. CacheGen: KV cache compression and streaming for fast large language model serving. In: Proceedings of the ACM SIGCOMM Conference; 2024.

[40] Gao B, He Z, Sharma P, Kang Q, Jevdjic D, Deng J, et al. Cost-efficient large language model serving for multi-turn conversations with CachedAttention. In: Proceedings of the USENIX Annual Technical Conference (ATC); 2024.

[41] Lee W, Lee J, Seo J, Sim J. InfiniGen: Efficient generative inference of large language models with dynamic KV cache management. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2024.

[42] Zhang Z, Sheng Y, Zhou T, Chen T, Zheng L, Cai R, et al. H2O: Heavy-hitter oracle for efficient generative inference of large language models. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS); 2023.

[43] Sheng Y, Zheng L, Yuan B, Li Z, Ryabinin M, Chen B, et al. FlexGen: High-throughput generative inference of large language models with a single GPU. In: Proceedings of the International Conference on Machine Learning (ICML); 2023.

[44] Aminabadi RY, Rajbhandari S, Awan AA, Li C, Li D, Zheng E, et al. DeepSpeed-Inference: Enabling efficient inference of transformer models at unprecedented scale. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC); 2022.

[45] Lin J, Tang J, Tang H, Yang S, Chen WM, Wang WC, et al. AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. In: Proceedings of the Conference on Machine Learning and Systems (MLSys); 2024.

[46] Frantar E, Ashkboos S, Hoefler T, Alistarh D. GPTQ: Accurate post-training quantization for generative pre-trained transformers. In: Proceedings of the Interna-tional Conference on Learning Representations (ICLR); 2023.

[47] Fu Y, Xue L, Huang Y, Brabete AO, Ustiugov D, Patel Y, et al. ServerlessLLM: Low-latency serverless inference for large language models. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2024.

[48] Wu B, Liu S, Zhong Y, Sun P, Liu X, Jin X. LoongServe: Efficiently serving long-context large language models with elastic sequence parallelism. In: Proceedings of the ACM Symposium on Operating Systems Principles (SOSP); 2024.

[49] Stojkovic J, Zhang C, Goiri I, Torrellas J, Choukse E. DynamoLLM: Designing LLM inference clusters for per-formance and energy efficiency. In: Proceedings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA); 2025.

[50] Shapiro M, Preguica N, Baquero C, Zawirski M. Conflict-free replicated data types. In: Proceedings of the Symposium on Stabilization, Safety, and Security of Distributed Systems (SSS); 2011.

[51] Iyengar J, Thomson M. QUIC: A UDP-based multiplexed and secure transport. IETF; 2021. RFC 9000.

[52] Burns B, Grant B, Oppenheimer D, Brewer E, Wilkes J. Borg, Omega, and Kubernetes. Communications of the ACM. 2016;59(5):50–57.

[53] Calinescu G, Chekuri C, Pal M, Vondrak J. Maxi-mizing a monotone submodular function subject to a matroid constraint. SIAM Journal on Computing. 2011;40(6):1740–1766.

[54] Fisher ML, Nemhauser GL, Wolsey LA. An anal-ysis of approximations for maximizing submodular set functions—II. Mathematical Programming Study. 1978;8:73–87.

[55] Hu C, Huang H, Hu J, Xu J, Chen X, Xie T, et al. MemServe: Context caching for disaggregated LLM serving with elastic memory pool. arXiv preprint arXiv:2406.17565. 2024.

[56] Jin Y, Wang T, Lin H, Song M, Li P, Ma Y, et al. P/D-Serve: Serving disaggregated large language model at scale. arXiv preprint arXiv:2408.08147. 2024.

[57] Chen Y, Xu Y, Jin M, Liu J, Wu C. KVDirect: Dis-tributed disaggregated LLM inference. arXiv preprint arXiv:2501.14743. 2025.

[58] Cui Z, Zhao S, Liu Y, et al. FlowKV: Enhancing multi-turn conversational coherence and low-latency KV-cache scheduling in disaggregated inference. arXiv preprint arXiv:2504.03775. 2025.

[59] Liu Y, Yao J, Li H, Cheng Y, Du K, Lu S, et al. LMCache: A KV-cache connector layer for fast and reusable LLM serving. arXiv preprint arXiv:2510.09665. 2025.

[60] Qin R, He W, Wang Y, Li Z, Xu X, Wu Y, et al. Prefill-as-a-Service: KVCache of next-generation models could go cross-datacenter. arXiv preprint arXiv:2604.15039. 2026.

[61] Ainslie J, Lee-Thorp J, de Jong M, Zemlyanskiy Y, Lebron F, Sanghai S. GQA: Training generalized multi-query transformer models from multi-head checkpoints. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP); 2023.

[62] Lee J, Mirrokni VS, Nagarajan V, Sviridenko M. Non-monotone submodular maximization under matroid and knapsack constraints. In: Proceedings of the ACM Symposium on Theory of Computing (STOC); 2009. p. 323–332.

Downloads

Published

12-08-2026

How to Cite

1.
Aruna KB, Kaliraj V, Sudha I, Sureka V. Janus: Joint Prefill/Decode Disaggregation with KV-Cache-Aware Multi-Cloud Routing for Edge-Adjacent LLM Serving. EAI Endorsed Trans IoT [Internet]. 2026 Aug. 12 [cited 2026 Aug. 12];11. Available from: https://publications.eai.eu/index.php/IoT/article/view/13349