Janus: Joint Prefill/Decode Disaggregation with KV-Cache-Aware Multi-Cloud Routing for Edge-Adjacent LLM Serving
DOI:
https://doi.org/10.4108/eetiot.13349Keywords:
LLM inference, disaggregated serving, multi-cloud, edge serving, KV-cache, Lyapunov optimization, scheduling, submodular optimizationAbstract
INTRODUCTION: Disaggregated large language model (LLM) serving separates the compute-bound prefill phase from the memory-bound decode phase and is increasingly deployed across heterogeneous multi-cloud and edge-adjacent fleets serving geographically distributed (including IoT and edge) clients. The key-value (KV) cache that couples the two phases raises a stateful routing problem—migrate, recompute, or partially ship the cache across inter-cloud links of varying bandwidth, on hardware of varying capability, under spot prices that change every few minutes—that, to our knowledge, no published framework fully addresses.
OBJECTIVES: To jointly optimize prefill placement, decode placement, KV-cache transport policy, and slow-timescale pool sizing across clouds with heterogeneous link bandwidths, GPU capabilities, and volatile spot prices, with explicit provable guarantees.
METHODS: We present Janus, an online scheduler that formulates per-request scheduling as a constrained graph-routing problem with stateful edges and decomposes it into a monotone-submodular prefix-aware placement subproblem and a Lyapunov drift-plus-penalty control subproblem, with four KV-transport policies including a hybrid layer-pipelined policy admitting a closed-form layer-split optimum. A 17.1K-line prototype implements the scheduling logic; evaluation uses a trace-driven, discrete-event simulator whose timing and cost models are calibrated against measured single-pod microbenchmarks, configured to model a 96-pod (512-GPU) three-cloud, six-region fleet.
RESULTS: We prove a (1 1/e) approximation for prefix reuse under continuous greedy (with a 1/2 guarantee for the deployed combinatorial greedy under slack capacity, degrading to 1/3 when heterogeneous KV capacity binds), an O(1/V ) gap to the best policy in the decomposed class with O(V ) queue bound stated with its explicit additive constants, a sample-path robustness guarantee under adversarially time-varying prices and bandwidth, and a hybrid-transport optimality theorem. In simulation, versus the strongest multi-cloud baseline we construct, Janus attains 3.8 median and 4.6 P99 time-to-first-token reduction, 2.1 goodput, a 71% reuse-capture rate, and 38% cost reduction, with graceful degradation under spot-preemption, WAN-bandwidth-collapse, and region-failure scenarios.
CONCLUSION: Janus is, to our knowledge, the first scheduler to treat the KV-transport decision as a first-class scheduling variable jointly with prefill and decode routing across heterogeneous multi-cloud fleets, with provable guarantees; physical multi-cloud deployment and hardware validation of the simulated results are explicitly left as future work.
Downloads
References
[1] Agrawal A, Kedia N, Panwar A, Mohan J, Kwatra N,Gulavani B, et al. Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2024.
[2] Kwon W, Li Z, Zhuang S, Sheng Y, Zheng L, Yu CH, et al. Efficient memory management for large language model serving with PagedAttention. In: Proceedings of the ACM Symposium on Operating Systems Principles (SOSP); 2023.
[3] Red Hat, IBM, Google, NVIDIA, CoreWeave, and contributors. llm-d: A Kubernetes-native distributed inference framework; 2025. Available from: https:// github.com/llm-d/llm-d.
[4] Zheng L, Yin L, Xie Z, Huang J, Sun C, Yu CH, et al. SGLang: Efficient execution of structured language model programs. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS); 2024.
[5] Zhong Y, Liu S, Chen J, Hu J, Zhu Y, Liu X, et al. DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2024.
[6] Patel P, Choukse E, Zhang C, Shah A, Goiri I, Maleki S, et al. Splitwise: Efficient generative LLM inference using phase splitting. In: Proceedings of the International Symposium on Computer Architecture (ISCA); 2024.
[7] Qin R, Li Z, He W, Zhang M, Wu Y, Zheng W, et al. Mooncake: A KVCache-centric disaggregated architecture for LLM serving. In: Proceedings of the USENIX Conference on File and Storage Technologies (FAST); 2025.
[8] Dao T, Fu DY, Ermon S, Rudra A, Re C. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS); 2022.
[9] Hu J, Wu X, Liu D, Yang F, Yang Z, Zhou L. Bench-marking multi-cloud LLM inference under heteroge-neous accelerators and pricing. In: Proceedings of the Conference on Machine Learning and Systems (MLSys); 2025.
[10] Pope R, Douglas S, Chowdhery A, Devlin J, Bradbury J, Heek J, et al. Efficiently scaling transformer inference. In: Proceedings of the Conference on Machine Learning and Systems (MLSys); 2023.
[11] Kwon W, Li Z, et al. vLLM: Easy, fast, and cheap LLM serving with PagedAttention; 2023. Available from: https://github.com/vllm-project/vllm.
[12] Nemhauser GL, Wolsey LA, Fisher ML. An analysis of approximations for maximizing submodular set functions—I. Mathematical Programming. 1978;14(1):265–294.
[13] Neely MJ. Stochastic Network Optimization with Application to Communication and Queueing Systems. Morgan & Claypool; 2010.
[14] Dao T. FlashAttention-2: Faster attention with better parallelism and work partitioning. In: Proceedings of the International Conference on Learning Representations (ICLR); 2024.
[15] Sun B, Huang Z, Zhao H, Xiao W, Zhang X, Li Y, et al. Llumnix: Dynamic scheduling for large language model serving. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2024.
[16] Yu GI, Jeong JS, Kim GW, Kim S, Chun BG. Orca: A distributed serving system for transformer-based generative models. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2022.
[17] Kosaian J, Rashmi KV, Venkataraman S. Parity models: Erasure-coded resilience for prediction serving systems. In: Proceedings of the ACM Symposium on Operating Systems Principles (SOSP); 2019.
[18] Stoica I, Shenker S. From cloud computing to sky computing. In: Proceedings of the Workshop on Hot Topics in Operating Systems (HotOS); 2021.
[19] Yang Z, Wu Z, Luo M, Chiang WL, Bhardwaj R, Kwon W, et al. SkyPilot: An intercloud broker for sky computing. In: Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI); 2023.
[20] Dehghan M, Jiang B, Seetharam A, He T, Salonidis T, Kurose J, et al. On the complexity of optimal request routing and content caching in heterogeneous cache networks. IEEE/ACM Transactions on Networking. 2017;25(3).
[21] Mei Y, Zhuang Y, Miao X, Yang J, Jia Z, Vinayak R. Helix: Distributed serving of large language models via max-flow on heterogeneous GPUs. In: Proceedings of the ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS); 2025.
[22] Hsieh K, Harlap A, Vijaykumar N, Konomis D, Ganger GR, Gibbons PB, et al. Gaia: Geo-distributed machine learning approaching LAN speeds. In: Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI); 2017.
[23] Ren J, Rajbhandari S, Aminabadi RY, Ruwase O, Yang S, Zhang M, et al. ZeRO-Offload: Democratizing billion-scale model training. In: Proceedings of the USENIX Annual Technical Conference (ATC); 2021.
[24] Ongaro D, Ousterhout J. In search of an understandable consensus algorithm. In: Proceedings of the USENIX Annual Technical Conference (ATC); 2014.
[25] Verma A, Pedrosa L, Korupolu M, Oppenheimer D, Tune E, Wilkes J. Large-scale cluster management at Google with Borg. In: Proceedings of the European Conference on Computer Systems (EuroSys); 2015.
[26] Grandl R, Ananthanarayanan G, Kandula S, Rao S, Akella A. Multi-resource packing for cluster schedulers. In: Proceedings of the ACM SIGCOMM Conference; 2014.
[27] Delgado P, Dinu F, Kermarrec AM, Zwaenepoel W. Hawk: Hybrid datacenter scheduling. In: Proceedings of the USENIX Annual Technical Conference (ATC); 2015.
[28] Ousterhout K, Wendell P, Zaharia M, Stoica I. Sparrow: Distributed, low-latency scheduling. In: Proceedings of the ACM Symposium on Operating Systems Principles (SOSP); 2013.
[29] Nishihara R, Moritz P, Wang S, Tumanov A, Paul W, Schleier-Smith J, et al. Real-time machine learning: The missing pieces. In: Proceedings of the Workshop on Hot Topics in Operating Systems (HotOS); 2017.
[30] Nishtala R, Fugal H, Grimm S, Kwiatkowski M, Lee H, Li HC, et al. Scaling Memcache at Facebook. In: Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI); 2013.
[31] Shoeybi M, Patwary M, Puri R, LeGresley P, Casper J, Catanzaro B. Megatron-LM: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053. 2019.
[32] Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, et al. Language models are few-shot learners. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS); 2020.
[33] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS); 2017.
[34] Touvron H, Lavril T, Izacard G, Martinet X, Lachaux MA, Lacroix T, et al. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971. 2023.
[35] Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. 2023.
[36] Jiang AQ, Sablayrolles A, Roux A, Mensch A, Savary B, Bamford C, et al. Mixtral of experts. arXiv preprint arXiv:2401.04088. 2024.
[37] Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A, Letman A, et al. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783. 2024.
[38] Leviathan Y, Kalman M, Matias Y. Fast inference from transformers via speculative decoding. In: Proceedings of the International Conference on Machine Learning (ICML); 2023.
[39] Liu Y, Li H, Cheng Y, Ray S, Huang Y, Zhang Q, et al. CacheGen: KV cache compression and streaming for fast large language model serving. In: Proceedings of the ACM SIGCOMM Conference; 2024.
[40] Gao B, He Z, Sharma P, Kang Q, Jevdjic D, Deng J, et al. Cost-efficient large language model serving for multi-turn conversations with CachedAttention. In: Proceedings of the USENIX Annual Technical Conference (ATC); 2024.
[41] Lee W, Lee J, Seo J, Sim J. InfiniGen: Efficient generative inference of large language models with dynamic KV cache management. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2024.
[42] Zhang Z, Sheng Y, Zhou T, Chen T, Zheng L, Cai R, et al. H2O: Heavy-hitter oracle for efficient generative inference of large language models. In: Proceedings of the Conference on Neural Information Processing Systems (NeurIPS); 2023.
[43] Sheng Y, Zheng L, Yuan B, Li Z, Ryabinin M, Chen B, et al. FlexGen: High-throughput generative inference of large language models with a single GPU. In: Proceedings of the International Conference on Machine Learning (ICML); 2023.
[44] Aminabadi RY, Rajbhandari S, Awan AA, Li C, Li D, Zheng E, et al. DeepSpeed-Inference: Enabling efficient inference of transformer models at unprecedented scale. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC); 2022.
[45] Lin J, Tang J, Tang H, Yang S, Chen WM, Wang WC, et al. AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration. In: Proceedings of the Conference on Machine Learning and Systems (MLSys); 2024.
[46] Frantar E, Ashkboos S, Hoefler T, Alistarh D. GPTQ: Accurate post-training quantization for generative pre-trained transformers. In: Proceedings of the Interna-tional Conference on Learning Representations (ICLR); 2023.
[47] Fu Y, Xue L, Huang Y, Brabete AO, Ustiugov D, Patel Y, et al. ServerlessLLM: Low-latency serverless inference for large language models. In: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI); 2024.
[48] Wu B, Liu S, Zhong Y, Sun P, Liu X, Jin X. LoongServe: Efficiently serving long-context large language models with elastic sequence parallelism. In: Proceedings of the ACM Symposium on Operating Systems Principles (SOSP); 2024.
[49] Stojkovic J, Zhang C, Goiri I, Torrellas J, Choukse E. DynamoLLM: Designing LLM inference clusters for per-formance and energy efficiency. In: Proceedings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA); 2025.
[50] Shapiro M, Preguica N, Baquero C, Zawirski M. Conflict-free replicated data types. In: Proceedings of the Symposium on Stabilization, Safety, and Security of Distributed Systems (SSS); 2011.
[51] Iyengar J, Thomson M. QUIC: A UDP-based multiplexed and secure transport. IETF; 2021. RFC 9000.
[52] Burns B, Grant B, Oppenheimer D, Brewer E, Wilkes J. Borg, Omega, and Kubernetes. Communications of the ACM. 2016;59(5):50–57.
[53] Calinescu G, Chekuri C, Pal M, Vondrak J. Maxi-mizing a monotone submodular function subject to a matroid constraint. SIAM Journal on Computing. 2011;40(6):1740–1766.
[54] Fisher ML, Nemhauser GL, Wolsey LA. An anal-ysis of approximations for maximizing submodular set functions—II. Mathematical Programming Study. 1978;8:73–87.
[55] Hu C, Huang H, Hu J, Xu J, Chen X, Xie T, et al. MemServe: Context caching for disaggregated LLM serving with elastic memory pool. arXiv preprint arXiv:2406.17565. 2024.
[56] Jin Y, Wang T, Lin H, Song M, Li P, Ma Y, et al. P/D-Serve: Serving disaggregated large language model at scale. arXiv preprint arXiv:2408.08147. 2024.
[57] Chen Y, Xu Y, Jin M, Liu J, Wu C. KVDirect: Dis-tributed disaggregated LLM inference. arXiv preprint arXiv:2501.14743. 2025.
[58] Cui Z, Zhao S, Liu Y, et al. FlowKV: Enhancing multi-turn conversational coherence and low-latency KV-cache scheduling in disaggregated inference. arXiv preprint arXiv:2504.03775. 2025.
[59] Liu Y, Yao J, Li H, Cheng Y, Du K, Lu S, et al. LMCache: A KV-cache connector layer for fast and reusable LLM serving. arXiv preprint arXiv:2510.09665. 2025.
[60] Qin R, He W, Wang Y, Li Z, Xu X, Wu Y, et al. Prefill-as-a-Service: KVCache of next-generation models could go cross-datacenter. arXiv preprint arXiv:2604.15039. 2026.
[61] Ainslie J, Lee-Thorp J, de Jong M, Zemlyanskiy Y, Lebron F, Sanghai S. GQA: Training generalized multi-query transformer models from multi-head checkpoints. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP); 2023.
[62] Lee J, Mirrokni VS, Nagarajan V, Sviridenko M. Non-monotone submodular maximization under matroid and knapsack constraints. In: Proceedings of the ACM Symposium on Theory of Computing (STOC); 2009. p. 323–332.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 K. B. Aruna, V. Kaliraj, I. Sudha, V. Sureka

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
This is an open-access article distributed under the terms of the Creative Commons Attribution CC BY 4.0 license, which permits unlimited use, distribution, and reproduction in any medium so long as the original work is properly cited.
