A Communication-Reduced Privacy-Preserving ViT Inference Framework for Distributed Edge Intelligence

Authors

DOI:

https://doi.org/10.4108/eetsis.13362

Keywords:

privacy-preserving inference, secret sharing, Vision Transformer, fixed-point arithmetic, distributed edge intelligence, communication reduction

Abstract

To address privacy leakage from user images, intermediate representations, and outputs during Vision Transformer (ViT) inference in distributed edge services, this paper presents SViT, a two-server secret-sharing framework with offline correlated randomness. The revised design specifies fixed-point arithmetic over the ring Z_(2^64) with 16 fractional bits, fresh one-time masks, probabilistic truncation, numerical ranges, and complete input-output procedures for SExp, SDiv, SSqrt, SVar, SLayerNorm, SSoftmax, and SGeLU. The protocols use fixed-depth range reduction, lookup-assisted initialization, and a constant number of Newton updates, so their online depth is independent of numerical convergence tolerances. Under the semi-honest, non-colluding-server model, the revised security analysis defines approximate ideal functionalities, public leakage, simulator inputs, and sequential composition. Existing microbenchmarks show 2.28-6.50 times lower runtime and 4.00-14.20 times lower online communication than CrypTen for the reported core operators; for SDiv, runtime decreases from 9.1 ms to 1.4 ms and communication from 10.8337 MB to 0.7629 MB. Synthetic Q16 numerical checks report maximum absolute errors of 5.10×10−5 for SExp and 4.73×10−4 for SGeLU, maximum relative errors of 1.27×10−5 for reciprocal and 3.72×10−5 for square root, and 100% top-1 consistency over 10,000 randomly generated Softmax vectors. The reported end-to-end result remains 3799.744 ms and 2.85 GB per inference; therefore, SViT is described as communication-reduced relative to the evaluated MPC baselines rather than universally lightweight, and its present practical scope is primarily high-bandwidth LAN or provider-edge deployments.

References

[1] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C]//Advances in Neural Information Processing Systems. 2017, 30: 5998-6008.

[2] Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale[C]//International Conference on Learning Representations. 2021.

[3] Devlin J, Chang M W, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics. 2019: 4171-4186.

[4] Zhou C, Li Q, Li C, et al. A comprehensive survey on pretrained foundation models: A history from BERT to ChatGPT[J]. International Journal of Machine Learning and Cybernetics, 2025, 16(12): 9851-9915. doi:10.1007/s13042-024-02443-6.

[5] Shi W, Cao J, Zhang Q, Li Y, Xu L. Edge computing: Vision and challenges[J]. IEEE Internet of Things Journal, 2016, 3(5): 637-646.

[6] Zhou Z, Chen X, Li E, Zeng L, Luo K, Zhang J. Edge intelligence: Paving the last mile of artificial intelligence with edge computing[J]. Proceedings of the IEEE, 2019, 107(8): 1738-1762.

[7] Teerapittayanon S, McDanel B, Kung H T. Distributed deep neural networks over the cloud, the edge and end devices[C]//IEEE International Conference on Distributed Computing Systems. 2017: 328-339.

[8] Mohassel P, Zhang Y. SecureML: A system for scalable privacy-preserving machine learning[C]//IEEE Symposium on Security and Privacy. 2017: 19-38.

[9] Liu J, Juuti M, Lu Y, Asokan N. Oblivious neural network predictions via MiniONN transformations[C]//ACM SIGSAC Conference on Computer and Communications Security. 2017: 619-631.

[10] Rouhani B D, Riazi M S, Koushanfar F. DeepSecure: Scalable provably-secure deep learning[C]//Proceedings of the 55th Annual Design Automation Conference. New York: ACM, 2018: 2:1-2:6. doi:10.1145/3195970.3196023.

[11] Juvekar C, Vaikuntanathan V, Chandrakasan A. GAZELLE: A low latency framework for secure neural network inference[C]//USENIX Security Symposium. 2018: 1651-1669.

[12] Riazi M S, Samragh M, Chen H, et al. XONN: XNOR-based oblivious deep neural network inference[C]//USENIX Security Symposium. 2019: 1501-1518.

[13] Wagh S, Gupta D, Chandran N. SecureNN: 3-party secure computation for neural network training[J]. Proceedings on Privacy Enhancing Technologies, 2019(3): 26-49.

[14] Wagh S, Tople S, Benhamouda F, et al. FALCON: Honest-majority maliciously secure framework for private deep learning[J]. Proceedings on Privacy Enhancing Technologies, 2021(1): 188-208.

[15] Mishra P, Lehmkuhl R, Srinivasan A, Zheng W, Popa R A. Delphi: A cryptographic inference service for neural networks[C]//USENIX Security Symposium. 2020: 2505-2522.

[16] Kumar N, Rathee M, Chandran N, Gupta D, Rastogi A, Sharma R. CrypTFlow: Secure TensorFlow inference[C]//IEEE Symposium on Security and Privacy. 2020: 336-353.

[17] Rathee D, Rathee M, Kumar N, et al. CrypTFlow2: Practical 2-party secure inference[C]//ACM SIGSAC Conference on Computer and Communications Security. 2020: 325-342.

[18] Rathee D, Rathee M, Goli R K K, Gupta D, Sharma R, Chandran N, Rastogi A. SiRnn: A math library for secure RNN inference[C]//2021 IEEE Symposium on Security and Privacy. IEEE, 2021: 1003-1020. doi:10.1109/SP40001.2021.00086.

[19] Knott B, Venkataraman S, Hannun A, et al. CrypTen: Secure multi-party computation meets machine learning[C]//Advances in Neural Information Processing Systems. 2021, 34: 4961-4973.

[20] Demmler D, Schneider T, Zohner M. ABY - A framework for efficient mixed-protocol secure two-party computation[C]//Network and Distributed System Security Symposium. 2015.

[21] Patra A, Schneider T, Suresh A, Yalame H. ABY2.0: Improved mixed-protocol secure two-party computation[C]//USENIX Security Symposium. 2021: 2165-2182.

[22] Keller M. MP-SPDZ: A versatile framework for multi-party computation[C]//ACM SIGSAC Conference on Computer and Communications Security. 2020: 1575-1590.

[23] Li D, Shao R, Wang H, Guo H, Xing E P, Zhang H. MPCFormer: Fast, performant and private Transformer inference with MPC[C]//International Conference on Learning Representations. 2023.

[24] Chen T, Bao H, Huang S, et al. THE-X: Privacy-preserving Transformer inference with homomorphic encryption[C]//Findings of the Association for Computational Linguistics: ACL 2022. 2022: 3510-3520.

[25] Hao M, Li H, Chen H, Xing P, Xu G, Zhang T. Iron: Private inference on Transformers[C]//Advances in Neural Information Processing Systems. 2022, 35: 15718-15731.

[26] Luo J, Zhang Y, Zhang Z, et al. SecFormer: Fast and accurate privacy-preserving inference for Transformer models via SMPC[C]//Findings of the Association for Computational Linguistics: ACL 2024. 2024: 13333-13348. doi:10.18653/v1/2024.findings-acl.790.

[27] Dong Y, Lu W J, Zheng Y, et al. PUMA: Secure inference of LLaMA-7B in five minutes[J]. Security and Safety, 2025, 4: 2025014. doi:10.1051/sands/2025014.

[28] Pang Q, Zhu J, Möllering H, Zheng W, Schneider T. BOLT: Privacy-preserving, accurate and efficient inference for Transformers[C]//2024 IEEE Symposium on Security and Privacy. San Francisco, CA, USA: IEEE, 2024: 4753-4771. doi:10.1109/SP54263.2024.00130.

[29] Diaa A, Fenaux L, Humphries T, et al. Fast and private inference of deep neural networks by co-designing activation functions[C]//33rd USENIX Security Symposium. 2024: 2191-2208.

[30] Yuan B, Yang S, Zhang Y, Ding N, Gu D, Sun S F. MD-ML: Super fast privacy-preserving machine learning for malicious security with a dishonest majority[C]//33rd USENIX Security Symposium. 2024: 2227-2244.

[31] Lu W J, Huang Z, Gu Z, et al. BumbleBee: Secure two-party inference framework for large Transformers[C]//Network and Distributed System Security Symposium. 2025.

[32] Xu T, Lu W J, Yu J, et al. Breaking the layer barrier: Remodeling private Transformer inference with hybrid CKKS and MPC[C]//34th USENIX Security Symposium. 2025: 2653-2672.

[33] Pang Z, Feng B, Luo M, Liu C, Xu S, Zhao K, Wu Y, Zeng B. Privacy-friendly adaptation of Vision Transformers for communication and latency-efficient private inference[C]//Proceedings of the ACM Web Conference 2026. New York: ACM, 2026: 2695-2706. doi:10.1145/3774904.3792211.

[34] Zhang J, Yang X P, He L, Chen K, Lu W J, Wang Y, Hou X, Liu J, Ren K, Yang X H. Secure Transformer inference made non-interactive[C]//Network and Distributed System Security Symposium. San Diego, CA, USA: Internet Society, 2025.

[35] Chenthara S, Ahmed K, Whittaker F. Privacy-Preserving Data Sharing using Multi-layer Access Control Model in Electronic Health Environment[J]. EAI Endorsed Transactions on Scalable Information Systems, 2019, 6(22): e6. doi:10.4108/eai.13-7-2018.159356.

[36] Zeng Y, Kang Z, Shi Z. Secure Data Processing Technology of Distribution Network OPGW Line with Edge Computing[J]. EAI Endorsed Transactions on Scalable Information Systems, 2023, 10(3): e7. doi:10.4108/eetsis.v10i3.2837.

[37] Beaver D. Efficient multiparty protocols using circuit randomization[C]//Advances in Cryptology-CRYPTO '91. Lecture Notes in Computer Science, vol. 576. Springer, 1992: 420-432. doi:10.1007/3-540-46766-1_34.

[38] Catrina O, Saxena A. Secure computation with fixed-point numbers[C]//Financial Cryptography and Data Security-FC 2010. Lecture Notes in Computer Science, vol. 6052. Springer, 2010: 35-50. doi:10.1007/978-3-642-14577-3_6.

[39] Canetti R. Security and composition of multiparty cryptographic protocols[J]. Journal of Cryptology, 2000, 13(1): 143-202. doi:10.1007/s001459910006.

Downloads

Published

20-08-2026

Issue

Section

Data Security and Privacy Protection in New Distributed Networks and System

How to Cite

1.
Chen T. A Communication-Reduced Privacy-Preserving ViT Inference Framework for Distributed Edge Intelligence. EAI Endorsed Scal Inf Syst [Internet]. 2026 Aug. 20 [cited 2026 Aug. 21];13(2). Available from: https://publications.eai.eu/index.php/sis/article/view/13362