Reference-Assisted Learner-Aligned Distributed Stacking for Edge Intrusion Detection in the Industrial Internet of Things
DOI:
https://doi.org/10.4108/eetsis.14107Keywords:
Industrial Internet of Things, edge intrusion detection, distributed stacking, non-IID data, reference data, model fusionAbstract
Networked manufacturing couples sensors, programmable logic controllers, and industrial gateways to production that cannot simply be paused. An intrusion detector in this setting must control false alarms and edge-resource use as well as detect attacks. Industrial Internet of Things (IIoT) edge nodes add three practical constraints: training traffic cannot be centralized, local statistical distributions differ, and the deployed models may be structurally heterogeneous. We address this setting with Reference-Assisted Learner-Aligned Distributed Stacking (RA-LADS). It evaluates node models on mutually exclusive reference pools, groups LightGBM and XGBoost responses by learner family, and represents each family by its mean, standard deviation, and five quantiles in a fixed, permutation-invariant 7M-dimensional vector. Across 20 matched partitioning and training seeds, the full 14-dimensional summary attained a mean F1 of 0.984514. Its gain over a global 7-dimensional summary without family semantics was 0.000756 (95% CI: 0.000559–0.000962; Holm-adjusted p = 1.08 × 10−5). By comparison, the difference from node-wise raw prediction concatenation was −0.000009, with an interval spanning zero. Random balanced grouping was not significantly different from true learner-family grouping after multiplicity correction; family identity is therefore a useful grouping prior, but not the only one. Mean F1 scores for a compact reference-pool LightGBM, a 41-parameter Tiny Deep Sets model, and a heterogeneous one-model-per-node summary were 0.983887, 0.983659, and 0.983865. The full method delivered small yet reproducible paired gains over all three controls. At fixed false-alarm-rate (FAR) budgets, no stable advantage over a single LightGBM appeared between 0.1% and 2% FAR; at 5% FAR, F1 increased by 0.000417 (Holm-adjusted p = 0.017). Five-seed retraining on two external NetFlow datasets did not show a general advantage. The claim is therefore limited to settings in which node responses carry exploitable distributional structure. Within that boundary, RA-LADS offers a lightweight interface for heterogeneous tree collaboration when labeled reference data and explicit alert operating points are available, although system cost still grows linearly.
References
[1] Sarhan M, Layeghy S, Moustafa N, Portmann M. NetFlow Datasets for Machine Learning-Based Network Intrusion Detection Systems. In: Big Data Technologies and Applications. Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering; 2021. p. 117-35.
[2] Sarhan M, Layeghy S, Moustafa N, Gallagher M, Portmann M. Feature Extraction for Machine Learning-Based Intrusion Detection in IoT Networks. Digital Communications and Networks. 2024;10(1):205-16.
[3] Mirsky Y, Doitshman T, Elovici Y, Shabtai A. Kitsune: An Ensemble of Autoencoders for Online Network Intrusion Detection. In: Network and Distributed System Security Symposium; 2018.
[4] Barradas D, Santos N, Rodrigues L, Signorello S, Ramos FMV, Madeira A. FlowLens: Enabling Efficient Flow Classification for ML-Based Network Security Applications. In: Network and Distributed System Security Symposium; 2021.
[5] Xu H, Wu D, Lu Y, Lu J, Zeng H. Models on the Move: Towards Feasible Embedded AI for Intrusion Detection on Vehicular CAN Bus. In: 2024 USENIX Annual Technical Conference; 2024. p. 1049-63.
[6] McMahan HB, Moore E, Ramage D, Hampson S, Ag¨uera y Arcas B. Communication-Efficient Learning of Deep Networks from Decentralized Data. In: Proceedings of the 20th International Conference on Artificial Intelligence and Statistics. vol. 54 of Proceedings of Machine Learning Research; 2017. p. 1273-82.
[7] Li T, Sahu AK, Zaheer M, Sanjabi M, Talwalkar A, Smith V. Federated Optimization in Heterogeneous Networks. In: Proceedings of Machine Learning and Systems. vol. 2; 2020. p. 429-50.
[8] Karimireddy SP, Kale S, Mohri M, Reddi S, Stich S, Suresh AT. SCAFFOLD: Stochastic Controlled Averaging for Federated Learning. In: Proceedings of the 37th International Conferenceon Machine Learning. vol. 119 of Proceedings of Machine Learning Research; 2020. p. 5132-43.
[9] Wang J, Liu Q, Liang H, Joshi G, Poor HV. Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization. In: Advances in Neural Information Processing Systems 33; 2020. p. 7611-23.
[10] Kairouz P, McMahan HB, Avent B, et al. Advances and Open Problems in Federated Learning. Foundations and Trends in Machine Learning. 2021;14(1–2):1-210.
[11] Khan LU, Saad W, Han Z, Hossain E, Hong CS. Federated Learning for Internet of Things: Recent Advances, Taxonomy, and Open Challenges. IEEE Communications Surveys & Tutorials. 2021;23(3):1759-99.
[12] Lai F, Dai Y, Singapuram S, Liu J, Zhu X, Madhyastha H, et al. FedScale: Benchmarking Model and System Performance of Federated Learning at Scale. In: Proceedings of the 39th International Conference on Machine Learning. vol. 162 of Proceedings of Machine Learning Research; 2022. p. 11814-27.
[13] Wolpert DH. Stacked Generalization. Neural Networks. 1992;5(2):241-59.
[14] Breiman L. Random Forests. Machine Learning. 2001;45(1):5- 32.
[15] Chen T, Guestrin C. XGBoost: A Scalable Tree Boosting System. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; 2016. p. 785-94.
[16] Ke G, Meng Q, Finley T, Wang T, Chen W, Ma W, et al. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In: Advances in Neural Information Processing Systems 30; 2017. p. 3146-54.
[17] Zaheer M, Kottur S, Ravanbakhsh S, Poczos B, Salakhutdinov RR, Smola A. Deep Sets. In: Advances in Neural Information Processing Systems 30. vol. 30; 2017.
[18] Lee J, Lee Y, Kim J, Kosiorek A, Choi S, Teh YW. Set Transformer: A Framework for Attention-Based Permutation- Invariant Neural Networks. In: Proceedings of the 36th International Conference on Machine Learning. vol. 97 of Proceedings of Machine Learning Research; 2019. p. 3744-53.
[19] Guo C, Pleiss G, Sun Y, Weinberger KQ. On Calibration of Modern Neural Networks. In: Proceedings of the 34th International Conference on Machine Learning. vol. 70 of Proceedings of Machine Learning Research; 2017. p. 1321-30.
[20] Li Q, He B, Song D. Model-Contrastive Federated Learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2021. p. 10713-22.
[21] Li T, Hu S, Beirami A, Smith V. Ditto: Fair and Robust Federated Learning Through Personalization. In: Proceedings of the 38th International Conference on Machine Learning. vol. 139 of Proceedings of Machine Learning Research; 2021. p. 6357-68.
[22] Li X, Jiang M, Zhang X, Kamp M, Dou Q. FedBN: Federated Learning on Non-IID Features via Local Batch Normalization. In: International Conference on Learning Representations; 2021.
[23] Zhang J, Li Z, Li B, Xu J, Wu S, Ding S, et al. Federated Learning with Label Distribution Skew via Logits Calibration. In: Proceedings of the 39th International Conference on Machine Learning. vol. 162 of Proceedings of Machine Learning Research; 2022. p. 26311-29.
[24] Diao E, Ding J, Tarokh V. HeteroFL: Computation and Communication Efficient Federated Learning for Heterogeneous Clients. In: International Conference on Learning Representations; 2021.
[25] Tan Y, Long G, Liu L, Zhou T, Lu Q, Jiang J, et al. FedProto: Federated Prototype Learning across Heterogeneous Clients. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 36; 2022. p. 8432-40.
[26] Cao X, Fang M, Liu J, Gong NZ. FLTrust: Byzantine-Robust Federated Learning via Trust Bootstrapping. In: Network and Distributed System Security Symposium; 2021.
[27] Lin T, Kong L, Stich SU, Jaggi M. Ensemble Distillation for Robust Model Fusion in Federated Learning. In: Advances in Neural Information Processing Systems 33. vol. 33; 2020. p. 2351-63.
[28] Zhu Z, Hong J, Zhou J. Data-Free Knowledge Distillation for Heterogeneous Federated Learning. In: Proceedings of the 38th International Conference on Machine Learning. vol. 139 of Proceedings of Machine Learning Research; 2021. p. 12878-89.
[29] Shen J, Yang W, Chu Z, Fan J, Niyato D, Lam KY. Effective Intrusion Detection in Heterogeneous Internet-of- Things Networks via Ensemble Knowledge Distillation-Based Federated Learning. In: 2024 IEEE International Conference on Communications; 2024. p. 2034-9.
[30] Li B, Wu Y, Song J, Lu R, Li T, Zhao L. DeepFed: Federated Deep Learning for Intrusion Detection in Industrial Cyber–Physical Systems. IEEE Transactions on Industrial Informatics. 2021;17(8):5615-24.
[31] Liu Y, Kumar N, Xiong Z, Lim WYB, Kang J, Niyato D. Communication-Efficient Federated Learning for Anomaly Detection in Industrial Internet of Things. In: 2020 IEEE Global Communications Conference; 2020. p. 1-6.
[32] Nguyen TD, Marchal S, Miettinen M, Fereidooni H, Asokan N, Sadeghi AR. D¨IoT: A Federated Self-Learning Anomaly Detection System for IoT. In: 2019 IEEE 39th International Conference on Distributed Computing Systems; 2019. p. 756- 67.
[33] Liu T, Dib O. Advanced Intrusion Detection for IoT Devices Using Federated Deep Learning. International Journal of Information Security. 2026;25(1):19.
[34] Shokri R, Stronati M, Song C, Shmatikov V. Membership Inference Attacks Against Machine Learning Models. In: 2017 IEEE Symposium on Security and Privacy; 2017. p. 3-18.
[35] Bagdasaryan E, Veit A, Hua Y, Estrin D, Shmatikov V. How To Backdoor Federated Learning. In: Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics. vol. 108 of Proceedings of Machine Learning Research; 2020. p. 2938-48.
[36] Demšar J. Statistical Comparisons of Classifiers over Multiple Data Sets. Journal of Machine Learning Research. 2006;7:1-30.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Jing Li, Shuhao Shen, Kangrui Xu, Lin Cui

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
This is an open access article distributed under the terms of the CC BY-NC-SA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.