A Scalable Machine Learning Generalization Analysis Model for Intelligent Information Systems Incorporating PAC-Bayes Bounds
DOI:
https://doi.org/10.4108/eetsis.14496Keywords:
PAC-Bayes bound, machine learning generalization ability, intelligent information system, knowledge management, data drift, knowledge update, model reliability, scalable learningAbstract
This study proposes SGAM-PB, a scalable machine-learning generalization analysis model that incorporates PAC-Bayes bounds. The experimental results show that the SGAM-PB achieves the best overall performance across five intelligent information system tasks. Compared with ERM, the generalization gap is reduced from 0.118 to 0.036, a decrease of approximately 69.5%. The PAC-Bayes bound is reduced to approximately 0.15 after 100 training epochs, showing a smoother and tighter convergence trend than the baseline methods. Under 50% drift intensity, the SGAM-PB maintains an accuracy of approximately 0.77, outperforming the ERM by nearly 0.10. In dynamic knowledge update experiments, the recovered accuracy reached approximately 0.86 after 30 update steps, which was higher than that of the ERM and traditional PAC-Bayes baseline. In scalability testing, SGAM-PB maintains the runtime cost below that of the deep ensemble and BNN methods when the data volume increases to 1000 K samples. These results demonstrate that the SGAM-PB can effectively reduce deployment risk, improve drift robustness, accelerate update recovery, and provide interpretable generalization monitoring for intelligent information systems.
References
[1] Viallard, P., Germain, P., Habrard, A., & Morvant, E. A general framework for the practical disintegration of PAC-Bayesian bounds. Machine Learning. 2024; 113(2): 519–604. https://doi.org/10.1007/s10994-023-06391-0.
[2] Biggs, F., & Guedj, B. Differentiable PAC-Bayes objectives with partially aggregated neural networks. Entropy. 2021; 23(10): 1280. https://doi.org/10.3390/e23101280.
[3] Mai, T. T. Misclassification bounds for PAC-Bayesian sparse deep learning. Machine Learning. 2025; 114:18. https://doi.org/10.1007/s10994-024-06690-0.
[4] Sun, S., Yu, M., Shawe-Taylor, J., & Mao, L. Stability-based PAC-Bayes analysis for multi-view learning algorithms. Information Fusion. 2022; 86–87, 76–92. https://doi.org/10.1016/j.inffus.2022.06.006.
[5] Flynn, H., Reeb, D., Kandemir, M., & Peters, J. PAC-Bayesian lifelong learning for multi-armed bandits. Data Mining and Knowledge Discovery. 2022; 36(2): 841–876. https://doi.org/10.1007/s10618-022-00825-4.
[6] Dong, S., Wang, Q., Sahri, S., Palpanas, T., & Srivastava, D. Efficiently mitigating the impact of data drift on machine learning pipelines. Proceedings of the VLDB Endowment. 2024; 17(11): 3072–3081. https://doi.org/10.14778/3681954.3681984.
[7] Bayram, F., Ahmed, B. S., & Kassler, A. From concept drift to model degradation: An overview on performance-aware drift detectors. Knowledge-Based Systems. 2022; 245: 108632. doi: 10.1016/j.knosys.2022.108632.
[8] Karimian, M., & Beigy, H. Concept drift handling: A domain adaptation perspective. Expert Systems with Applications. 2023; 224: 119946. doi: 10.1016/j.eswa.2023.119946.
[9] Aguiar, G. J., & Cano, A. A comprehensive analysis of concept drift locality in data streams. Knowledge-Based Systems. 2024; 289: 111535. doi: 10.1016/j.knosys.2024.111535.
[10] Hinder, F., Vaquet, V., & Hammer, B. One or two things we know about concept drift—a survey on monitoring in evolving environments. Part B: Locating and explaining concept drift. Frontiers in Artificial Intelligence. 2024; 7: 1330258. doi: 10.3389/frai.2024.1330258.
[11] Chen, Y., & Dai, H.-L. Concept drift adaptation with continuous kernel learning. Information Sciences. 2024; 670: 120649. doi: 10.1016/j.ins.2024.120649.
[12] Suárez-Cetrulo, A. L., Quintana, D., & Cervantes, A. A survey on machine learning for recurring concept drifting data streams. Expert Systems with Applications. 2023; 213, 118934. https://doi.org/10.1016/j.eswa.2022.118934.
[13] Han, M., Chen, Z., Li, M., Wu, H., & Zhang, X. A survey of active and passive concept drift handling methods. Computational Intelligence. 2022; 38(4): 1492–1535. https://doi.org/10.1111/coin.12520.
[14] Chugg, B., Wang, H., & Ramdas, A. A Unified Recipe for Deriving (Time-Uniform) PAC-Bayes Bounds. Journal of Machine Learning Research. 2023; 24(372): 1–61. http://jmlr.org/papers/v24/23-0401.html.
[15] Rodríguez-Gálvez, B., Thobaben, R., & Skoglund, M. More PAC-Bayes bounds: From bounded losses, to losses with general tail behaviors, to anytime validity. Journal of Machine Learning Research. 2024; 25(110): 1–43. http://jmlr.org/papers/v25/23-1360.html.
[16] Dupuis, B., Viallard, P., Deligiannidis, G., & Simsekli, U. Uniform Generalization Bounds on Data-Dependent Hypothesis Sets via PAC-Bayesian Theory on Random Sets. Journal of Machine Learning Research. 2024; 25(409): 1–55. http://jmlr.org/papers/v25/24-0605.html.
[17] Tang, H., & Liu, Y. Information-Theoretic Generalization Bounds for Transductive Learning and its Applications. Journal of Machine Learning Research. 2024; 25(407): 1–69. http://jmlr.org/papers/v25/23-1368.html.
[18] Hou, S., Kassraie, P., Kratsios, A., Krause, A., & Rothfuss, J. Instance-Dependent Generalization Bounds via Optimal Transport. Journal of Machine Learning Research. 2023; 24(349): 1–51. http://jmlr.org/papers/v24/22-1293.html.
[19] Rothfuss, J., Josifoski, M., Fortuin, V., & Krause, A. Scalable PAC-Bayesian Meta-Learning via the PAC-Optimal Hyper-Posterior: From Theory to Practice. Journal of Machine Learning Research. 2023; 24(386): 1–62. http://jmlr.org/papers/v24/22-1254.html.
[20] Pelosi, D., Cacciagrano, D., & Piangerelli, M. Explainability and interpretability in concept and data drift: A systematic literature review. Algorithms. 2025; 18(7): 443. https://doi.org/10.3390/a18070443.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Lina Guo, Zijing Zhang, Yujia Zhang, Le Ji

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
This is an open access article distributed under the terms of the CC BY-NC-SA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.