CrossVerify: low-payload external validation signals for heterogeneous black-box LLM cascades
DOI:
https://doi.org/10.4108/eetsis.14497Keywords:
LLM cascade, cross-verification, black-box routing, selective prediction, confidence calibrationAbstract
INTRODUCTION: Heterogeneous black-box LLM cascades need routing signals that are reliable, calibratable, and computable without hidden-state access.
OBJECTIVES: This study evaluates whether external verifier-support signals can complement proposer self-confidence in low-payload cascade routing.
METHODS: CrossVerify is designed and tested on six datasets, three proposer-verifier pairs, and five random seeds under both signal-analysis and fixed-fallback escalation settings.
RESULTS: Proposer confidence is the strongest average single-feature baseline, while verifier support is the strongest external signal; their combination yields the best overall routing performance, with clear task dependence and a 28-byte black-box-compatible interface.
CONCLUSION: CrossVerify offers a deployment-compatible signal design whose main value is complementing self-confidence rather than replacing it.
References
[1] Varangot-Reille C, Bouvard C, Ciancone M, et al. Doing more with less: A survey on routing strategies for resource optimisation in large language model-based systems. J. Artif. Intell. Res. 2026;86:865-893. https://doi.org/10.1613/jair.1.19801.
[2] Dong Y, Ito T. From consensus theory to LLM agents: Practical consensus-building for multi-issue negotiation. Expert Syst. Appl. 2026;322:132250. https://doi.org/10.1016/j.eswa.2026.132250.
[3] Huang J, Han T, Yu N, Yi X. Socratic elenchus-inspired multi-agent debate for mitigating hallucinations in large language models. Expert Syst. Appl. 2026;320:132208. https://doi.org/10.1016/j.eswa.2026.132208.
[4] Valkanas A, Pal S, Rumiantsev P, Zhang Y, Coates M. C3PO: Optimized large language model cascades with probabilistic cost constraints for reasoning. In: Advances in Neural Information Processing Systems. 2025;38:84481-84522.
[5] Luo H, Liu Y, Zhang R, et al. Toward edge general intelligence with multiple-large language model (Multi-LLM): Architecture, trust, and orchestration. IEEE Trans. Cogn. Commun. Netw. 2025;11(6):3563-3585. https://doi.org/10.1109/TCCN.2025.3612760.
[6] Hendrickx K, Perini L, Van der Plas D, Meert W, Davis J. Machine learning with a reject option: A survey. Mach. Learn. 2024;113(5):3073–3110.
[7] Silva Filho TM, Song H, Perello-Nieto M, Santos-Rodriguez R, Kull M, Flach P. Classifier calibration: A survey on how to assess and improve predicted class probabilities. Mach. Learn. 2023;112(9):3211–3260.
[8] Hullermeier E, Waegeman W. Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Mach. Learn. 2021;110(3):457–506.
[9] La Malfa E, et al. Language-models-as-a-service: Overview of a new paradigm and its challenges. J. Artif. Intell. Res. 2024;80:1497–1523.
[10] Dong Y, et al. Safeguarding large language models: A survey. Artif. Intell. Rev. 2025;58(12):382. https://doi.org/10.1007/s10462-025-11389-2.
[11] Coussement K, Abedin MZ, Kraus M, Maldonado S, Topuz K. Explainable AI for enhanced decision-making. Decis. Support Syst. 2024;184:114276.
[12] Chen J, Mueller J. Quantifying uncertainty in answers from any language model and enhancing their trustworthiness. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. p. 5186–5200.
[13] Lin Z, Trivedi S, Sun J. Generating with confidence: Uncertainty quantification for black-box large language models. Trans. Mach. Learn. Res. 2024. https://openreview.net/forum?id=DWkJCSxKU5.
[14] Kapoor S, et al. Large language models must be taught to know what they don’t know. In: Advances in Neural Information Processing Systems. 2024;37:85932-85972.
[15] Wang X, et al. MixLLM: Dynamic routing in mixed large language models. In: Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2025. p. 10001–10018.
[16] Shnitzer T, Ou A, Silva M, et al. Large language model routing with benchmark datasets. In: Proceedings of the First Conference on Language Modeling (COLM). 2024.
[17] Shen Y, Liu Y, Huang Z, Yin R, Zheng X, Huang X. SATER: A self-aware and token-efficient approach to routing and cascading. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. p. 10515–10529.
[18] Chen B-W, Chen C-C, Yen A-Z. Confidence-driven multi-scale model selection for cost-efficient inference. In: Findings of the Association for Computational Linguistics: EACL 2026. 2026. p. 1760–1770.
[19] Shirkavand R, Gao S, Yu P, Huang H. Cost-aware contrastive routing for LLMs. In: Advances in Neural Information Processing Systems. 2025;38. https://arxiv.org/abs/2508.12491.
[20] Felicioni N, Maystre L, Ghiassian S, Ciosek K. On the importance of uncertainty in decision-making with large language models. Trans. Mach. Learn. Res. 2024. https://openreview.net/forum?id=YfPzUX6DdO.
[21] Bhattacharjya D, et al. SIMBA UQ: Similarity-based aggregation for uncertainty quantification in large language models. In: Findings of the Association for Computational Linguistics: EMNLP 2025. 2025. p. 15880–15894.
[22] Mahaut M, Aina L, Czarnowska P, Hardalov M, Müller T, Marquez L. Factual confidence of LLMs: On reliability and robustness of current estimators. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. p. 4554–4570.
[23] Khanmohammadi R, et al. Calibrating LLM confidence by probing perturbed representation stability. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. p. 10448–10514.
[24] Hong R, Zhang H, Pang X, Yu D, Zhang C. A closer look at the self-verification abilities of large language models in logical reasoning. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. p. 900–925.
[25] Zhang Y, et al. Small language models need strong verifiers to self-correct reasoning. In: Findings of the Association for Computational Linguistics: ACL 2024. 2024. p. 15637–15653.
[26] Xiong M, et al. Can LLMs express their uncertainty? An empirical evaluation of confidence elicitation in LLMs. In: The Twelfth International Conference on Learning Representations. 2024. https://openreview.net/forum?id=gjeQKFxFpZ.
[27] Trapeznikov K, Saligrama V, Castanon DA. Multi-stage classifier design. Mach. Learn. 2013;92(2–3):479–502.
[28] Wang Z, Wang Q, Zhang Y, Chen T, Zhu X, Shi X, Xu K. SConU: Selective Conformal Uncertainty in Large Language Models. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 19052–19075). Association for Computational Linguistics. 2025. https://doi.org/10.18653/v1/2025.acl-long.934.
[29] Zoppi T, Popov P. Confidence ensembles: Tabular data classifiers on steroids. Inf. Fusion. 2025;120:103126.
[30] Haas S, Hüllermeier E. Conformalized prescriptive machine learning for uncertainty-aware automated decision making: The case of goodwill requests. Int. J. Data Sci. Anal. 2024;20(3):2061–2077.
[31] Longo L, et al. Explainable artificial intelligence (XAI) 2.0: A manifesto of open challenges and interdisciplinary research directions. Inf. Fusion. 2024;106:102301.
[32] Lyu Q, et al. Calibrating large language models with sample consistency. Proc. AAAI Conf. Artif. Intell. 2025;39(18):19260–19268.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Li Luo, Keng Yap Ng, Wei Chong Choo

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
This is an open access article distributed under the terms of the CC BY-NC-SA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.