CrossVerify: low-payload external validation signals for heterogeneous black-box LLM cascades

Authors

DOI:

https://doi.org/10.4108/eetsis.14497

Keywords:

LLM cascade, cross-verification, black-box routing, selective prediction, confidence calibration

Abstract

INTRODUCTION: Heterogeneous black-box LLM cascades need routing signals that are reliable, calibratable, and computable without hidden-state access.
OBJECTIVES: This study evaluates whether external verifier-support signals can complement proposer self-confidence in low-payload cascade routing.
METHODS: CrossVerify is designed and tested on six datasets, three proposer-verifier pairs, and five random seeds under both signal-analysis and fixed-fallback escalation settings.
RESULTS: Proposer confidence is the strongest average single-feature baseline, while verifier support is the strongest external signal; their combination yields the best overall routing performance, with clear task dependence and a 28-byte black-box-compatible interface.
CONCLUSION: CrossVerify offers a deployment-compatible signal design whose main value is complementing self-confidence rather than replacing it.

References

[1] Varangot-Reille C, Bouvard C, Ciancone M, et al. Doing more with less: A survey on routing strategies for resource optimisation in large language model-based systems. J. Artif. Intell. Res. 2026;86:865-893. https://doi.org/10.1613/jair.1.19801.

[2] Dong Y, Ito T. From consensus theory to LLM agents: Practical consensus-building for multi-issue negotiation. Expert Syst. Appl. 2026;322:132250. https://doi.org/10.1016/j.eswa.2026.132250.

[3] Huang J, Han T, Yu N, Yi X. Socratic elenchus-inspired multi-agent debate for mitigating hallucinations in large language models. Expert Syst. Appl. 2026;320:132208. https://doi.org/10.1016/j.eswa.2026.132208.

[4] Valkanas A, Pal S, Rumiantsev P, Zhang Y, Coates M. C3PO: Optimized large language model cascades with probabilistic cost constraints for reasoning. In: Advances in Neural Information Processing Systems. 2025;38:84481-84522.

[5] Luo H, Liu Y, Zhang R, et al. Toward edge general intelligence with multiple-large language model (Multi-LLM): Architecture, trust, and orchestration. IEEE Trans. Cogn. Commun. Netw. 2025;11(6):3563-3585. https://doi.org/10.1109/TCCN.2025.3612760.

[6] Hendrickx K, Perini L, Van der Plas D, Meert W, Davis J. Machine learning with a reject option: A survey. Mach. Learn. 2024;113(5):3073–3110.

[7] Silva Filho TM, Song H, Perello-Nieto M, Santos-Rodriguez R, Kull M, Flach P. Classifier calibration: A survey on how to assess and improve predicted class probabilities. Mach. Learn. 2023;112(9):3211–3260.

[8] Hullermeier E, Waegeman W. Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods. Mach. Learn. 2021;110(3):457–506.

[9] La Malfa E, et al. Language-models-as-a-service: Overview of a new paradigm and its challenges. J. Artif. Intell. Res. 2024;80:1497–1523.

[10] Dong Y, et al. Safeguarding large language models: A survey. Artif. Intell. Rev. 2025;58(12):382. https://doi.org/10.1007/s10462-025-11389-2.

[11] Coussement K, Abedin MZ, Kraus M, Maldonado S, Topuz K. Explainable AI for enhanced decision-making. Decis. Support Syst. 2024;184:114276.

[12] Chen J, Mueller J. Quantifying uncertainty in answers from any language model and enhancing their trustworthiness. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. p. 5186–5200.

[13] Lin Z, Trivedi S, Sun J. Generating with confidence: Uncertainty quantification for black-box large language models. Trans. Mach. Learn. Res. 2024. https://openreview.net/forum?id=DWkJCSxKU5.

[14] Kapoor S, et al. Large language models must be taught to know what they don’t know. In: Advances in Neural Information Processing Systems. 2024;37:85932-85972.

[15] Wang X, et al. MixLLM: Dynamic routing in mixed large language models. In: Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2025. p. 10001–10018.

[16] Shnitzer T, Ou A, Silva M, et al. Large language model routing with benchmark datasets. In: Proceedings of the First Conference on Language Modeling (COLM). 2024.

[17] Shen Y, Liu Y, Huang Z, Yin R, Zheng X, Huang X. SATER: A self-aware and token-efficient approach to routing and cascading. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. p. 10515–10529.

[18] Chen B-W, Chen C-C, Yen A-Z. Confidence-driven multi-scale model selection for cost-efficient inference. In: Findings of the Association for Computational Linguistics: EACL 2026. 2026. p. 1760–1770.

[19] Shirkavand R, Gao S, Yu P, Huang H. Cost-aware contrastive routing for LLMs. In: Advances in Neural Information Processing Systems. 2025;38. https://arxiv.org/abs/2508.12491.

[20] Felicioni N, Maystre L, Ghiassian S, Ciosek K. On the importance of uncertainty in decision-making with large language models. Trans. Mach. Learn. Res. 2024. https://openreview.net/forum?id=YfPzUX6DdO.

[21] Bhattacharjya D, et al. SIMBA UQ: Similarity-based aggregation for uncertainty quantification in large language models. In: Findings of the Association for Computational Linguistics: EMNLP 2025. 2025. p. 15880–15894.

[22] Mahaut M, Aina L, Czarnowska P, Hardalov M, Müller T, Marquez L. Factual confidence of LLMs: On reliability and robustness of current estimators. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. p. 4554–4570.

[23] Khanmohammadi R, et al. Calibrating LLM confidence by probing perturbed representation stability. In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. p. 10448–10514.

[24] Hong R, Zhang H, Pang X, Yu D, Zhang C. A closer look at the self-verification abilities of large language models in logical reasoning. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. p. 900–925.

[25] Zhang Y, et al. Small language models need strong verifiers to self-correct reasoning. In: Findings of the Association for Computational Linguistics: ACL 2024. 2024. p. 15637–15653.

[26] Xiong M, et al. Can LLMs express their uncertainty? An empirical evaluation of confidence elicitation in LLMs. In: The Twelfth International Conference on Learning Representations. 2024. https://openreview.net/forum?id=gjeQKFxFpZ.

[27] Trapeznikov K, Saligrama V, Castanon DA. Multi-stage classifier design. Mach. Learn. 2013;92(2–3):479–502.

[28] Wang Z, Wang Q, Zhang Y, Chen T, Zhu X, Shi X, Xu K. SConU: Selective Conformal Uncertainty in Large Language Models. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 19052–19075). Association for Computational Linguistics. 2025. https://doi.org/10.18653/v1/2025.acl-long.934.

[29] Zoppi T, Popov P. Confidence ensembles: Tabular data classifiers on steroids. Inf. Fusion. 2025;120:103126.

[30] Haas S, Hüllermeier E. Conformalized prescriptive machine learning for uncertainty-aware automated decision making: The case of goodwill requests. Int. J. Data Sci. Anal. 2024;20(3):2061–2077.

[31] Longo L, et al. Explainable artificial intelligence (XAI) 2.0: A manifesto of open challenges and interdisciplinary research directions. Inf. Fusion. 2024;106:102301.

[32] Lyu Q, et al. Calibrating large language models with sample consistency. Proc. AAAI Conf. Artif. Intell. 2025;39(18):19260–19268.

Downloads

Published

16-09-2026

How to Cite

1.
Luo L, Yap Ng K, Chong Choo W. CrossVerify: low-payload external validation signals for heterogeneous black-box LLM cascades. EAI Endorsed Scal Inf Syst [Internet]. 2026 Sep. 16 [cited 2026 Sep. 25];13(3). Available from: https://publications.eai.eu/index.php/sis/article/view/14497