Temporal-Structural Stress Testing and Cross-Granularity Robustness Evaluation of Graph AI for Cryptocurrency AML Detection

Authors

DOI:

https://doi.org/10.4108/airo.13658

Keywords:

cryptocurrency AML, graph AI, graph neural networks, financial crime, temporal validation

Abstract

Cryptocurrency anti-money-laundering (AML) is typically framed as a graph-learning problem, since suspicious value flows rarely appear as isolated records. This study examines whether graph AI architectures retain an advantage over strong non-GNN baselines when evaluation is temporal, structurally explicit, and aligned with investigator triage. We answer that question using two public Bitcoin AML datasets: Elliptic, a transaction- node dataset, and Elliptic2, a subgraph-level dataset. Our comparison spans classical models, tree ensembles, graph-derived features, graph embeddings, neural baselines, and GNN variants, evaluated across AUPRC, AUROC, accuracy, precision and recall at an investigation budget, calibration, operating points, and bootstrap uncertainty. On the temporal Elliptic test period, Extra Trees achieves AUPRC 0.570 and AUROC 0.869, while GraphSAGE, the strongest standard GNN baseline, reaches AUPRC 0.311. A strict-temporal GraphSAGE variant improves to AUPRC 0.381 but still falls behind the leading tree ensemble model. The paired bootstrap difference between Extra Trees and GraphSAGE is 0.259 AUPRC, with a 95% interval of [0.222, 0.297]. On the full Elliptic2 subgraph benchmark, Random Forest achieves AUPRC 0.494 and AUROC 0.923 in the main split and remains the strongest model by mean AUPRC across repeated splits. These findings indicate that graph AI for cryptocurrency AML should be stress-tested against capable tree ensemble models and operational metrics before architectural complexity is treated as deployment evidence.

Downloads

Download data is not yet available.

References

[1] He H, Garcia EA. Learning from Imbalanced Data. IEEE Transactions on Knowledge and Data Engineering. 2009;21(9):1263-84.

[2] Weber M, Domeniconi G, Chen J, Weidele DKI, Bellei C, Robinson T, et al. Anti-Money Laundering in Bitcoin: Experimenting with Graph Convolutional Networks for Financial Forensics. In: KDD ’19 Workshop on Anomaly Detection in Finance. Anchorage, AK, USA; 2019. Available from: https://arxiv.org/abs/1908.02591.

[3] Bellei C, Xu M, Phillips R, Robinson T, Weber M, Kaler T, et al.. The Shape of Money Laundering: Subgraph Representation Learning on the Blockchain with the Elliptic2 Dataset; 2024. ArXiv preprint. Available from: https://arxiv.org/abs/2404.19109.

[4] Kaufman S, Rosset S, Perlich C, Stitelman O. Leakage in Data Mining: Formulation, Detection, and Avoidance. ACM Transactions on Knowledge Discovery from Data. 2012;6(4):15.

[5] Ron D, Shamir A. Quantitative Analysis of the Full Bitcoin Transaction Graph. In: Financial Cryptography and Data Security. vol. 7859 of Lecture Notes in Computer Science. Springer; 2013. p. 6-24.

[6] Meiklejohn S, Pomarole M, Jordan G, Levchenko K, McCoy D, Voelker GM, et al. A Fistful of Bitcoins: Characterizing Payments Among Men with No Names. In: Proceedings of the 2013 Internet Measurement Conference. ACM; 2013. p. 127-39.

[7] Alarab I, Prakoonwit S. Graph-Based LSTM for Anti- Money Laundering: Experimenting Temporal Graph Convolutional Network with Bitcoin Data. Neural Processing Letters. 2023;55:689-707.

[8] Bolton RJ, Hand DJ. Statistical Fraud Detection: A Review. Statistical Science. 2002;17(3):235-55.

[9] Lessmann S, Baesens B, Seow HV, Thomas LC. Benchmarking State-of-the-Art Classification Algorithms for Credit Card Fraud Detection. European Journal of Operational Research. 2016;249(2):666-81.

[10] Breiman L. Random Forests. Machine Learning. 2001;45(1):5-32.

[11] Geurts P, Ernst D, Wehenkel L. Extremely Randomized Trees. Machine Learning. 2006;63(1):3-42.

[12] Friedman JH. Greedy Function Approximation: A Gradient Boosting Machine. The Annals of Statistics. 2001;29(5):1189-232.

[13] Chen T, Guestrin C. XGBoost: A Scalable Tree Boosting System. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; 2016. p. 785-94.

[14] Rudin C. Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. Nature Machine Intelligence. 2019;1:206-15.

[15] Kipf TN, Welling M. Semi-Supervised Classification with Graph Convolutional Networks. In: International Conference on Learning Representations; 2017. Available from: https://openreview.net/forum?id=SJU4ayYgl.

[16] Hamilton WL, Ying R, Leskovec J. Inductive Representation Learning on Large Graphs. In: Advances in Neural Information Processing Systems. vol. 30; 2017. Available from: https://proceedings.neurips.cc/paper/2017/hash/5dd9db5e033da9c6fb5ba837a7ebea9-Abstract.html.

[17] Velickovic P, Cucurull G, Casanova A, Romero A, Lio P, Bengio Y. Graph Attention Networks. In: International Conference on Learning Representations; 2018. Available from: https://openreview.net/forumid=rJXMpikCZ.

[18] Wu Z, Pan S, Chen F, Long G, Zhang C, Yu PS. A Comprehensive Survey on Graph Neural Networks. IEEE Transactions on Neural Networks and Learning Systems. 2021;32(1):4-24.

[19] Pareja A, Domeniconi G, Chen J, Ma T, Suzumura T, Kanezashi H, et al. EvolveGCN: Evolving Graph Convolutional Networks for Dynamic Graphs. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34; 2020. p. 5363-70.

[20] Rossi E, Chamberlain B, Frasca F, Eynard D, Monti F, Bronstein M. Temporal Graph Networks for Deep Learning on Dynamic Graphs; 2020.

[21] Grover A, Leskovec J. node2vec: Scalable Feature Learning for Networks. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; 2016. p. 855-64.

[22] Dou Y, Liu Z, Sun L, Deng Y, Peng H, Yu PS. Enhancing Graph Neural Network-Based Fraud Detectors against Camouflaged Fraudsters. In: Proceedings of the 29th ACM International Conference on Information and Knowledge Management; 2020. p. 315-24.

[23] Ma X, Wu J, Xue S, Yang J, Zhou C, Sheng QZ, et al. A Comprehensive Survey on Graph Anomaly Detection With Deep Learning. IEEE Transactions on Knowledge and Data Engineering. 2023;35(12):12012-38.

[24] Egressy B, von Niederhäusern L, Blanuša J, Altman E, Wattenhofer R, Atasu K. Provably Powerful Graph Neural Networks for Directed Multigraphs. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 38; 2024. p. 11838-46.

[25] Lin Z, Luo Q, Wu D, Shen J, Li L, Nong X, et al. Detecting Illicit Transactions in Bitcoin: A Wavelet- Temporal Graph Transformer Approach for Anti-Money Laundering. Scientific Reports. 2026;16(1):1548.

[26] Li E, Chen M, Xiang S, Chen L. Graph Learning- Empowered Financial Fraud Detection: Progress and Future Directions. Intelligent Computing. 2025;4:0146.

[27] Hu W, Fey M, Zitnik M, Dong Y, Ren H, Liu B, et al. Open Graph Benchmark: Datasets for Machine Learning on Graphs. In: Advances in Neural Information Processing Systems. vol. 33; 2020. p. 22118-33. Available from: https://proceedings.neurips.cc/paper/2020/hash/fb60d411a5c5b72b2e7d3527cfc84fd0-Abstract.html.

[28] Saito T, Rehmsmeier M. The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLOS ONE. 2015;10(3):e0118432.

[29] Guo C, Pleiss G, Sun Y, Weinberger KQ. On Calibration of Modern Neural Networks. In: Proceedings of the 34th International Conference on Machine Learning. vol. 70 of Proceedings of Machine Learning Research. PMLR; 2017. p. 1321-30. Available from: https://proceedings.mlr.press/v70/guo17a.html.

[30] Yuan H, Yu H, Gui S, Ji S. Explainability in Graph Neural Networks: A Taxonomic Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2023;45(5):5782-99.

[31] Efron B, Tibshirani RJ. An Introduction to the Bootstrap. Chapman and Hall/CRC; 1994.

[32] Cao X, Li S, Katsikis V, Khan AT, He H, Liu Z, et al. Empowering Financial Futures: Large Language Models in the Modern Financial Landscape. EAI Endorsed Transactions on AI and Robotics. 2024;3(1).

[33] Hossan MZ, Riipa MB, Hossain MA, Dhar SR, Zaman AM, Hossain M, et al. AI-Powered Predictive Analytics for Financial Risk Management in U.S. Markets. EAI Endorsed Transactions on AI and Robotics. 2025 August;4(1). Available from: https://publications. eai.eu/index.php/airo/article/view/9532.

[34] Dang QV, Nguyen PL, Le D, Dinh MN. XHBot: eXplainable Heterophily-Aware Graph Neural Networks for Social Bot Detection. EAI Endorsed Transactions on AI and Robotics. 2026 July;5. Available from: https://publications.eai.eu/index.php/airo/article/view/12969.

[35] Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research. 2011;12:2825-30. Available from: https://www.jmlr.org/papers/v12/pedregosa11a.html.

[36] Cox DR. The Regression Analysis of Binary Sequences. Journal of the Royal Statistical Society: Series B (Methodological). 1958;20(2):215-32.

[37] Rumelhart DE, Hinton GE, Williams RJ. Learning Representations by Back-Propagating Errors. Nature. 1986;323(6088):533-6.

[38] Page L, Brin S, Motwani R, Winograd T. The PageRank Citation Ranking: Bringing Order to the Web. Stanford InfoLab; 1999. 1999-66. Available from: http://ilpubs.stanford.edu:8090/422/1/1999-66.pdf.

[39] Seidman SB. Network Structure and Minimum Degree. Social Networks. 1983;5(3):269-87.

[40] Fawcett T. An Introduction to ROC Analysis. Pattern Recognition Letters. 2006;27(8):861-74.

[41] Brier GW. Verification of Forecasts Expressed in Terms of Probability. Monthly Weather Review. 1950;78(1):1-3.

[42] Tabassi E. Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology; 2023. NIST AI 100-1.

Downloads

Published

02-09-2026

How to Cite

1.
Sadia Akter, Sabab Islam, Md Faysal Ahmed, Md Hossain Jamil, Md Fokrul Islam Khan, Partha Singha, et al. Temporal-Structural Stress Testing and Cross-Granularity Robustness Evaluation of Graph AI for Cryptocurrency AML Detection. EAI Endorsed Trans AI Robotics [Internet]. 2026 Sep. 2 [cited 2026 Sep. 2];5. Available from: https://publications.eai.eu/index.php/airo/article/view/13658

Most read articles by the same author(s)