Relative Position Encoding-Enhanced Vision Transformer for Image-Level Defect Classification in Manufacturing Textured-Surface Inspection
DOI:
https://doi.org/10.4108/eetsis.13596Keywords:
Vision transformer, relative position encoding, surface defect classification, industrial visual inspection, manufacturing quality inspectionAbstract
INTRODUCTION: In textile and textured-surface manufacturing, visual defects in woven fabrics, leather-like materials, tiles, and grid products directly influence product grading, downstream cutting, assembly reliability, after-sales repair, and supply-chain delivery stability. Weak defects embedded in repetitive textures are still difficult for manual inspection and conventional handcrafted vision methods.
OBJECTIVES: This study constructs an image-level defective/non-defective classification model for manufacturing quality inspection by enhancing Vision Transformer with relative spatial modeling, so that subtle texture disruption can be identified before products enter subsequent processing or distribution stages.
METHODS: A ViT-Small backbone was combined with learnable relative position encoding. Input images were resized, normalized, divided into 16 × 16 patches, and transformed into token embeddings. Relative spatial bias was inserted into multi-head self-attention to encode patch-to-patch displacement while retaining absolute positional information. The model was evaluated with precision, accuracy, recall, F1-score, and AUROC on a supervised MVTec-derived texture subset.
RESULTS: ViT-RPE outperformed ResNet-18, an activation-embedded CNN, EfficientFormer-L1, Swin-Tiny, and the baseline ViT under the same supervised image-level classification protocol. It achieved an accuracy of 0.954 ± 0.004 and an AUROC of 0.977 ± 0.003 on DAGM, and an accuracy of 0.939 ± 0.005 and an AUROC of 0.969 ± 0.004 on the MVTec-derived subset. Ablation results showed that combining absolute and relative positional information produced the best performance, with limited additional computational cost.
CONCLUSION: The method provides a compact image-level pre-screening solution for automated product inspection in manufacturing lines, especially where repetitive texture, material appearance consistency, and rapid quality sorting are required.
References
[1] Kahraman Y, Durmuşoğlu A. Deep learning-based fabric defect detection: A review[J]. Textile Research Journal, 2023, 93(5-6): 1485-1503.
[2] Smith A D, Du S, Kurien A. Vision transformers for anomaly detection and localisation in leather surface defect classification based on low-resolution images and a small dataset[J]. Applied sciences, 2023, 13(15): 8716.
[3] Hassan S A, Beliatis M J, Radziwon A, et al. Textile fabric defect detection using enhanced deep convolutional neural network with safe human–robot collaborative interaction[J]. Electronics, 2024, 13(21): 4314.
[4] Carrilho R, Yaghoubi E, Lindo J, et al. Toward automated fabric defect detection: A survey of recent computer vision approaches[J]. Electronics, 2024, 13(18): 3728.
[5] Tabernik D, Šela S, Skvarč J, et al. Segmentation-based deep-learning approach for surface-defect detection. Journal of Intelligent Manufacturing. 2020;31:759–776. doi:10.1007/s10845-019-01476-x.
[6] Mewada H, Pires I M, Engineer P, et al. Fabric surface defect classification and systematic analysis using a cuckoo search optimized deep residual network[J]. Engineering Science and Technology, an International Journal, 2024, 53: 101681.
[7] Aksakalli I K, Demir K, Sokmen O. A hybrid PatchNet-Attention based deep learning architecture for multi-type fabric defect classification in textile manufacturing and quality control[J]. Engineering Science and Technology, an International Journal, 2025, 72: 102231.
[8] Machado R, Barros L A M, Vieira V, et al. Textile defect detection using artificial intelligence and computer vision—a preliminary deep learning approach[J]. Electronics, 2025, 14(18): 3692.
[9] Erdogan M, Dogan M. Enhanced curvature-based fabric defect detection: a experimental study with gabor transform and deep learning[J]. Algorithms, 2024, 17(11): 506.
[10] Geze R A, Akbaş A. Detection and classification of fabric defects using deep learning algorithms[J]. Politeknik Dergisi, 2023, 27(1): 371-378.
[11] Liu Z, Lin Y, Cao Y, et al. Swin Transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021:10012–10022. doi:10.1109/ICCV48922.2021.00986.
[12] Carion N, Massa F, Synnaeve G, et al. End-to-end object detection with transformers. In: Vedaldi A, Bischof H, Brox T, Frahm J M, eds. Computer Vision – ECCV 2020. Cham: Springer; 2020:213–229. doi:10.1007/978-3-030-58452-8_13.
[13] Batzner K, Heckler L, König R. Efficientad: Accurate visual anomaly detection at millisecond-level latencies[C]//Proceedings of the IEEE/CVF winter conference on applications of computer vision. 2024: 128-138.
[14] Hyun J, Kim S, Jeon G, et al. Reconpatch: Contrastive patch representation learning for industrial anomaly detection[C]//Proceedings of the IEEE/CVF winter conference on applications of computer vision. 2024: 2052-2061.
[15] Bae J, Lee J H, Kim S. PNI: Industrial anomaly detection using position and neighborhood information[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023: 6373-6383.
[16] Ruan J, He J, Tong Y, et al. Knowledge Embedding Relation Network for Small Data Defect Detection[J]. Applied Sciences, 2024, 14(17): 7922.
[17] Ozek A, Seckin M, Demircioglu P, et al. Artificial intelligence driving innovation in textile defect detection[J]. Textiles, 2025, 5(2): 12.
[18] Cauteruccio F, Marchetti M, Traini D, et al. Adaptive patch selection to improve Vision Transformers through Reinforcement Learning: F. Cauteruccio et al[J]. Applied Intelligence, 2025, 55(7): 607.
[19] Wu K, Peng H, Chen M, et al. Rethinking and improving relative position encoding for vision transformer[C]//Proceedings of the IEEE/CVF international conference on computer vision. 2021: 10033-10041.
[20] Ma X, Zhang Z, Yu R, et al. SAVE: Encoding spatial interactions for vision transformers[J]. Image and Vision Computing, 2024, 152: 105312.
[21] Heo B, Park S, Han D, et al. Rotary position embedding for vision transformer[C]//European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024: 289-305.
[22] Jeong J, Zou Y, Kim T, et al. Winclip: Zero-/few-shot anomaly classification and segmentation[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023: 19606-196
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Jiamin Liu, Yuhui Sun

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
This is an open access article distributed under the terms of the CC BY-NC-SA 4.0, which permits copying, redistributing, remixing, transformation, and building upon the material in any medium so long as the original work is properly cited.