Abstract
Crack segmentation plays an important role in intelligent infrastructure inspection and safety assessment. However, accurate pixel-level crack segmentation remains challenging due to irregular crack morphology, complex backgrounds, and noise interference, particularly when processing high-resolution inspection images under computational constraints. To address these challenges, this paper proposes a computationally efficient and computationally efficient crack segmentation network based on adaptive Mamba enhancement and multi-level feature fusion. The proposed framework aims to achieve accurate crack representation while maintaining a favorable balance between segmentation performance and computational cost. Specifically, an adaptive Mamba enhancement (AME) module is introduced as the fundamental building block of the U-shaped architecture, which integrates selective state-space modeling with local convolutional refinement to effectively capture long-range dependencies while reducing redundant feature responses. A fine-grained perception feedforward network (FPN) is further developed to enhance local structural representation through multi-scale directional convolution, and a multi-scale progressive fusion (MPF) module is embedded into skip connections to alleviate information loss during hierarchical feature reconstruction. Extensive experiments conducted on three public crack datasets, including CrackTree260, CFD, and CrackLS315, demonstrate that the proposed network achieves mIoU scores of 84.25%, 82.07%, and 70.63%, respectively, outperforming eight representative segmentation methods while maintaining competitive computational efficiency. The results indicate that the proposed method provides an effective solution for high-resolution crack analysis and demonstrates potential for computationally constrained infrastructure inspection applications.
Data availability
All data generated or analyzed during this study are included in this published article (and its supplementary information files).
Code availability
Code and models will be available at https://github.com/wyogMg/CM-UNet
References
Barisin T, Jung C, Müsebeck F et al (2022) Methods for segmenting cracks in 3D images of concrete: a comparison based on semi-synthetic images. Pattern Recognit 129:108747
Fang F, Li L, Gu Y et al (2020) A novel hybrid approach for crack detection. Pattern Recognit 107:107474
Flah M, Suleiman AR, Nehdi ML (2020) Classification and quantification of cracks in concrete structures using deep learning image-based techniques. Cem Concr Compos 114:103781
Akagic A, Buza E, Omanovic S, et al. (2018) Pavement crack detection using Otsu thresholding for image segmentation. In: 41st international convention on information and communication technology, electronics and microelectronics (MIPRO). IEEE, pp 1092–1097
Su M, Wan J, Zhou Q et al (2024) Utilizing pretrained convolutional neural networks for crack detection and geometric feature recognition in concrete surface images. J Build Eng 98:111386
Sari Y, Prakoso PB, Baskara AR (2019) Road crack detection using support vector machine (SVM) and OTSU algorithm. In: 2019 6th International Conference on Electric Vehicular Technology (ICEVT). IEEE, pp 349–354
Yuan C, Zhang H, Wang L et al (2024) Research on strength prediction of crack rock mass based on random forest algorithm. Bull Eng Geol Environ 83(4):128
Hacıefendioğlu K, Başağa HB (2022) Concrete road crack detection using deep learning-based faster R-CNN method. Iran J Sci Technol Trans Civ Eng 46(2):1621–1633
Jiang T, Huang Y, Hu C et al (2025) Bridge component segmentation for health monitoring an enhanced DeepLabV3+ model with lightweight network and multi-scale channel attention mechanism. Adv Struct Eng 28(5):939–951
Ronneberger, Olaf, Philipp Fischer, Thomas Brox (2015) U-net: convolutional networks for biomedical image segmentation. In: International Conference on Medical Image Computing and Computer-assisted Intervention. Cham: Springer international publishing
Cheng J, Xiong W, Chen W, et al. (2018) Pixel-level crack detection using U-Net. In: TENCON 2018-2018 IEEE Region 10 Conference. IEEE, pp 0462–0466
Chen J, He Y (2022) A novel U‐shaped encoder–decoder network with attention mechanism for detection and evaluation of road cracks at pixel level. Comput-Aided Civ Infrastruct Eng 37(13):1721–1736
Dosovitskiy A (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929
Liu H, Miao X, Mertz C, et al. (2021) Crackformer: transformer network for fine-grained crack detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp 3783–3792
Xiang C, Guo J, Cao R et al (2023) A crack-segmentation algorithm fusing transformers and convolutional neural networks for complex detection scenarios. Autom Constr 152:104894
Sun Z, Zhai J, Pei L et al (2023) Automatic pavement crack detection transformer based on convolutional and sequential feature fusion. Sensors (Basel) 23(7):3772
Liu Z, Lin Y, Cao Y, et al. (2021) Swin transformer: Hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision pp 10012–10022
Li K, Wang Y, Zhang J et al (2023) Uniformer: unifying convolution and self-attention for visual recognition. IEEE Trans Pattern Anal Mach Intell 45(10):12581–12600
Gu A, Dao T (2023) Mamba: linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752
Wang Z, Zheng J Q, Zhang Y, et al. (2024) Mamba-unet: Unet-like pure visual mamba for medical image segmentation. arXiv preprint arXiv:2402.05079
Ma J, Li F, Wang B (2024) U-mamba: enhancing long-range dependency for biomedical image segmentation[J]. arXiv preprint arXiv:2401.04722
Ma X, Zhang X, Pun MO (2024) RS 3 Mamba: visual state space model for remote sensing image semantic segmentation. IEEE Geosci Remote Sens Lett. https://doi.org/10.1109/lgrs.2024.3414293
Zhu Q, Cai Y, Fang Y et al (2024) Samba: semantic segmentation of remotely sensed images with state space model. Heliyon. https://doi.org/10.1016/j.heliyon.2024.e38495
Li Y, Hou Q, Zheng Z, et al. (2023) Large selective kernel network for remote sensing object detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp 16794–16805
Chu H, Wang W, Deng L (2022) Tiny-Crack-Net: a multiscale feature fusion network with attention mechanisms for segmentation of tiny cracks[J]. Computer-Aided Civil Infrastructure Eng 37(14):1914–1931
Wen X, Li S, Yu H et al (2024) Multi-scale context feature and cross-attention network-enabled system and software-based for pavement crack detection. Eng Appl Artif Intell 127:107328
Wang J, Zeng Z, Wang J et al (2024) Automatic crack segmentation model based on multi-branch aggregation transformer. Adv Struct Eng 27(13):2289–2302
Guo MH, Lu CZ, Liu ZN et al (2023) Visual attention network. Comput Vis Media 9(4):733–752
Gu A, Goel K, Ré C (2021) Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396
Liu Y, Tian Y, Zhao Y, Yu H, Xie L, Wang Y, Ye Q, Liu Y (2024) Vmamba: visual state space model. arXiv preprint arXiv:2401.10166
Cheng C, Wang H, Sun H (2024) Activating wider areas in image super-resolution. arXiv preprint arXiv:2403.08330
Guo H, Li J, Dai T, et al. (2025) Mambair: a simple baseline for image restoration with state-space model. In: European Conference on Computer Vision. Springer, Cham, pp 222–241
Ba J L (2016) Layer normalization. arXiv preprint arXiv:1607.06450
Ma M, Yang L, Liu Y et al (2024) An attention-based progressive fusion network for pixelwise pavement crack detection. Measurement 226:114159
Elfwing S, Uchibe E, Doya K (2018) Sigmoid-weighted linear units for neural network function approximation in reinforcement learning. Neural Netw 107:3–11
Wang B, Deng F, Jiang P et al (2024) WiTUnet: a U-shaped architecture integrating CNN and transformer for improved feature alignment and local information fusion. Sci Rep 14(1):25525
Lou M, Zhang S, Zhou HY, Yang S, Wu C, Yu Y (2025) TransXNet: learning both global and local dynamics with a dual dynamic token mixer for visual recognition. IEEE Trans Neural Netw Learn Syst. https://doi.org/10.1109/tnnls.2025.3550979
Lau KW, Po LM, Rehman YAU (2024) Large separable kernel attention: rethinking the large kernel attention design in CNN. Expert Syst Appl 236:121352
Milletari F, Navab N, Ahmadi SA (2016) V-net: fully convolutional neural networks for volumetric medical image segmentation. In: 2016 Fourth International Conference on 3D Vision (3DV). IEEE, pp 565–571
Zou Q, Cao Y, Li Q et al (2012) CrackTree: automatic crack detection from pavement images. Pattern Recognit Lett 33(3):227–238
Zou Q, Zhang Z, Li Q et al (2018) Deepcrack: learning hierarchical convolutional features for crack detection. IEEE Trans Image Process 28(3):1498–1512
Wang J, Zeng Z, Sharma PK et al (2024) Dual-path network combining CNN and transformer for pavement crack segmentation. Autom Constr 158:105217
Kingma DP (2014) Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980
Badrinarayanan V, Kendall A, Cipolla R (2017) Segnet: a deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans Pattern Anal Mach Intell 39(12):2481–2495
Chen L C, Zhu Y, Papandreou G, et al. (2018) Encoder-decoder with atrous separable convolution for semantic image segmentation. In: Proceedings of the European Conference on Computer Vision (ECCV). pp 801–818
Liu Y, Yao J, Lu X et al (2019) Deepcrack: a deep hierarchical feature learning architecture for crack segmentation. Neurocomputing 338:139–153
Cao H, Wang Y, Chen J, et al. (2022) Swin-unet: Unet-like pure transformer for medical image segmentation. In: European Conference on Computer Vision. Cham: Springer Nature Switzerland, pp 205–218
Pang J, Zhang H, Zhao H, et al. (2022) DcsNet: a real-time deep network for crack segmentation. Signal, Image and Video Proc 1–9
Liu, Hui, et al. (2025) SCSegamba: lightweight structure-aware vision mamba for crack segmentation in structures. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE
Luo W, Li Y, Urtasun R, Zemel R (2016) Understanding the effective receptive field in deep convolutional neural networks. Adv Neural Inf Proc Sys 29
Funding
This work was supported by Natural Science Foundation of Hunan Province (2025JJ70638), and Scientific Research Fund of Hunan Provincial Education Department (24C0233).
Author information
Authors and Affiliations
Contributions
Yongming Wang involved in conceptualization, methodology, data curation, formal analysis, visualization, writing—original draft, writing review and editing; Siqi Liang performed conceptualization, validation, visualization, funding acquisition, project administration, writing review and editing; Shigang Hu contributed to methodology, investigation, resources, software, validation, visualization, supervision, writing—review and editing.
Corresponding author
Ethics declarations
Conflict of interests
The authors declare that they have no conflict of interest.
Additional information
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Wang, Y., Liang, S. & Hu, S. The U-shaped crack segmentation network based on adaptive Mamba enhancement and multi-scale aggregation. J Supercomput 82, 732 (2026). https://doi.org/10.1007/s11227-026-08860-4
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s11227-026-08860-4
Facts Only
* The proposed framework uses an adaptive Mamba enhancement (AME) module and multi-level feature fusion for crack segmentation.
* The AME module integrates selective state-space modeling with local convolutional refinement.
* A fine-grained perception feedforward network (FPN) is developed using multi-scale directional convolution to enhance local structural representation.
* A multi-scale progressive fusion (MPF) module is embedded in skip connections for hierarchical feature reconstruction.
* Experiments were conducted on CrackTree260, CFD, and CrackLS315 datasets.
* The network achieved mIoU scores of 84.25%, 82.07%, and 70.63% on these respective datasets.
* The results outperformed eight representative segmentation methods.
* Code and models are available at https://github.com/wyogMg/CM-UNet
Executive Summary
Full Take
Sentinel — Human
This text exhibits the high structural density and technical specificity characteristic of a peer-reviewed scientific publication rather than general synthetic content.
