Vision Transformers for Multi‑Modal Damage Detection in Reinforced Concrete Structures: A Systematic Review of Visible and Infrared Image Fusion
سال انتشار: 1405
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 11
فایل این مقاله در 11 صفحه با فرمت PDF قابل دریافت می باشد
- صدور گواهی نمایه سازی
- من نویسنده این مقاله هستم
استخراج به نرم افزارهای پژوهشی:
شناسه ملی سند علمی:
CAUCONG05_100
تاریخ نمایه سازی: 18 مرداد 1405
چکیده مقاله:
Automated damage detection in reinforced concrete (RC) structures is a critical challenge in structural health monitoring (SHM), directly affecting infrastructure safety, maintenance costs, and life-cycle management [۱, ۲]. Conventional convolutional neural networks (CNNs), despite their success in many computer vision tasks, suffer significant performance degradation in real-world inspection environments characterized by poor illumination, dust, shadows, occlusion, and adverse weather [۳]. Furthermore, CNN-based detectors operating solely on visible (RGB) imagery cannot identify subsurface anomalies such as moisture ingress, delamination, thermal damage, or fire-induced residual heat, all of which are key indicators of RC deterioration [۴, ۵]. This systematic review investigates the application of Vision Transformers (ViTs) combined with multi-modal fusion of visible and infrared (IR) thermal images for automated detection of cracks, corrosion, spalling, and fire damage in RC structures [۶]. Following PRISMA ۲۰۲۰ guidelines [۷], a systematic search of Web of Science, Scopus, and IEEE Xplore (۲۰۲۰–۲۰۲۵) retrieved ۳۴۱ records, from which ۴۲ high-quality papers were selected after rigorous screening. The quantitative synthesis reveals that ViTs, owing to their self-attention mechanism, consistently outperform state-of-the-art CNNs in modeling long-range spatial dependencies and fusing complementary visible-thermal information, achieving an average improvement of approximately ۱۲ percent in mean average precision (mAP) under challenging conditions [۸, ۹]. Among the three fusion strategies evaluated—early, late, and cross-attention—cross-attention achieves the highest performance with a mean mAP of ۸۸.۷ percent, due to its ability to dynamically align multi-modal features at multiple hierarchical levels [۱۰]. Despite these promising results, critical obstacles remain: the absence of a large-scale, standardized, publicly available multi-modal benchmark dataset with pixel-level annotations [۱۱]; the high computational cost and inference latency of ViTs limiting real-time and edge deployment [۱۲]; and limited model interpretability undermining practitioner trust [۱۳]. To address these barriers, future research should prioritize four directions: diffusion-based generative data augmentation to expand limited training data [۱۴]; spatio-temporal ViTs for continuous monitoring of damage evolution [۱۵]; lightweight ViT architectures suitable for resource-constrained edge devices such as drones [۱۶]; and integration of ViTs with large language models (LLMs) for automated, intelligent inspection report generation [۱۷]. This review serves as a comprehensive reference for researchers and practitioners working on ViT-based multi-modal damage detection for reinforced concrete infrastructure.
کلیدواژه ها:
نویسندگان
Evaz Tajik
PhD student in civil engineering, structural engineering