Vision Transformers in Object Detection; A Survey
سال انتشار: 1405
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 41
فایل این مقاله در 8 صفحه با فرمت PDF قابل دریافت می باشد
- صدور گواهی نمایه سازی
- من نویسنده این مقاله هستم
استخراج به نرم افزارهای پژوهشی:
شناسه ملی سند علمی:
ECMCONF11_005
تاریخ نمایه سازی: 13 مرداد 1405
چکیده مقاله:
Transformer was applied to the field of natural language processing, which has numerous advantages over Convolution Neural Networks (CNN). The outstanding performance of Transformer models in NLP tasks encouraged a lot of researchers to explore new fields for applying Transformers which led to Computer Vision tasks. Vision Transformer (ViT) is an image recognition model and It uses transformer architecture. Vision Transformer is broadly used in Image Classification, Semantic Segmentation, Object Detection and other fields. DEtection Transformer (DETR) reframes detection as a set prediction problem. However, DETR has several disadvantages like slow training convergence and not being effective for smaller objects. To mitigate these issues, several researchers have proposed different methods. In this survey our goal is to cover eight Transformer-based object detection methods. We found it essential to classify and compare these methods. Both the foundational modules of DETR and its recent advancements are covered.
کلیدواژه ها:
Vision Transformer (VIT) ، Computer Vision ، Deep Neural Networks ، Convolutional Neural Networks (CNN) ، Object Detection ، DETR ، Deformable-DETR
نویسندگان
Sara Mohseni
Department of Computer Engineering, Islamic Azad University of Qods
Vahid Ahadi Moghadam
Department of Computer Engineering, Islamic Azad University of Qods