Vision Transformers in Object Detection; A Survey

سال انتشار: 1405
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 41

فایل این مقاله در 8 صفحه با فرمت PDF قابل دریافت می باشد

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

ECMCONF11_005

تاریخ نمایه سازی: 13 مرداد 1405

چکیده مقاله:

Transformer was applied to the field of natural language processing, which has numerous advantages over Convolution Neural Networks (CNN). The outstanding performance of Transformer models in NLP tasks encouraged a lot of researchers to explore new fields for applying Transformers which led to Computer Vision tasks. Vision Transformer (ViT) is an image recognition model and It uses transformer architecture. Vision Transformer is broadly used in Image Classification, Semantic Segmentation, Object Detection and other fields. DEtection Transformer (DETR) reframes detection as a set prediction problem. However, DETR has several disadvantages like slow training convergence and not being effective for smaller objects. To mitigate these issues, several researchers have proposed different methods. In this survey our goal is to cover eight Transformer-based object detection methods. We found it essential to classify and compare these methods. Both the foundational modules of DETR and its recent advancements are covered.

کلیدواژه ها:

نویسندگان

Sara Mohseni

Department of Computer Engineering, Islamic Azad University of Qods

Vahid Ahadi Moghadam

Department of Computer Engineering, Islamic Azad University of Qods