Multimodal Transformer for Integrating Clinical Text, Imaging, and Genomics in Precision Oncology

سال انتشار: 1404
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 164

فایل این مقاله در 13 صفحه با فرمت PDF قابل دریافت می باشد

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

AICNF03_057

تاریخ نمایه سازی: 3 اسفند 1404

چکیده مقاله:

Recent advances in artificial intelligence have highlighted the potential of multimodal learning to enhance precision diagnostics in oncology. This study proposes a novel transformer-based framework that integrates three critical data modalities—radiological images, genomic profiles, and clinical text records—to achieve comprehensive cancer diagnosis and subtype classification. The proposed architecture employs modality-specific encoders (Vision Transformer for MRI, BioBERT for clinical text, and a one-dimensional transformer for genomic data) and a cross-attention fusion transformer to learn shared representations across heterogeneous sources. The model was trained and validated on subsets of the TCGA and MIMIC-IV datasets, demonstrating superior performance compared to unimodal and dual-modality baselines. Quantitative results show that the proposed tri-modal transformer improves diagnostic accuracy by up to ۱۱.۳% and enhances interpretability through modality-specific attention visualization. The findings confirm that integrating imaging, molecular, and textual data significantly strengthens diagnostic reliability and supports precision oncology applications. This work establishes a foundation for unified, explainable, and data-efficient diagnostic systems in future AI-driven healthcare.

نویسندگان

Mohammad Khooshebast Baghsangani

MCs student Engineering Faculty Islamic Azad University of Mashhad

Jafar Khooshebast Baghsangani

BCs student Engineering Faculty Ferdowsi University of Mashhad

Mostafa Khosheh BastBaghsangani

PhD candidate Electrical Department Hakim Sabzevari University