Enhanced ۳D Scene Reconstruction Using Neural Radiance Fields and Vision Transformers
سال انتشار: 1405
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 24
فایل این مقاله در 10 صفحه با فرمت PDF قابل دریافت می باشد
- صدور گواهی نمایه سازی
- من نویسنده این مقاله هستم
استخراج به نرم افزارهای پژوهشی:
شناسه ملی سند علمی:
AAIEH02_066
تاریخ نمایه سازی: 22 شهریور 1405
چکیده مقاله:
Recent progress in Neural Radiance Fields (NeRFs) has substantially improved ۳D scene reconstruction and novel view generation by modeling environments as continuous volumetric functions learned from multi-view imagery. Despite these advances, NeRFs still face persistent challenges—most notably their reliance on dense camera coverage, limited global scene understanding, and considerable computational overhead. In contrast, Vision Transformers (ViTs) excel at capturing long-range dependencies and holistic image context using patch-based token representations. This work proposes a unified framework that brings together NeRF and ViT-based feature extraction to enhance the robustness and fidelity of ۳D reconstruction. By embedding global transformer-derived features into the radiance field, the system is better equipped to reason about geometry and appearance under difficult conditions such as sparse viewpoints, occlusions, noise, or disordered image input. The paper first revisits foundational concepts in NeRFs and Vision Transformers, then reviews influential developments in both neural rendering and attention-based architectures. Building on this background, we introduce a hybrid model composed of a per-view ViT encoder, a multi-view transformer for cross-view aggregation, and a NeRF-inspired MLP that predicts density and color using both positional information and aggregated tokens. We outline a training strategy based on photometric, depth, and regularization losses, and propose an evaluation protocol using established benchmarks. We conclude with a discussion of advantages, limitations, and future directions, including potential integrations with ۳D Gaussian Splatting and applications in autonomous robotics.
کلیدواژه ها:
نویسندگان
Yeganeh Ahmadi
BSc Student of Computer Software Engineering, Islamic Azad University, Najafabad Branch, Isfahan, Iran.
Ehsan Narimani
PhD in Computer Software Engineering, Islamic Azad University, Najafabad Branch, Isfahan, Iran.
Nasim Ahmadi
PhD in Computer Science (Soft Computing and Artificial Intelligence), University of Tehran Kish International Campus, Tehran, Iran.
Aram Ghalavand
MSc in Computer Software Engineering, Islamic Azad University, Khorramabad Branch, Khorramabad, Iran.