The Fragile Grammar of Latent Trust

سال انتشار: 1405
نوع سند: مقاله کنفرانسی
زبان: انگلیسی
مشاهده: 110

فایل این مقاله در 47 صفحه با فرمت PDF قابل دریافت می باشد

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

DMECONF11_078

تاریخ نمایه سازی: 26 شهریور 1405

چکیده مقاله:

Large language models (LLMs) have become core components of contemporary artificial intelligence systems, enabling capabilities in reasoning, code synthesis, information retrieval, and autonomous tool utilization. However, their deployment in safety-critical and adversarial settings has introduced a broad and evolving attack surface spanning training, inference, and system integration stages. This paper presents a structured survey of the security landscape of LLMS, consolidating recent developments in attack methodologies, vulnerability classes, and defensive strategies. Threats are systematically categorized into model-centric and system-centric vectors, including model extraction, training data poisoning, back door insertion, membership inference, model inversion, adversarial prompt injection, jailbreak techniques, and inference-time manipulation in agentic workflows. Emerging risks associated with retrieval-augmented generation pipelines, indirect prompt injection, supply-chain compromise of model components, and automated adversarial exploitation are also examined. The analysis further highlights structural vulnerabilities arising from the dual role of natural language as both input data and executable instruction, which undermines conventional security boundaries. Defensive mechanisms are reviewed, encompassing alignment methodologies, adversarial and robust fine-tuning, input-output filtering, secure tool-use orchestration, and formal verification approaches. Finally, the paper discusses governance frameworks and evaluation benchmarks, emphasizing persistent asymmetries between rapidly advancing attack techniques and comparatively immature defensive measures, and motivating unified threat models and security-by-design principles for future LLM systems.