Parameter Efficiency, Not Scale, Is the True Measure of Intelligence

فایل این در 25 صفحه با فرمت PDF قابل دریافت می باشد

  • من نویسنده این مقاله هستم

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این :

چکیده :

Abstract The field of artificial intelligence has spent the last several years in a parameter arms race. The tacit assumption—that raw parameter count is a reliable proxy for intelligence—has shaped funding, publication incentives, and public perception. This paper argues that the assumption is both scientifically misleading and practically costly. Intelligence per parameter (IPP), the ratio of capability and generalization to model size (particularly active parameters), offers a more rigorous measure of progress. We develop a framework centered on Pareto efficiency in the performance-parameter space, formalize IPP through explicit mathematical definitions, and show how it can be operationalized with existing benchmark suites. Recent empirical developments through mid-2026, including the Chinchilla scaling laws, the Llama 3 family, DeepSeek’s sparse architectures, the rise of strong sub-10-billion-parameter models, and the July 2026 release of Moonshot’s Kimi K3 (2.8 trillion total parameters with only 104 billion active), demonstrate that the efficiency frontier has shifted sharply leftward. Smaller or sparsely activated systems now match or exceed models with far higher nominal parameter counts on many tasks. We examine the architectural changes driving this shift—Multi-Latent Attention, Mixture-of-Experts routing, extreme quantization, reasoning distillation, and newer mechanisms such as Kimi Delta Attention—and ground the argument in scaling theory and algorithmic information theory. The paper ends with a concrete proposal: treat parameter count as a cost, not a credential, and redesign leaderboards to reward efficiency as aggressively as absolute performance.

نویسندگان

مراجع و منابع این :

لیست زیر مراجع و منابع استفاده شده در این را نمایش می دهد. این مراجع به صورت کاملا ماشینی و بر اساس هوش مصنوعی استخراج شده اند و لذا ممکن است دارای اشکالاتی باشند که به مرور زمان دقت استخراج این محتوا افزایش می یابد. مراجعی که مقالات مربوط به آنها در سیویلیکا نمایه شده و پیدا شده اند، به خود لینک شده اند :
  • Abdin, M., Aneja, J., Behl, H., et al. (2024). Phi-4 ...
  • Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & ...
  • DeepSeek-AI. (2024). DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language ...
  • DeepSeek-AI. (2025). DeepSeek-V3 technical report. arXiv preprint arXiv:2412.19437. ...
  • DeepSeek-AI. (2025). DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement ...
  • Dettmers, T., Lewis, M., Belkada, Y., & Zettlemoyer, L. (2022). ...
  • Dubey, A., Jauhri, A., Pandey, A., et al. (2024). The ...
  • Fedus, W., Zoph, B., & Shazeer, N. (2022). Switch transformers: ...
  • Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2023). ...
  • Frantar, E., & Alistarh, D. (2023). SparseGPT: Massive language models ...
  • Frankle, J., & Carbin, M. (2019). The lottery ticket hypothesis: ...
  • Frankle, J., Dziugaite, G. K., Roy, D., & Carbin, M. ...
  • Gemma Team. (2024). Gemma 2: Improving open language models at ...
  • Gu, A., & Dao, T. (2023). Mamba: Linear-time sequence modeling ...
  • Gunasekar, S., Zhang, Y., Aneja, J., et al. (2023). Textbooks ...
  • Han, S., Mao, H., & Dally, W. (2016). Deep compression: ...
  • Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the ...
  • Hoffmann, J., Borgeaud, S., Mensch, A., et al. (2022). Training ...
  • Hu, E. J., Shen, Y., Wallis, P., et al. (2022). ...
  • Jiang, A. Q., Sablayrolles, A., Mensch, A., et al. (2024). ...
  • Kaplan, J., McCandlish, S., Henighan, T., et al. (2020). Scaling ...
  • Li, Y., Bubeck, S., Eldan, R., et al. (2023). Textbooks ...
  • Lin, J., Tang, J., Tang, H., et al. (2023). AWQ: ...
  • Ma, X., Fang, G., & Wang, X. (2024). BitNet b1.58: ...
  • Moonshot AI. (2026). Kimi K3 technical report and model card ...
  • Qwen Team. (2024). Qwen2.5 technical report. arXiv preprint arXiv:2412.15115. ...
  • Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2020). ...
  • Schaeffer, R., Miranda, B., & Koyejo, S. (2023). Are emergent ...
  • Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and ...
  • Touvron, H., Lavril, T., Izacard, G., et al. (2023). LLaMA: ...
  • Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention ...
  • نمایش کامل مراجع