Unlocking Book Genre from Covers: A Multimodal Approach to Book Genre Prediction

سال انتشار: 1404
نوع سند: مقاله ژورنالی
زبان: انگلیسی
مشاهده: 182

فایل این مقاله در 9 صفحه با فرمت PDF قابل دریافت می باشد

استخراج به نرم افزارهای پژوهشی:

لینک ثابت به این مقاله:

شناسه ملی سند علمی:

JR_IJWR-8-3_006

تاریخ نمایه سازی: 11 شهریور 1404

چکیده مقاله:

In today’s visually driven market, book cover design plays a crucial role in conveying a work’s narrative and thematic essence. A book cover is a multimodal entity, consisting of various visual and textual elements. While conventional recommendation systems have often overlooked the semantic richness of cover imagery, prior work attempting to incorporate textual information relied on OCR to extract text from covers. However, these raw tokens capture only a fraction of the cover's meaning and often miss deeper thematic and narrative cues. Recognizing these limitations, we leverage the advanced knowledge accumulated in VLMs to derive a more comprehensive representation, using this knowledge to add it as an additional feature to the system. In this paper, we use VLM-generated descriptions and integrate these rich descriptions as a new textual feature. Our enhanced corpus comprises ۵۷,۰۰۰ book covers across ۳۰ genres (۱,۹۰۰ per genre), each annotated with both raw imagery and VLM-generated narrative summaries. We fuse two state-of-the-art vision encoders (ViT and VisionMamba) with a text encoder that processes these VLM descriptions. Experimental results demonstrate a Top ۱ accuracy of ۶۳.۳۱% and a Top ۳ accuracy of ۸۳.۰۳%, marking a substantial improvement over the state-of-the-art variant and underscoring the value of VLM-derived context in multimodal genre classification.

نویسندگان

Reza Toosi

Department of Computer Engineering, Faculty of Engineering, Golestan University, Gorgan, Iran.

Alireza Hosseini

School of Electrical and Computer Engineering, College of Engineering, University of Tehran, Tehran, Iran.

Ramin Toosi

School of Electrical and Computer Engineering, College of Engineering, University of Tehran, Tehran, Iran.

Mohammad Ali Akhaee

School of Electrical and Computer Engineering, College of Engineering, University of Tehran, Tehran, Iran.

مراجع و منابع این مقاله:

لیست زیر مراجع و منابع استفاده شده در این مقاله را نمایش می دهد. این مراجع به صورت کاملا ماشینی و بر اساس هوش مصنوعی استخراج شده اند و لذا ممکن است دارای اشکالاتی باشند که به مرور زمان دقت استخراج این محتوا افزایش می یابد. مراجعی که مقالات مربوط به آنها در سیویلیکا نمایه شده و پیدا شده اند، به خود مقاله لینک شده اند :
  • C. A. Kratz, “On Telling/Selling a Book by Its Cover,” ...
  • Oramas, O. Nieto, F. Barbieri, and X. Serra, “Multi-label music ...
  • Dorochowicz and B. Kostek, “Relationship between album cover design and ...
  • Barney and K. Kaya, “Predicting genre from movie posters”, Machine ...
  • Behrouzi, R. Toosi, and M. A. Akhaee, “Multimodal movie genre ...
  • Zeng, Y. Gong, and X. Zeng, “Controllable digital restoration of ...
  • Cetinic and S. Grgic, “Genre classification of paintings,” ۲۰۱۶ International ...
  • Agarwal, H. Karnick, N. Pant, and U. Patel, “Genre and ...
  • K. Iwana, S. T. R. Rizvi, S. Ahmed, A. Dengel, ...
  • Hatamizadeh and J. Kautz, “MambaVision: A hybrid Mamba-Transformer vision backbone”, ...
  • Gu and T. Dao, ‘Mamba: Linear-time sequence modeling with selective ...
  • M. Rahman, A. A. Tutul, A. Nath, L. Laishram, S. ...
  • Buczkowski, A. Sobkowicz, and M. Kozlowski, “Deep Learning Approaches towards ...
  • Pye, “Content-based methods for the management of digital music,” ۲۰۰۰ ...
  • Z. Afzal et al., “Deepdocclassifier: Document classification with deep Convolutional ...
  • Gupta, M. Agarwal, and S. Jain, “Automated Genre Classification of ...
  • R. Doyle and P. A. Bottomley, “Dressed for the Occasion: ...
  • W. Henderson, J. L. Giese, and J. A. Cote, “Impression ...
  • Chiang, Y. Ge, and C. Wu, “Classification of book genres ...
  • Haraguchi, B. K. Iwana, and S. Uchida, “What Text Design ...
  • H. Patel and D. Aggarwal, “BGCNet: A novel deep visual ...
  • Dosovitskiy et al., “An image is worth ۱۶x۱۶ words: Transformers ...
  • Kjartansson and A. Ashavsky, “Can you judge a book by ...
  • نمایش کامل مراجع