株式会社極東書店トップ商品一覧Recent Advances in Multimodal Hallucination.

商品詳細

Recent Advances in Multimodal Hallucination.

Recent Advances in Multimodal Hallucination.

・ISBN 978-3-032-36320-6 hard EUR 49.99

¥13,361.- (税込) (※)価格はご注文時の参考価格となります。
納品価格につきましては書籍の入荷時点で確定となります。
版元の原価改定、外国為替の変動等により異なる場合がございますので、予めご了承下さい。

お気に入り
著者・編者Jing, Liqiang / Zhang, Yue / Du, Xinya,
出版社 (Springer Nature Switzerland AG, SZ)
出版年月2026
ページ数109 pp.
言語ENG
ニュース番号<A05-86352>

解説

Hallucination in Multimodal Models is the first comprehensive research monograph dedicated to the growing challenge of hallucinations in large-scale multimodal AI systems, particularly vision-language models (VLMs) and multimodal large language models (MLLMs). The book systematically defines, categorizes, evaluates, and mitigates hallucinations - cases where models generate content that is factually inconsistent, visually unsupported, or commonsensically implausible. These hallucinations have become increasingly problematic in real-world applications of AI, including robotics, autonomous systems, and AI-generated media, where the consequences of inaccurate outputs can be severe.

The purpose of this book is threefold:

(1) to formalize the types and causes of hallucination in multimodal models;

(2) to present state-of-the-art evaluation frameworks, such as FaithScore and FIHA, for quantifying hallucination at a fine-grained level; and

(3) to introduce a unified mitigation framework.

The book presents recent research results, including FaithScore, FIHA, FIFA, FGAIF, and Dentist, together with empirical studies across leading large vision-language models. Of particular interest are its fine-grained approaches to atomic fact verification, semantic dependency modeling, unified text-video hallucination evaluation, reward-based alignment, and training-free hallucination mitigation. The book also examines how these methods can improve the faithfulness and reliability of multimodal model outputs, offering practical tools for both researchers and engineers.

This book complements and extends the existing literature on multimodal model evaluation (e.g., MME, SEED-Bench, LAMM) by moving beyond surface-level metrics to offer deeper, interpretable, and automated hallucination analysis. Unlike survey papers or benchmarks that only diagnose the problem, our monograph provides a cohesive solution path from diagnosis to mitigation, built upon novel technical contributions and real-world implementations.

As this is the first edition, it introduces original theoretical frameworks, algorithms, benchmarks, and design paradigms. It is intended to serve as a reference for graduate students, academic researchers, and industry practitioners working in natural language processing, computer vision, embodied AI, and trustworthy AI systems.