株式会社極東書店トップ商品一覧Explorations in Automatic Thesaurus Discovery. Softcover reprint of the original 1st ed. 1994

商品詳細

Explorations in Automatic Thesaurus Discovery. Softcover reprint of the original 1st ed. 1994

Explorations in Automatic Thesaurus Discovery. Softcover reprint of the original 1st ed. 1994

・ISBN 978-1-4613-6167-1 paper EUR 149.99

¥40,091.- (税込) (※)価格はご注文時の参考価格となります。
納品価格につきましては書籍の入荷時点で確定となります。
版元の原価改定、外国為替の変動等により異なる場合がございますので、予めご了承下さい。

お気に入り
著者・編者Grefenstette, Gregory,
シリーズ (The Springer International Series in Engineering and Computer Science)
出版社 (Springer-Verlag New York Inc., US)
出版年月2012
ページ数305 pp.
言語ENG
ニュース番号<A04-84159>

解説

Explorations in Automatic Thesaurus Discovery presents an automated method for creating a first-draft thesaurus from raw text. It describes natural processing steps of tokenization, surface syntactic analysis, and syntactic attribute extraction. From these attributes, word and term similarity is calculated and a thesaurus is created showing important common terms and their relation to each other, common verb--noun pairings, common expressions, and word family members.
The techniques are tested on twenty different corpora ranging from baseball newsgroups, assassination archives, medical X-ray reports, abstracts on AIDS, to encyclopedia articles on animals, even on the text of the book itself. The corpora range from 40,000 to 6 million characters of text, and results are presented for each in the Appendix.
The methods described in the book have undergone extensive evaluation. Their time and space complexity are shown to be modest. The results are shown to converge to a stable state as the corpus grows. The similarities calculated are compared to those produced by psychological testing. A method of evaluation using Artificial Synonyms is tested. Gold Standards evaluation show that techniques significantly outperform non-linguistic-based techniques for the most important words in corpora.
Explorations in Automatic Thesaurus Discovery includes applications to the fields of information retrieval using established testbeds, existing thesaural enrichment, semantic analysis. Also included are applications showing how to create, implement, and test a first-draft thesaurus.