Sciweavers

SMC
2010
IEEE

Semantic enrichment of text representation with wikipedia for text classification

13 years 3 months ago
Semantic enrichment of text representation with wikipedia for text classification
—Text classification is a widely studied topic in the area of machine learning. A number of techniques have been developed to represent and classify text documents. Most of the techniques try to achieve good classification performance while taking a document only by its words (e.g. statistical analysis on word frequency and distribution patterns). One of the recent trends in text classification research is to incorporate more semantic interpretation in text classification, especially by using Wikipedia. This paper introduces a technique for incorporating the vast amount of human knowledge accumulated in Wikipedia into text representation and classification. The aim is to improve classification performance by transforming general terms into a set of related concepts grouped around semantic themes. In order to achieve this goal, this paper proposes a unique method for breaking the enormous amount of extracted Wikipedia knowledge (concepts) into smaller pieces (subsets of concepts). The...
Hiroki Yamakawa, Jing Peng, Anna Feldman
Added 30 Jan 2011
Updated 30 Jan 2011
Type Journal
Year 2010
Where SMC
Authors Hiroki Yamakawa, Jing Peng, Anna Feldman
Comments (0)