Sciweavers

ACL
2009

Employing Topic Models for Pattern-based Semantic Class Discovery

13 years 2 months ago
Employing Topic Models for Pattern-based Semantic Class Discovery
A semantic class is a collection of items (words or phrases) which have semantically peer or sibling relationship. This paper studies the employment of topic models to automatically construct semantic classes, taking as the source data a collection of raw semantic classes (RASCs), which were extracted by applying predefined patterns to web pages. The primary requirement (and challenge) here is dealing with multi-membership: An item may belong to multiple semantic classes; and we need to discover as many as possible the different semantic classes the item belongs to. To adopt topic models, we treat RASCs as "documents", items as "words", and the final semantic classes as "topics". Appropriate preprocessing and postprocessing are performed to improve results quality, to reduce computation cost, and to tackle the fixed-k constraint of a typical topic model. Experiments conducted on 40 million web pages show that our approach could yield better results than a...
Huibin Zhang, Mingjie Zhu, Shuming Shi, Ji-Rong We
Added 16 Feb 2011
Updated 16 Feb 2011
Type Journal
Year 2009
Where ACL
Authors Huibin Zhang, Mingjie Zhu, Shuming Shi, Ji-Rong Wen
Comments (0)