Sciweavers

TSD
2009
Springer

Improving the Clustering of Blogosphere with a Self-term Enriching Technique

13 years 11 months ago
Improving the Clustering of Blogosphere with a Self-term Enriching Technique
The analysis of blogs is emerging as an exciting new area in the text processing field which attempts to harness and exploit the vast quantity of information being published by individuals. However, their particular characteristics (shortness, vocabulary size and nature, etc.) make it difficult to achieve good results using automated clustering techniques. Moreover, the fact that many blogs may be considered to be narrow domain means that exploiting external linguistic resources can have limited value. In this paper, we present a methodology to improve the performance of clustering techniques on blogs, which does not rely on external resources. Our results show that this technique can produce significant improvements in the quality of clusters produced.
Fernando Perez-Tellez, David Pinto, John Cardiff,
Added 27 May 2010
Updated 27 May 2010
Type Conference
Year 2009
Where TSD
Authors Fernando Perez-Tellez, David Pinto, John Cardiff, Paolo Rosso
Comments (0)