Sciweavers

ECIR
2010
Springer

Text Clustering for Peer-to-Peer Networks with Probabilistic Guarantees

13 years 6 months ago
Text Clustering for Peer-to-Peer Networks with Probabilistic Guarantees
Text clustering is an established technique for improving quality in information retrieval, for both centralized and distributed environments. However, for highly distributed environments, such as peer-topeer networks, current clustering algorithms fail to scale. Our algorithm for peer-to-peer clustering achieves high scalability by using a probabilistic approach for assigning documents to clusters. It enables a peer to compare each of its documents only with very few selected clusters, without significant loss of clustering quality. The algorithm offers probabilistic guarantees for the correctness of each document assignment to a cluster. Extensive experimental evaluation with up to 100000 peers and 1 million documents demonstrates the scalability and effectiveness of the algorithm.
Odysseas Papapetrou, Wolf Siberski, Norbert Fuhr
Added 29 Oct 2010
Updated 29 Oct 2010
Type Conference
Year 2010
Where ECIR
Authors Odysseas Papapetrou, Wolf Siberski, Norbert Fuhr
Comments (0)