Sciweavers

AIRS
2008
Springer

On the Construction of a Large Scale Chinese Web Test Collection

13 years 11 months ago
On the Construction of a Large Scale Chinese Web Test Collection
The lack of a large scale Chinese test collection is an obstacle to the Chinese information retrieval development. In order to address this issue, we built such a collection composed of millions of Chinese web pages, known as the Chinese Web Test collection with 100 gigabyte (CWT100g) in data volume, which is the largest Chinese web test collection as of this writing, and has been used by several dozen research groups besides being adopted in the evaluation of the SEWM-2004 Chinese Web Track[1] and the HTRDPE-2004[2]. We present the total solution for constructing a large scale test collection like the CWT100g. Further, we found that: 1) the distribution of the number of pages within sites obeys a Zipf-like law instead of a power law proposed by Adamic and Huberman [3, 4]; 2) and an appropriate filtering method on host alias will economize resources for about 25% while crawling pages. The Zipf-like law and the method of filtering host alias proposed in the paper will facilitate both to...
Hongfei Yan, Chong Chen, Bo Peng, Xiaoming Li
Added 01 Jun 2010
Updated 01 Jun 2010
Type Conference
Year 2008
Where AIRS
Authors Hongfei Yan, Chong Chen, Bo Peng, Xiaoming Li
Comments (0)