Wikicorpus: A Word-Sense Disambiguated Multilingual Wikipedia Corpus

15 years 6 months ago

Download www.lsi.upc.edu

This article presents a new freely available trilingual corpus (Catalan, Spanish, English) that contains large portions of the Wikipedia and has been automatically enriched with linguistic information. To our knowledge, this is the largest such corpus that is freely available to the community: In its present version, it contains over 750 million words. The corpora have been annotated with lemma and part of speech information using the open source library FreeLing. Also, they have been sense annotated with the state of the art Word Sense Disambiguation algorithm UKB. As UKB assigns WordNet senses, and WordNet has been aligned across languages via the InterLingual Index, this sort of annotation opens the way to massive explorations in lexical semantics that were not possible before. We present a first attempt at creating a trilingual lexical resource from the sense-tagged Wikipedia corpora, namely, WikiNet. Moreover, we present two by-products of the project that are of use for the NLP ...

Samuel Reese, Gemma Boleda, Montse Cuadros, Llu&ia

Real-time Traffic

Algorithm Ukb | Available Trilingual Corpus | Disambiguation Algorithm Ukb | Education | LREC 2010 |

claim paper

» Projecting Parameters for Multilingual Word Sense Disambiguation

» Construction of a Benchmark Data Set for Crosslingual Word Sense Disambiguation

» Evaluating the Impact of Some Linguistic Information on the Performances of a Similarityba...

» Overview of the Clef 2008 Multilingual Question Answering Track

» WikiTranslate Query Translation for CrossLingual Information Retrieval Using Only Wikipedi...

» Mapping Lexical Entries in a Verbs Database to WordNet Senses

Post Info
More Details (n/a)

Added	29 Oct 2010
Updated	29 Oct 2010
Type	Conference
Year	2010
Where	LREC
Authors	Samuel Reese, Gemma Boleda, Montse Cuadros, Lluís Padró, German Rigau

Comments (0)

Sciweavers

Wikicorpus: A Word-Sense Disambiguated Multilingual Wikipedia Corpus

Algorithm Ukb | Available Trilingual Corpus | Disambiguation Algorithm Ukb | Education | LREC 2010 |

Explore & Download

Productivity Tools

Sciweavers