Sciweavers

AAAI
2015

Relating Romanized Comments to News Articles by Inferring Multi-Glyphic Topical Correspondence

8 years 1 months ago
Relating Romanized Comments to News Articles by Inferring Multi-Glyphic Topical Correspondence
Commenting is a popular facility provided by news sites. Analyzing such user-generated content has recently attracted research interest. However, in multilingual societies such as India, analyzing such user-generated content is hard due to several reasons: (1) There are more than 20 official languages but linguistic resources are available mainly for Hindi. It is observed that people frequently use romanized text as it is easy and quick using an English keyboard, resulting in multiglyphic comments, where the texts are in the same language but in different scripts. Such romanized texts are almost unexplored in machine learning so far. (2) In many cases, comments are made on a specific part of the article rather than the topic of the entire article. Off-the-shelf methods such as correspondence LDA are insufficient to model such relationships between articles and comments. In this paper, we extend the notion of correspondence to model multi-lingual, multi-script, and inter-lingual top...
Goutham Tholpadi, Mrinal Kanti Das, Trapit Bansal,
Added 27 Mar 2016
Updated 27 Mar 2016
Type Journal
Year 2015
Where AAAI
Authors Goutham Tholpadi, Mrinal Kanti Das, Trapit Bansal, Chiranjib Bhattacharyya
Comments (0)