Sciweavers

Free Online Productivity Tools i2Speak i2Symbol i2OCR iTex2Img iWeb2Print iWeb2Shot i2Type iPdf2Split iPdf2Merge i2Bopomofo i2Arabic i2Style i2Image i2PDF iLatex2Rtf Sci2ools

141

ACL
2003

89views Computational Linguistics» more ACL 2003»

Parametric Models of Linguistic Count Data

15 years 7 months ago

Parametric Models of Linguistic Count Data

Download acl.ldc.upenn.edu

It is well known that occurrence counts of words in documents are often modeled poorly by standard distributions like the binomial or Poisson. Observed counts vary more than simple models predict, prompting the use of overdispersed models like Gamma-Poisson or Beta-binomial mixtures as robust alternatives. Another deﬁciency of standard models is due to the fact that most words never occur in a given document, resulting in large amounts of zero counts. We propose using zeroinﬂated models for dealing with this, and evaluate competing models on a Naive Bayes text classiﬁcation task. Simple zero-inﬂated models can account for practically relevant variation, and can be easier to work with than overdispersed models.

Martin Jansche

Real-time Traffic

ACL 2003 | ACL 2007 | Overdispersed Models | Simple Models | Simple Zero-inﬂated Models |

claim paper

Related Content

» Parametric and nonparametric Bayesian model specification A case study involving models fo...

» Comparison of Poisson Mixture Models for Count Data Clusterization

» On the Parametrization of Clapping

» A Logistic Regression Model of Determiner Omission in PPs

» Connectionist Modeling of Linguistic Quantifiers

» Word usage and posting behaviors modeling blogs with unobtrusive data collection methods

» Twoway translation of compound sentences and arm motions by recurrent neural networks

» Language identification of names with SVMs

» Training Phrase Translation Models with LeavingOneOut

Post Info
More Details (n/a)

Added	31 Oct 2010
Updated	31 Oct 2010
Type	Conference
Year	2003
Where	ACL
Authors	Martin Jansche

Comments (0)