Exploring a Few Good Tuples from Text Databases

14 years 6 months ago

Download www1.cs.columbia.edu

Information extraction from text databases is a useful paradigm to populate relational tables and unlock the considerable value hidden in plain-text documents. However, information extraction can be expensive, due to various complex text processing steps necessary in uncovering the hidden data. There are a large number of text databases available, and not every text database is necessarily relevant to every relation. Hence, it is important to be able to quickly explore the utility of running an extractor for a specific relation over a given text database before carrying out the expensive extraction task. In this paper, we present a novel exploration methodology of finding a few good tuples for a relation that can be extracted from a database which allows for judging the relevance of the database for the relation. Specifically, we propose the notion of a good(k, ) query as one that can return any k tuples for a relation among the top- fraction of tuples ranked by their aggregated confid...

Alpa Jain, Divesh Srivastava

Real-time Traffic