Sciweavers

Free Online Productivity Tools i2Speak i2Symbol i2OCR iTex2Img iWeb2Print iWeb2Shot i2Type iPdf2Split iPdf2Merge i2Bopomofo i2Arabic i2Style i2Image i2PDF iLatex2Rtf Sci2ools

197

EMNLP
2006

73views Natural Language Processing» more EMNLP 2006»

Learning Field Compatibilities to Extract Database Records from Unstructured Text

15 years 9 months ago

Learning Field Compatibilities to Extract Database Records from Unstructured Text

Download www.cs.umass.edu

Named-entity recognition systems extract entities such as people, organizations, and locations from unstructured text. Rather than extract these mentions in isolation, this paper presents a record extraction system that assembles mentions into records (i.e. database tuples). We construct a probabilistic model of the compatibility between field values, then employ graph partitioning algorithms to cluster fields into cohesive records. We also investigate compatibility functions over sets of fields, rather than simply pairs of fields, to examine how higher representational power can impact performance. We apply our techniques to the task of extracting contact records from faculty and student homepages, demonstrating a 53% error reduction over baseline approaches.

Michael L. Wick, Aron Culotta, Andrew McCallum

Real-time Traffic

Cohesive Records | EMNLP 2006 | EMNLP 2007 | Graph Partitioning Algorithms | Systems Extract Entities |

claim paper

Related Content

» Extracting Data Records from Unstructured Biomedical Full Text

» Generalized Expectation Criteria for Bootstrapping Extractors using RecordText Alignment

» Integrating Unstructured Data into Relational Databases

» Segmentation of Publication Records of Authors from the Web

» Information Extraction

» Automatic Segmentation of Text into Structured Records

» Feature Extraction for Massive Data Mining

» From Field Notes towards a Knowledge Base

» Clinical and financial outcomes analysis with existing hospital patient records

Post Info
More Details (n/a)

Added	30 Oct 2010
Updated	30 Oct 2010
Type	Conference
Year	2006
Where	EMNLP
Authors	Michael L. Wick, Aron Culotta, Andrew McCallum

Comments (0)