Sciweavers

IJMMS
2008

Ontology-based information extraction and integration from heterogeneous data sources

13 years 4 months ago
Ontology-based information extraction and integration from heterogeneous data sources
In this paper we present the design, implementation and evaluation of SOBA, a system for ontology-based information extraction from heterogeneous data resources, including plain text, tables and image captions. SOBA is capable of processing structured information, text and image captions to extract information and integrate it into a coherent knowledge base. To establish coherence, SOBA interlinks the information extracted from different sources and detects duplicate information. The knowledge base produced by SOBA can then be used to query for information contained in the different sources in an integrated and seamless manner. Overall, this allows for advanced retrieval functionality by which questions can be answered precisely. A further distinguishing feature of the SOBA system is that it straightforwardly integrates deep and shallow natural language processing to increase robustness and accuracy. We discuss the implementation and application of the SOBA system within the SmartWeb ...
Paul Buitelaar, Philipp Cimiano, Anette Frank, Mat
Added 12 Dec 2010
Updated 12 Dec 2010
Type Journal
Year 2008
Where IJMMS
Authors Paul Buitelaar, Philipp Cimiano, Anette Frank, Matthias Hartung, Stefania Racioppa
Comments (0)