Sciweavers

DGO
2010

Digital sustainable publication of legacy parliamentary proceedings

13 years 5 months ago
Digital sustainable publication of legacy parliamentary proceedings
We address the problem of publishing parliamentary proceedings in a digital sustainable manner. We give an extensive requirements analysis, and based on that propose a uniform XML format. We evaluated our approach by collecting and automatically processing proceedings from six parliaments spanning almost 200 years in total. Most of this data is real legacy data consisting of scanned and OCRed documents. The approach scales very well and produces high quality data. All documents are transformed into UTF-8 encoded XML files with extensive metadata in Dublin Core standard. The text itself is divided into pages which are divided into paragraphs. Every document, page and paragraph has a unique URN which resolves to a web page. Every page element in the XML files is connected to a facsimile image of that page in PDF or JPEG format. We created a viewer in which both versions can be inspected simultaneously. A search-engine for the complete collection is available online. Categories and Subje...
Maarten Marx, Nelleke Aders, Anne Schuth
Added 29 Oct 2010
Updated 29 Oct 2010
Type Conference
Year 2010
Where DGO
Authors Maarten Marx, Nelleke Aders, Anne Schuth
Comments (0)