Skip to main navigation Skip to search Skip to main content

Semantics-based content extraction in typewritten historical documents

A. Antonacopoulos*, D. Karatzas

*Corresponding author for this work

Research output: Chapter in BookChapterResearchpeer-review

Abstract

This paper presents a flexible approach to extracting content from scanned historical documents using semantic information. The final electronic document is the result of a "digital historical document lifecycle" process, where the expert knowledge of the historian/archivist user is incorporated at different stages. Results show that such a conversion strategy aided by (expert) user-specified semantic information and which enables the processing of individual parts of the document in a specialised way, produces superior (in a variety of significant ways) results than document analysis and understanding techniques devised for contemporary documents.

Original languageEnglish
Title of host publicationProceedings of the Eighth International Conference on Document Analysis and Recognition
Pages48-53
Number of pages6
DOIs
Publication statusPublished - 2005

Publication series

NameProceedings of the International Conference on Document Analysis and Recognition, ICDAR
Volume2005
ISSN (Print)1520-5363

Fingerprint

Dive into the research topics of 'Semantics-based content extraction in typewritten historical documents'. Together they form a unique fingerprint.

Cite this