Skip to main navigation Skip to search Skip to main content

Can one deep learning model learn script-independent multilingual word-spotting?

Research output: Chapter in BookChapterResearchpeer-review

Abstract

Word spotting has gained increased attention lately as it can be used to extract textual information from handwritten documents and scene-text images. Current word spotting approaches are designed to work on a single language and/or script. Building intelligent models that learn script-independent multilingual word-spotting is challenging due to the large variability of multilingual alphabets and symbols. We used ResNet-152 and the Pyramidal Histogram of Characters (PHOC) embedding to build a one-model script-independent multilingual word-spotting and we tested it on Latin, Arabic, and Bangla (Indian) languages. The one-model we propose performs on par with the multi-model language-specific word-spotting system, and thus, reduces the number of models needed for each script and/or language.

Original languageEnglish
Title of host publicationProceedings - 15th IAPR International Conference on Document Analysis and Recognition, ICDAR 2019
Pages260-267
Number of pages8
ISBN (Electronic)9781728128610
DOIs
Publication statusPublished - Sept 2019

Publication series

NameProceedings of the International Conference on Document Analysis and Recognition, ICDAR
ISSN (Print)1520-5363

Keywords

  • Handwriting
  • Histogram of Characters
  • Multitasking
  • PHOC
  • ResNet
  • Scene text images

Fingerprint

Dive into the research topics of 'Can one deep learning model learn script-independent multilingual word-spotting?'. Together they form a unique fingerprint.

Cite this