Ir directamente a la navegación principal Ir directamente a la búsqueda Ir directamente al contenido principal

Scene text visual question answering

Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marcal Rusinol, C. V. Jawahar, Ernest Valveny, DImosthenis Karatzas

Producción científica: Capítulo de libroCapítuloInvestigaciónrevisión exhaustiva

Resumen

Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image. In this work, we present a new dataset, ST-VQA, that aims to highlight the importance of exploiting high-level semantic information present in images as textual cues in the Visual Question Answering process. We use this dataset to define a series of tasks of increasing difficulty for which reading the scene text in the context provided by the visual information is necessary to reason and generate an appropriate answer. We propose a new evaluation metric for these tasks to account both for reasoning errors as well as shortcomings of the text recognition module. In addition we put forward a series of baseline methods, which provide further insight to the newly released dataset, and set the scene for further research.

Idioma originalInglés
Título de la publicación alojadaProceedings - 2019 International Conference on Computer Vision, ICCV 2019
EditorialInstitute of Electrical and Electronics Engineers Inc.
Páginas4290-4300
Número de páginas11
ISBN (versión digital)9781728148038
DOI
EstadoPublicada - oct 2019

Serie de la publicación

NombreProceedings of the IEEE International Conference on Computer Vision
Volumen2019-October
ISSN (versión impresa)1550-5499

Huella

Profundice en los temas de investigación de 'Scene text visual question answering'. En conjunto forman una huella única.

Citar esto