Mostrar el registro sencillo del ítem

Artículo

dc.creatorRoldán Salvador, Juan Carloses
dc.creatorJiménez Aguirre, Patriciaes
dc.creatorCorchuelo Gil, Rafaeles
dc.date.accessioned2022-04-08T07:17:04Z
dc.date.available2022-04-08T07:17:04Z
dc.date.issued2020
dc.identifier.citationRoldán Salvador, J.C., Jiménez Aguirre, P. y Corchuelo Gil, R. (2020). On Extracting Data from Tables that are Encoded using HTML. Knowledge-Based Systems, 190 (February 2020, art. nº 105157)
dc.identifier.issn0950-7051es
dc.identifier.urihttps://hdl.handle.net/11441/131963
dc.description.abstractTables are a common means to display data in human-friendly formats. Many authors have worked on proposals to extract those data back since this has many interesting applications. In this article, we summarise and compare many of the proposals to extract data from tables that are encoded using HTML and have been published between 2000 and 2018. We first present a vocabulary that homogenises the terminology used in this field; next, we use it to summarise the proposals; finally, we compare them side by side. Our analysis highlights several challenges to which no proposal provides a conclusive solution and a few more that have not been addressed sufficiently; simply put, no proposal provides a complete solution to the problem, which seems to suggest that this research field shall keep active in the near future. We have also realised that there is no consensus regarding the datasets and the methods used to evaluate the proposals, which hampers comparing the experimental results.es
dc.description.sponsorshipMinisterio de Economía y Competitividad TIN2013-40848-Res
dc.description.sponsorshipMinisterio de Economía y Competitividad TIN2016-75394-Res
dc.formatapplication/pdfes
dc.format.extent43es
dc.language.isoenges
dc.publisherElsevieres
dc.relation.ispartofKnowledge-Based Systems, 190 (February 2020, art. nº 105157)
dc.rightsAttribution-NonCommercial-NoDerivatives 4.0 Internacional*
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/4.0/*
dc.subjectHTML documentses
dc.subjectWeb tableses
dc.subjectTable mininges
dc.subjectData extractiones
dc.titleOn Extracting Data from Tables that are Encoded using HTMLes
dc.typeinfo:eu-repo/semantics/articlees
dcterms.identifierhttps://ror.org/03yxnpp24
dc.type.versioninfo:eu-repo/semantics/submittedVersiones
dc.rights.accessRightsinfo:eu-repo/semantics/openAccesses
dc.contributor.affiliationUniversidad de Sevilla. Departamento de Lenguajes y Sistemas Informáticoses
dc.relation.projectIDTIN2013-40848-Res
dc.relation.projectIDTIN2016-75394-Res
dc.relation.publisherversionhttps://www.sciencedirect.com/science/article/pii/S095070511930509X?via%3Dihubes
dc.identifier.doi10.1016/j.knosys.2019.105157es
dc.contributor.groupUniversidad de Sevilla. TIC258: Data-centric Computing Research Hubes
dc.journaltitleKnowledge-Based Systemses
dc.publication.volumen190es
dc.publication.issueFebruary 2020, art. nº 105157es
dc.identifier.sisius21888241es
dc.contributor.funderMinisterio de Economía y Competitividad (MINECO). Españaes

FicherosTamañoFormatoVerDescripción
1903.08305.pdf1.116MbIcon   [PDF] Ver/Abrir  

Este registro aparece en las siguientes colecciones

Mostrar el registro sencillo del ítem

Attribution-NonCommercial-NoDerivatives 4.0 Internacional
Excepto si se señala otra cosa, la licencia del ítem se describe como: Attribution-NonCommercial-NoDerivatives 4.0 Internacional