Article
Roller: A novel approach to web information extraction
Author/s | Jiménez Aguirre, Patricia
Corchuelo Gil, Rafael |
Department | Universidad de Sevilla. Departamento de Lenguajes y Sistemas Informáticos |
Publication Date | 2016 |
Deposit Date | 2022-04-08 |
Published in |
|
Abstract | The research regarding web information extraction focuses
on learning rules to extract some selected information from web documents.
Many proposals are ad-hoc and cannot benefit from the advances
in machine learning; ... The research regarding web information extraction focuses on learning rules to extract some selected information from web documents. Many proposals are ad-hoc and cannot benefit from the advances in machine learning; furthermore, they are likely to fade away as theWeb evolves and their intrinsic assumptions are not satisfied. Some authors have explored transforming web documents into relational data and then using techniques that got inspiration from inductive logic programming. In theory, such proposals should be easier to adapt as the Web evolves because they build on catalogues of features that can be adapted without changing the proposals themselves. Unfortunately, they are difficult to scale as the number of documents or features increases. In the general field of machine learning, there are propositio-relational proposals that attempt to provide effective and efficient means to learn from relational data using propositional techniques, but they have seldom been explored regarding web information extraction. In this article, we present a new proposal called Roller: it relies on a search procedure that uses a dynamic flattening technique to explore the context of the nodes that provide the information to be extracted; it is configured with an open catalogue of features, so that it can adapt to the evolution of the Web; it also requires a base learner and a rule scorer, which helps it benefit from the continuous advances in machine learning. Our experiments confirm that it outperforms other state-of-the-art proposals in terms of effectiveness and that it is very competitive in terms of efficiency; we have also confirmed that our conclusions are solid from a statistical point of view. |
Funding agencies | Ministerio de Educación y Ciencia (MEC). España Junta de Andalucía Ministerio de Ciencia e Innovación (MICIN). España Ministerio de Economia, Industria y Competitividad (MINECO). España Ministerio de Economía y Competitividad (MINECO). España |
Project ID. | TIN2007-64119
P07-TIC-2602 P08-TIC-4100 TIN2008-04718-E TIN2010-21744 TIN2010-09809-E TIN2010-10811-E TIN2010-09988-E TIN2011-15497-E TIN2013-40848-R |
Citation | Jiménez Aguirre, P. y Corchuelo Gil, R. (2016). Roller: A novel approach to web information extraction. Knowledge and Information Systems, 49 (1), 197-241. |
Files | Size | Format | View | Description |
---|---|---|---|---|
ROLLER A novel approach to web ... | 1.049Mb | [PDF] | View/ | |