ProVe: A Pipeline for Automated Provenance Verification of Knowledge Graphs against Textual Sources

by   Gabriel Amaral, et al.

Knowledge Graphs are repositories of information that gather data from a multitude of domains and sources in the form of semantic triples, serving as a source of structured data for various crucial applications in the modern web landscape, from Wikipedia infoboxes to search engines. Such graphs mainly serve as secondary sources of information and depend on well-documented and verifiable provenance to ensure their trustworthiness and usability. However, their ability to systematically assess and assure the quality of this provenance, most crucially whether it properly supports the graph's information, relies mainly on manual processes that do not scale with size. ProVe aims at remedying this, consisting of a pipelined approach that automatically verifies whether a Knowledge Graph triple is supported by text extracted from its documented provenance. ProVe is intended to assist information curators and consists of four main steps involving rule-based methods and machine learning models: text extraction, triple verbalisation, sentence selection, and claim verification. ProVe is evaluated on a Wikidata dataset, achieving promising results overall and excellent performance on the binary classification task of detecting support from provenance, with 87.5 accuracy and 82.9 scripts used in this paper are available on GitHub and Figshare.


page 18

page 19

page 20

page 21

page 22


Entity Context Graph: Learning Entity Representations fromSemi-Structured Textual Sources on the Web

Knowledge is captured in the form of entities and their relationships an...

Creating Knowledge Graphs for Geographic Data on the Web

Geographic data plays an essential role in various Web, Semantic Web and...

Growing and Serving Large Open-domain Knowledge Graphs

Applications of large open-domain knowledge graphs (KGs) to real-world p...

CBIM: A Graph-based Approach to Enhance Interoperability Using Semantic Enrichment

Interoperability remains a challenge in the construction industry. In th...

Towards Knowledge Graphs Validation through Weighted Knowledge Sources

The performance of applications, such as personal assistants, search eng...

MapSDI: A Scaled-up Semantic Data Integration Framework for Knowledge Graph Creation

Semantic web technologies have significantly contributed with effective ...

Toward the Automated Construction of Probabilistic Knowledge Graphs for the Maritime Domain

International maritime crime is becoming increasingly sophisticated, oft...

Please sign up or login with your details

Forgot password? Click here to reset