SoMeSci- A 5 Star Open Data Gold Standard Knowledge Graph of Software Mentions in Scientific Articles

by   David Schindler, et al.

Knowledge about software used in scientific investigations is important for several reasons, for instance, to enable an understanding of provenance and methods involved in data handling. However, software is usually not formally cited, but rather mentioned informally within the scholarly description of the investigation, raising the need for automatic information extraction and disambiguation. Given the lack of reliable ground truth data, we present SoMeSci (Software Mentions in Science) a gold standard knowledge graph of software mentions in scientific articles. It contains high quality annotations (IRR: κ=.82) of 3756 software mentions in 1367 PubMed Central articles. Besides the plain mention of the software, we also provide relation labels for additional information, such as the version, the developer, a URL or citations. Moreover, we distinguish between different types, such as application, plugin or programming environment, as well as different types of mentions, such as usage or creation. To the best of our knowledge, SoMeSci is the most comprehensive corpus about software mentions in scientific articles, providing training samples for Named Entity Recognition, Relation Extraction, Entity Disambiguation, and Entity Linking. Finally, we sketch potential use cases and provide baseline results.


page 1

page 2

page 3

page 4


Investigating Software Usage in the Social Sciences: A Knowledge Graph Approach

Knowledge about the software used in scientific investigations is necess...

MORTY: Structured Summarization for Targeted Information Extraction from Scholarly Articles

Information extraction from scholarly articles is a challenging task due...

Automatic Generation of Benchmarks for Entity Recognition and Linking

The velocity dimension of Big Data plays an increasingly important role ...

Constructing a Knowledge Graph from Textual Descriptions of Software Vulnerabilities in the National Vulnerability Database

Knowledge graphs have shown promise for several cybersecurity tasks, suc...

Slot Filling for Extracting Reskilling and Upskilling Options from the Web

Disturbances in the job market such as advances in science and technolog...

KGEA: A Knowledge Graph Enhanced Article Quality Identification Dataset

With so many articles of varying quality being produced at every moment,...

Better Call the Plumber: Orchestrating Dynamic Information Extraction Pipelines

In the last decade, a large number of Knowledge Graph (KG) information e...

Please sign up or login with your details

Forgot password? Click here to reset