Overview of STEM Science as Process, Method, Material, and Data Named Entities

05/24/2022
by   Jennifer D'Souza, et al.
0

We are faced with an unprecedented production in scholarly publications worldwide. Stakeholders in the digital libraries posit that the document-based publishing paradigm has reached the limits of adequacy. Instead, structured, machine-interpretable, fine-grained scholarly knowledge publishing as Knowledge Graphs (KG) is strongly advocated. In this work, we develop and analyze a large-scale structured dataset of STEM articles across 10 different disciplines, viz. Agriculture, Astronomy, Biology, Chemistry, Computer Science, Earth Science, Engineering, Material Science, Mathematics, and Medicine. Our analysis is defined over a large-scale corpus comprising 60K abstracts structured as four scientific entities process, method, material, and data. Thus our study presents, for the first-time, an analysis of a large-scale multidisciplinary corpus under the construct of four named entity labels that are specifically defined and selected to be domain-independent as opposed to domain-specific. The work is then inadvertently a feasibility test of characterizing multidisciplinary science with domain-independent concepts. Further, to summarize the distinct facets of scientific knowledge per concept per discipline, a set of word cloud visualizations are offered. The STEM-NER-60k corpus, created in this work, comprises over 1M extracted entities from 60k STEM articles obtained from a major publishing platform and is publicly released https://github.com/jd-coderepos/stem-ner-60k.

READ FULL TEXT

page 7

page 8

page 9

page 10

research
03/28/2022

Computer Science Named Entity Recognition in the Open Research Knowledge Graph

Domain-specific named entity recognition (NER) on Computer Science (CS) ...
research
10/14/2022

Self-Adaptive Named Entity Recognition by Retrieving Unstructured Knowledge

Although named entity recognition (NER) helps us to extract various doma...
research
07/08/2022

Lessons from Deep Learning applied to Scholarly Information Extraction: What Works, What Doesn't, and Future Directions

Understanding key insights from full-text scholarly articles is essentia...
research
03/09/2023

Position Paper on Dataset Engineering to Accelerate Science

Data is a critical element in any discovery process. In the last decades...
research
08/30/2022

Large-scale Multi-granular Concept Extraction Based on Machine Reading Comprehension

The concepts in knowledge graphs (KGs) enable machines to understand nat...
research
09/10/2021

WikiCSSH: Extracting and Evaluating Computer Science Subject Headings from Wikipedia

Hierarchical domain-specific classification schemas (or subject heading ...

Please sign up or login with your details

Forgot password? Click here to reset