Improving reference mining in patents with BERT

01/04/2021
by   Ken Voskuil, et al.
0

References in patents to scientific literature provide relevant information for studying the relation between science and technological inventions. These references allow us to answer questions about the types of scientific work that leads to inventions. Most prior work analysing the citations between patents and scientific publications focussed on the front-page citations, which are well structured and provided in the metadata of patent archives such as Google Patents. In the 2019 paper by Verberne et al., the authors evaluate two sequence labelling methods for extracting references from patents: Conditional Random Fields (CRF) and Flair. In this paper we extend that work, by (1) improving the quality of the training data and (2) applying BERT-based models to the problem. We use error analysis throughout our work to find problems in the dataset, improve our models and reason about the types of errors different models are susceptible to. We first discuss the work by Verberne et al. and other related work in Section2. We describe the improvements we make in the dataset, and the new models proposed for this task. We compare the results of our new models with previous results, both on the labelled dataset and a larger unlabelled corpus. We end with a discussion on the characteristics of the results of our new models, followed by a conclusion. Our code and improved dataset are released under an open-source license on github.

READ FULL TEXT

page 7

page 8

research
03/16/2021

Two tales of science technology linkage: Patent in-text versus front-page references

There is recurrent debate about how useful science is for technological ...
research
03/05/2023

Are self-citations a normal feature of knowledge accumulation?

Science is a cumulative activity, which can manifest itself through the ...
research
05/29/2023

Do Language Models Know When They're Hallucinating References?

Current state-of-the-art language models (LMs) are notorious for generat...
research
12/21/2020

A qualitative and quantitative analysis of open citations to retracted articles: the Wakefield et al.'s case

In this article, we show the results of a quantitative and qualitative a...
research
09/19/2022

Relationship between incidence of breathing obstruction and degree of muzzle shortness in pedigree dogs

There has been much concern about health issues associated with the bree...
research
02/07/2023

The Effect of Metadata on Scientific Literature Tagging: A Cross-Field Cross-Model Study

Due to the exponential growth of scientific publications on the Web, the...
research
08/07/2023

What has ChatGPT read? The origins of archaeological citations used by a generative artificial intelligence application

The public release of ChatGPT has resulted in considerable publicity and...

Please sign up or login with your details

Forgot password? Click here to reset