Identifying Historical Travelogues in Large Text Corpora Using Machine Learning

by   Jan Rörden, et al.

Travelogues represent an important and intensively studied source for scholars in the humanities, as they provide insights into people, cultures, and places of the past. However, existing studies rarely utilize more than a dozen primary sources, since the human capacities of working with a large number of historical sources are naturally limited. In this paper, we define the notion of travelogue and report upon an interdisciplinary method that, using machine learning as well as domain knowledge, can effectively identify German travelogues in the digitized inventory of the Austrian National Library with F1 scores between 0.94 and 1.00. We applied our method on a corpus of 161,522 German volumes and identified 345 travelogues that could not be identified using traditional search methods, resulting in the most extensive collection of early modern German travelogues ever created. To our knowledge, this is the first time such a method was implemented for the bibliographic indexing of a text corpus on this scale, improving and extending the traditional methods in the humanities. Overall, we consider our technique to be an important first step in a broader effort of developing a novel mixed-method approach for the large-scale serial analysis of travelogues.


page 1

page 2

page 3

page 4


A Corpus for Automatic Readability Assessment and Text Simplification of German

In this paper, we present a corpus for use in automatic readability asse...

German Parliamentary Corpus (GerParCor)

Parliamentary debates represent a large and partly unexploited treasure ...

Compiling and Processing Historical and Contemporary Portuguese Corpora

This technical report describes the framework used for processing three ...

Robin: A Novel Online Suicidal Text Corpus of Substantial Breadth and Scale

Suicide is a major public health crisis. With more than 20,000,000 suici...

Factuality Detection using Machine Translation – a Use Case for German Clinical Text

Factuality can play an important role when automatically processing clin...

Diachronic Analysis of German Parliamentary Proceedings: Ideological Shifts through the Lens of Political Biases

We analyze bias in historical corpora as encoded in diachronic distribut...

Applications of Machine Learning in Document Digitisation

Data acquisition forms the primary step in all empirical research. The a...

Please sign up or login with your details

Forgot password? Click here to reset