Suum Cuique: Studying Bias in Taboo Detection with a Community Perspective

by   Osama Khalid, et al.

Prior research has discussed and illustrated the need to consider linguistic norms at the community level when studying taboo (hateful/offensive/toxic etc.) language. However, a methodology for doing so, that is firmly founded on community language norms is still largely absent. This can lead both to biases in taboo text classification and limitations in our understanding of the causes of bias. We propose a method to study bias in taboo classification and annotation where a community perspective is front and center. This is accomplished by using special classifiers tuned for each community's language. In essence, these classifiers represent community level language norms. We use these to study bias and find, for example, biases are largest against African Americans (7/10 datasets and all 3 classifiers examined). In contrast to previous papers we also study other communities and find, for example, strong biases against South Asians. In a small scale user study we illustrate our key idea which is that common utterances, i.e., those with high alignment scores with a community (community classifier confidence scores) are unlikely to be regarded taboo. Annotators who are community members contradict taboo classification decisions and annotations in a majority of instances. This paper is a significant step toward reducing false positive taboo decisions that over time harm minority communities.


page 1

page 2

page 3

page 4


WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models

We present WinoQueer: a benchmark specifically designed to measure wheth...

Discovering and Categorising Language Biases in Reddit

We present a data-driven approach using word embeddings to discover and ...

Towards WinoQueer: Developing a Benchmark for Anti-Queer Bias in Large Language Models

This paper presents exploratory work on whether and to what extent biase...

Enriching Abusive Language Detection with Community Context

Uses of pejorative expressions can be benign or actively empowering. Whe...

TalkDown: A Corpus for Condescension Detection in Context

Condescending language use is caustic; it can bring dialogues to an end ...

Sentence level estimation of psycholinguistic norms using joint multidimensional annotations

Psycholinguistic normatives represent various affective and mental const...

Altruistic and Profit-oriented: Making Sense of Roles in Web3 Community from Airdrop Perspective

Regardless of which community, incentivizing users is a necessity for we...

Please sign up or login with your details

Forgot password? Click here to reset