Thesis icon

Thesis

Improving the evaluation and effectiveness of hate speech detection models

Abstract:

Online hate speech is a widespread and deeply harmful problem. To tackle hate at scale, we need models that can automatically detect it. This has motivated research in Natural Language Processing (NLP) to develop text-based hate speech detection models. In recent years, these models have improved substantially, following general advances in language modelling.

In my thesis, I show that impressive headline results paint an incomplete picture of model quality. I argue that much progress in hate speech detection so far has rested on simplifying assumptions, which, while useful in some settings, we need to move past in order to develop more truly effective models. In particular, I argue that current standards for model evaluation tend to be overly aggregated, static, monolithic and English language-centric, because of four common simplifying assumptions, which I use to structure my thesis.

I discuss core concepts in an introduction and literature review. Then, I present four Chapters, which each challenge one of the four simplifying assumptions. Assumption 1 is that model accuracy equals model quality. I introduce a suite of functional tests for hate speech detection models, which enables fine-grained diagnostic insights and reveals critical weaknesses in seemingly accurate models. Assumption 2 is that hate speech today equals hate speech tomorrow. I find that model performance degrades over time and explore temporal adaptation as a remedy. Assumption 3 is that hate speech for me equals hate speech for you. I evidence subjectivity in labelling hate speech and introduce two contrasting data annotation paradigms for managing subjectivity. Assumption 4 is that hate speech in English equals all hate speech. I explore data-efficient strategies for expanding detection into more under-resourced languages.

Overall, my thesis seeks to 1) work towards better, more comprehensive quality standards for hate speech detection models, and 2) improve models along these standards.

Actions

Access Document

Files:

Authors

More by this author
Institution:
University of Oxford
Division:
SSD
Role:
Author

Contributors

Institution:
University of Oxford
Role:
Supervisor
ORCID:
0000-0002-5989-3574
Institution:
University of Oxford
Division:
SSD
Department:
Oxford Internet Institute
Role:
Supervisor
ORCID:
0000-0003-4597-8283
Institution:
University of Oxford
Division:
SSD
Department:
Oxford Internet Institute
Role:
Examiner
ORCID:
0000-0002-6894-4951
Role:
Examiner


More from this funder
Funder identifier:
https://ror.org/05xwwfy96
Programme:
Promotionsstipendium


DOI:
Type of award:
DPhil
Level of award:
Doctoral
Awarding institution:
University of Oxford

Terms of use


Views and Downloads






If you are the owner of this record, you can report an update to it here: Report update to this record

TO TOP