Thesis
Improving the evaluation and effectiveness of hate speech detection models
- Abstract:
-
Online hate speech is a widespread and deeply harmful problem. To tackle hate at scale, we need models that can automatically detect it. This has motivated research in Natural Language Processing (NLP) to develop text-based hate speech detection models. In recent years, these models have improved substantially, following general advances in language modelling.
In my thesis, I show that impressive headline results paint an incomplete picture of model quality. I argue that much progress in hate speech detection so far has rested on simplifying assumptions, which, while useful in some settings, we need to move past in order to develop more truly effective models. In particular, I argue that current standards for model evaluation tend to be overly aggregated, static, monolithic and English language-centric, because of four common simplifying assumptions, which I use to structure my thesis.
I discuss core concepts in an introduction and literature review. Then, I present four Chapters, which each challenge one of the four simplifying assumptions. Assumption 1 is that model accuracy equals model quality. I introduce a suite of functional tests for hate speech detection models, which enables fine-grained diagnostic insights and reveals critical weaknesses in seemingly accurate models. Assumption 2 is that hate speech today equals hate speech tomorrow. I find that model performance degrades over time and explore temporal adaptation as a remedy. Assumption 3 is that hate speech for me equals hate speech for you. I evidence subjectivity in labelling hate speech and introduce two contrasting data annotation paradigms for managing subjectivity. Assumption 4 is that hate speech in English equals all hate speech. I explore data-efficient strategies for expanding detection into more under-resourced languages.
Overall, my thesis seeks to 1) work towards better, more comprehensive quality standards for hate speech detection models, and 2) improve models along these standards.
Actions
Access Document
- Files:
-
-
(Preview, Dissemination version, pdf, 3.3MB, Terms of use)
-
Authors
Contributors
- Institution:
- University of Oxford
- Role:
- Supervisor
- ORCID:
- 0000-0002-5989-3574
- Institution:
- University of Oxford
- Division:
- SSD
- Department:
- Oxford Internet Institute
- Role:
- Supervisor
- ORCID:
- 0000-0003-4597-8283
- Institution:
- University of Oxford
- Division:
- SSD
- Department:
- Oxford Internet Institute
- Role:
- Examiner
- ORCID:
- 0000-0002-6894-4951
- Role:
- Examiner
- Funder identifier:
- https://ror.org/05xwwfy96
- Programme:
- Promotionsstipendium
- DOI:
- Type of award:
- DPhil
- Level of award:
- Doctoral
- Awarding institution:
- University of Oxford
- Language:
-
English
- Keywords:
- Subjects:
- Deposit date:
-
2024-09-09
- ARK identifier:
Terms of use
- Copyright holder:
- Röttger, P
- Copyright date:
- 2023
- Licence:
- CC Attribution (CC BY)
If you are the owner of this record, you can report an update to it here: Report update to this record