freq_list_wales.csv
A frequency list of word forms in the corpora. Word forms with a frequency below 1.00 per million words in the Wales corpus (equivalent to 36 occurrences) are omitted, as are word forms beginning with numerals (1, 1989, 12pm, 4th etc.).

Columns are:
word: the word form in the corpus
n_south: number of tokens in the southern England subcorpus
freq_south: frequency per million words in the southern England subcorpus
n_north: number of tokens in the norther England subcorpus
freq_north: frequency per million words in the northern England subcorpus
n_wales: number of tokens in the Wales subcorpus
freq_wales: frequency per million words in the Wales subcorpus
n_total: number of tokens across all three subcorpora
ratio_south: ratio of frequency in the Wales subcorpora to the frequency in the southern England subcorpora
ratio_north: ratio of frequency in the Wales subcorpora to the frequency in the northern England subcorpora
ratio_england: ratio of frequency in the Wales subcorpora to the overall frequency in the two England subcorpora

fig_4.1.png
Geospatial distribution of the word 'lush' in the corpus as compared to the frequency of other adjectives of positive evaluation listed below.

fig_4.2-11.png
Geospatial distribution of terms of address in clause-final position in the corpus.

supp1–6.png
Geospatial distribution of additional terms of address in clause-final position in the corpus.

vocatives_16variants_en.csv
Raw data by location underlying fig_4.2-11 and supp1-6.

Columns:
x,y: longitude and latitude of the location
var1–16 frequency of each variant at that location, expressed as the sum of proportions of use for each Twitter account localized to that location. 

Example:
If, at a given location,
User 1 uses variant 1 once and variant 2 four times = variant 1 is used 0.2 of the time; and
User 2 uses variant 1 twice and variant 2 zero times = variant 1 is used 1.0 of the time;
and there are no other users localized to that location, then
var1 for that location is 1.2

total: total number of observations at that location; this would be 7 in the example above.

vocatives_16variants_en_kde.csv
Underlying data smoothed using kernelPhil (https://cran.r-project.org/src/contrib/Archive/kernelPhil/)

Columns:
x,y: longitude and latitude of the location
bw: bandwidth
relative_density_var1-16: relative smoothed density of the given variant at this location
effective_sample_size: effective sample size
weight_at_point: number of actual observations at this location
best: variant with the highest smoothed frequency

In both files:
var1 = 'boy/boi'
var2 = 'boyo/boio'
var3 = 'bro/bruv/blud/blad'
var4 = 'bud(dy)'
var5 = 'butt(y)'
var6 = 'dude'
var7 = 'fella'
var8 = 'girl'
var9 = 'hun'
var10 = 'lad'
var11 = 'man'
var12 = 'mate(y)'
var13 = 'mucka/mukka/mucker/mukker'
var14 = 'mun'
var15 = 'mush'
var16 = 'pal'

lush_en.csv and lush_en_kde.csv
The same data in the same format as above for adjectives of positive evaluation.

In both files:
var1 = 'amazing'
var2 = 'brilliant'
var3 = 'fantastic'
var4 = 'lovely'
var5 = 'lush'

For details of the corpus, see Willis, David. 2026. Welsh–English social-media lexicon in comparative context: Adjectives of positive evaluation and terms of address. In Natalie Braber & Rhys Sandow (eds.), Sociolinguistic approaches to lexical variation in English. London: Routledge.