What a Language Keeps of the Word It Will Not Say
English says gosh for God and heck for hell, and every one of those keeps the first sound and throws the rest away. That is a fact about English until somebody checks another language. English Wiktionary carries a minced-oath category for 33 languages; 273 pairs from 16 of them were measured here with one symmetric ruler, the first phone against the last phone, so neither end gets a bigger target than the other. The first phone survives 81.3 per cent of the time and the last 56.9, and the gap holds in all three transcription systems separately.
PRE-REGISTERED · 273 PAIRS · 16 LANGUAGES · THREE TRANSCRIPTION SYSTEMS · EVERY NUMBER RE-DERIVABLE FROM THE COMMITTED SNAPSHOT
A minced oath is a word a language builds so that a word it cannot say need not be said. English has hundreds: gosh for God, heck for hell, darn for damn, shoot for shit. Each one is a small natural experiment, because a speech community took a word it was not allowed to utter, deformed it until it was allowed, and the deformation is still recoverable from the two words side by side.
A page on this site measured 76 of the English ones and found that what survives is the beginning. The consonants before the first vowel came through 52 times; the vowel and everything after it came through never. That result has one obvious way to be wrong, and the page could not test it: it may be a fact about English rather than a fact about mincing.
English Wiktionary carries a minced oaths category for 33 languages. This page is what happens when you point the same question at all of them.
One ruler, and it is symmetric
The English measurement compared a short thing against a long one. An onset is a consonant or two; a rime is a vowel and every phone after it to the end of the word. Even with no phonology at all, the short target is easier to hit, so some of that 52-to-0 was the ruler and not the language. And the ruler itself does not travel: the syllabifier it used cuts onsets at the longest cluster English law allows, and pointing that at Polish or Welsh would be a lie about Polish and Welsh.
So this page uses a ruler with no phonological theory in it at all. Write each word as a string of phones. Ask two questions:
- Did the first phone survive? One phone, at the left edge.
- Did the last phone survive? One phone, at the right edge.
The two targets are exactly the same size. Neither end is favoured. Then the same question for the first and last two phones, three phones, and so on out to six, which gives two curves that can be laid on top of each other and read at a glance.
The instrument
Every pair that survived the audit is here. Green marks the phones that match counting in from the left; blue marks the phones that match counting in from the right; grey is where the two words have parted company. The play button speaks the pair, source first.
The audio is espeak-ng synthesis of the exact strings the analysis measured. It is a way to hear the shape of the pair, not a recording of a speaker and not a claim about how anyone pronounces these words. espeak-ng is a grapheme-to-phoneme program and it is wrong sometimes, most often where a language's spelling is least phonemic; that is why the same result is computed twice more, once from the IPA Wiktionary itself prints and once from the plain letters, and all three are published below.
What the two curves do
Pooled across 14 languages, a minced oath keeps the first phone of the word it replaces 81.3 per cent of the time and the last phone 56.9 per cent of the time. 82 pairs keep the first and lose the last; 31 do the reverse; McNemar's exact test on those discordant pairs gives p = 0.0000017.
Reading from the plain letters instead, which requires trusting no phonetic software at all and adds the two languages espeak-ng cannot speak (Cebuano and Tagalog), it is 83.3 against 54.5 over 233 pairs in 16 languages, p = 2.3e-10. Reading from the IPA Wiktionary prints on the entries themselves, where it prints any, 88.2 against 52.0 over 127 pairs, p = 1.3e-8.
Chance is not 50 per cent, so the page computes it
A language with few possible initial consonants will match at the left edge by luck more often than one with many. Quoting a raw percentage across languages with different inventories would therefore say more about phonotactics than about mincing. So for every language the null is built from that language's own words: take its euphemisms and its source words, shuffle which goes with which 100,000 times with a seeded generator, and see how often the first phones happen to agree. That is the by chance column.
| language | pairs | first kept | last kept | first, by chance | perm p | b/c | McNemar p |
|---|---|---|---|---|---|---|---|
| Czech | 9 | 89 | 78 | 24.7 | 0.00023 | 2/1 | 1.0 |
| Dutch | 10 | 80 | 50 | 25.0 | 0.00044 | 5/2 | 0.45 |
| English | 77 | 70 | 57 | 9.5 | <10-4 | 26/16 | 0.16 |
| Finnish | 6 | 100 | 83 | 38.9 | 0.017 | 1/0 | 1.0 |
| French | 21 | 90 | 52 | 15.6 | <10-4 | 10/2 | 0.039 |
| German | 7 | 71 | 43 | 14.3 | 0.00065 | 4/2 | 0.69 |
| Italian | 10 | 100 | 50 | 23.9 | <10-4 | 5/0 | 0.063 |
| Polish | 11 | 100 | 55 | 35.5 | <10-4 | 5/0 | 0.063 |
| Portuguese | 7 | 86 | 100 | 30.6 | 0.0047 | 0/1 | 1.0 |
| Russian | 18 | 83 | 44 | 18.5 | <10-4 | 10/3 | 0.092 |
| Spanish | 16 | 94 | 44 | 13.7 | <10-4 | 9/1 | 0.021 |
| Swedish | 7 | 86 | 71 | 63.3 | 0.14 | 1/0 | 1.0 |
| Ukrainian | 5 | 60 | 60 | 31.8 | 0.16 | 2/2 | 1.0 |
| Vietnamese | 5 | 80 | 60 | 32.0 | 0.033 | 2/1 | 1.0 |
| pooled, 14 languages | 209 | 81.3 | 56.9 | 82/31 | 0.0000017 |
Red in the first kept column marks a language where the last phone did at least as well as the first. Languages with fewer than 5 usable pairs are not shown here and are not pooled; their counts are in the method section, and they are not silently dropped.
The prediction I registered, and lost
While designing this study, before a single statistic had been computed, I read the Wiktionary entry for sacrebleu and noticed a whole closed family of French oaths built by replacing Dieu with bleu: sacrebleu, morbleu, parbleu, corbleu, palsambleu, maugrebleu. In that family what survives looks like the rhyme, and the onset of the replaced word, the /dj/ of Dieu, is precisely what is destroyed. So I wrote the prediction into the pre-registration: French would show a weaker head advantage than English, and the -bleu family taken alone would show a tail advantage.
It is on the record because it is wrong.
| replaced | euphemism | source phones | euphemism phones | first | last |
|---|---|---|---|---|---|
| corps de Dieu | corbleu | kˈɔʁ də- djˈø | kɔʁblˈø | ✓ | ✓ |
| malgré Dieu | maugrebleu | malɡʁˈe djˈø | moɡʁəblˈø | ✓ | ✓ |
| mort Dieu | morbleu | mˈɔʁ djˈø | mɔʁblˈø | ✓ | ✓ |
| par le sang de Dieu | palsambleu | paʁ lə- sˈɑ̃ də- djˈø | palsɑ̃blˈø | ✓ | ✓ |
| par Dieu | parbleu | paʁ djˈø | paʁblˈø | ✓ | ✓ |
| sacre Dieu | sacrebleu | sˈakʁ djˈø | sakʁəblˈø | ✓ | ✓ |
All 6 of them keep both ends. The reason is a thing I had not thought about: these oaths deform a phrase, not a word. Par Dieu begins with par, which nobody touches, so the phrase's first phone is safe no matter what happens to Dieu; and bleu was chosen to rhyme with Dieu, so the phrase's last phone is safe too. What the family destroys is the middle. Far from being weak, French turns out to have one of the highest head rates in the study at 90 per cent, against a chance rate of 15.6 per cent.
This is what a pre-registration is for. Had I noticed the -bleu family only after running the numbers, the honest-looking move would have been to leave it out, and no reader could have known. Written down first, a wrong prediction costs one paragraph and buys the rest of the page its credibility.
Where the tail wins, which the pre-registration called a finding
2 eligible languages reverse: Portuguese and Ukrainian. Portuguese does it because most of its category is the rhyming-phrase kind, ponte que partiu for puta que pariu and fruta que caiu for the same, where the whole tail of the phrase is the joke and is kept deliberately. Below the threshold there is a cleaner reversal still: Old Norse has exactly two pairs, ragr for argr and regi for ergi, and both work by deleting the initial vowel and prefixing a consonant, which keeps the tail and destroys the head in both. Two pairs prove nothing. They are here because the pre-registration said a reversal gets a section rather than a footnote, and because they show the mechanism is not a law of nature.
The published claim this was supposed to test
Lev-Ari and McKay asked whether profanity has universal sound patterns, and their best candidate was that swear words under-use the approximants, the sonorous sounds /l/, /r/, /w/, /j/. Their Study 2 turned that around on English minced oaths: if approximants are un-swear-like, adding one should soften a word. They found approximants more frequent in the minced oaths than in the originals, 29 against 12, across 67 minced oaths drawn from 24 source swear words in the OED and the Wikipedia list. That study is English only, of necessity, because those are English sources.
Here is the same comparison on 209 pairs across 14 languages that are mostly not English. Approximants rose in 54 pairs and fell in 23, a mean of +0.163 phones per pair, sign test p = 0.00054. Counting every rhotic as an approximant, which is what their own coding does and which matters for French uvular /ʁ/, it is 55 against 28, p = 0.0040. It replicates.
And now the check that was registered in advance to stop it meaning more than it should. The same paired test was run for every phoneme class, not only the one the hypothesis is about, and all of them ship together.
| class | rose | fell | mean count Δ | p | mean rate Δ (pp) | p |
|---|---|---|---|---|---|---|
| stop | 62 | 51 | +0.062 | 0.35 | -1.42 | 0.21 |
| fricative | 33 | 55 | -0.115 | 0.025 | -2.27 | 0.0017 |
| affricate | 0 | 0 | +0.000 | 1.0 | +0.00 | 1.0 |
| nasal | 48 | 14 | +0.196 | 0.000017 | +2.10 | 0.012 |
| liquid | 51 | 18 | +0.167 | 0.000088 | +1.45 | 0.20 |
| glide | 20 | 21 | -0.005 | 1.0 | -0.18 | 0.47 |
| vowel | 62 | 29 | +0.268 | 0.00071 | +0.32 | 1.0 |
| approximant | 54 | 23 | +0.163 | 0.00054 | +1.27 | 0.41 |
Read the count columns and approximants rise. Read the rate columns, which divide by word length and therefore ask whether a euphemism is proportionally more approximant, and the approximant effect goes away: 1.27 percentage points, p = 0.41. Most of the raw rise is that euphemisms are simply longer words. What does survive the length control is a different shape: fricatives fall (-2.27 pp, p = 0.0017) and nasals rise (+2.10 pp, p = 0.012).
So the replication is real and the interpretation needs care, and the only reason this page can tell you which is which is that the all-classes table was promised before the numbers existed.
Every way I could think of to break it
Each row below drops a category of pair that a sceptic might object to, and recomputes the headline from scratch on what is left.
| subset | pairs | first kept | last kept | McNemar p |
|---|---|---|---|---|
| all | 209 | 81.3 | 56.9 | 0.0000017 |
| single words only | 136 | 79.4 | 55.9 | 0.00026 |
| no chains (source not itself in the category) | 194 | 81.4 | 56.7 | 0.0000027 |
| no swaps (deformations only) | 171 | 78.9 | 57.3 | 0.00019 |
| audit-untouched only (machine verdict kept) | 188 | 85.1 | 54.8 | 1.0e-8 |
| no case-duplicates | 207 | 81.2 | 56.5 | 0.0000017 |
| strictest: single, no chain, no swap, no dup | 97 | 76.3 | 56.7 | 0.013 |
| excluding English | 132 | 87.9 | 56.8 | 0.0000010 |
The strictest row throws out phrases, chains where the source word is itself a minced oath, swaps where the euphemism is an ordinary word of the language pressed into service, and case-variant duplicates, and it still holds. Removing English entirely raises the head rate, which is the single most useful line in the table: the effect is not an English artefact leaking into a pooled average.
How the pairs were built, and what was thrown away
Every pair starts from one mechanical rule applied identically to all 33 languages, with no human judgement anywhere in it: take the first cue-carrying etymology or definition line, and its first linked word in the same language that is neither the entry title nor a meta-term. That rule produced 288 pairs from 540 category pages.
Two bugs in the rule were found by reading its output and were fixed before any statistic was computed. A bare wiki-link inside a non-English section is almost always the English gloss, which is how the first pass paired Polish cholewa with dagnabbit; and the templates der, inh, bor carry two language codes, so matching on the entry's own code captures a language code rather than a word, which is how Polish kurdebalans came out paired with the string fr. Restricting to language-tagged term templates fixes both. The looser rule is still computed and is in the data: it yields 422 pairs, and the extra ones are mostly English glosses.
A third bug was found later, by looking at this page's own instrument rather than at any table. espeak-ng writes a French clitic with a trailing hyphen, də- for de, and writes Vietnamese tone as a digit inside the string; Wiktionary writes Chinese and Vietnamese tone as superscript numerals. None of those is a phone, and all of them were being counted as one. The rule now applied to all three arms is stated once: a character that is not a segment of the IPA is not counted as one. Fixing it moved the pooled last-phone rate by about one point and removed a spurious hundred-per-cent tail rate for Vietnamese. It is written down here because the bug was visible only because every pair is on the page with its phones showing, which is most of the argument for putting them there.
All 288 machine pairs were then hand-audited against their own evidence text, one line each with a reason, and that audit was committed before the analysis script was written. 15 were dropped and 24 were corrected where the etymology plainly named a different word than the rule had grabbed. Here is every drop:
the 15 dropped pairs, with reasons
- Albanian bretk against bretëkë: the etymology gives the word's own history (a frog), not a taboo source; the gloss 'dreq' is a definition, and no etymology states the substitution
- Chinese Richard against 碌柒: the euphemism is an English name, so the pair crosses languages
- English B against BCPL: the etymology is about a programming language and names no taboo source
- English blimey against blind#Verb: two competing derivations in the same sentence ('(God) blind me ... or blame me') and an optional element
- English by jingo against crikey: three competing theories are listed; 'compare crikey' is a comparison, not a source
- English consarn it against consarnation: the taboo source is given as 'possibly ... confound it'; explicitly uncertain
- English holy buckets against holy cow: 'possibly ... alternatively', two competing derivations, one of them a pun on holey (spelled holey)
- English let's go, Brandon against fuck: no etymology; the relation is a mishearing of a chanted phrase, not a deformation of the linked word
- English pooh against poo: the line 'poo: a minced oath for shit' leaves it ambiguous whether the source is poo or shit
- English Sam Hill against hell: the etymology opens 'Its exact etymology is uncertain' and lists rival named individuals
- Finnish jumalauta against jumankauta: the etymology is Jumala + auttaa; 'possibly originally a minced oath. Compare jumankauta' names no source
- Icelandic bévítans against bé: a blend of two named sources (bé from bölvaður, the ending from ansvítans); no single source word
- Irish muise against más ea: 'Possibly from más ea ... or a euphemistic alteration of Muire', two competing sources
- Portuguese caçamba against cacete,caralho: the template names two sources, cacete and caralho, and they differ at the phone that matters
- Vietnamese cần lời giải thích against clgt: the source clgt is an initialism, so there is no comparable phone string on the source side
Of the 273 that remain, 91 have a phrase on one side or the other, 16 are chains whose source is itself in the same minced-oath category, and 44 are swaps where the euphemism is an ordinary existing word. All three are flagged per pair and every headline is recomputed without them in the table above.
Limits, stated rather than buried
- Wiktionary is not the OED. It is a tertiary source, CC BY-SA, and its etymologies vary in quality. The category is whatever its editors have tagged, so no claim about all the minced oaths of any language is made or can be.
- Category size measures Wiktionary, not swearing. English has 77 usable pairs and Hungarian has 1, because English Wiktionary is written by English speakers. Comparing counts across languages would be a statement about the encyclopedia.
- espeak-ng is a program, not a dictionary. Its error rate differs by language and is worst where orthography is least phonemic. It agrees with Wiktionary's own printed IPA on the first-phone verdict for 97.7 per cent of the 128 pairs where both exist, and on the last-phone verdict for 95.3 per cent, while producing an identical phone string only 33.6 per cent of the time. The verdicts are robust where the transcriptions are not.
- Chinese is reported and never pooled. A tonal syllable written in a romanisation is not commensurable with a segmental phone string from Polish, so it is excluded from every pooled figure by the pre-registration, not by the result.
- Eleven languages have a category but fewer than 5 usable pairs and appear nowhere in the pooled numbers: Chinese (10), Galician (3), Hindi (1), Hungarian (1), Icelandic (4), Ilocano (1), Indonesian (1), Irish (2), Norwegian Bokmål (4), Norwegian Nynorsk (4), Old Norse (2), Scottish Gaelic (2), Turkish (3), Welsh (2).
- This is a study of what a dictionary records, not of what mouths do. Nothing here was heard. It measures written etymological claims about pairs of written words.
What was already known, and by whom
The generalisation on this page appears to be unclaimed, and the near misses are worth naming precisely because they are near.
- Allan and Burridge, Forbidden Words (Cambridge, 2006), is the standard treatment, and their term for the phenomenon is remodelling. They lay out Shit! against Sugar!, Shoot!, Shivers!, Shucks!, which is a perfect four-for-four demonstration of onset preservation, and they do not remark on it. Their discussion of how a hearer recovers a remodelled word turns instead on a scrambled-letters anecdote.
- Rossi, in the Treccani Enciclopedia dell'Italiano (2011), documents Italian deformation by phoneme substitution, cacchio for cazzo and cribbio for Cristo, and explicitly identifies no systematic rule about which part is kept.
- Vallery's 2024 doctoral thesis at Lille surveys this literature and names the gap directly, writing that minced oaths "must be different from, but very close to the swear word they replace, so that speakers understand what swear word is replaced", and that "a more exhaustive, systematic study of minced oaths would allow to confirm or disconfirm this provisional hypothesis". His hypothesis is about sonority. The question of which end he does not ask.
- Pimentel, Cotterell and Roark (EACL 2021) show across hundreds of languages that lexicons front-load their disambiguating information, and Marslen-Wilson's cohort work shows listeners use word onsets to narrow candidates. Between them they make the head advantage a prediction rather than a curiosity, which is why the interesting number here is not that the head wins but the size of the margin over the permutation baseline.
- Wong (2016), "Flip the Bleeping Page", an undergraduate dissertation supervised by Andrew Nevins, is the one piece of possible prior art on cross-linguistic taboo avoidance that I could not read: it is deposited behind a login wall and has no DOI or index record. Its indexed abstract describes lexical neighbourhood density, frequency counts and native-speaker judgements, and does not say whether it measures which part of the word survives. Named here as an unresolved gap rather than quietly omitted.
One thing the data here cannot settle is why. There are two clean stories and this page distinguishes neither. The hearer's: the onset survives because that is where a listener recovers the word you are pointedly not saying, so mincing that destroyed the beginning would fail at its whole job. The speaker's: the onset survives because it is already out of the mouth before the flinch arrives, so only the rest is available to repair. Both predict exactly what is on this page.
Provenance
Category memberships and entry text are from English Wiktionary, CC BY-SA 4.0, snapshotted 2026-08-05 with revision ids in research/minced-oaths-crosslinguistic/data/. Pronunciation is espeak-ng (GPL), Wiktionary's own printed IPA, and the plain orthography. The pre-registration is research/minced-oaths-crosslinguistic/PREREGISTRATION.md, committed before the analysis script existed; the audit is data/audit.json, committed before it was run. Lev-Ari, S. and McKay, R., "The sound of swearing: are there universal patterns in profanity?", Psychonomic Bulletin & Review 30(3), 1103 to 1114, issue dated 2023, online 6 December 2022, doi:10.3758/s13423-022-02202-0. Vallery, R., Tark, shperlack, burfip, and other alien bad words, PhD thesis, Université de Lille, 2024. Allan, K. and Burridge, K., Forbidden Words, Cambridge University Press, 2006. Pimentel, T., Cotterell, R. and Roark, B., "Disambiguatory signals are stronger in word-initial positions", EACL 2021, arXiv:2102.02183.
Every figure on this page is regenerated from data/results.json by build-page.mjs, and verify-what-a-language-keeps.mjs at the repository root re-derives the results from the committed snapshot with its own arithmetic and then asserts the bytes that shipped.