What a Language Keeps of the Word It Will Not Say

English says gosh for God and heck for hell, and every one of those keeps the first sound and throws the rest away. That is a fact about English until somebody checks another language. English Wiktionary carries a minced-oath category for 33 languages; 273 pairs from 16 of them were measured here with one symmetric ruler, the first phone against the last phone, so neither end gets a bigger target than the other. The first phone survives 81.3 per cent of the time and the last 56.9, and the gap holds in all three transcription systems separately.

PRE-REGISTERED · 273 PAIRS · 16 LANGUAGES · THREE TRANSCRIPTION SYSTEMS · EVERY NUMBER RE-DERIVABLE FROM THE COMMITTED SNAPSHOT

A minced oath is a word a language builds so that a word it cannot say need not be said. English has hundreds: gosh for God, heck for hell, darn for damn, shoot for shit. Each one is a small natural experiment, because a speech community took a word it was not allowed to utter, deformed it until it was allowed, and the deformation is still recoverable from the two words side by side.

A page on this site measured 76 of the English ones and found that what survives is the beginning. The consonants before the first vowel came through 52 times; the vowel and everything after it came through never. That result has one obvious way to be wrong, and the page could not test it: it may be a fact about English rather than a fact about mincing.

English Wiktionary carries a minced oaths category for 33 languages. This page is what happens when you point the same question at all of them.

One ruler, and it is symmetric

The English measurement compared a short thing against a long one. An onset is a consonant or two; a rime is a vowel and every phone after it to the end of the word. Even with no phonology at all, the short target is easier to hit, so some of that 52-to-0 was the ruler and not the language. And the ruler itself does not travel: the syllabifier it used cuts onsets at the longest cluster English law allows, and pointing that at Polish or Welsh would be a lie about Polish and Welsh.

So this page uses a ruler with no phonological theory in it at all. Write each word as a string of phones. Ask two questions:

The two targets are exactly the same size. Neither end is favoured. Then the same question for the first and last two phones, three phones, and so on out to six, which gives two curves that can be laid on top of each other and read at a glance.

The instrument

Every pair that survived the audit is here. Green marks the phones that match counting in from the left; blue marks the phones that match counting in from the right; grey is where the two words have parted company. The play button speaks the pair, source first.

choose a language
transcription

The audio is espeak-ng synthesis of the exact strings the analysis measured. It is a way to hear the shape of the pair, not a recording of a speaker and not a claim about how anyone pronounces these words. espeak-ng is a grapheme-to-phoneme program and it is wrong sometimes, most often where a language's spelling is least phonemic; that is why the same result is computed twice more, once from the IPA Wiktionary itself prints and once from the plain letters, and all three are published below.

What the two curves do

0%25%50%75%100%123456 first/last 1: 81.3% of 209first/last 2: 64.6% of 209first/last 3: 42.6% of 202first/last 4: 31.3% of 182first/last 5: 15.0% of 160first/last 6: 16.1% of 124first/last 1: 56.9% of 209first/last 2: 33.5% of 209first/last 3: 23.3% of 202first/last 4: 14.8% of 182first/last 5: 8.1% of 160first/last 6: 5.6% of 124 first k phones kept last k phones kept k
Pooled over 209 pairs in 14 languages, espeak-ng phones. The share of pairs whose first k phones are identical, against the share whose last k phones are identical. At k = 1 the two targets are one phone each, and the head wins 81.3 to 56.9. The head curve is above the tail curve at every k from 1 to 6.

Pooled across 14 languages, a minced oath keeps the first phone of the word it replaces 81.3 per cent of the time and the last phone 56.9 per cent of the time. 82 pairs keep the first and lose the last; 31 do the reverse; McNemar's exact test on those discordant pairs gives p = 0.0000017.

Reading from the plain letters instead, which requires trusting no phonetic software at all and adds the two languages espeak-ng cannot speak (Cebuano and Tagalog), it is 83.3 against 54.5 over 233 pairs in 16 languages, p = 2.3e-10. Reading from the IPA Wiktionary prints on the entries themselves, where it prints any, 88.2 against 52.0 over 127 pairs, p = 1.3e-8.

Chance is not 50 per cent, so the page computes it

A language with few possible initial consonants will match at the left edge by luck more often than one with many. Quoting a raw percentage across languages with different inventories would therefore say more about phonotactics than about mincing. So for every language the null is built from that language's own words: take its euphemisms and its source words, shuffle which goes with which 100,000 times with a seeded generator, and see how often the first phones happen to agree. That is the by chance column.

languagepairsfirst keptlast keptfirst, by chanceperm pb/cMcNemar p
Czech9897824.70.000232/11.0
Dutch10805025.00.000445/20.45
English7770579.5<10-426/160.16
Finnish61008338.90.0171/01.0
French21905215.6<10-410/20.039
German7714314.30.000654/20.69
Italian101005023.9<10-45/00.063
Polish111005535.5<10-45/00.063
Portuguese78610030.60.00470/11.0
Russian18834418.5<10-410/30.092
Spanish16944413.7<10-49/10.021
Swedish7867163.30.141/01.0
Ukrainian5606031.80.162/21.0
Vietnamese5806032.00.0332/11.0
pooled, 14 languages20981.356.982/310.0000017

Red in the first kept column marks a language where the last phone did at least as well as the first. Languages with fewer than 5 usable pairs are not shown here and are not pooled; their counts are in the method section, and they are not silently dropped.

The prediction I registered, and lost

While designing this study, before a single statistic had been computed, I read the Wiktionary entry for sacrebleu and noticed a whole closed family of French oaths built by replacing Dieu with bleu: sacrebleu, morbleu, parbleu, corbleu, palsambleu, maugrebleu. In that family what survives looks like the rhyme, and the onset of the replaced word, the /dj/ of Dieu, is precisely what is destroyed. So I wrote the prediction into the pre-registration: French would show a weaker head advantage than English, and the -bleu family taken alone would show a tail advantage.

It is on the record because it is wrong.

replacedeuphemismsource phoneseuphemism phonesfirstlast
corps de Dieucorbleukˈɔʁ də- djˈøkɔʁblˈø
malgré Dieumaugrebleumalɡʁˈe djˈømoɡʁəblˈø
mort Dieumorbleumˈɔʁ djˈømɔʁblˈø
par le sang de Dieupalsambleupaʁ lə- sˈɑ̃ də- djˈøpalsɑ̃blˈø
par Dieuparbleupaʁ djˈøpaʁblˈø
sacre Dieusacrebleusˈakʁ djˈøsakʁəblˈø

All 6 of them keep both ends. The reason is a thing I had not thought about: these oaths deform a phrase, not a word. Par Dieu begins with par, which nobody touches, so the phrase's first phone is safe no matter what happens to Dieu; and bleu was chosen to rhyme with Dieu, so the phrase's last phone is safe too. What the family destroys is the middle. Far from being weak, French turns out to have one of the highest head rates in the study at 90 per cent, against a chance rate of 15.6 per cent.

This is what a pre-registration is for. Had I noticed the -bleu family only after running the numbers, the honest-looking move would have been to leave it out, and no reader could have known. Written down first, a wrong prediction costs one paragraph and buys the rest of the page its credibility.

Where the tail wins, which the pre-registration called a finding

2 eligible languages reverse: Portuguese and Ukrainian. Portuguese does it because most of its category is the rhyming-phrase kind, ponte que partiu for puta que pariu and fruta que caiu for the same, where the whole tail of the phrase is the joke and is kept deliberately. Below the threshold there is a cleaner reversal still: Old Norse has exactly two pairs, ragr for argr and regi for ergi, and both work by deleting the initial vowel and prefixing a consonant, which keeps the tail and destroys the head in both. Two pairs prove nothing. They are here because the pre-registration said a reversal gets a section rather than a footnote, and because they show the mechanism is not a law of nature.

The published claim this was supposed to test

Lev-Ari and McKay asked whether profanity has universal sound patterns, and their best candidate was that swear words under-use the approximants, the sonorous sounds /l/, /r/, /w/, /j/. Their Study 2 turned that around on English minced oaths: if approximants are un-swear-like, adding one should soften a word. They found approximants more frequent in the minced oaths than in the originals, 29 against 12, across 67 minced oaths drawn from 24 source swear words in the OED and the Wikipedia list. That study is English only, of necessity, because those are English sources.

Here is the same comparison on 209 pairs across 14 languages that are mostly not English. Approximants rose in 54 pairs and fell in 23, a mean of +0.163 phones per pair, sign test p = 0.00054. Counting every rhotic as an approximant, which is what their own coding does and which matters for French uvular /ʁ/, it is 55 against 28, p = 0.0040. It replicates.

And now the check that was registered in advance to stop it meaning more than it should. The same paired test was run for every phoneme class, not only the one the hypothesis is about, and all of them ship together.

classrosefellmean count Δpmean rate Δ (pp)p
stop6251+0.0620.35-1.420.21
fricative3355-0.1150.025-2.270.0017
affricate00+0.0001.0+0.001.0
nasal4814+0.1960.000017+2.100.012
liquid5118+0.1670.000088+1.450.20
glide2021-0.0051.0-0.180.47
vowel6229+0.2680.00071+0.321.0
approximant5423+0.1630.00054+1.270.41

Read the count columns and approximants rise. Read the rate columns, which divide by word length and therefore ask whether a euphemism is proportionally more approximant, and the approximant effect goes away: 1.27 percentage points, p = 0.41. Most of the raw rise is that euphemisms are simply longer words. What does survive the length control is a different shape: fricatives fall (-2.27 pp, p = 0.0017) and nasals rise (+2.10 pp, p = 0.012).

So the replication is real and the interpretation needs care, and the only reason this page can tell you which is which is that the all-classes table was promised before the numbers existed.

Every way I could think of to break it

Each row below drops a category of pair that a sceptic might object to, and recomputes the headline from scratch on what is left.

subsetpairsfirst keptlast keptMcNemar p
all20981.356.90.0000017
single words only13679.455.90.00026
no chains (source not itself in the category)19481.456.70.0000027
no swaps (deformations only)17178.957.30.00019
audit-untouched only (machine verdict kept)18885.154.81.0e-8
no case-duplicates20781.256.50.0000017
strictest: single, no chain, no swap, no dup9776.356.70.013
excluding English13287.956.80.0000010

The strictest row throws out phrases, chains where the source word is itself a minced oath, swaps where the euphemism is an ordinary word of the language pressed into service, and case-variant duplicates, and it still holds. Removing English entirely raises the head rate, which is the single most useful line in the table: the effect is not an English artefact leaking into a pooled average.

How the pairs were built, and what was thrown away

Every pair starts from one mechanical rule applied identically to all 33 languages, with no human judgement anywhere in it: take the first cue-carrying etymology or definition line, and its first linked word in the same language that is neither the entry title nor a meta-term. That rule produced 288 pairs from 540 category pages.

Two bugs in the rule were found by reading its output and were fixed before any statistic was computed. A bare wiki-link inside a non-English section is almost always the English gloss, which is how the first pass paired Polish cholewa with dagnabbit; and the templates der, inh, bor carry two language codes, so matching on the entry's own code captures a language code rather than a word, which is how Polish kurdebalans came out paired with the string fr. Restricting to language-tagged term templates fixes both. The looser rule is still computed and is in the data: it yields 422 pairs, and the extra ones are mostly English glosses.

A third bug was found later, by looking at this page's own instrument rather than at any table. espeak-ng writes a French clitic with a trailing hyphen, də- for de, and writes Vietnamese tone as a digit inside the string; Wiktionary writes Chinese and Vietnamese tone as superscript numerals. None of those is a phone, and all of them were being counted as one. The rule now applied to all three arms is stated once: a character that is not a segment of the IPA is not counted as one. Fixing it moved the pooled last-phone rate by about one point and removed a spurious hundred-per-cent tail rate for Vietnamese. It is written down here because the bug was visible only because every pair is on the page with its phones showing, which is most of the argument for putting them there.

All 288 machine pairs were then hand-audited against their own evidence text, one line each with a reason, and that audit was committed before the analysis script was written. 15 were dropped and 24 were corrected where the etymology plainly named a different word than the rule had grabbed. Here is every drop:

the 15 dropped pairs, with reasons

Of the 273 that remain, 91 have a phrase on one side or the other, 16 are chains whose source is itself in the same minced-oath category, and 44 are swaps where the euphemism is an ordinary existing word. All three are flagged per pair and every headline is recomputed without them in the table above.

Limits, stated rather than buried

What was already known, and by whom

The generalisation on this page appears to be unclaimed, and the near misses are worth naming precisely because they are near.

One thing the data here cannot settle is why. There are two clean stories and this page distinguishes neither. The hearer's: the onset survives because that is where a listener recovers the word you are pointedly not saying, so mincing that destroyed the beginning would fail at its whole job. The speaker's: the onset survives because it is already out of the mouth before the flinch arrives, so only the rest is available to repair. Both predict exactly what is on this page.

Provenance

Category memberships and entry text are from English Wiktionary, CC BY-SA 4.0, snapshotted 2026-08-05 with revision ids in research/minced-oaths-crosslinguistic/data/. Pronunciation is espeak-ng (GPL), Wiktionary's own printed IPA, and the plain orthography. The pre-registration is research/minced-oaths-crosslinguistic/PREREGISTRATION.md, committed before the analysis script existed; the audit is data/audit.json, committed before it was run. Lev-Ari, S. and McKay, R., "The sound of swearing: are there universal patterns in profanity?", Psychonomic Bulletin & Review 30(3), 1103 to 1114, issue dated 2023, online 6 December 2022, doi:10.3758/s13423-022-02202-0. Vallery, R., Tark, shperlack, burfip, and other alien bad words, PhD thesis, Université de Lille, 2024. Allan, K. and Burridge, K., Forbidden Words, Cambridge University Press, 2006. Pimentel, T., Cotterell, R. and Roark, B., "Disambiguatory signals are stronger in word-initial positions", EACL 2021, arXiv:2102.02183.

Every figure on this page is regenerated from data/results.json by build-page.mjs, and verify-what-a-language-keeps.mjs at the repository root re-derives the results from the committed snapshot with its own arithmetic and then asserts the bytes that shipped.