Pun research ranges from carefully classified Japanese sentences to large joke-recognition datasets and multilingual generation benchmarks. The numbers show how researchers define a pun, assemble examples, measure humor, and compare generated wordplay.
Table of contents
- How large are pun corpora?
- Which pun types appear most often?
- Where do pun examples come from?
- What do annotation and humor ratings show?
- How do multilingual datasets compare?
- What do generation benchmarks measure?
How large are pun corpora?
One Japanese pun-corpus study gives a useful view of the difference between collected material and a cleaned research resource. Its corpus builder first extracted 52,995 pun sentences. After duplicates and near-duplicates were removed, the final corpus contained 45,970 pun sentences. More than 95% of the extracted data was shared with the Japanese pun database created by Araki et al. in 2017. These figures describe that study’s corpus-building process; they are not an estimate of all Japanese puns.
The same research project used a 600-sentence sample for detailed classification. Of those sentences, 563 were juxtaposed puns, representing 93.8% of the sample. The sample also included 42 perfect puns, or 7.0%, 521 imperfect puns, or 86.8%, 33 superposed puns, or 5.5%, and 4 sentences classified as other, or 0.7% (Comparison of Pun Detection Methods Using Japanese Pun Corpus).
These categories should not be added together as though they were necessarily one exclusive list: the study reports several classification views of the same 600-sentence sample. The practical lesson is that a corpus can become much smaller through deduplication while still retaining tens of thousands of examples. It also shows why a pun statistic needs its sample definition beside it.
For comparison, Puntuguese contains 4,903 manually gathered punning one-liners in Brazilian and European Portuguese. Its reported humor-recognition F1-score is 68.9% (Puntuguese: A Corpus of Puns in Portuguese with Micro-edits). A smaller, manually assembled corpus and a larger extracted corpus answer different research questions, so their counts should not be treated as a quality ranking.
Which pun types appear most often?
The Japanese sample is especially informative for writers interested in how wordplay is structured. Juxtaposed puns accounted for 563 of 600 sampled sentences. Perfect puns accounted for 42, while imperfect puns accounted for 521. Superposed puns accounted for 33, and the other category accounted for 4. In the study’s summary, juxtaposed and imperfect puns together accounted for more than 85.0% of targeted puns (Comparison of Pun Detection Methods Using Japanese Pun Corpus).
The difference between the broad and narrower labels matters. “Juxtaposed” describes one reported structural grouping, while “perfect” and “imperfect” describe another distinction used in the same sample. A writer comparing pun statistics should therefore preserve the original label and denominator instead of presenting every percentage as a single taxonomy.
The Portuguese resource takes a different approach: it is defined by 4,903 punning one-liners and micro-edits rather than the Japanese study’s sentence-level classification counts. The available result for that corpus is a 68.9% humor-recognition F1-score, which summarizes recognition performance rather than the proportion of one-liners judged funny.
Where do pun examples come from?
The Japanese extracted data came from several named sources. Dajare Nabi contributed 39,120 sentences, Dajare Suteshon contributed 8,795, and Dajare Netto contributed 1,621. Hitokuchi Dajare Daishuugou contributed 1,067 sentences, while Dajare Shu Dajare Jiten contributed 982.
Smaller sources were still part of the assembled resource. Dajare No Kanzume contributed 572 sentences, Dajare Kurabu contributed 428, Dajare Hiroba contributed 303, and Dajare o Itta no ha Dareja? contributed 107 (Comparison of Pun Detection Methods Using Japanese Pun Corpus).
The largest named source was therefore Dajare Nabi, with 39,120 sentences. The smallest named source contributed 107. That spread is useful context when reading the final total of 45,970 cleaned sentences: the corpus is not a uniform sample from equally sized sources. It is an aggregation in which some repositories contribute far more material than others.
Another benchmark used a non-pun corpus of 20,000 sentences from Wikipedia and the Gutenberg BookCorpus (Pun Generation with Surprise). A non-pun comparison set changes the task: instead of only asking whether a sentence contains wordplay, researchers can compare puns with ordinary text.
What do annotation and humor ratings show?
The Pun Generation with Surprise research reported several evaluation sets. The SemEval development set contained 33 puns, 33 swap-puns, and 64 non-puns. Human funniness ratings were collected on 130 development sentences. Forty-eight workers participated in the Amazon Mechanical Turk rating task, each sentence was rated by 5 workers, and 10 workers were removed for low agreement. The average Spearman correlation among the remaining workers was 0.3.
The reported inter-annotator Spearman correlations were 0.57 for success, 0.36 for funniness, and 0.32 for grammaticality. These values distinguish different judgments: people can agree more about whether a task succeeded than about how funny or grammatical a line feels (Pun Generation with Surprise).
The study also reported correlations between surprisal and labels. In SemEval, the pun-versus-non-pun surprisal correlation was 0.46 with p=0.00, while the pun-versus-swap-pun correlation was 0.48 with p=0.00. In KAO, the corresponding values were 0.58 with p=0.00 and 0.26 with p=0.15. Within puns, the SemEval surprisal correlation was 0.08 with p=0.37.
Other measured relationships were similarly specific. In SemEval, ambiguity correlated 0.40 with pun-versus-non-pun funniness and 0.18 with pun-versus-swap-pun funniness. In KAO, ambiguity correlated 0.59 with pun-versus-non-pun funniness. SemEval distinctiveness correlations were -0.17 for pun versus non-pun funniness, 0.15 for pun versus swap-pun funniness, and 0.41 within puns. KAO distinctiveness correlated 0.29 with pun versus non-pun funniness and 0.27 within puns.
Unusualness correlated 0.37 with pun-versus-non-pun funniness in SemEval and 0.36 in KAO. Its pun-versus-swap-pun correlation was 0.19 in SemEval. These are reported research correlations, not rules that guarantee a funnier pun.
How do multilingual datasets compare?
Several datasets show how strongly corpus size depends on the task and language.
| Dataset or task | Quantified composition | Source label |
|---|---|---|
| SemEval-2017 Task 7 | About 4,000 contexts; 71% were puns | Large Dataset and Language Model Fun-Tuning for Humor Recognition |
| STIERLITZ | 46,608 jokes and 46,608 non-jokes; 93,216 total | Large Dataset and Language Model Fun-Tuning for Humor Recognition |
| PUNS | 213 jokes and 0 non-jokes; 213 total | Large Dataset and Language Model Fun-Tuning for Humor Recognition |
| FUN | 156,605 jokes and 156,605 non-jokes; 313,210 total | Large Dataset and Language Model Fun-Tuning for Humor Recognition |
| GOLD | 899 jokes and 978 non-jokes; 1,877 total | Large Dataset and Language Model Fun-Tuning for Humor Recognition |
| ChinesePun | 1,049 homophonic and 1,057 homographic puns; 2,106 total | Are U a Joke Master? Pun Generation via Multi-Stage Curriculum |
STIERLITZ was also reported with 37,447 jokes and 37,447 non-jokes in training, or 65,530 total examples as reported in the paper; validation contained 4,682 of each, or 9,364 total; and test contained 9,361 of each, or 18,722 total. The reported training total does not match the sum of the two displayed class counts, so the paper’s stated total should remain labeled as reported rather than silently recalculated.
FUN training contained 125,708 jokes and 125,708 non-jokes, or 251,416 total examples. Its test set contained 30,897 jokes and 30,897 non-jokes, or 61,794 total. PUNS had no non-jokes in the reported composition, unlike the balanced STIERLITZ and FUN structures.
ChinesePun contained 187,315 words in total. Its main composition counted 1,049 homophonic puns and 1,057 homographic puns. An earlier split table reported 879 training and 219 test homophonic examples, for 1,098 total, plus 1,039 training and 259 test homographic examples, for 1,298 total. A later split reported 839 homophonic training examples and 210 test examples, alongside 845 homographic training examples and 212 test examples. The study also reported a generation pipeline using 1,684 training samples and 422 test samples.
What do generation benchmarks measure?
The ChinesePun work used 10,000 preference-data pairs or triplets after sampling. Its PGCL experiments ran on 2 RTX 3090 GPUs, with a learning rate of 5e-5, fine-tuning batch size of 4, DPO beta of 0.5, decoding temperature of 0.95, top-p of 0.95, and top-k of 5 (Are U a Joke Master? Pun Generation via Multi-Stage Curriculum). The SemEval generation pipeline used 1,918 training samples and 478 test samples.
The reported ChinesePun results compare ChatGPT, AmbiPun, Baichuan2-7B, Baichuan2sft, Baichuan2dpo, and PGCL. ChatGPT had an average length of 53.94, corpus-diversity Dist-1 of 5.51, corpus-diversity Dist-2 of 41.88, sentence-diversity Dist-1 of 77.87, sentence-diversity Dist-2 of 95.17, humor A/B win of 93.07, and humor A/B lose of 28.00.
AmbiPun reported average length 79.31, corpus-diversity Dist-1 2.68, corpus-diversity Dist-2 15.50, sentence-diversity Dist-1 64.50, sentence-diversity Dist-2 85.06, structure success 28.00, pun success 42.00, humor A/B win 53.84, and humor A/B lose 16.00. Baichuan2-7B reported average length 91.12, corpus-diversity Dist-1 3.00, corpus-diversity Dist-2 26.92, sentence-diversity Dist-1 70.84, sentence-diversity Dist-2 92.36, structure success 18.00, pun success 44.00, humor A/B win 78.69, and humor A/B lose 14.00.
Baichuan2sft reported average length 85.57, corpus-diversity Dist-1 3.16, corpus-diversity Dist-2 30.94, sentence-diversity Dist-1 62.61, sentence-diversity Dist-2 85.27, structure success 20.00, pun success 46.00, humor A/B win 74.30, and humor A/B lose 16.00. Baichuan2dpo reported average length 79.26, corpus-diversity Dist-1 3.22, corpus-diversity Dist-2 30.25, sentence-diversity Dist-1 60.04, sentence-diversity Dist-2 82.74, structure success 30.00, pun success 52.00, humor A/B win 80.41, and humor A/B lose 20.00.
PGCL reported average length 85.20, corpus-diversity Dist-1 2.78, corpus-diversity Dist-2 27.52, sentence-diversity Dist-1 58.73, sentence-diversity Dist-2 81.39, structure success 64.00, pun success 22.00, humor A/B win 89.10, and humor A/B lose 44.00 (Are U a Joke Master? Pun Generation via Multi-Stage Curriculum).
Taken together, these pun statistics measure different things: collection scale, linguistic form, human agreement, classifier performance, and generated-output preferences. Keeping those denominators and source labels visible is essential when using the numbers to understand or write better wordplay.