Skip to main content

Names: Poetry and Probability

A name generator is only as Irish as the pool you let it learn from. Most of the week was deciding what --region ireland is allowed to mean.

Names: Poetry and Probability
Software11 min readNLPNamesEtymology

Written by

Kartik Jha

Published on

August 17, 2026

Saorleea. Crionne. Kaikiyotsu. Fusahide.

None of those strings were in the training lists - every three-letter fragment inside them was. Which is both the trick and the danger: a letter model will happily reproduce whatever culture you hand it, wrong one included.

import { generateNames } from '@alistairheus/name-generator';
 
generateNames({
  strategy: 'markov',
  gender: 'female',
  region: 'ireland',
  pool: 'exclusive',
  count: 10,
});

The call is small because the hard part moved into the data. --region ireland has to mean something you can defend. Most of the week was spent deciding what that something is.

This runs on Node 20+, reading JSON with node:fs and sampling with node:crypto - fine for a worldbuilding server, not something you drop into a React bundle.

The Takeshi problem

Language isn't a bag of letters. Start a name with c and English will hand you vowels far more often than b or f - the same is true everywhere, except now the probabilities belong to a culture, not to "letters in general."

So the training set doesn't sit upstream of the product - it is the product, and a bad one produces a model that's wrong with total confidence.

Sweden's first SCB import pulled the 2-bearer tilltalsnamn stock as of 31 December 2022 - about 49,376 male names and 57,768 female names, close to every calling name registered in the country, Takeshi included.

--region sweden then meant nothing more than "SCB listed this string somewhere." True, but useless as a training set if the goal is Swedish-shaped invention - Markov or morph run on that list would learn a global mix and call it Swedish spelling.

The cutoff moved to 200+ bearers, matching the threshold Norway's SSB table 10467 already used. Sweden shrank to 1,711 male and 1,787 female names. Norway, on the same 200-person rule, has 932 male and 1,042 female names. Takeshi lost the Sweden tag. He's tagged japan now.

Ireland's CSO tables produced the mismatch in the opposite direction - Nusaybah and Hadassah both appear in Irish birth records, and drawing them with --region ireland --pool register is entirely faithful to the register. Feeding them to a letter model as "Irish phonology," though, is a different claim, and not one the register data actually backs up.

PyPI's names-dataset looked like a shortcut and wasn't one - it derives from the 2021 Facebook leak, and we never downloaded the 3.3GB person-level CSV to find out how bad it was.

Where the names came from

The 12-13 August push built the unique files: football and player CSVs first, then japanese-names-dataset.csv plus some romanized leftover CJK (941aadd), French prenoms (ba97ea5), a mixed Nordic txt/csv (41d7170), and a small A Song of Ice and Fire extract (ca6985f). Last names got ingested in the same pass and then never wired into generation - they're still sitting in data/last_names.json, gitignored, waiting.

The source CSVs were gitignored too, under a global *.csv / *.txt rule, which meant that by the time we actually needed provenance, the files themselves were just gone. Japanese, French, and Nordic tags had to be reconstructed from git deltas instead of from the source data - not a great way to spend an afternoon - and those tags are correspondingly lower-confidence than the official tables that arrived later. The Japanese delta also mixed the CSV in with romanized CJK from the sports dump, and missed Takeshi entirely, since he'd already made it into the unique list before 941aadd landed.

Japan was rebuilt as an intersection: romanizations already in the unique files that also show up in shuheilocale/japanese-personal-name-dataset. That lands at 3,803 male and 3,008 female names, all romaji - there's no national Japanese register behind any of it.

Official lists that did land, with source metadata, live in data/regions/:

TagMaleFemaleSource
ireland28443759CSO VSA50/VSA60, 3+ births, 1964-2025. Count is summed births.
scotland25603701NRS babies' first names, 1974-2025. Count is summed births.
norway9321042SSB 10467, 200+ bearers. No per-name count in the dump.
sweden17111787SCB tilltalsnamn, 31 Dec 2022, 200+ bearers. Count is bearers.
germany1117211202Cologne births, 2010-2017 and 2023. Count is summed births.

Destatis doesn't publish a national given-name file at all, so germany is really just Köln. The Cologne copies carry two different licenses - CC BY 3.0 DE for fxnn/vornamen's 2010-2017 data, and Datenlizenz Deutschland - Zero - 2.0 for 2023 - and the counts are annual births, not a running total. 2023 alone accounts for about 2,657 male names; the union across all the years comes to roughly 11,178. We kept every year rather than collapsing them.

Denmark, Finland, and Iceland are still open. Denmark's data sits behind DST's search UI with no bulk export, Finland's Avoindata endpoint 403s from this environment for reasons I never tracked down, and Iceland's Mannanafnaskrá has no bulk CSV at all. All three just need someone to sit down and drop a file into data/raw/.

Data storage

The unique JSON under data/unique/ is a cleaned merged index: 83,610 male names, 90,067 female names. A record looks like this:

{"name": "Aoife", "regions": ["germany", "ireland", "scotland"], "counts": {"germany": 2, "ireland": 22063, "scotland": 699}}

Aoife's Irish count of 22,063 is summed births over decades, with annual suppression baked in; the Cologne 2 is births in a handful of selected years and name positions. These aren't the same unit, and using them to weight a model only makes sense after per-source scaling - log counts, or probabilities computed within a single file. Add them straight across countries and you get a number that means nothing.

loadNameData() doesn't touch that index anymore. Generation instead unions data/regions/ (or the minified copies in pack/data/regions/), which comes to 26,606 male names and 30,219 female names, every one of them tagged. The untagged sports-data strings stay behind in the unique files for future rebuilds but never make it into the published tarball.

regions still mixes together kinds of label that don't really belong in one field - jurisdictions (Ireland, Scotland, Norway, Sweden), one city standing in for a country, reconstructed collections (Japan, France, Nordic), and fiction (ASOIAF). It's an unfinished schema more than a deliberate design; registers, origins, and collections should probably be split into separate fields eventually.

JSON stayed, for now. Parsing took about 12ms on the 3MB unique file; the slow part was Zod plus cleanName running over every string, at roughly 545ms on the unique dump and 230ms after the regional loader trimmed things down. SQLite would be the better store the day last names, SQL filters, and frequency queries become daily work - that day hasn't come yet, so we didn't switch.

--region and --pool

--region keeps names that carry that tag. On top of that, --pool picks a view:

ModeWhat you get
registerEvery name tagged with the region (default)
exclusiveOther tags stay inside a small overlap set: Ireland/Scotland/Cologne/France, or Norway/Sweden/nordic. Japan exclusive allows only japan
morphRegister names that pass European stem-and-suffix eligibility

exclusive means, precisely, "no conflicting tag in the datasets we currently have" - a narrower and more honest claim than "no conflicting tag, period." Missing coverage just leaves an origin unknown rather than resolved; Oluwaseun survives Irish exclusive not because the dataset has confirmed anything about him, but because it doesn't tag him as pan-regional either. Reach for --pool morph instead when the claim you actually want to make is about European given-name shape.

morph is that shape, concretely: Latin names carrying productive endings like -ina, -elle, -ian, -o. It skips Japanese morphemes (suke, tsu, ichi) and prefixes like abd or moham, so it won't isolate Gaelic and it won't emit anything Japanese-shaped.

random samples the selected view without replacement, via crypto.randomInt. Twenty-one tests cover parsing, filters, and generation - mostly just checking that the software does what the flags claim it does.

Invented names

morph and markov both train in memory, on whatever view you've already selected, and nothing gets saved to disk. "Train" here just means: count statistics from that pool, then sample from them.

Markov here is order-3 over letters: take the last two characters, ask the pool what actually followed them in real names, sample from that distribution, append, slide the window, repeat.

Order 1 is a bag of letters. Order 2 is "what follows c?" Order 3 is "what follows cr?" Higher orders start copying the training set. We stop at 3.

Each a-z name (accents folded, hyphens dropped) gets one vote; source counts don't factor in. The walker starts from ^^^, so what it learns as "the first letters of a name" are actually the first letters of names, not some slice out of the middle of one. It stops at $, throws out anything that copies the training set outright, and keeps results between 3 and 16 characters. Training on Irish exclusive female took about 7ms - cheap enough that writing the triples to disk would cost roughly as much as just rebuilding them, and the cached file would go stale the moment the lists or the exclusive rules changed anyway.

scripts/eval-markov.ts holds out 10% of a view, generates 100 names from the rest, and reports collisions plus how many generated triples already existed in train:

ViewPool sizeHoldout hitsSample
Ireland exclusive female22381Ryl, Crionne, Saorleea, Eshaelynn
Ireland morph female23141Eina, Yevanna, Antoinelle, Hilde
Japan exclusive male36770Fusahide, Kaikiyotsu, Tenshusuke
Sweden exclusive female4210Irjo, Gretheres

Every run produced 100 strings that don't exist in training, built entirely out of triples that do. Kaikiyotsu is attested Japanese-looking fragments recombined in a new order; Saorleea is the Irish exclusive pool doing the identical job. None of the cultural specificity here comes from the algorithm - Japan exclusive stays inside romanized Japanese letter patterns and Ireland exclusive stays inside Irish fragments because the pool enforces it, not because Markov chains know anything about language.

Run the same algorithm on a mixed world list and you get the opposite outcome - fragments from everywhere, names that belong nowhere. That's exactly what the 2-bearer Swedish stock would have taught the model to call "Swedish." No amount of clever sampling fixes that after the fact. The pool has to do that work, because nothing downstream of it will.

A reject list for illegal letter clusters would be the next cheap experiment to try. Count-weighted Markov is further out - it needs per-source scaling worked out first. And if we ever train something larger, it should be looking at --pool exclusive or --pool morph on the regional union, not the raw unique index.

Packaging and publishing

Unscoped name-generator and fantasy-name-generator were both taken, unsurprisingly, so the package lives at @alistairheus/name-generator. prepack compiles to dist/, minifies the regional JSON into pack/data/regions/, and rewrites dist/cli.js with an LF shebang - npm had been silently dropping bin whenever that file shipped with CRLF line endings on Windows, which took a minute to track down.

The tarball comes out to about 329 kB packed, 1.9 MB unpacked, across 60 files - Cologne alone accounts for most of that JSON. The unique index stays in git rather than in the package. 1.0.0 went out on 17 August 2026.

The layout after b316426: src/ and test/ unchanged, rebuild scripts moved under scripts/, the one-shot August 12-13 Python moved to scripts/legacy/ (though the paths inside those files still point at the old root), bulk downloads in data/raw/ (ignored), and ASOIAF extracts in data/sources/. The global CSV/txt ignore rule is gone now, so the next raw dump can actually be committed if that turns out to be useful.

Open gaps, plainly stated: no Denmark, Finland, or Iceland; no Destatis national file; last names still unwired; mixed regions semantics; counts that can't be compared across sources; Markov running unweighted; no phonotactic filter. But the flags already commit you to a claim you can actually stand behind. --region ireland --pool exclusive --strategy markov means: invent from letter triples seen in names this dataset tags as Irish, and doesn't also tag outside a defined overlap set. Narrower than "an Irish name" - and unlike that claim, checkable.

Let's get to work!