From d98d4766777f5bb3c3ec05ad1bf3aa77e7efdd95 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sun, 4 Oct 2026 02:54:32 -0700 Subject: [PATCH 1/2] Add salutations, offices and ranks in other languages to TITLES (#606) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TITLES had no salutation for Finnish, Estonian, Romanian, Croatian/Serbian, Icelandic, Czech, Lithuanian, Malay, Filipino, Georgian and others, so "Herra Väinö Johansson" read given 'Herra'. 41 words join TITLES, grouped by language in titles.py, and knt (Knight) joins SUFFIX_ACRONYMS beside kt with a 'Knt' case mask. Thirteen of the words are borne as names, though rarely as the first word of one (greve, graaf, knight, herra, hrabia, ...). They ship on the graf precedent (Derek's call): a name that begins with one gives it to the title, as "Graf Steffi" already did. rules.md#H1 gains an Accepted clause for that cost, and decisions.md#salutation-titles records the per-word counts plus the Excluded (TITLES) rows (ông/bà, thiru, puan, pan, rodina, gróf, marshal, justice, ...). The one corpus mover is the new H1 example "Greve Anna", classified in all five ledgers; three comma rules' claims grow by "Greve, Anna". Co-Authored-By: Claude Opus 5.5 --- AGENTS.md | 2 +- docs/design/decisions.md | 18 +++++ docs/design/rules.md | 11 +++ docs/release_log.rst | 2 + nameparser/config/capitalization.py | 1 + nameparser/config/suffixes.py | 1 + nameparser/config/titles.py | 74 ++++++++++++++++++++ tests/v2/cases.py | 49 +++++++++++++ tests/v2/test_facade_cases.py | 3 + tests/v2/test_ledger_guards.py | 36 ++++++++-- tools/differential/corpus_rules.jsonl | 3 + tools/differential/expected_since_1.4.0.toml | 19 +++++ tools/differential/expected_since_2.0.0.toml | 19 +++++ tools/differential/expected_since_2.1.0.toml | 19 +++++ tools/differential/expected_since_2.2.0.toml | 19 +++++ tools/differential/expected_since_2.3.0.toml | 19 +++++ 16 files changed, 289 insertions(+), 6 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 0bae9fba..25727bf4 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -367,7 +367,7 @@ Add a dedicated `copy.deepcopy()` round-trip test for it too (see `test_regexes_ **`HumanName.C` is a property backed by `_C`, but pickles under the public key `'C'`** — `__init__`/direct assignment route through the `C` setter, which calls the shared `_validate_constants` staticmethod (also used by `__init__`) so an invalid value raises `TypeError` immediately instead of surfacing later as an unrelated `AttributeError` deep in parsing (#239). `__getstate__`/`__setstate__` deliberately translate `self._C` ↔ a `'C'` key in the pickled dict (with the usual `CONSTANTS`-singleton-becomes-`None` sentinel) rather than pickling `_C` directly, so the on-disk pickle format hasn't changed across this fix — don't "simplify" that translation away or old pickles/tests that hand-build a state dict with a `'C'` key will break. -**Titles permanently shadow first names — be conservative** — any word in `TITLES` is always consumed as a title and can never be parsed as a first name. `"Dean"` is the canonical example: it's a common academic title *and* a common given name, so it is intentionally absent from the default titles (see `docs/customize.rst` — users who need it add it via opt-in `Constants`). Before adding a word to `TITLES`, ask: "Could this plausibly be someone's given name in any culture?" If yes, don't add it globally; it belongs in caller-supplied `Constants` instead. This same caution applies to international honorifics — `Prince`, `Sheikh`, `Frau` are all first names in some contexts. It also applies to any prefix sub-set gated on "never a first name": obscure-looking foreign particles are surprisingly often real given names — `Von` (Von Miller), `Vander` (Brazilian, also the Arcane character). When unsure, exclude — a missing member just means that name isn't auto-handled, whereas a wrong member misparses a real person. +**Titles permanently shadow first names — be conservative** — any word in `TITLES` is always consumed as a title and can never be parsed as a first name. `"Dean"` is the canonical example: it's a common academic title *and* a common given name, so it is intentionally absent from the default titles (see `docs/customize.rst` — users who need it add it via opt-in `Constants`). Before adding a word to `TITLES`, ask: "Could this plausibly be someone's given name in any culture?" If yes, don't add it globally; it belongs in caller-supplied `Constants` instead. This same caution applies to international honorifics — `Prince`, `Sheikh`, `Frau` are all first names in some contexts. It also applies to any prefix sub-set gated on "never a first name": obscure-looking foreign particles are surprisingly often real given names — `Von` (Von Miller), `Vander` (Brazilian, also the Arcane character). When unsure, exclude — a missing member just means that name isn't auto-handled, whereas a wrong member misparses a real person. The decided exception is a word borne as a name only RARELY in the leading position — mostly a surname written last (`graf`, `greve`, `knight`), a few with small given-name or surname-first counts (`herra`, `hrabia`): the title claim acts only in front, so it ships, and a name that begins with it (`Greve Anna`, `Herra Wijaya` → title) is the accepted cost. That is a judgment by incidence, made per word with the counts recorded in `decisions.md#salutation-titles` (rules.md#H1's Accepted clause states the cost), not a rule a sweep can apply: a word that commonly leads a real name (Vietnamese `Ông`, the given names `thiru`, `puan`, `marshal`) is still out. **The period-abbreviation title inference runs at the head of the GIVEN-NAME part, not the head of the name** — an unrecognized multi-letter word ending in a single trailing period (`_pieces._PERIOD_ABBREV`, a hand copy of the `period_abbreviation` regex, `{2,}` letters — it was assign's until #424 and group's until #439) is treated as a title in the leading title run, e.g. `"Insp. Jane Morse"` → `title='Insp.'`. "Leading" is per SEGMENT: `_peel_leading_titles` is called for NO_COMMA segment 0, SUFFIX_COMMA segment 0, and FAMILY_COMMA **segment 1**, so `"Morse, Det. Insp. Jane"` → `title='Det. Insp.'` and a lone `"Smith, Xyz."` → `title='Xyz.'` — long-standing, verified against 1.4.0, and the mechanism behind #296 (`"Smith, Jr."` → title, which the shape rule claims even once `jr` leaves `TITLES`). The docs said "leading word" until 2026-08-01 and were wrong for every comma path. It does not mutate `C.titles`, so the periodless form (`"Insp"`) is unaffected elsewhere. The `{2,}` length requirement — not a separate initials check — is what excludes single-letter initials like `"J."`; the same word after the given name is left as a middle name. **The inference OUTRANKS vocabulary where it runs**: `"Esq. Smith"` → `title='Esq.'` even though `esq` is suffix-only vocabulary, because the shape rule fires before anything consults the suffix sets. **The INFERENCE still runs in one direction only, and that is what the two slots share and where they part** (rewritten 2026-09-08, #316): a period-marked word is claimed by SHAPE at the front and by VOCABULARY at the back, so an unlisted abbreviation opening a name is a title while an unlisted abbreviation ending one is a NAME word (`"John Smith Xyz."` → `family='Xyz.'`). A trailing period-marked word the vocabulary knows as a title now IS one (`"John Smith Prof."` → `title='Prof.'`, `rules.md#H5`) — the comma path always read it that way and the two agree now — and a trailing period-marked word the SUFFIX vocabulary knows is still a suffix, the suffix run being peeled first (`"John Smith Esq."` → `suffix='Esq.'`). A BARE trailing title word stays a name word (`"John Smith Sir"`, `"Mary Jane King"`), which is the doctrine line below. Meanwhile `period_joined_vocab` resolves INTERIOR-period tokens (`Lt.Gov.`, `Msc.Ed.`) to title-or-suffix by vocabulary, and `_extract._suffix_shaped` treats any period-final delimited content as not-a-nickname. Four sites, and the trailing one now resolves by two vocabularies in order rather than one; unifying the rest is open design work, not settled. (#109; see `docs/usage.rst` "Titles you didn't configure") diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 477f807e..3d20c9b6 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -739,6 +739,15 @@ Closes #346, #344 and #343 as one bundle: each CLASS is decided once, and each s - **Not recorded as excluded:** the Tier 4 candidates — transliterated military ranks, Advocate, Engineer, Marhum, and the kinship forms — are skipped pending a corpus check, not declined, and मां/माता/মা and the hafiz/syed/shahid forms are left alone the same way. An Excluded entry is a standing prohibition, and none of these has been argued to that standard. One of them carries a duty anyway: প্রকৌশলী (Engineer) is rules.md#H2's executable example of an unlisted abugida honorific, so whoever ships it must move the example to another unlisted word first, or the doc test and the contract-tier gate fail together. - **Measurement (2026-09-06).** Four corpus names move, all under the renunciate fold, all in one direction (the family empties and the given fills): Swami Vivekananda, Guru Nanak, Baba Ramdev and Lama Zopa, every one from the radar-tier corpus_issues.jsonl. Recompute by parsing every name in the tools/differential/corpus*.jsonl glob twice — once with the shipped lexicon, once with `Lexicon.default().remove(given_name_titles={"swami","guru","baba","lama"})` — and diffing the seven name fields. The multi-name spellings are the boundary and do not move: Swami Vivekananda Saraswati and Guru Gobind Singh keep families Saraswati and Singh, rules.md#H1's fold requiring exactly one name word. The Devanagari and Bengali additions move no corpus name at all, the corpora holding no name written in either script apart from rules.md's own H2 example, which this bundle rewrote. The gate's intentional counts rose by four at every baseline; read today's from the `corpus:` line of `uv run python tools/differential/compare.py --baseline X`. +### salutation-titles — salutations, offices and ranks for the languages TITLES had none for (2026-10-04, #606) + +Closes #606. No parser code moves: 41 words join TITLES and one, knt (Knight, beside kt), joins SUFFIX_ACRONYMS with a `Knt` case mask. The class is decided once — a word that addresses or ranks a person and is written in FRONT of the name — and each word is then argued on #vocabulary-collisions C-i, the position test, because TITLES has no ambiguous subset and the only two answers available are ship and do not ship. The per-word lists, grouped by language, are the #606 block in nameparser/config/titles.py. + +- **The graf class, decided (Derek, 2026-10-04).** A word borne as a name, but RARELY as the first word of one, ships anyway, on the precedent of graf, which shipped though Steffi Graf exists. This is a judgment on each word's incidence in the leading position — the position C-i asks about — and not a rule: it relaxes C-i's "under uncertainty, default to ambiguous", which for TITLES means do not ship, for words whose counts were measured and found small. The title claim acts only in front, so most bearers are untouched in the forms they write: `Gladys Knight` and `Knight, Gladys` keep family Knight. The cost falls on a name that BEGINS with the word, the cost `Graf Steffi` already paid: `Greve Anna` reads title Greve, family Anna, and a given name is lost the same way, `Herra Wijaya` reading title Herra, family Wijaya. Declaring `FAMILY_FIRST` does not help, since the title claim is read before the order (`Hrabia Anna` reads title Hrabia under it); the family comma does. Two further limits, both pre-existing: a trailing period makes the word the title by rules.md#H5 (`Gladys Knight.` reads title `Knight.`, as `Mary Jane King.` does), and every word this entry adds widens that rule's reach, H5 claiming any listed title behind a period. Accepted rather than prevented, and stated at rules.md#H1's Accepted clause. The issue proposed the class for three words — greve (Scandinavian Count), graaf (Dutch Count), knight. The spot-check the issue asked for (forebears.io, 2026-10-04, worldwide counts rounded) found ten more of its "no known collision" words borne, and Derek extended the decision to all of them rather than excluding them: herra (~3,000 surnames, Costa Rica and the Dominican Republic; ~670 given-name uses, Indonesia and the Philippines), batoni (~1,100, the painter Pompeo Batoni; ~300 given-name uses, Tanzania), ponas (~4,000, Iran), counsel (~730, Australia, the US, England), ponia (~560), the Filipino ginoo, ginang and binibini (~500 each, Philippine lists writing Surname, Given with a comma), and two whose registers often write the surname first, hrabia (~750, Poland) and kreivi (~270, Finland). Those last two, and the given-name uses of herra and batoni, are the leading-position bearers this class accepts, and were decided on with their counts in hand. Where the line falls is the counts and nothing finer: thiru and puan are out as given names COMMON in the leading position, rodina as a surname common in surname-first Russian lists, and gróf for want of any count at all. The comment `# borne` in titles.py marks each member, so a sweep meets the decision where it meets the word. +- **Accents are not folded, and two words depend on it.** `_lexicon._normalize` lowercases and composes NFC and does nothing else to letters, so paní and härra cannot match Pani (~76,000 surnames and ~65,000 given names, India and Indonesia) or Harra (~1,100). Accent-folding the lexicon would turn both into collisions; whoever proposes it re-argues this entry. frú depends on nothing: the unaccented fru (Swedish/Danish Mrs) already ships, so the Cameroonian surname Fru (~10,800, where the surname is often written first) already reads as a title in `Fru Anna` and did before this entry; that pre-existing collision is recorded here rather than argued, being no part of #606. +- **gravos was dropped, not excluded.** The issue listed it as the Latin-script Greek for Count; the spot-check could not confirm it is the Greek word at all, which is κόμης (kómis). Nothing was shown, so there is nothing to keep out — it is recorded here so that a later sweep starts from that rather than from the issue's table. κόμης itself is NOT added and NOT excluded: it is a Greek surname (~420 bearers in Greece, official lists writing the surname first), which by its counts sits inside the class above beside hrabia, and it was left out of this change for having no use in current Greek address rather than for its collision. Adding it is a decision still to make. +- **Measurement (2026-10-04).** One corpus name moves, `Greve Anna`, given Greve to title Greve: the rules.md#H1 example this entry added, and the only one of its three examples that moves. Recompute: take the 41 entries of the #606 block in titles.py — the lines that are a quoted string, not the words quoted inside its comments, which name existing entries and excluded ones — parse every distinct name in the tools/differential/corpus*.jsonl glob (1,518 on the day) with the shipped parser and with `Lexicon.default().remove(titles=, suffix_acronyms={"knt"})`, and diff `as_dict()`. Before the rules.md examples were added, the same diff moved nothing over the 1,809 corpus lines of the tree before #604 merged, gravos still in the word list. The old readings the release-log bullet quotes were checked against the 1.4.0 and 2.3.0 wheels. + ### cjk-comma-demotion — the script shapes are pure, the wrappers are tolerated (2026-09-01, #469) Closes #469, and continues the corpus-tier arc below rather than standing apart from it: the tier split gave the differential somewhere to WATCH a name without promising it, and this is the first doctrine narrowing to spend that. No parser behavior moves anywhere in it. Counts are this session's and every one is recomputable from the checked-in tree — `wc -l` over tools/differential/, `uv run python tools/differential/build_shapes_corpus.py --coverage`, and one `uv run python tools/differential/compare.py --baseline X` run per baseline, read off its `corpora:` and `corpus:` lines. @@ -857,6 +866,15 @@ Excluded (TITLES): - शेख / শেখ (Sheikh) — a clan name and a family name in both scripts ("শেখ হাসিনা"). The divergence from Latin is deliberate and is the sri/shri case run backwards: the Latin sheikh/sheik/shaykh/shaikh cluster SHIPS as given-name titles for the Arabic addressing form, while the Indic spellings name the family (2026-09-06, #344/#343). - आचार्य (Acharya) — a Brahmin surname (2026-09-06, #344). - राजा / रानी (Raja/Rani) — common given names (2026-09-06, #344). +- ông / bà (Vietnamese Mr / Mrs) — Ông is a Vietnamese surname and Vietnamese names lead with the surname, the position the title claim acts on; pinned by the case row salutation_that_leads_a_surname_stays_a_name (2026-10-04, #606, #salutation-titles). +- thiru (Tamil Mr) — a Tamil given name (2026-10-04, #606). +- puan (Malay Mrs) — an Indonesian given name, Puan Maharani (2026-10-04, #606). +- pan (Polish/Czech Mr) — the Chinese surname Pan, which leads (2026-10-04, #606). The feminine paní ships, its accent keeping it off the name Pani; the unaccented pani does not. +- rodina (Czech "family") — a Russian surname, often written surname-first (2026-10-04, #606). familie ships. +- gróf (Hungarian Count) — possibly a surname in a surname-first culture; unverified, and under uncertainty the answer is do not ship (2026-10-04, #606). +- cik, paron, tikin, janab (Latin spellings) — unverified, so not shipped. Latin janab follows the sri/shri pattern: Devanagari जनाब ships, and the Latin spelling needs its own check (2026-10-04, #606). +- marshal — a given name (Marshal Yanda). The accepted cost is that `Field Marshal Bernard Montgomery` keeps reading title Field, given Marshal (2026-10-04, #606). +- justice — a given name (Justice Smith). Likewise `Chief Justice John Roberts` keeps title Chief, given Justice (2026-10-04, #606). - The trailing-position rule that must NOT be adopted, and 2026-09-08 says what the line is: TITLES holds hundreds of words in no suffix set, at least nineteen of them ordinary English surnames (king, judge, bishop, baron, sheriff, ...), so a blanket "vocabulary outranks position in the trailing slot" reading would turn "Mary Jane King" into title="King" with the family name gone. The leading half of this argument is AGENTS.md's "Dean is deliberately absent" gotcha; this is the trailing half, and it shadows the family name rather than the given (#316). What #316 settled is that the prohibition is on BARE words: "Mary Jane King" still reads family King, while "John Smith Prof." reads title Prof. **The second half of that argument is RETRACTED 2026-09-09**, in review of the docs commit: "a period-marked trailing word being one nobody writes as a surname" claimed the period keeps the collision out of the trailing slot, and it does not. Measured on the branch tree — "Mary Jane King." reads title "King.", given Mary, family Jane, the family name shadowed exactly as the bare reading would have shadowed it, and "John Smith Judge." reads title "Judge." The reach is the vocabulary rather than a handful: of the 746 titles in no suffix set, 627 read as a trailing title once a period is written behind them, and every one of the 119 that do not is kept out by the abbreviation SHAPE — a digit, a hyphen, an apostrophe, or a script whose letters carry combining marks (rules.md#H2 records that half) — no plain ASCII-letter title missing at all. So king, judge, bishop, baron and sheriff ARE reachable in the trailing slot; what the prohibition above buys is the bare spelling and nothing more. The period is a WRITING convention the rule can read, not evidence about the word, and reading it is an ACCEPTED COST under the premise that the input is a name (rules.md H Background): someone who ends a name with a period-marked word has written an abbreviation. rules.md#H5 carries it as an Accepted clause with "Mary Jane King." as the executable example. Recompute the reach with `L = Parser().lexicon; ws = [w for w in L.titles - L.suffix_acronyms - L.suffix_words if " " not in w]; len([w for w in ws if parse("Mary Jane %s." % w).title])` against `len(ws)` — 627 of 746 on 2026-09-09 — the digits move as the vocabulary grows and neither argument does. See #H5. Excluded (Policy.script_orders defaults): Script.KATAKANA is deliberately absent — a pure-katakana token is predominantly a transcribed foreign name kept in its source order, so nothing defaults on it (rule W4's boundary). Noted 2026-08-15: of the three Script-keyed axes, this is the one with no force-a-decision guard (mechanisms.md#FORCE-A-DECISION-TABLE), so a new Script member silently gets no order. diff --git a/docs/design/rules.md b/docs/design/rules.md index 60628924..35892d88 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -85,6 +85,17 @@ H1. Rationale: a title normally addresses by surname, so a title title or surname for every peer and every wife (`Lord Byron`, `Lady Thatcher`), and the text does not say which the bearer is, which the list cannot express; `prince` and `princess` are in it. + Accepted: a title word that is also borne as a name, though + rarely as the first word of one (`graf`, `greve`, `herra`), is + the title in front of a name, so a name that BEGINS with it -- a + surname written first without a comma, or a given name -- gives + that word to the title, under a declared family-first order too. + Written last without a period, or before a family comma, it is + the name; a period behind it makes it the title by H5, as for + `Mary Jane King.` (#606). + "Greve Anna" → family="Anna" + "Anna Greve" → family="Greve" + "Greve, Anna" → family="Greve" history: decisions.md#H1 · interacts: H3, H5, P2, P3, P5, M2, S1, S2, N1, N3 · implemented: nameparser/_pipeline/_post_rules.py H2. Rationale: before a name, an abbreviation is almost always a diff --git a/docs/release_log.rst b/docs/release_log.rst index 057cafa4..b3020b77 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -76,6 +76,8 @@ Release Log - **Fix the Irish particles Ó, Ní and Ua and the Malay binti being read as a middle name.** ``HumanName("Liam Ó Murchú")`` gives first ``Liam``, last ``Ó Murchú``, where 1.4.0 through 2.3.0 gave middle ``Ó``, last ``Murchú``; ``Sinéad Ní Mhurchú``, ``Seán Ua Buachalla`` and ``Ina binti Navalamar`` (and the Singapore spelling ``binte``) move the same way. ``Ó`` and ``Ní`` are never given names, so ``Ó Murchú`` alone is all last name, where every release gave first ``Ó``. ``Ua`` and ``binti`` can be, so a leading one stays the first name and ``parse()`` reports ``particle-or-given``: ``Ua Buachalla`` gives first ``Ua``, last ``Buachalla``, as before. ``Ó.`` written with a period is still an initial: ``Juan Ó. Pérez`` keeps middle ``Ó.``, while ``Juan Ó Pérez`` gives last ``Ó Pérez``. Case repair writes the Irish particles capitalized, ``SEÁN Ó MURCHÚ`` repairing to ``Seán Ó Murchú``, and ``binti`` in lowercase: ``INA BINTI NAVALAMAR`` repairs to ``Ina binti Navalamar``, where 2.3.0 gave ``Ina Binti Navalamar``. The Irish casing comes from new ``capitalization_exceptions`` entries, which now outrank the lowercase case repair gives a particle, and that holds for your own entries too: with ``constants.capitalization_exceptions['van'] = 'Van'``, ``ludwig van beethoven`` repairs to ``Ludwig Van Beethoven``, where 1.4.0 through 2.3.0 kept ``van``. See ``P7`` and the #604 entries under ``vocabulary-collisions`` and ``R4`` in ``docs/design/decisions.md`` (closes #604) + - **Fix salutations in Finnish, Estonian, Romanian, Croatian/Serbian, Icelandic, Czech, Lithuanian, Malay, Filipino and other languages being read as a first name.** ``HumanName("Herra Väinö Johansson")`` gives title ``Herra``, first ``Väinö``, last ``Johansson``, where 1.4.0 through 2.3.0 gave first ``Herra``, middle ``Väinö``. The titles gain Mr/Mrs/Miss forms such as ``rouva``, ``proua``, ``doamna``, ``gospođa``, ``frú``, ``meneer``, ``paní``, ``ponas``, ``encik`` and ``ginang``; the word for Count in several of them (``kreivi``, ``krahv``, ``hrabia``, ``greve``, ``graaf``); ``familie`` (``Familie Hansen`` gives title ``Familie``, last ``Hansen``); ``knight``; and the offices ``commissioner``, ``counsel`` and ``administrator``, which complete titles that half-worked: ``Police Commissioner James Gordon`` gives title ``Police Commissioner``, first ``James``, where it gave title ``Police``, first ``Commissioner``. Some of the new titles are also names, nearly always the last name (``Greve``, ``Knight``), and those still read as the last name there: ``Gladys Knight`` and ``Knight, Gladys`` are unchanged (``Gladys Knight.``, with a trailing period, gives title ``Knight.``, as ``Mary Jane King.`` already does). The cost falls on a name that begins with one of them, the cost ``Graf`` already pays: ``Greve Anna`` gives title ``Greve``, last ``Anna``, and so does ``Hrabia Anna`` under ``FAMILY_FIRST``; the few borne as first names (``Herra``, ``Batoni``) lose them the same way. Words that commonly lead a real name stay out, among them Vietnamese ``Ông``, Polish and Czech ``Pan``, and ``Marshal`` and ``Justice``. ``Knt`` (Knight) joins the post-nominals beside ``Kt``: ``Sir John Smith Knt`` gives suffix ``Knt``, where it gave middle ``Smith``, last ``Knt``. See the ``salutation-titles`` entry of ``docs/design/decisions.md`` (closes #606) + **Additions** - **Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials.** A subset of ``conjunctions`` holding ``e`` and ``i`` by default; it is the knob for the change above rather than a switch. Portuguese data, where ``e`` links surnames the way ``y`` does in Spanish, takes it out: ``Lexicon.default().remove(conjunctions_ambiguous={"e"})`` restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one: ``Lexicon.default().add(conjunctions_ambiguous={"y"})``. A v1 ``Constants`` has no manager of its own for it -- deleting the word from ``conjunctions`` is what turns the marking off, the same rule the glued-honorific tails follow. See ``docs/customize.rst`` (#383, #479) diff --git a/nameparser/config/capitalization.py b/nameparser/config/capitalization.py index 2249b017..9c41b23e 100644 --- a/nameparser/config/capitalization.py +++ b/nameparser/config/capitalization.py @@ -17,6 +17,7 @@ 'dmin': 'DMin', 'drph': 'DrPH', 'dsc': 'DSc', + 'knt': 'Knt', 'kt': 'Kt', 'mdiv': 'MDiv', 'pharmd': 'PharmD', diff --git a/nameparser/config/suffixes.py b/nameparser/config/suffixes.py index 0cf723b3..f8d8c73e 100644 --- a/nameparser/config/suffixes.py +++ b/nameparser/config/suffixes.py @@ -714,6 +714,7 @@ 'kcvo', 'kg', 'khs/dhs', + 'knt', # #606: Knight, the older spelling of 'kt' 'kp', 'kt', 'lac', diff --git a/nameparser/config/titles.py b/nameparser/config/titles.py index c9ecec4d..80b63d38 100644 --- a/nameparser/config/titles.py +++ b/nameparser/config/titles.py @@ -766,6 +766,80 @@ 'writer', 'zoologist', + # #606: salutations and offices for the languages the block above + # had none for, so "Herra Väinö Johansson" stops reading given + # 'Herra'. TITLES has no ambiguous subset, so a word ships only if + # the leading claim misreads nobody often enough to matter + # (decisions.md#vocabulary-collisions C-i). Those marked "borne" + # are names somewhere, but rarely the FIRST word of one -- mostly + # surnames written last. They ship on the 'graf' precedent above, + # paying the cost "Graf Steffi" already pays: a name that begins + # with one gives it to the title (rules.md#H1). That is a + # judgment on each word's counts, not a rule to sweep with. Counts + # and sources are in + # decisions.md#salutation-titles, and so are the words that failed: + # among them Vietnamese ông/bà, Tamil thiru and pl/cs pan, all in + # its Excluded (TITLES) block. + # English offices, completing runs like "Police Commissioner". + # 'marshal' and 'justice' stay out: both are given names. + 'administrator', + 'commissioner', + 'counsel', # borne + # Finnish: Mr, Mrs, Miss, Count. + 'herra', # borne + 'rouva', + 'neiti', + 'kreivi', # borne + # Estonian: Mr, Mrs, Miss, Count. + 'härra', + 'proua', + 'preili', + 'krahv', + # Romanian: Mr, Mrs (with and without the breve), Miss. + 'domnul', + 'doamna', + 'doamnă', + 'domnișoara', + # Croatian/Serbian/Bosnian, Latin script: Mr, Mrs, Miss. + 'gospodin', + 'gospođa', + 'gospođica', + # Icelandic: Mrs, Miss, Count ('herra', Mr, is the Finnish entry). + 'frú', + 'ungfrú', + 'greifi', + # Dutch: Mr, Miss, Count ('mevrouw' is above). + 'meneer', + 'juffrouw', + 'graaf', # borne + # Scandinavian Count ('fru' and 'herr' are above), Swedish Miss. + 'greve', # borne + 'fröken', + # Czech: Mrs, Miss. Mr is 'pan', the Chinese surname Pan, so out; + # 'paní' keeps its accent, so the Indian name Pani is untouched. + 'paní', + 'slečna', + # Polish: Count. + 'hrabia', # borne + # Lithuanian: Mr, Mrs, Count. + 'ponas', # borne + 'ponia', # borne + 'grafas', + # Malay: Mr. Mrs is 'puan', a given name (Puan Maharani), so out. + 'encik', + # Filipino: Mr, Mrs, Miss. + 'ginoo', # borne + 'ginang', # borne + 'binibini', # borne + # Georgian, Latin script: Mr, Mrs, Count. + 'batoni', # borne + 'kalbatoni', + 'grafi', + # Danish/Norwegian/German "Familie Hansen", the family addressed. + 'familie', + # English Knight, the rank the post-nominal 'knt' names. + 'knight', # borne + # #269: Cyrillic (ru/uk) -- mr/mrs/dr/prof/academician/pan(i) # honorifics, same title-then-family convention as 'mr'/'dr'/'prof' # above (not GIVEN_NAME_TITLES: "г-н Петров" families the surname diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 4eec6a97..4ac2a3de 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -1544,6 +1544,55 @@ def _check_cjk_shape_purity(self) -> None: "1b folds the name into the family. 1.4.0 and 2.1 gave " "first 'de Mesnil' with no family, because the chain " "left 1b nothing standing alone to fire on"), + # #606: salutations in languages TITLES had none for, against a + # salutation left out because it leads a real name. + Case("salutation_in_a_new_language_is_a_title", + "Herra Väinö Johansson", + {"title": "Herra", "given": "Väinö", "family": "Johansson"}, + classification="fix(#606)", + notes="Finnish Mr; every release read given 'Herra', middle " + "'Väinö'. The contrast row is the one below"), + Case("salutation_that_leads_a_surname_stays_a_name", + "Ông Văn Tùng", + {"given": "Tùng", "middle": "Văn", "family": "Ông"}, + policy=Policy(name_order=FAMILY_FIRST_GIVEN_LAST), + notes="Vietnamese ông (Mr) is excluded from TITLES: Ông is a " + "surname and Vietnamese names lead with the surname, " + "the position the title claim acts on " + "(decisions.md#salutation-titles)"), + Case("office_completes_a_title_run", + "Police Commissioner James Gordon", + {"title": "Police Commissioner", "given": "James", + "family": "Gordon"}, + classification="fix(#606)", + notes="'police' was a title and 'commissioner' was not, so the " + "office read as the given name"), + Case("surname_borne_title_claims_the_front", "Greve Anna", + {"title": "Greve", "family": "Anna"}, + classification="fix(#606)", + notes="greve is a Scandinavian count AND a surname; shipped on " + "the graf precedent, so the uncommaed family-first form " + "loses its surname as 'Graf Steffi' does " + "(rules.md#H1's Accepted clause). The two rows below are " + "the readings it keeps"), + Case("surname_borne_title_trailing_is_the_surname", "Anna Greve", + {"given": "Anna", "family": "Greve"}, + notes="the title claim acts only in front (#606)"), + Case("surname_borne_title_before_a_comma_is_the_surname", + "Greve, Anna", + {"given": "Anna", "family": "Greve"}, + notes="the family comma fixes the surname (#606)"), + Case("knt_is_a_post_nominal", "Sir John Smith Knt", + {"title": "Sir", "given": "John", "family": "Smith", + "suffix": "Knt"}, + classification="fix(#606)", + notes="knt (Knight) joined the suffix acronyms beside kt; " + "every release read family 'Knt', middle 'Smith'"), + Case("knt_after_a_comma_is_a_suffix", "John Smith, Knt.", + {"given": "John", "family": "Smith", "suffix": "Knt."}, + classification="fix(#606)", + notes="every release read title 'Knt.', the period shape " + "claiming an unlisted word in front of the given part"), Case("title_plus_one_word_with_maiden", "Dr. Smith née Jones", {"title": "Dr.", "family": "Smith", "maiden": "Jones"}, classification="fix(#410)", diff --git a/tests/v2/test_facade_cases.py b/tests/v2/test_facade_cases.py index ce2a474f..e750017f 100644 --- a/tests/v2/test_facade_cases.py +++ b/tests/v2/test_facade_cases.py @@ -102,6 +102,9 @@ # v1 spelling, so the row is core-only. The default-order twin # ("Lord Chancellor") is an ordinary row and runs here. "all_titles_input_family_first", + # #606: the excluded salutation under the order Vietnamese names + # use, which has no v1 spelling. + "salutation_that_leads_a_surname_stays_a_name", # #518 review round: the `by_script` scope's positive controls. # An emptied script table has no v1 spelling -- v1 has no script # orders to empty -- so both rows are core-only, though the roles diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 1b47378f..ae5c0698 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -3880,7 +3880,9 @@ def _claim(rule: dict) -> _Claim: # comma name. # 2026-10-04, #604: 442 -> 443, 'Pérez, Juan Ó.', #604's # rules.md#P7 comma example. Reach. - _Claim(443, ('given', 'suffix', 'title'), "e00fc6c8266b", None), + # 2026-10-04, #606: 443 -> 444, 'Greve, Anna', #606's + # rules.md#H1 comma example. Reach. + _Claim(444, ('given', 'suffix', 'title'), "f2ad75955602", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3917,7 +3919,8 @@ def _claim(rule: dict) -> _Claim: # for both, so there is no diff here to explain. "fix(#296) a lone post-comma credential is a suffix": # 2026-10-01, #564: 23 -> 24, 'Smith, XYZ'. Reach. - _Claim(24, ('family', 'given', 'suffix', 'title'), "eae9a2bb02b0", None), + # 2026-10-04, #606: 24 -> 25, 'Greve, Anna'. Reach. + _Claim(25, ('family', 'given', 'suffix', 'title'), "b2bc2b830168", None), # 2026-09-27, #544: 6 -> 7; gains 'Smith, PhD MEng'. # 2026-09-28, #544: 7 -> 8; gains 'Smith, PhD Ma'. "fix(#325) a split credential followed by another suffix after a one-word family comma reads as suffixes": @@ -3994,7 +3997,8 @@ def _claim(rule: dict) -> _Claim: # 2026-10-03, #549: 439 -> 442, the same three #549 examples. # Reach. # 2026-10-04, #604: 442 -> 443, 'Pérez, Juan Ó.'. Reach. - _Claim(443, ('family', 'given'), "e00fc6c8266b", None), + # 2026-10-04, #606: 443 -> 444, 'Greve, Anna'. Reach. + _Claim(444, ('family', 'given'), "f2ad75955602", None), # 2026-10-01, #575: new, 4; 'De La Cruz, Ed', 'Freiherr von # Berg, Ed', 'Van Buren, Ed', 'de la Cruz, Ma'. "fix(#575) a particle surname before a comma is one name word": @@ -4663,6 +4667,10 @@ def _claim(rule: dict) -> _Claim: # two rules.md#R4 case-repair examples. "fix(#604) the Irish and Malay patronymic particles": _Claim(3, ('family', 'middle'), "cfcfd91f9d58", None), + # 2026-10-04, #606: new, 1; 'Greve Anna', rules.md#H1's + # Accepted example for a surname-borne title. + "fix(#606) salutations and titles in other languages": + _Claim(1, ('given', 'title'), "0c78dbac8a6b", None), }, "expected_since_2.0.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -4977,7 +4985,8 @@ def _claim(rule: dict) -> _Claim: # `given`, a field outside this rule's own ('suffix', 'title'). "fix(#296) a lone post-comma credential is a suffix": # 2026-10-01, #564: 23 -> 24, 'Smith, XYZ'. Reach. - _Claim(24, ('suffix', 'title'), "eae9a2bb02b0", None), + # 2026-10-04, #606: 24 -> 25, 'Greve, Anna'. Reach. + _Claim(25, ('suffix', 'title'), "b2bc2b830168", None), # 2026-09-27, #544: 6 -> 7; gains 'Smith, PhD MEng'. # 2026-09-28, #544: 7 -> 8; gains 'Smith, PhD Ma'. "fix(#325) a split credential followed by another suffix after a one-word family comma reads as suffixes": @@ -5450,6 +5459,10 @@ def _claim(rule: dict) -> _Claim: # two rules.md#R4 case-repair examples. "fix(#604) the Irish and Malay patronymic particles": _Claim(3, ('family', 'middle'), "cfcfd91f9d58", None), + # 2026-10-04, #606: new, 1; 'Greve Anna', rules.md#H1's + # Accepted example for a surname-borne title. + "fix(#606) salutations and titles in other languages": + _Claim(1, ('given', 'title'), "0c78dbac8a6b", None), }, # The 2.3 cycle's first rule, and a facade-only render fix: every # role is identical, so `_initials` alone. Reach and digest as in @@ -5940,6 +5953,10 @@ def _claim(rule: dict) -> _Claim: # two rules.md#R4 case-repair examples. "fix(#604) the Irish and Malay patronymic particles": _Claim(3, ('family', 'middle'), "cfcfd91f9d58", None), + # 2026-10-04, #606: new, 1; 'Greve Anna', rules.md#H1's + # Accepted example for a surname-borne title. + "fix(#606) salutations and titles in other languages": + _Claim(1, ('given', 'title'), "0c78dbac8a6b", None), }, "expected_since_2.1.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -6207,7 +6224,8 @@ def _claim(rule: dict) -> _Claim: # `given`, a field outside this rule's own ('suffix', 'title'). "fix(#296) a lone post-comma credential is a suffix": # 2026-10-01, #564: 23 -> 24, 'Smith, XYZ'. Reach. - _Claim(24, ('suffix', 'title'), "eae9a2bb02b0", None), + # 2026-10-04, #606: 24 -> 25, 'Greve, Anna'. Reach. + _Claim(25, ('suffix', 'title'), "b2bc2b830168", None), # 2026-09-27, #544: 6 -> 7; gains 'Smith, PhD MEng'. # 2026-09-28, #544: 7 -> 8; gains 'Smith, PhD Ma'. "fix(#325) a split credential followed by another suffix after a one-word family comma reads as suffixes": @@ -6675,6 +6693,10 @@ def _claim(rule: dict) -> _Claim: # two rules.md#R4 case-repair examples. "fix(#604) the Irish and Malay patronymic particles": _Claim(3, ('family', 'middle'), "cfcfd91f9d58", None), + # 2026-10-04, #606: new, 1; 'Greve Anna', rules.md#H1's + # Accepted example for a surname-borne title. + "fix(#606) salutations and titles in other languages": + _Claim(1, ('given', 'title'), "0c78dbac8a6b", None), }, "expected_since_2.3.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -7008,6 +7030,10 @@ def _claim(rule: dict) -> _Claim: # two rules.md#R4 case-repair examples. "fix(#604) the Irish and Malay patronymic particles": _Claim(3, ('family', 'middle'), "cfcfd91f9d58", None), + # 2026-10-04, #606: new, 1; 'Greve Anna', rules.md#H1's + # Accepted example for a surname-borne title. + "fix(#606) salutations and titles in other languages": + _Claim(1, ('given', 'title'), "0c78dbac8a6b", None), }, } diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index dc94125c..f9ba1ea3 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -19,6 +19,7 @@ "Anh Van Do" "Anh van Do" "Anna () Smith" +"Anna Greve" "Anna z (domu) Nowak" "Anna z Nowak" "Asst. Vice Chancellor John Smith" @@ -101,6 +102,8 @@ "García Márquez, MJ JK" "García Márquez, MJ PhD" "García Márquez, Ms G.J." +"Greve Anna" +"Greve, Anna" "Hans „Erster“ und “Zweiter” Müller" "Hassan Mohamad Ali" "Hassan, Mohamad Ahmad Ali" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 786c3020..b8dc85f9 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -5015,3 +5015,22 @@ issue = "fix(#604) the Irish and Malay patronymic particles" # 'Ó MURCHÚ', 'binti navalamar'. name_regex = "^(?:Juan Ó Pérez|SEÁN Ó MURCHÚ|ina binti navalamar)$" fields = ["middle", "family"] + +# --------------------------------------------------------------- +# #606: SALUTATIONS AND TITLES IN OTHER LANGUAGES. +# TITLES gains salutations and offices for languages it had none +# for, among them three that are also surnames written last +# wherever they are borne (graaf, greve, knight), shipped on the +# graf precedent (rules.md#H1). The one corpus mover is the +# H1 Accepted example for that cost: the surname-first form +# without a comma gives the surname-borne word to the title. +# +# Literal-anchored to the name: it was reported unexplained on the +# run before this rule existed (2026-10-04). +# --------------------------------------------------------------- + +[[change]] +issue = "fix(#606) salutations and titles in other languages" +# The given name becomes the title: 'Greve'. +name_regex = "^Greve Anna$" +fields = ["title", "given"] diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index ed828c56..9133273f 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -4082,3 +4082,22 @@ issue = "fix(#604) the Irish and Malay patronymic particles" # 'Ó MURCHÚ', 'binti navalamar'. name_regex = "^(?:Juan Ó Pérez|SEÁN Ó MURCHÚ|ina binti navalamar)$" fields = ["middle", "family"] + +# --------------------------------------------------------------- +# #606: SALUTATIONS AND TITLES IN OTHER LANGUAGES. +# TITLES gains salutations and offices for languages it had none +# for, among them three that are also surnames written last +# wherever they are borne (graaf, greve, knight), shipped on the +# graf precedent (rules.md#H1). The one corpus mover is the +# H1 Accepted example for that cost: the surname-first form +# without a comma gives the surname-borne word to the title. +# +# Literal-anchored to the name: it was reported unexplained on the +# run before this rule existed (2026-10-04). +# --------------------------------------------------------------- + +[[change]] +issue = "fix(#606) salutations and titles in other languages" +# The given name becomes the title: 'Greve'. +name_regex = "^Greve Anna$" +fields = ["title", "given"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 041388d7..9caff2e6 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -4027,3 +4027,22 @@ issue = "fix(#604) the Irish and Malay patronymic particles" # 'Ó MURCHÚ', 'binti navalamar'. name_regex = "^(?:Juan Ó Pérez|SEÁN Ó MURCHÚ|ina binti navalamar)$" fields = ["middle", "family"] + +# --------------------------------------------------------------- +# #606: SALUTATIONS AND TITLES IN OTHER LANGUAGES. +# TITLES gains salutations and offices for languages it had none +# for, among them three that are also surnames written last +# wherever they are borne (graaf, greve, knight), shipped on the +# graf precedent (rules.md#H1). The one corpus mover is the +# H1 Accepted example for that cost: the surname-first form +# without a comma gives the surname-borne word to the title. +# +# Literal-anchored to the name: it was reported unexplained on the +# run before this rule existed (2026-10-04). +# --------------------------------------------------------------- + +[[change]] +issue = "fix(#606) salutations and titles in other languages" +# The given name becomes the title: 'Greve'. +name_regex = "^Greve Anna$" +fields = ["title", "given"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index d28d0e0d..e9beaefc 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2423,3 +2423,22 @@ issue = "fix(#604) the Irish and Malay patronymic particles" # 'Ó MURCHÚ', 'binti navalamar'. name_regex = "^(?:Juan Ó Pérez|SEÁN Ó MURCHÚ|ina binti navalamar)$" fields = ["middle", "family"] + +# --------------------------------------------------------------- +# #606: SALUTATIONS AND TITLES IN OTHER LANGUAGES. +# TITLES gains salutations and offices for languages it had none +# for, among them three that are also surnames written last +# wherever they are borne (graaf, greve, knight), shipped on the +# graf precedent (rules.md#H1). The one corpus mover is the +# H1 Accepted example for that cost: the surname-first form +# without a comma gives the surname-borne word to the title. +# +# Literal-anchored to the name: it was reported unexplained on the +# run before this rule existed (2026-10-04). +# --------------------------------------------------------------- + +[[change]] +issue = "fix(#606) salutations and titles in other languages" +# The given name becomes the title: 'Greve'. +name_regex = "^Greve Anna$" +fields = ["title", "given"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 954dca0e..916c42a8 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1672,3 +1672,22 @@ issue = "fix(#604) the Irish and Malay patronymic particles" # 'Ó MURCHÚ', 'binti navalamar'. name_regex = "^(?:Juan Ó Pérez|SEÁN Ó MURCHÚ|ina binti navalamar)$" fields = ["middle", "family"] + +# --------------------------------------------------------------- +# #606: SALUTATIONS AND TITLES IN OTHER LANGUAGES. +# TITLES gains salutations and offices for languages it had none +# for, among them three that are also surnames written last +# wherever they are borne (graaf, greve, knight), shipped on the +# graf precedent (rules.md#H1). The one corpus mover is the +# H1 Accepted example for that cost: the surname-first form +# without a comma gives the surname-borne word to the title. +# +# Literal-anchored to the name: it was reported unexplained on the +# run before this rule existed (2026-10-04). +# --------------------------------------------------------------- + +[[change]] +issue = "fix(#606) salutations and titles in other languages" +# The given name becomes the title: 'Greve'. +name_regex = "^Greve Anna$" +fields = ["title", "given"] From c789c9e85ce594d6d74e7c6686925bea15101037 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sun, 4 Oct 2026 03:09:35 -0700 Subject: [PATCH 2/2] Name the #606 borne-word class as an exception to C-i Second review round: the class overrides C-i's position test for measured leading-position bearers, not its uncertainty fallback, so C-i's entry now carries a dated pointer to the one exception and its scope (TITLES, these words). The entry also says the excluded side of the line was judged without counts, so the line is not reproducible from the record. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 3 ++- nameparser/config/titles.py | 15 ++++++++------- 2 files changed, 10 insertions(+), 8 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 3d20c9b6..d4f51e86 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -278,6 +278,7 @@ Recomputed 2026-09-07 with that recipe, after this section's own #342 decision l - 2026-08-16 (collision keystone; #348, #360, #342, #385) — the 58%-vs-0.65% gap is BASE RATE, not disagreement. Both sets apply the same test; most particles are short words that double as names (van, bin, le, do, bar, mac) while most credential acronyms are not (abpp, acp). Recorded because the gap reads as an inconsistency and is not one — a reviewer who "harmonizes" the two shares will break one of them. - **C-i, vocabulary vs. name.** A word belongs in its set's ambiguous subset iff it is borne as an ordinary name IN THE POSITION THE VOCABULARY CLAIM ACTS ON. Existence of a bearer anywhere is not the test, and the first draft of this criterion (recorded 2026-08-16, corrected 2026-08-17) got that wrong. Under uncertainty, default to AMBIGUOUS. This generalizes the evidence standard already stated in NON_GIVEN_NAME_PARTICLES' docstring — a wrong unambiguous claim misparses a real person, a wrong ambiguous marking only adds a flag — from "which set" to "which subset". Applies uniformly to particles, suffix_acronyms and titles. +- 2026-10-04 #606 — C-i has ONE recorded exception, for TITLES only: #salutation-titles ships words measured as borne in the leading position at small incidence (herra, batoni, hrabia, kreivi), a weighing C-i itself does not make. It extends to no other set and no other word without its own counts. - **C-ii, vocabulary vs. vocabulary.** Where two sets claim a word and NEITHER reading is a name, precedence is a frequency judgment recorded per word. vd is the live case: never-given particle AND credential acronym (the British Volunteer Decoration), neither of them a name. Decision: the Dutch van der reading, as the more common. That is what unblocks #380, whose trailing-orphan half is a separate decision recorded under its own rule. - 2026-08-17 — C-i CORRECTED: the position qualifier. Writing the per-word records for the never-given particles falsified the first @@ -743,7 +744,7 @@ Closes #346, #344 and #343 as one bundle: each CLASS is decided once, and each s Closes #606. No parser code moves: 41 words join TITLES and one, knt (Knight, beside kt), joins SUFFIX_ACRONYMS with a `Knt` case mask. The class is decided once — a word that addresses or ranks a person and is written in FRONT of the name — and each word is then argued on #vocabulary-collisions C-i, the position test, because TITLES has no ambiguous subset and the only two answers available are ship and do not ship. The per-word lists, grouped by language, are the #606 block in nameparser/config/titles.py. -- **The graf class, decided (Derek, 2026-10-04).** A word borne as a name, but RARELY as the first word of one, ships anyway, on the precedent of graf, which shipped though Steffi Graf exists. This is a judgment on each word's incidence in the leading position — the position C-i asks about — and not a rule: it relaxes C-i's "under uncertainty, default to ambiguous", which for TITLES means do not ship, for words whose counts were measured and found small. The title claim acts only in front, so most bearers are untouched in the forms they write: `Gladys Knight` and `Knight, Gladys` keep family Knight. The cost falls on a name that BEGINS with the word, the cost `Graf Steffi` already paid: `Greve Anna` reads title Greve, family Anna, and a given name is lost the same way, `Herra Wijaya` reading title Herra, family Wijaya. Declaring `FAMILY_FIRST` does not help, since the title claim is read before the order (`Hrabia Anna` reads title Hrabia under it); the family comma does. Two further limits, both pre-existing: a trailing period makes the word the title by rules.md#H5 (`Gladys Knight.` reads title `Knight.`, as `Mary Jane King.` does), and every word this entry adds widens that rule's reach, H5 claiming any listed title behind a period. Accepted rather than prevented, and stated at rules.md#H1's Accepted clause. The issue proposed the class for three words — greve (Scandinavian Count), graaf (Dutch Count), knight. The spot-check the issue asked for (forebears.io, 2026-10-04, worldwide counts rounded) found ten more of its "no known collision" words borne, and Derek extended the decision to all of them rather than excluding them: herra (~3,000 surnames, Costa Rica and the Dominican Republic; ~670 given-name uses, Indonesia and the Philippines), batoni (~1,100, the painter Pompeo Batoni; ~300 given-name uses, Tanzania), ponas (~4,000, Iran), counsel (~730, Australia, the US, England), ponia (~560), the Filipino ginoo, ginang and binibini (~500 each, Philippine lists writing Surname, Given with a comma), and two whose registers often write the surname first, hrabia (~750, Poland) and kreivi (~270, Finland). Those last two, and the given-name uses of herra and batoni, are the leading-position bearers this class accepts, and were decided on with their counts in hand. Where the line falls is the counts and nothing finer: thiru and puan are out as given names COMMON in the leading position, rodina as a surname common in surname-first Russian lists, and gróf for want of any count at all. The comment `# borne` in titles.py marks each member, so a sweep meets the decision where it meets the word. +- **The graf class, decided (Derek, 2026-10-04).** A word borne as a name, but RARELY as the first word of one, ships anyway, on the precedent of graf, which shipped though Steffi Graf exists. This is a judgment on each word's incidence in the leading position — the position C-i asks about — and not a rule, and it is an EXCEPTION to C-i rather than an application of it: C-i is an iff with no threshold, and herra, batoni, hrabia and kreivi were measured as borne in that position, so C-i says do not ship them (TITLES having no ambiguous subset to mark them in). What the exception adds is a weighing C-i does not make, a small measured incidence against the benefit of the title reading, and it is scoped to TITLES and to the words this entry names. It licenses nothing in the particle or suffix sets, and a later TITLES word that wants it is argued on its own counts. The title claim acts only in front, so most bearers are untouched in the forms they write: `Gladys Knight` and `Knight, Gladys` keep family Knight. The cost falls on a name that BEGINS with the word, the cost `Graf Steffi` already paid: `Greve Anna` reads title Greve, family Anna, and a given name is lost the same way, `Herra Wijaya` reading title Herra, family Wijaya. Declaring `FAMILY_FIRST` does not help, since the title claim is read before the order (`Hrabia Anna` reads title Hrabia under it); the family comma does. Two further limits, both pre-existing: a trailing period makes the word the title by rules.md#H5 (`Gladys Knight.` reads title `Knight.`, as `Mary Jane King.` does), and every word this entry adds widens that rule's reach, H5 claiming any listed title behind a period. Accepted rather than prevented, and stated at rules.md#H1's Accepted clause. The issue proposed the class for three words — greve (Scandinavian Count), graaf (Dutch Count), knight. The spot-check the issue asked for (forebears.io, 2026-10-04, worldwide counts rounded) found ten more of its "no known collision" words borne, and Derek extended the decision to all of them rather than excluding them: herra (~3,000 surnames, Costa Rica and the Dominican Republic; ~670 given-name uses, Indonesia and the Philippines), batoni (~1,100, the painter Pompeo Batoni; ~300 given-name uses, Tanzania), ponas (~4,000, Iran), counsel (~730, Australia, the US, England), ponia (~560), the Filipino ginoo, ginang and binibini (~500 each, Philippine lists writing Surname, Given with a comma), and two whose registers often write the surname first, hrabia (~750, Poland) and kreivi (~270, Finland). Those last two, and the given-name uses of herra and batoni, are the leading-position bearers this class accepts, and were decided on with their counts in hand. The line is NOT reproducible from this record, and that is said rather than hidden: the accepted side was counted, while the excluded side — thiru, puan, pan, rodina, ông/bà — was judged from general knowledge of each word, as #606 drafted it, with no count taken (the count source refused further lookups the same day). Each was judged a COMMON leading-position name, against counts in the hundreds on the accepted side; gróf was excluded for want of any evidence at all. Whoever decides κόμης or the next salutation should count the excluded words first, and is free to find the line in a different place. The comment `# borne` in titles.py marks each member, so a sweep meets the decision where it meets the word. - **Accents are not folded, and two words depend on it.** `_lexicon._normalize` lowercases and composes NFC and does nothing else to letters, so paní and härra cannot match Pani (~76,000 surnames and ~65,000 given names, India and Indonesia) or Harra (~1,100). Accent-folding the lexicon would turn both into collisions; whoever proposes it re-argues this entry. frú depends on nothing: the unaccented fru (Swedish/Danish Mrs) already ships, so the Cameroonian surname Fru (~10,800, where the surname is often written first) already reads as a title in `Fru Anna` and did before this entry; that pre-existing collision is recorded here rather than argued, being no part of #606. - **gravos was dropped, not excluded.** The issue listed it as the Latin-script Greek for Count; the spot-check could not confirm it is the Greek word at all, which is κόμης (kómis). Nothing was shown, so there is nothing to keep out — it is recorded here so that a later sweep starts from that rather than from the issue's table. κόμης itself is NOT added and NOT excluded: it is a Greek surname (~420 bearers in Greece, official lists writing the surname first), which by its counts sits inside the class above beside hrabia, and it was left out of this change for having no use in current Greek address rather than for its collision. Adding it is a decision still to make. - **Measurement (2026-10-04).** One corpus name moves, `Greve Anna`, given Greve to title Greve: the rules.md#H1 example this entry added, and the only one of its three examples that moves. Recompute: take the 41 entries of the #606 block in titles.py — the lines that are a quoted string, not the words quoted inside its comments, which name existing entries and excluded ones — parse every distinct name in the tools/differential/corpus*.jsonl glob (1,518 on the day) with the shipped parser and with `Lexicon.default().remove(titles=, suffix_acronyms={"knt"})`, and diff `as_dict()`. Before the rules.md examples were added, the same diff moved nothing over the 1,809 corpus lines of the tree before #604 merged, gravos still in the word list. The old readings the release-log bullet quotes were checked against the 1.4.0 and 2.3.0 wheels. diff --git a/nameparser/config/titles.py b/nameparser/config/titles.py index 80b63d38..1430731a 100644 --- a/nameparser/config/titles.py +++ b/nameparser/config/titles.py @@ -769,13 +769,14 @@ # #606: salutations and offices for the languages the block above # had none for, so "Herra Väinö Johansson" stops reading given # 'Herra'. TITLES has no ambiguous subset, so a word ships only if - # the leading claim misreads nobody often enough to matter - # (decisions.md#vocabulary-collisions C-i). Those marked "borne" - # are names somewhere, but rarely the FIRST word of one -- mostly - # surnames written last. They ship on the 'graf' precedent above, - # paying the cost "Graf Steffi" already pays: a name that begins - # with one gives it to the title (rules.md#H1). That is a - # judgment on each word's counts, not a rule to sweep with. Counts + # it is borne as no name in the LEADING position + # (decisions.md#vocabulary-collisions C-i) -- with one recorded + # exception. Those marked "borne" are names somewhere, but rarely + # the FIRST word of one, mostly surnames written last, and ship + # anyway on the 'graf' precedent above, paying the cost "Graf + # Steffi" already pays: a name that begins with one gives it to + # the title (rules.md#H1). That is a judgment on each word's + # counts, scoped to these words, not a rule to sweep with. Counts # and sources are in # decisions.md#salutation-titles, and so are the words that failed: # among them Vietnamese ông/bà, Tamil thiru and pl/cs pan, all in