diff --git a/AGENTS.md b/AGENTS.md index 0a1b7d99..e15f80b4 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -269,7 +269,7 @@ The library has two layers: `nameparser/config/` (data) and `nameparser/parser.p **A constant's membership is a question you may reopen.** Proposing that a word be ADDED, REMOVED or MOVED between vocabulary sets is ordinary design work — a shipped entry is not evidence that anyone judged it. `SUFFIX_ACRONYMS` arrived in `af5bdab` as a bulk Wikipedia import never reviewed against surname collisions: 572 of its 577 alphabetic entries leave `family` empty in `"John "` against five ambiguous-gated exceptions (recomputed 2026-09-07; the fifth is `ba`; 568 of 575 against seven, measured 2026-09-25 after #540), and `sa`, `se` and `om` are borne as surnames (measured 2026-08-23), as were `rai`, `cha`, `ba` and `mc` before `decisions.md#suffix-acronym-collisions` decided all four — `rai` and `cha` removed, `ba` marked ambiguous, `mc` left alone as no borne name at all. The same entry marked `meng` and `lac` ambiguous on 2026-09-25 (#540). When a fix starts to look like new machinery, check the vocabulary first. Criterion: `decisions.md#vocabulary-collisions`, with #360's positional qualifier. -**Sweep the forms a change can reach: comma shapes, then `name_order`.** No comma, a FULL name before the comma, a ONE-WORD name before the comma and NOTHING before the comma are four paths, and the last two are the misses — #429, #430 and #432 are all the one-word path disagreeing with the full-name path on inputs the full-name path reads correctly, and #613's PR review found `, PhD` and `, Jr.` (an empty surname field, as CSV data writes it) read as given and title by a decision that returned early on an empty head, after two review rounds whose every grid had a name before the comma. Put an empty head and a nickname-only head (`(Bob), PhD`) in every comma grid. Count the one-word path the way C1 counts, which is neither tokens nor pieces: a particle run and the one word it attaches to are one word (#575), a connective join and a bound given-name pair count as their words (#613, so `Ortega y Gasset, Dr.` is a whole name), and a suffix beside the surname counts as no name word at all. #603's review found the suffix half (`Vega y Lopez Jr., PhD Jones` lost its family where `Vega y Lopez, PhD Jones` kept it); #613 then reversed the connective half, and both now read given 'Vega', middle 'y', family 'Lopez' — so test a particle surname, a connective surname and a surname with a suffix, each before a comma. For orders the sweep already exists (`tests/v2/test_cases.py` runs every row under all three) but carries one assertion, R2's family partition, so it checks nothing a new change moves; coverage there has been vacuous before (PR #394's review found the suite passed with `name_order` discarded from grouping). Parse your change's names in each comma shape and each order, and read the ones you did not predict. And read a generated grid's moves by SHAPE, not by corpus membership: #614's first design passed the suite, the five gates and a fingerprint of every corpus name, because its grid moves were sorted into 'corpus name moved' and 'garbage', and the review found the moves that mattered among the 'garbage' -- a bare salutation (`Mr. and Mrs.` lost its connective to the family name) and a name with one trailing word (`Freiherr von Berg X.Y.Z.`). Pull out the head-plus-one-word slice of the grid and a salutation with nothing after it, and read every move there. Two more were the second review's (#614): no grid put TWO words that are both a title and a particle at the head (`Freiherr St van Berg MA`, where an index into the pieces was applied to a merged view of them -- `TITLES ∩ particles_ambiguous` is `{freiherr, st}`, so put both in every head list), and the grids compared ambiguity KINDS, so a report telling the caller the wrong field (`Freiherr von Berg Ma`, 'read as a given name' of a word in the family) passed every pin. Compare each report's `detail` against the field its tokens land in. And #620's: an exclusion list in code being replaced is a set of rules, not noise -- #614's merged copy excluded every word the trailing read weighs from its units, the first #620 prototype dropped the list as a symptom of the copy, and its review found the list had been keeping positional counting's rule that a member in front counts as a word (`Freiherr von Berg MA X.Y.Z.` lost its suffix). Before deleting one, mutate each entry out and read what moves. +**Sweep the forms a change can reach: comma shapes, then `name_order`.** No comma, a FULL name before the comma, a ONE-WORD name before the comma and NOTHING before the comma are four paths, and the last two are the misses — #429, #430 and #432 are all the one-word path disagreeing with the full-name path on inputs the full-name path reads correctly, and #613's PR review found `, PhD` and `, Jr.` (an empty surname field, as CSV data writes it) read as given and title by a decision that returned early on an empty head, after two review rounds whose every grid had a name before the comma. Put an empty head and a nickname-only head (`(Bob), PhD`) in every comma grid. Count the one-word path the way C1 counts, which is neither tokens nor pieces: a particle run and the one word it attaches to are one word (#575), a connective join and a bound given-name pair count as their words (#613, so `Ortega y Gasset, Dr.` is a whole name), and a suffix beside the surname counts as no name word at all. #603's review found the suffix half (`Vega y Lopez Jr., PhD Jones` lost its family where `Vega y Lopez, PhD Jones` kept it); #613 then reversed the connective half, and both now read given 'Vega', middle 'y', family 'Lopez' — so test a particle surname, a connective surname and a surname with a suffix, each before a comma. For orders the sweep already exists (`tests/v2/test_cases.py` runs every row under all three) but carries one assertion, R2's family partition, so it checks nothing a new change moves; coverage there has been vacuous before (PR #394's review found the suite passed with `name_order` discarded from grouping). Parse your change's names in each comma shape and each order, and read the ones you did not predict. And read a generated grid's moves by SHAPE, not by corpus membership: #614's first design passed the suite, the five gates and a fingerprint of every corpus name, because its grid moves were sorted into 'corpus name moved' and 'garbage', and the review found the moves that mattered among the 'garbage' -- a bare salutation (`Mr. and Mrs.` lost its connective to the family name) and a name with one trailing word (`Freiherr von Berg X.Y.Z.`). Pull out the head-plus-one-word slice of the grid and a salutation with nothing after it, and read every move there. Two more were the second review's (#614): no grid put TWO words that are both a title and a particle at the head (`Freiherr St van Berg MA`, where an index into the pieces was applied to a merged view of them -- `TITLES ∩ particles_ambiguous` is `{freiherr, st}`, so put both in every head list), and the grids compared ambiguity KINDS, so a report telling the caller the wrong field (`Freiherr von Berg Ma`, 'read as a given name' of a word in the family) passed every pin. Compare each report's `detail` against the field its tokens land in -- since #626 `test_a_report_names_the_field_its_word_lands_in` (tests/v2/test_cases.py) does it over every non-locale case row under all three orders (a row carrying its own policy only as declared), for the phrasings its `_FIELD_CLAIMS` lists -- so a row with the shape is the test only where the report's wording is listed: a new emitter's wording goes in that list, and segment's "consumed as suffix" is not in it (#629). #626's own first grid then missed an emitter for want of one comma shape (`, Ma Dr.`: an empty head, an ambiguous word, a trailing title, which H1 moves to the family): a trailing title belongs in every empty-head comma grid too. And #620's: an exclusion list in code being replaced is a set of rules, not noise -- #614's merged copy excluded every word the trailing read weighs from its units, the first #620 prototype dropped the list as a symptom of the copy, and its review found the list had been keeping positional counting's rule that a member in front counts as a word (`Freiherr von Berg MA X.Y.Z.` lost its suffix). Before deleting one, mutate each entry out and read what moves. **Then sweep the spellings: a class reached by vocabulary is also reached by SHAPE.** A test that uses only the listed spelling of a word class walks only the vocabulary path, and the shape path is a separate branch that can be missing while every test passes. Roman numerals are the sharpest case: `ii`, `iii` and `iv` are suffix vocabulary; `i` and `v` are vocabulary too but initial-shaped, so `is_suffix_piece` vetoes them and only the numeral fork reads them; and `vi` through `x` are in no list at all, read by the fork's shape test alone (checked 2026-10-04 against `Lexicon.default()`) -- which stops at `x`: `_vocab._ROMAN` matches I to X, so `John Smith XI` reads `XI` as a name word (re-checked 2026-10-07, #614's /simplify, which found a docstring claiming "any length"). So a change that handles `V` and `III` can still miss `VI` — #610 found exactly that, the maiden take's candidate loop admitting a shape-only numeral in one position where the peel reads it in two (`Jane Doe nee Smith VI Prof.` kept maiden `Smith VI` while `V` and `III` worked). The same split runs through the credentials (`MA` listed, `X.Y.Z.` by the dotted shape, `XYZ` by the caps shape, and `Ph. D.` merged by group into one flagged piece the peel's walk never holds -- #603 found #602's run-start predicate refusing it too, so `Smith, John Ph. D. Jones` kept middle 'Jones' where `Smith, John PhD Jones` read suffix 'PhD Jones') and the titles (`Prof.` period-marked and read at the end of a name, `Prof` bare and a name word there). When a change touches a word class, test one spelling from each path. @@ -320,7 +320,7 @@ The 2.0 rewrite lands as underscore-private modules alongside the v1 code. These - **Method organization**, fixed section order in every class: fields + `__post_init__` validation → alternative constructors → dunders (construction/equality → protocol → operators) → properties → public methods by concern (access → editing → comparison → rendering delegates) → private helpers last, except a helper serving exactly one section may sit at that section's head. Sanctioned deviation, facade layer only: `HumanName` and the shim `Constants` organize by v1 concern groups (`# -- render defaults --`, `# -- config / parsing --`, `# -- fields --`, ..., dunders and pickle last) — the classes mirror v1's own surface and die in 3.0; the canonical order still binds every core type. - **Validation is eager and fail-loud**: every `raise` states the offending value, the expected form, and the fix. Exception taxonomy: wrong type — including wrong element type inside a collection, bare `str` where an iterable of strings is expected, or a `Mapping` where a plain iterable is expected — raises `TypeError`; well-typed but unacceptable values raise `ValueError`; failed enum lookups stay `ValueError` for any input (stdlib `EnumType` precedent). **When the message hands the reader code to paste, that code has to survive a type checker** — nameparser ships `py.typed`. #337's segmenterless warning offered `Policy(segment_scripts=())`, an `arg-type` error, because these fields are annotated with what they STORE rather than everything the constructor accepts. Prefer the `frozenset()` / `()` spellings in messages and docstrings, and pin the offered spelling in a test — the warning tests matched on `ja_segmenter` and never checked the actionable half of the message. **A warning emitted in `Parser.__post_init__` needs `parser_for` to re-emit it from its own frame** (the `catch_warnings(record=True)` block at its return): `__post_init__`'s `stacklevel` is sized for direct `Parser(...)` construction, and through `parser_for`'s extra frame the default one-line rendering attributes the warning to the library's own `return Parser(...)` — the exact call the message tells the user to change becomes invisible. No single stacklevel serves both entry points; a new construction warning gets the re-emission for free, but a new CONSTRUCTION SITE for `Parser` inside this package needs its own re-emission or its callers get library-attributed warnings (#337 review). - **Guard, hint, and emit for the WHOLE family, and parametrize the test over it**: a check added to one member of a set belongs on all of it, and the test must sweep the family, not one example. This session shipped `_reject_str_and_mapping` on `Policy` but not `PolicyPatch`, the bytes decode hint on three of five config entry points, and a regex-sync roster missing four of its copies — each a separate follow-up bug that a `{class} × {field} × {bad-value}` parametrization would have caught and a per-example test hid. When you find you're guarding member N, grep for the other members first. -- **Ambiguities are emitted at the DECISION site**: an `Ambiguity` records a fork the parse had to call, not a token that sits in an ambiguous vocabulary. Emit where the branch is taken — the trailing-suffix peel (read once in `_group` for a main segment since #614, its picks reported by `_assign` once the roles they took are known, except a pick the particle chain takes into the name, which the chain reports and the read then gives up), the delimiter escape's follow-up in `classify` — never by scanning for a `vocab:*-ambiguous` tag. The same tagged token is a genuine fork in one position and unremarkable in another (`do` mid-name in "Joao da Silva do Amaral de Souza" chooses nothing). **A branch that runs but changes nothing is not a decision either** -- the prefix chain's `merge_pieces(pieces, ptags, k, j)` executes even when `j == k + 1`, folding a piece into itself, and keying on "the code got here" reported a fork for all ambiguous particles on "Do Van Jr." (`Dr.` when that was written, before #367 made a plain title transparent and put the shape out of the loop's reach entirely), where the particle stayed a lone leading name piece — the GIVEN name under the default order, the family name under `FAMILY_FIRST` — and `_assign` reported the same token again. Check that the branch actually claimed something (`j > k + 1`) before recording. Structure often settles the question before it arises, which is why `PARTICLE_OR_GIVEN` is not emitted on the `FAMILY_COMMA` path's WHOLLY-FAMILY read -- the comma fixed which piece is the family -- and `SUFFIX_OR_NAME` is not emitted for "Ma, Jack". A tail segment is the same case from the other side: assign reads it as suffixes, a maiden marker there included since #601, and since #603 a title word past the second comma as a title (rules.md#C2, so `Freiherr` below reads title) (`Jane Doe, PhD, Jr nee van Ma` reads suffix `PhD, Jr nee van Ma`, the marker an ordinary word, rules.md#M2), so group's two chain emitters are handed no report list there, after either comma — `John Smith, Jr., Freiherr von Richthofen` still chains `von` and reports only `comma-structure` (rules.md#C2, 2026-09-28). Read that scope narrowly: the comma settles nothing about a particle trailing the given name, so P6's attachment decides that fork on the same path and reports it (#405; in `post_rules` until #613, in `_assign` since, right after the walk that reads the given part), in the kind naming the reading it OVERRODE, which is the reading assign made and not the word's vocabulary: `SUFFIX_OR_NAME` where assign had read the run as a post-nominal (`vd`, `mc`), else `PARTICLE_OR_GIVEN` where the run holds an ambiguous particle (`van`, and `do`, which is in the suffix vocabulary too but in its AMBIGUOUS half, so no credential reading was overridden), else silence. The decision site also has the token index and the detail text in hand, which the tag scan would have to reconstruct. **If a fork's two branches are taken in DIFFERENT stages, every one of them needs the emitter** -- `PARTICLE_OR_GIVEN` is decided in `_assign` when the ambiguous particle stays a lone leading piece, in `_group` when something shifts it off the name's leading piece and the prefix chain claims it, and wherever P6's attachment takes a trailing particle into the family -- in `_assign` after a comma that names a family (since #613), in `post_rules` at the end of a name read family-first (#467) and after a comma with nothing before it, where H1 and M4 must read the name first -- so every one of them reports; for two years only the first did. What can still do the shifting is narrow, and #367 is why: a plain title no longer can (`Dr. Van Johnson` reads as `Van Johnson` does and reports from `_assign`), so the `_group` emitter needs a word that is BOTH a title and a particle — measured, `TITLES ∩ particles_ambiguous` is `{freiherr, st}` in the default vocabulary (`do` left TITLES in #296's audit; decisions.md's Excluded block records the before and after), plus any overlap a caller's config creates — standing ahead of the chained particle as the LEADING NAME word. Titles may precede it, so `Dr. St van Johnson` reaches the emitter and `St van Johnson` does too; a given name may not, so `Jan Freiherr von Richthofen` does not reach it while `Freiherr von Richthofen` and `Dr. Freiherr von Richthofen` do. Two shapes that look like they should reach it and do NOT, both measured by stepping `STAGES` and watching where `ambiguities` grows: `Dr. Do van Johnson` and `Do St Johnson` report from `assign`, not `group`, because `do` is no longer a title and so stays the leading name piece assign reports on — a both-vocabulary word CHAINED (`Jan St Johnson`) reports nothing at all. When checking whether that emitter is dead, a both-vocabulary word in the leading name position is the thing to look for, and the answer is that it is not dead. The stage-ownership map in `tests/v2/pipeline/test_state.py` must list `ambiguities` for each such stage, and it passes vacuously until a case row exercises the path, so add the row too. Report BOTH directions of a two-way fork — "John Smith MA" (read as a suffix) and "Jack MA" (read as the family name) are equally guesses. Every kind needs a trigger in `tests/v2/test_contracts.py::_AMBIGUITY_TRIGGERS` (an explicit `None`, strict-xfail, while reserved), and case-table rows pin expected kinds exactly, so a new emitter shows up in both immediately. **Pin the decision, not the vocabulary**: the only titled-particle test used an UNAMBIGUOUS particle, so it walked the right code path and proved nothing about the branch under test -- two criticals passed 1539 tests. A row contrasting the two readings ("John Smith V" against "John Smith B") is what makes an emitter's absence meaningful. +- **Ambiguities are emitted at the DECISION site**: an `Ambiguity` records a fork the parse had to call, not a token that sits in an ambiguous vocabulary. Emit where the branch is taken — the trailing-suffix peel (read once in `_group` for a main segment since #614, its picks reported by `_assign` once the roles they took are known, except a pick the particle chain takes into the name, which the chain reports and the read then gives up), the delimiter escape's follow-up in `classify` — never by scanning for a `vocab:*-ambiguous` tag. The same tagged token is a genuine fork in one position and unremarkable in another (`do` mid-name in "Joao da Silva do Amaral de Souza" chooses nothing). **A branch that runs but changes nothing is not a decision either** -- the prefix chain's `merge_pieces(pieces, ptags, k, j)` executes even when `j == k + 1`, folding a piece into itself, and keying on "the code got here" reported a fork for all ambiguous particles on "Do Van Jr." (`Dr.` when that was written, before #367 made a plain title transparent and put the shape out of the loop's reach entirely), where the particle stayed a lone leading name piece — the GIVEN name under the default order, the family name under `FAMILY_FIRST` — and `_assign` reported the same token again. Check that the branch actually claimed something (`j > k + 1`) before recording. Structure often settles the question before it arises, which is why `PARTICLE_OR_GIVEN` is not emitted on the `FAMILY_COMMA` path's WHOLLY-FAMILY read -- the comma fixed which piece is the family -- and `SUFFIX_OR_NAME` is not emitted for "Ma, Jack". A tail segment is the same case from the other side: assign reads it as suffixes, a maiden marker there included since #601, and since #603 a title word past the second comma as a title (rules.md#C2, so `Freiherr` below reads title) (`Jane Doe, PhD, Jr nee van Ma` reads suffix `PhD, Jr nee van Ma`, the marker an ordinary word, rules.md#M2), so group's two chain emitters are handed no report list there, after either comma — `John Smith, Jr., Freiherr von Richthofen` still chains `von` and reports only `comma-structure` (rules.md#C2, 2026-09-28). Read that scope narrowly: the comma settles nothing about a particle trailing the given name, so P6's attachment decides that fork on the same path and reports it (#405; in `post_rules` until #613, in `_assign` since, right after the walk that reads the given part), in the kind naming the reading it OVERRODE, which is the reading assign made and not the word's vocabulary: `SUFFIX_OR_NAME` where assign had read the run as a post-nominal (`vd`, `mc`), else `PARTICLE_OR_GIVEN` where the run holds an ambiguous particle (`van`, and `do`, which is in the suffix vocabulary too but in its AMBIGUOUS half, so no credential reading was overridden), else silence. The decision site also has the token index and the detail text in hand, which the tag scan would have to reconstruct -- all but the FIELD a report names, which a rule in post_rules can still change (H1's move behind a title, P1's family-first fold, P6's attachment, and under opt-in policies the patronymic rotations and `middle_as_family`'s fold): an emitter naming one passes the rest of its sentence as `field_tail` and assemble words the field from the final role (#626, where `Kim Min Do` under `FAMILY_FIRST` was told 'Do' was a middle name with 'Do' in the family). **If a fork's two branches are taken in DIFFERENT stages, every one of them needs the emitter** -- `PARTICLE_OR_GIVEN` is decided in `_assign` when the ambiguous particle stays a lone leading piece, in `_group` when something shifts it off the name's leading piece and the prefix chain claims it, and wherever P6's attachment takes a trailing particle into the family -- in `_assign` after a comma that names a family (since #613), in `post_rules` at the end of a name read family-first (#467) and after a comma with nothing before it, where H1 and M4 must read the name first -- so every one of them reports; for two years only the first did. What can still do the shifting is narrow, and #367 is why: a plain title no longer can (`Dr. Van Johnson` reads as `Van Johnson` does and reports from `_assign`), so the `_group` emitter needs a word that is BOTH a title and a particle — measured, `TITLES ∩ particles_ambiguous` is `{freiherr, st}` in the default vocabulary (`do` left TITLES in #296's audit; decisions.md's Excluded block records the before and after), plus any overlap a caller's config creates — standing ahead of the chained particle as the LEADING NAME word. Titles may precede it, so `Dr. St van Johnson` reaches the emitter and `St van Johnson` does too; a given name may not, so `Jan Freiherr von Richthofen` does not reach it while `Freiherr von Richthofen` and `Dr. Freiherr von Richthofen` do. Two shapes that look like they should reach it and do NOT, both measured by stepping `STAGES` and watching where `ambiguities` grows: `Dr. Do van Johnson` and `Do St Johnson` report from `assign`, not `group`, because `do` is no longer a title and so stays the leading name piece assign reports on — a both-vocabulary word CHAINED (`Jan St Johnson`) reports nothing at all. When checking whether that emitter is dead, a both-vocabulary word in the leading name position is the thing to look for, and the answer is that it is not dead. The stage-ownership map in `tests/v2/pipeline/test_state.py` must list `ambiguities` for each such stage, and it passes vacuously until a case row exercises the path, so add the row too. Report BOTH directions of a two-way fork — "John Smith MA" (read as a suffix) and "Jack MA" (read as the family name) are equally guesses. Every kind needs a trigger in `tests/v2/test_contracts.py::_AMBIGUITY_TRIGGERS` (an explicit `None`, strict-xfail, while reserved), and case-table rows pin expected kinds exactly, so a new emitter shows up in both immediately. **Pin the decision, not the vocabulary**: the only titled-particle test used an UNAMBIGUOUS particle, so it walked the right code path and proved nothing about the branch under test -- two criticals passed 1539 tests. A row contrasting the two readings ("John Smith V" against "John Smith B") is what makes an emitter's absence meaningful. - **A kind is worth adding only if a reader would hesitate too**: the test is not "does the code take a branch" but whether a person reading that input would genuinely be unsure. "Smith, John V" reads as a middle initial to anyone -- the comma settles it -- so reporting it would be noise that teaches callers to ignore the field, which costs more than the missing report. Reachability of the second branch is necessary, not sufficient. Prefer leaving a fork silent and documenting the omission over emitting on input nobody finds ambiguous. - **Parser owns config-dependent conveniences**: `Parser.matches`/`Parser.capitalized`/`Parser.revise` exist because the `ParsedName` equivalents fall back to DEFAULT config for str/omitted arguments (documented loudly in both docstrings). `revise` harvests tokens from a full sub-parse of each replacement value (tags kept minus `FOLDED_TAG`, roles forced, the R1 entry pass `suffix_entries` re-run over the forced state so a suffix value's entries follow its own commas, ambiguities discarded); the merge tail is shared with `replace()` via `ParsedName._with_field_tokens`. `Parser.capitalized` delegates through `name.capitalized(self.lexicon)` specifically so `_parser` never imports `_render` — keep it that way. - **Per-word vocabulary fields warn on multi-word entries** (`_normset`/`_normpairs` via `_warn_dead_entry`, UserWarning, never a raise — see the given_name_titles Gotcha for why raising is wrong). `given_name_titles` is the one multi-word-matched field and is exempt; `_edit` passes `warn=False` (add() warns once via the new instance's `__post_init__`; remove() stores nothing). The default vocabulary and every locale pack must stay warning-free (`test_default_lexicon_builds_warning_free`, `test_pack_vocabulary_entries_are_single_words`). diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 333799a6..25ba09eb 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -18,6 +18,12 @@ Entry conventions: - 2026-09-20 #397 (second review) — A REPORT MUST NOT CONTRADICT THE READING BESIDE IT, and the shape could not arise before this cycle. `classify`'s `conjunction-or-initial` fork offers exactly two readings and its detail says which it took ("it is read as an initial"). Where the rest of the parse then roles that letter SUFFIX, the fork resolved to NEITHER branch: `JOHN QUINCY SMITH I` carried the connective-or-initial beside a `suffix-or-name` saying the same token reads as a generational suffix, and `john smith i`, `HENRY I`, `SMITH, JOHN I`, `JOSEP CAROD I ROVIRA III` and `abdul e i` were the same. `i` is the first word that is both a MARKED connective and suffix vocabulary, so no parse before #397 could reach it — which is why the rule is written down now rather than having been. DECIDED: the report is WITHDRAWN where the role is SUFFIX or TITLE, in `_pipeline/_assemble.py`, the one place that sees final roles and the ambiguity list together — and beside a withdrawal that was already there, for a report whose referents did not survive assembly. That is the narrow reading of mechanisms.md#AMBIGUITY-AT-THE-DECISION-SITE rather than an exception to it: classify still owns the fork and still takes it where it is taken; what goes is a report whose subject the parse went on to read as something else. Emitting later was the alternative and was rejected as more invasive for the same answer — the fork's inputs (the case class, the own-words span) live in classify and moving the emitter moves them. MEASURED: over the review's 43,925-name grid under eight configurations, 9,806 rows on 1,637 distinct names lose exactly one `conjunction-or-initial` and nothing else — no report of any other kind moves and none is gained. On the deduped corpus-union-cases population the only name affected is `JOHN QUINCY SMITH I`, a row this round added, so the ledger's three-name `fix(#397) the Catalan link reports a connective-or-initial in a one-case name` rule is untouched: all three of its names role the letter MIDDLE and all three still report. The invariant is `tests/v2/test_properties.py::test_no_report_contradicts_the_reading_beside_it`, which states it twice — no connective-or-initial on a token roled SUFFIX or TITLE, and no token named by two reports whose details assert different readings — and fails on 8,204 rows at `dc3bdf9c`, 0 here. +- 2026-10-08 #626 — A REPORT NAMES THE FIELD ITS WORD LANDS IN. Six of assign's reports worded the field from the role the word held at assign (`read as a given name`), and rules in post_rules still move a word after that: H1 moves the name behind a title to the family, P1's family-first fold takes the particle run and one unit and lays the rest out by the order, P6's attachment takes a trailing name word into the family at the end of a name read family-first, and under opt-in policies the patronymic rotations (O1, O2) and `middle_as_family`'s fold (O3) move one too. So `Attorney General of Minnesota` reported 'General of Minnesota' "read as a given name by convention" with the unit in the family (since 2.3.0, #491's join clause); `Kim Min Do` under FAMILY_FIRST told the caller 'Do' was "read as a middle name" with 'Do' in the family, and `de Kim Ma` under FAMILY_FIRST said 'Ma' was a middle name with 'Ma' the given name (both this cycle, since #289 made the Title-case word a name word); `, Ma Dr.` said 'Ma' after the comma was "read as the given name" with 'Ma' in the family (this cycle, since #289 added that comma report; 2.3.0 reads the same fields and reports nothing); and `Van Ivan Petrovich` under the East Slavic patronymic rule said the leading 'Van' was "read as a given name" after O1 rotated it into the family (at 2.3.0 too). #614's second review had fixed one instance of the class (`Freiherr von Berg Ma`) by moving its report to the stage that knows the outcome. + DECIDED: the DECISION stays at its site and the WORDING of the field moves to assemble, the one place that knows final roles (the same reasoning as the 2026-09-20 withdrawal above). An emitter whose detail names a field a later rule could change writes the sentence up to the field and hands the rest as `PendingAmbiguity.field_tail` (`""` where the field ends the sentence); assemble joins them around `FIELD_READING` of the first referent's final role. All six such emitters in assign use it -- H4's all-suffix name, title-or-name's join clause, O5's lone word, the trailing peel's picks, the leading ambiguous particle and the comma report on the word after a family comma -- including those right on every measured input, so no later rule can make them wrong. A report in a name emptied for want of content is withdrawn rather than shipped half-worded (no emitter reaches that today). A placeholder in the detail was declined: token text can carry any characters, and a name containing the placeholder would corrupt its own report. Reports worded where the field is settled keep their literal wording: a comma's credential reading, the suffix run's and the numeral fork's, H4's "the name" and P6's own "joins the family". The comma report's given-side wording moves from "the given name" to "a given name" with the rest. + MISSED, then caught in review: the first draft converted five emitters and its grid had no empty-head comma with a trailing title, the shape AGENTS.md's comma sweep already asks for; four of the review's five readers found the sixth emitter on `, Ma Dr.`. The draft also named only H1 and P6 as movers; P1's fold was found by stepping the stages, and the review of the fix found the opt-in movers O1 to O3, which no grid over the three orders reaches -- reverting the leading-particle emitter passed the whole suite until rows under those policies pinned it. + MEASURED 2026-10-08, each report's field phrase compared against its tokens' final roles. RECIPE: texts are the 1,536 names of `tools/differential/corpus*.jsonl` and 347 `tests/v2/cases.py` texts (the table at 0c54662b, the parent of #614's first commit, less the corpus names; today's table adds rows this entry's review wrote for the shapes the grid missed, so it reads more wrong claims on master, 6 more as measured on the review), plus every head in {'John', 'John Smith', 'John Q. Smith', 'Mary Ann Smith', 'John van Smith', 'Juan de la Vega', 'John van', 'anh van', 'Jan Freiherr von Berg', 'Freiherr von Berg', 'Freiherr St', 'St', 'Dr. St', 'Freiherr', 'Dr.', 'Mr. and Mrs.', 'Attorney General of', 'abdul salam', 'Juan y Garcia', 'Jane Doe nee van der Berg', 'Dr. John Smith', 'JOHN SMITH', 'john smith', 'Smith', 'de Mesnil', 'Dr. de Mesnil', 'Kim Min', 'Kim Min Jun', 'Ortega y Gasset', 'van', 'Do', ''} followed by one or two words (repetition allowed) from the list [PhD, MA, Ma, DO, Do, do, vd, mc, Jr, III, V, V., VI, X.Y.Z., Ph. D., Esq., Prof., Sir, Ed, van, and, Jones, ba, John, Smith, Minnesota, de, St, Freiherr], space-joined and stripped; 20,000 draws after Python's `random.seed(626)`, each `random.choice(heads)`, then `random.randint(3, 4)` words each `random.choice(words)`, space-joined and stripped, both lists in the order written here; then, unstripped, `{head}, {word}`, `Smith, {head} {word}` with any double space replaced by one, and `{word}, {head}` for every head and word, and `, {w1} {w2}` and `(Bob), {w1} {w2}` for every ordered pair of words; the empty text dropped and duplicates removed -- 53,517 texts; each parsed under the three orders, so the opt-in movers are outside these figures; a claim is a match of `(reads? as|joins) ()` for the phrases `_FIELD_CLAIMS` in tests/v2/test_cases.py lists plus "the given name", and it is wrong when its tokens' roles are not among the fields the phrase names. On master (26cdb891) 2,725 of 89,092 claims were wrong, in three emitters: title-or-name's join clause (1,575), the trailing peel's picks (1,114) and the comma report (36); on the tree, 0. The comparator is master, which words the field at assign; the tree words it at assemble. Fingerprint against master over the differential corpora, master's case-table names and a seeded comma grid (13,090 names) under six policies (the three orders, `lenient_comma_suffixes=False`, a `/` suffix delimiter and `CapsSuffixes.OFF`): no role, kind, initials or rendering moves; the details that change are 6,395 comma reports on 1,086 names reworded from "the given name" to "a given name", 85 on 15 names now naming the family they land in, and 40 title-or-name join reports on 10 names now naming the family. Pinned by `test_a_report_names_the_field_its_word_lands_in` in tests/v2/test_cases.py, over every case row under all three orders except the locale rows (a row with its own policy only as declared), whose recorded control is 14 failures on master, and by `test_the_field_sweep_sees_the_claims_it_checks`, which records how many claims that sweep checks; the opt-in movers by the rows `Van Ivan Petrovich` (East Slavic), `Van Ali Veli oglu` (Turkic) and `Do Bishop Do` (FAMILY_FIRST with `middle_as_family`), which the sweep reads as declared. + +Open: [#629](https://github.com/derek73/python-nameparser/issues/629) a comma-structure report on a part past the second comma says the part was "consumed as suffix" where a title in it now reads as a title (`John Smith, Jr., Freiherr von Richthofen`): the one report naming a field that the #626 entry's mechanism does not reach. ### A2 — the empty name @@ -667,7 +673,7 @@ Closes #342 (a wordlist question) and #454 (a rules.md question) together, becau AND THE COUNT AT A COMMA IS OF NAME WORDS. For this class, "words to spare" in `X, Y` form means the part before the comma holds two or more NAME words: `Smith Jr., MA` has two tokens and one name word, and a token count hands its family to `given` (measured on the prototype). rules.md#C1 already said both things — "more than one word precedes the comma" in its opening clause and "a part before the comma with more than one name word" later — so the doc edit is a reconciliation inside C1 rather than a new claim. - 2026-09-15 (spec review) — THE COUNT REACHES THE LISTED SET TOO — one uniform rule for the whole ambiguous class rather than two. `John Smith, Ed` → given John, family Smith, suffix Ed; `john smith, ma` and `JOHN SMITH, MA` move the same way, each being a one-case name that takes the positional rule, which now reaches the comma structure. All three are 1.4.0 RESTORATIONS — measured on the wheel, v1 reads all three as suffixes — so today's tree is what deviates and this restores parity rather than opening distance from it. Corpus population: ZERO (verified by grepping `tools/differential/corpus*.jsonl` for all three; `Royce, Ed` is the corpus's only member of the sibling one-word shape), so no corpus mover count moves and `Davis Royce, Ed` is the case row that carries the decision. - 2026-09-17, corrected — A CASELESS SCRIPT IS INERT ONLY AS A LEAN, NOT AS THE NAME-WORD COUNT. `is_one_case` answers True trivially for text with no case to write a contrast in, so `ambiguous_lean` returns `None` for `毛泽东, MA`, `마틴 킹, MA` and `田中 太郎, MA` alike — the LEAN really is inert by construction, as first written. But the comma count above is of NAME words "whatever case the name is written in", which is orthogonal to case and reaches a caseless script the same as any other: `마틴 킹, MA` and `田中 太郎, MA` each hold TWO pre-comma name words, so the structure flips and both read suffix 'MA' — a PARITY RESTORATION, not a fresh deviation (1.4.0 read both as suffix on the wheel), with the Hangul/Han surname split then running positionally over what the comma leaves (`마틴 킹` → given '틴', middle '킹', family '마'; `田中 太郎` → given '太郎', family '田中'), unrelated to this design and #271/#272's pre-existing behavior. `毛泽东, MA` does not move — one pre-comma token, the count never reaches two. Pinned `a_caseless_script_wrote_no_contrast_hangul` and `a_caseless_script_wrote_no_contrast_japanese` (tests/v2/cases.py), both `tolerated=True`. -- 2026-09-17, corrected — THE COMMA-PATH REPORT TRACKS THE FORK CONSULTED, NOT THE LEAN. A first pass gated the family-comma report on whether `ambiguous_lean` fired, which wrongly silenced `Smith, ma` (all-lower, no lean) and `毛泽东, MA` (caseless, no lean) — the TRAILING slot has never gated its own report on the lean either (a bare ambiguous acronym reported before #289 too, pick declined or not), because the report is about a fork the parse CONSULTED, and the fork is consulted whenever a class member sits in the slot, lean or no lean. Fixed: the guard is class MEMBERSHIP on the first post-comma piece (`vocab:suffix-ambiguous` or `shape:acronym`, never the lean) — `Smith, ma` reports `"'ma' after the comma is also an ordinary name word; read as the given name"`, `毛泽东, MA` reports the same wording for `'MA'`, and `Smith, MA PhD` (a two-piece post-comma part) reports through its FIRST piece regardless of what follows it, exactly as a one-piece part would. The structure decision reports in `segment`, where that branch is taken, and the family-comma reading in `assign`, so one DECISION never reports twice — a second ambiguous token elsewhere is a second fork and reports on its own (`Smith MA, Ed` carries two `suffix-or-name` reports, one per fork, plus a `given-or-family` for the lone remaining word). +- 2026-09-17, corrected — THE COMMA-PATH REPORT TRACKS THE FORK CONSULTED, NOT THE LEAN. A first pass gated the family-comma report on whether `ambiguous_lean` fired, which wrongly silenced `Smith, ma` (all-lower, no lean) and `毛泽东, MA` (caseless, no lean) — the TRAILING slot has never gated its own report on the lean either (a bare ambiguous acronym reported before #289 too, pick declined or not), because the report is about a fork the parse CONSULTED, and the fork is consulted whenever a class member sits in the slot, lean or no lean. Fixed: the guard is class MEMBERSHIP on the first post-comma piece (`vocab:suffix-ambiguous` or `shape:acronym`, never the lean) — `Smith, ma` reports `"'ma' after the comma is also an ordinary name word; read as the given name"` (since #626, 2026-10-08: "read as a given name", the field worded from where the word lands), `毛泽东, MA` reports the same wording for `'MA'`, and `Smith, MA PhD` (a two-piece post-comma part) reports through its FIRST piece regardless of what follows it, exactly as a one-piece part would. The structure decision reports in `segment`, where that branch is taken, and the family-comma reading in `assign`, so one DECISION never reports twice — a second ambiguous token elsewhere is a second fork and reports on its own (`Smith MA, Ed` carries two `suffix-or-name` reports, one per fork, plus a `given-or-family` for the lone remaining word). - 2026-09-14 (Derek) — #516's DOTTED HALF IS A SWITCH, `Policy.unlisted_dotted_suffixes`, DEFAULT ON: an unlisted token of two or more period-separated chunks, any chunk length (`X.Y.Z.`, `B.Tech.`, `Q.W.E.R.T.`), joins the ambiguous class BY SHAPE; whole-token vocabulary still wins (`M.A.`, `Ph.D.`, `A.B.C.` — `abc` IS a suffix acronym). The positional rule then decides and reports either way: `John Smith X.Y.Z.` → suffix, `Jack X.Y.Z.` → family, `Smith, A.B.` → given, `John Smith, A.B.` → suffix. Case is irrelevant — the periods are the evidence — so `john smith x.y.z.` reads as its mixed-case twin. Switch OFF: name material everywhere as 2.3 read it, and the fork is STILL reported, the parser having chosen the name reading over a credential one. - 2026-09-15 (Derek) — THE ROMAN-CHUNK ACCIDENT RETIRES, NARROWLY. rules.md#S3's chunk rule survives except where every chunk the vocabulary matches is a SINGLE ASCII CHARACTER — the accident exactly. Measured, `{w for w in (L.suffix_acronyms | L.suffix_words) if len(w) == 1}` is `{i, v, 2}` plus the glued CJK honorific tails `{様, 殿, 氏, 군, 님, 씨, 양}`, so "single ASCII character" names the first three and only those — CHARACTER because `2` is a digit and is in the set, ASCII because `씨` is the one that must KEEP its chunk claim (`J.씨`). So `John Smith R.A.I.` and `John Smith J.u.n.i.o.r.` reach the shape class and read suffix by POSITION with the same fields plus the report, `Jack X.Y.I.` reads family where it read suffix (1.4.0's reading), and `Msc.Ed.`, `JD.CPA`, `Lt.Gov.`, `J.씨` keep their chunk-derived readings. The WIDE retirement — any chunk match yielding to the shape — was MEASURED AND REJECTED: it moves three more corpus names, `Doe, John Msc.Ed.` losing a real credential to `middle` and the two `김민준씨, J.씨` rows re-routing their honorific peel, for nothing this design wants. - 2026-09-17/18, #516 review rounds — THE SHAPE VERDICT NEEDS TWO FURTHER GATES, both found by measuring after the shape verdict shipped. (i) ALPHABETIC: `period_joined_vocab`'s "shape" branch now requires every chunk to be `.isalpha()` — a bare digit chunk is not an acronym letter by any reading, so `Smith, 1.4`, `John Smith 1.4` and `John Smith, 1.4` stay unclaimed (no report, no role move), where an unguarded shape verdict had wrongly admitted them; the delimited `Bridge (1.4)` control only proved digits reach extract's escape, saying nothing about the same digits undelimited, which is the gap this closes. (ii) INITIALLESS-SCRIPT: the same test `is_title_shaped` already makes for H2's own leading-abbreviation inference (#323) — `text.isascii() or not in_initialless_script(text)` — now gates the shape verdict too, so `John Smith 田.中.`, `김 민준 이.박.` and `John Smith たな.か.` stay family: a script with no period abbreviations at all has nothing for interior periods to abbreviate, so a CJK word glued into period-separated characters is not spelling an acronym either. `John Smith 田.中.` is pinned `tolerated=True` rather than `shape=` — a Latin period convention glued onto CJK characters is a composed form no writing system produces, and the case table's own validator refuses a Latin shape tag on CJK text; `이.박.` and `たな.か.` are covered by direct `period_joined_vocab` assertions instead of a second corpus row each. diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 7ef06e77..068715cc 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -114,7 +114,7 @@ Problem shape. A retired default rides in on old pickles, but a user may have de ## AMBIGUITY-AT-THE-DECISION-SITE — emit where the branch is taken -Problem shape. An ambiguity report should fire exactly when the parse chose between live readings — no more, no less — and the choosing happens in code, not in vocabulary. Contract statement. Emit at the site that takes the branch, not where an ambiguous tag sits; when a fork's two branches are decided in different stages, EVERY deciding stage carries an emitter; and a branch that runs but changes nothing is not a decision and must not report. How it works. PARTICLE_OR_GIVEN fires from assignment for a lone leading particle and from grouping when a title shifts it off the front — for two years only the first site emitted. And keying on +Problem shape. An ambiguity report should fire exactly when the parse chose between live readings — no more, no less — and the choosing happens in code, not in vocabulary. Contract statement. Emit at the site that takes the branch, not where an ambiguous tag sits; when a fork's two branches are decided in different stages, EVERY deciding stage carries an emitter; and a branch that runs but changes nothing is not a decision and must not report. The decision is the site's, but a FIELD the report names is not: a later rule can still move the word, so the emitter leaves the field to assemble (`PendingAmbiguity.field_tail`), which words it from the final role (#626; a comma-structure report on a part past the second comma still names "suffix" for a part holding a title, #629). How it works. PARTICLE_OR_GIVEN fires from assignment for a lone leading particle and from grouping when a title shifts it off the front — for two years only the first site emitted. And keying on "the code got here" instead of "the outcome differed" once reported a fork for all ambiguous particles on "Dr. Van Jr.". Lives in. The AmbiguityKind emitters across _pipeline/ (rule A1 is the observable contract). Reach for it when. Adding any ambiguous vocabulary or any new fork — count the deciding sites, then count the emitters. diff --git a/docs/design/rules.md b/docs/design/rules.md index 6f208463..3be3c5fc 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -2429,7 +2429,9 @@ A1. Rationale: a caller can only act on doubt that is reported. contradicts: a letter both a connective and an initial, read as the generation it also spells, is neither of the two the connective-or-initial fork offered, and only that fork's report - goes — the generation's own stands. + goes — the generation's own stands. And a report that names the + field its word was read into names the field the word holds in + the result, whatever rule moved it after the report was made. "Van Johnson" → ambiguities=("particle-or-given",) "JOHN QUINCY SMITH I" → ambiguities=("suffix-or-name",) "Jane „JD Smith" → ambiguities=("unbalanced-delimiter",) @@ -2439,7 +2441,7 @@ A1. Rationale: a caller can only act on doubt that is reported. segmenter's own error, which propagates — a user-code error is not a content error. (Needs the optional extra to demonstrate, so no example line.) - history: decisions.md#A1 · interacts: P3, S2 · implemented: nameparser/_pipeline/_assemble.py, nameparser/_pipeline/_state.py + history: decisions.md#A1 · interacts: H1, O1, O2, O3, P1, P3, P6, S2 · implemented: nameparser/_pipeline/_assemble.py, nameparser/_pipeline/_state.py A2. Rationale: an input with no name content names nobody, and saying so beats inventing fields from punctuation. diff --git a/docs/release_log.rst b/docs/release_log.rst index fb32432b..1097d6db 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -36,6 +36,8 @@ Release Log - **Fix a comma part read wholly as suffixes reporting a particle in it as chained onto a name.** ``parse("John Smith, Jr., Freiherr von Richthofen").ambiguities`` names ``comma-structure`` alone, where 2.0 through 2.3 also named ``particle-or-given`` for ``von`` -- a word the same parse had put in the suffix, so the report described a reading it never made. This fix moves no field; the title change below (#603) moves ``Freiherr`` to the title. A part the parser consumes as suffixes, after a suffix comma or past the second comma, reports what the part is and nothing about its words as names, and the particle chain's credential-acronym report this release adds (above) keeps the same bound: ``John Smith, PhD Do Ma`` reports the comma's decision once and not again for ``Ma``. See the 2026-09-28 bullet of the ``C1`` entry in ``docs/design/decisions.md`` + - **Fix an ambiguity report naming the field its word was first placed in rather than the one it ends up in.** ``parse("Attorney General of Minnesota").ambiguities`` says ``'General of Minnesota'`` was "read as a family name by convention", where 2.3.0 said "read as a given name" of a name whose family it was, and a leading ``Van`` that a patronymic rule rotates into the family name (``Van Ivan Petrovich`` under ``Policy(patronymic_rules=frozenset({PatronymicRule.EAST_SLAVIC}))``) is reported as read as a family name where 2.3.0 said a given one. A report that names the field one word was read into now names the field the word holds in the result, whatever rule moved it after the report was made. This fix moves no field and no report kind. See the #626 entry under ``A1`` in ``docs/design/decisions.md`` (closes #626) + - **A credential after the name core starts a suffix run to the end of its part.** ``HumanName("John Smith PhD Jones")`` gives suffix ``PhD Jones`` and reports ``suffix-or-name`` for the name word it took, where 2.3.0 gave middle ``Smith PhD``, last ``Jones``; ``Smith, John PhD Jones`` gives suffix ``PhD Jones`` where 2.3.0 gave middle ``Jones``, suffix ``PhD``. A title in the run is a title, so ``Eric H. Holder Jr. Attorney General`` gives title ``Attorney General``, suffix ``Jr.``, as the comma spelling ``Eric H. Holder, Jr. Attorney General`` already read, where 2.3.0 gave middle ``H. Holder Jr. Attorney``, last ``General``. Only an unambiguous credential or generational word starts the run, and only behind the name core -- two name words with no comma, or the given part after one. A credential that is also a name does not start one, nor does a surname particle or a single letter, so ``John Smith MA Jones`` and ``Mohamed Ali Abd Allah`` read as before, and a bare title word starts nothing: ``Mary Jane King Smith`` keeps middle ``Jane King``. See the ``S2`` entry of ``docs/design/decisions.md`` (closes #602) - **Fix a comma part that opens with a credential being read as a given name.** ``HumanName("John Smith, PhD Jones")`` gives first ``John``, last ``Smith``, suffix ``PhD Jones`` and reports ``suffix-or-name`` for the name word it took, where 2.3.0 gave first ``PhD``, middle ``Jones``, last ``John Smith`` (and 1.4.0 title ``PhD``, first ``Jones``). With one word before the comma the whole part is suffixes: ``Doe, PhD Jones`` gives last ``Doe``, suffix ``PhD Jones``. A title word in the part stays a title. A word that is a title as well as a credential (``MD``, ``Ms``) opens nothing, so ``Smith, Ms Jane`` keeps title ``Ms``, first ``Jane``. An undeclared delimiter becomes one more word of the part: ``Steven Hardman, RN - CRNA`` gives first ``Steven``, last ``Hardman``, suffix ``RN - CRNA``, where 1.4.0 through 2.3.0 gave first ``RN``, middle ``-``, last ``Steven Hardman``. The split spelling now counts like the joined one, so ``Smith, John Ph. D. Jones`` gives suffix ``Ph. D. Jones``, where 2.3.0 gave middle ``Jones``, suffix ``Ph. D.``. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #603) diff --git a/nameparser/_pipeline/_assemble.py b/nameparser/_pipeline/_assemble.py index e3aede20..a3921d4f 100644 --- a/nameparser/_pipeline/_assemble.py +++ b/nameparser/_pipeline/_assemble.py @@ -1,6 +1,8 @@ """Not a stage: converts the final ParseState into a public ParsedName. -Consumes: tokens (all roles set), dropped, ambiguities (by index). +Consumes: tokens (all roles set), dropped, ambiguities (by index; +one carrying ``field_tail`` is worded from its word's final role, +#626). Produces: a validated ParsedName -- the constructor re-checks every invariant (span order/bounds, ambiguity subset), so a pipeline bug that would produce an invalid result dies HERE, not in a renderer @@ -18,7 +20,7 @@ """ from __future__ import annotations -from nameparser._pipeline._state import ParseState +from nameparser._pipeline._state import FIELD_READING, ParseState from nameparser._types import ( Ambiguity, AmbiguityKind, ParsedName, Role, Token, ) @@ -112,8 +114,17 @@ def assemble(state: ParseState) -> ParsedName: and any(t.role in (Role.SUFFIX, Role.TITLE) for t in materialized)): continue - ambiguities.append( - Ambiguity(pending.kind, pending.detail, materialized)) + detail = pending.detail + if pending.field_tail is not None: + # A report naming a field, in a name emptied above (no + # emitter of one reaches that today: every referent is a + # name word): no field is left to name, so it goes rather + # than ship the sentence's first half. + if not materialized: + continue + detail = (f"{detail}{FIELD_READING[materialized[0].role]}" + f"{pending.field_tail}") + ambiguities.append(Ambiguity(pending.kind, detail, materialized)) return ParsedName(original=state.original, tokens=tuple(final.values()), ambiguities=tuple(ambiguities)) diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index ab43a276..eae0ffad 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -363,9 +363,10 @@ def _assign_main(seg_idx: int, state: ParseState, # # Every bare ambiguous acronym the FINAL peel had to resolve is one # coin-flip each, in either direction, so the report collects - # rather than overwrites. Deferred to after assignment because the - # wording reads the role back, and which role "not peeled" means - # depends on name_order. (The roman-numeral fork needs no such + # rather than overwrites. Deferred to after assignment because what + # a pick is declined as depends on the role it took, and which role + # "not peeled" means depends on name_order; the field itself is + # worded by assemble (#626). (The roman-numeral fork needs no such # deferral and is reported here.) # # Read once in group, after P3's connective joins and ahead of the @@ -435,8 +436,6 @@ def _assign_main(seg_idx: int, state: ParseState, # call budget on every parse otherwise (decisions.md#parse-cost). if peeled.names <= 1: head = pieces[name_pieces[0]] - token = tokens[head[0]] - assert token.role is not None # Both conventions here turn on a lone name word, so both # report at the site that places one # (mechanisms.md#AMBIGUITY-AT-THE-DECISION-SITE), and one @@ -458,16 +457,16 @@ def _assign_main(seg_idx: int, state: ParseState, # `suffix-or-name`" (history: decisions.md#H4) -- only the # word made into a name reports, and no name piece survived # the peel, so the carve-out above made the first post-nominal - # the name. The role comes off the token for the reason stated - # at the particle emitter below. + # the name. Its field is worded by assemble from the final role + # (`PendingAmbiguity.field_tail`, #626). if peeled.names == 0 and field_undecided: text = " ".join(tokens[i].text for i in head) ambiguities.append(PendingAmbiguity( AmbiguityKind.SUFFIX_OR_NAME, f"{text!r} is post-nominal vocabulary with no name word " - f"beside it; read as a {token.role.value} name rather than " - f"a post-nominal, nothing else being left to be the name", - tuple(head))) + f"beside it; read as ", tuple(head), + field_tail=" rather than a post-nominal, nothing else " + "being left to be the name")) # One name piece off the peel: the convention placed a lone # name word. A suffix beside it is not a decision -- 'Smith # Jr.' and "'Smitty' Jones Jr." ARE this convention -- and the @@ -489,16 +488,17 @@ def _assign_main(seg_idx: int, state: ParseState, # Wales'): the leading-title peel took the whole name. if any("vocab:title" in tokens[i].tags for i in head): text = " ".join(tokens[i].text for i in head) + lone = len(head) == 1 ambiguities.append(PendingAmbiguity( AmbiguityKind.TITLE_OR_NAME, f"{text!r} is title vocabulary and the only name word " f"the title peel left standing; read as the name by " f"convention rather than as more title" - if len(head) == 1 else + if lone else f"{text!r} is the only name unit and joins title " - f"vocabulary to a name word; read as a " - f"{token.role.value} name by convention", - tuple(head))) + f"vocabulary to a name word; read as ", + tuple(head), + field_tail=None if lone else " by convention")) # rules.md#O5: "a name of one name word that nothing else has # decided reads that word as the given name under the default # given-first order, and as the family name under a declared @@ -517,16 +517,17 @@ def _assign_main(seg_idx: int, state: ParseState, ambiguities.append(PendingAmbiguity( AmbiguityKind.GIVEN_OR_FAMILY, f"{text!r} is the only name word and nothing else " - f"decides it; read as a {token.role.value} name by " - f"convention, which follows the read order", - tuple(head))) + f"decides it; read as ", tuple(head), + field_tail=" by convention, which follows the read " + "order")) for piece in peeled.picks: - # every pick is in rest, so the loops above just gave it a role + # every pick is in rest, so the loops above just gave it a + # role; whether it was read as a suffix or a name settles what + # it was declined as, and which field assemble names (#626) token = tokens[piece[0]] assert token.role is not None - taken, declined = ( - ("a suffix", "a name part") if token.role is Role.SUFFIX - else (f"a {token.role.value} name", "a post-nominal")) + declined = ("a name part" if token.role is Role.SUFFIX + else "a post-nominal") # A word in the class by SHAPE (rules.md#S2, #S3) is in no # wordlist and may be written with its periods, so the listed # member's wording would misdescribe it on both counts (#563) @@ -537,27 +538,23 @@ def _assign_main(seg_idx: int, state: ParseState, "and an ordinary name") ambiguities.append(PendingAmbiguity( AmbiguityKind.SUFFIX_OR_NAME, - f"{token.text!r} {what}; read as {taken} rather than " - f"{declined}", - piece)) + f"{token.text!r} {what}; read as ", piece, + field_tail=f" rather than {declined}")) # leading ambiguous particle read as a name (#121 surfaced) if name_pieces: head = pieces[name_pieces[0]] if (len(head) == 1 and len(name_pieces) > 1 and "vocab:particle-ambiguous" in tokens[head[0]].tags): - # the loops above gave the head piece its role from - # `order`, which is _effective_order's answer and not - # necessarily name_order's -- a script_orders entry - # overrides it. So read the role off the token rather than - # assume given, or re-derive it here; same reason as - # SUFFIX_OR_NAME just above. - token = tokens[head[0]] - assert token.role is not None + # the field is assemble's to word from the final role + # (`PendingAmbiguity.field_tail`, #626), never assumed given: + # `order` above is _effective_order's answer, not + # necessarily name_order's, and a later rule may still move + # the word ('Van Ivan Petrovich' under the East Slavic rule + # rotates 'Van' into the family) ambiguities.append(PendingAmbiguity( AmbiguityKind.PARTICLE_OR_GIVEN, - f"leading {token.text!r} may be a family-name " - f"particle; read as a {token.role.value} name", - tuple(head))) + f"leading {tokens[head[0]].text!r} may be a family-name " + f"particle; read as ", tuple(head), field_tail="")) return order @@ -852,11 +849,13 @@ def assign(state: ParseState) -> ParseState: if state.pieces[1] and len(state.pieces[1][0]) == 1: i = state.pieces[1][0][0] if not tokens[i].tags.isdisjoint(_AMBIGUOUS_CREDENTIAL_TAGS): + # the field is assemble's to word: with nothing before + # the comma and a title after the word, H1 moves it to + # the family (', Ma Dr.', #626's review) ambiguities.append(PendingAmbiguity( AmbiguityKind.SUFFIX_OR_NAME, f"{tokens[i].text!r} after the comma is also an " - f"ordinary name word; read as the given name", - (i,))) + f"ordinary name word; read as ", (i,), field_tail="")) # Segment 1 is read before segment 0's wholly-family pass. It # consumes piece tags and text only -- nothing segment 0's # read writes. (A title standing after the comma behind a whole diff --git a/nameparser/_pipeline/_state.py b/nameparser/_pipeline/_state.py index d76c6261..dd37336d 100644 --- a/nameparser/_pipeline/_state.py +++ b/nameparser/_pipeline/_state.py @@ -130,12 +130,35 @@ class PendingAmbiguity: extract_delimited knows only a character offset, so it records that and tokenize resolves it to the containing token's index. Stages after tokenize set ``indices`` directly and leave ``origin`` None. + + ``field_tail`` is for a report that names the field its word was + read into. post_rules may still move the word (H1's move behind a + title, P1's family-first fold, P6's attachment, and under opt-in + policies the patronymic rotations and `middle_as_family`'s fold), + so the emitter + writes ``detail`` up to the field and the rest of its sentence + here -- ``""`` where the field ends the sentence, which is why + assemble tests ``is not None`` -- and assemble joins them around + the FIELD_READING of the first referent's final role, the one + place that knows it (#626). Until then ``detail`` is a fragment. """ kind: AmbiguityKind detail: str indices: tuple[int, ...] = () origin: int | None = None + field_tail: str | None = None + + +#: How a report names the field a word was read into, by final role. +FIELD_READING: Mapping[Role, str] = MappingProxyType({ + Role.TITLE: "a title", Role.GIVEN: "a given name", + Role.MIDDLE: "a middle name", Role.FAMILY: "a family name", + Role.SUFFIX: "a suffix", Role.NICKNAME: "a nickname", + Role.MAIDEN: "a maiden name"}) +# every role a referent can hold, or a new one is a KeyError at the +# first parse whose report lands in it +assert FIELD_READING.keys() == set(Role) @dataclass(frozen=True, slots=True) diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 936111c4..c20ad5f9 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -8221,6 +8221,65 @@ def _check_cjk_shape_purity(self) -> None: "whether `Prince` is a title rather than which field " "the unit takes -- one of the inputs measured to reach " "that branch, `prince` being in TITLES"), + Case("a_joined_unit_h1_moves_is_reported_in_the_field_it_lands_in", + "Attorney General of Minnesota", + {"title": "Attorney", "family": "General of Minnesota"}, + ambiguities=("title-or-name",), classification="parity", + notes="#626: assign places the lone unit as the given name and " + "H1 moves it to the family behind the title; the report " + "was worded at assign and said 'given'. The order sweep " + "in test_cases.py checks the field it names"), + Case("a_pick_the_attachment_moves_is_reported_in_its_family", + "Kim Min Do", + {"given": "Kim", "middle": "Min", "family": "Do"}, + ambiguities=("suffix-or-name",), classification="fix(#289)", + notes="#626: under FAMILY_FIRST the order sweep reads P6 " + "attaching the Title-case 'Do' to the family ('Do Kim') " + "after assign read it as a middle name; the report was " + "worded at assign and said 'middle'. Under " + "FAMILY_FIRST_GIVEN_LAST nothing moves it: given 'Do'. " + "1.4.0 read suffix 'Do'"), + Case("a_pick_the_family_first_fold_moves_is_reported_where_it_lands", + "de Kim Ma", {"family": "de Kim", "given": "Ma"}, + ambiguities=("suffix-or-name",), classification="fix(#289)", + policy=Policy(name_order=FAMILY_FIRST), + notes="#626's review: P1's family-first fold takes the particle " + "run and one name unit, and lays the rest out by the " + "order, so 'Ma', which assign read as a middle name, " + "lands in the given name -- a third rule moving a " + "reported word after assign. 2.3.0 read suffix 'Ma'"), + Case("a_leading_particle_the_east_slavic_rotation_moves_is_reported_in_the_family", + "Van Ivan Petrovich", + {"given": "Ivan", "middle": "Petrovich", "family": "Van"}, + ambiguities=("particle-or-given",), classification="parity", + policy=_ES, + notes="#626's second review: O1 rotates the leading 'Van', which " + "assign reported as a given name, into the family; the " + "report names where it lands. Reverting that emitter " + "passed every test until this row and its Turkic twin. " + "1.4.0 with patronymic_name_order reads the same fields"), + Case("a_leading_particle_the_turkic_rotation_moves_is_reported_in_the_family", + "Van Ali Veli oglu", + {"given": "Ali", "middle": "Veli oglu", "family": "Van"}, + ambiguities=("particle-or-given",), classification="parity", + policy=_TK, + notes="#626's second review: O2's twin of the row above"), + Case("a_pick_the_middle_as_family_fold_moves_is_reported_in_the_family", + "Do Bishop Do", {"given": "Bishop", "family": "Do Do"}, + ambiguities=("particle-or-given", "suffix-or-name"), + classification="fix(#289)", + policy=Policy(name_order=FAMILY_FIRST, middle_as_family=True), + notes="#626's second review: the trailing 'Do' is a middle name " + "at assign and O3's fold takes it into the family. 2.3.0 " + "read suffix 'Do'"), + Case("a_comma_report_h1_moves_is_reported_in_the_family", + ", Ma Dr.", {"title": "Dr.", "family": "Ma"}, + ambiguities=("suffix-or-name",), classification="fix(#316)", + notes="#626's review: with nothing before the comma, the comma " + "report on 'Ma' was worded 'the given name' at assign, " + "and H1 then moved it to the family behind the trailing " + "title. The empty-head comma shape AGENTS.md's sweep " + "paragraph asks for; 1.4.0 read first 'Ma', suffix 'Dr.'"), Case("lone_joined_unit_carrying_collision_set_title_vocabulary", "Smith and King", {"given": "Smith and King"}, ambiguities=("title-or-name",), diff --git a/tests/v2/pipeline/test_assemble.py b/tests/v2/pipeline/test_assemble.py index 749ef85d..79efbad2 100644 --- a/tests/v2/pipeline/test_assemble.py +++ b/tests/v2/pipeline/test_assemble.py @@ -1,7 +1,7 @@ from nameparser._lexicon import Lexicon from nameparser._pipeline import run from nameparser._pipeline._assemble import assemble -from nameparser._pipeline._state import ParseState, WorkToken +from nameparser._pipeline._state import ParseState, PendingAmbiguity, WorkToken from nameparser._policy import Policy from nameparser._types import AmbiguityKind, ParsedName, Role, Span @@ -197,3 +197,56 @@ def test_the_withdrawal_reads_the_role_and_not_the_word() -> None: "read as an initial", (dr,)),)) assert AK.CONJUNCTION_OR_INITIAL not in [ a.kind for a in assemble(poisoned).ambiguities] + + +def _with_reports(original: str, tokens: tuple[WorkToken, ...], + *reports: PendingAmbiguity) -> ParsedName: + return assemble(ParseState( + original=original, lexicon=Lexicon.default(), policy=Policy(), + tokens=tokens, ambiguities=reports)) + + +def test_a_field_tail_report_is_worded_from_the_final_role() -> None: + # #626: the emitter writes the sentence up to the field and the + # rest as field_tail; assemble names the field of the FIRST + # referent's final role, whatever the emitter saw. An empty tail is + # a real tail (the field ends the sentence), not an absent one. + tokens = (WorkToken("Kim", Span(0, 3), role=Role.FAMILY), + WorkToken("Do", Span(4, 6), role=Role.FAMILY)) + pn = _with_reports( + "Kim Do", tokens, + PendingAmbiguity(AmbiguityKind.SUFFIX_OR_NAME, "'Do' read as ", + (1,), field_tail=" rather than a post-nominal"), + PendingAmbiguity(AmbiguityKind.PARTICLE_OR_GIVEN, + "leading 'Kim' read as ", (0,), field_tail=""), + PendingAmbiguity(AmbiguityKind.ORDER, "a whole sentence", (0,))) + assert [a.detail for a in pn.ambiguities] == [ + "'Do' read as a family name rather than a post-nominal", + "leading 'Kim' read as a family name", + "a whole sentence"] + + +def test_a_field_tail_report_in_an_emptied_name_is_withdrawn() -> None: + # A name emptied for want of content keeps its reports, but one + # naming a field has no field left to name: it goes rather than + # ship the sentence's first half. No emitter reaches this today. + pn = _with_reports( + "-", (WorkToken("-", Span(0, 1), role=Role.GIVEN),), + PendingAmbiguity(AmbiguityKind.GIVEN_OR_FAMILY, "'-' read as ", + (0,), field_tail=" by convention"), + PendingAmbiguity(AmbiguityKind.UNBALANCED_DELIMITER, "kept", (0,))) + assert not pn + assert [a.detail for a in pn.ambiguities] == ["kept"] + + +def test_a_field_tail_report_names_its_first_referent_s_field() -> None: + # The field comes from the FIRST referent: every emitter's referents + # share one role today, so this pins the documented choice rather + # than a case the pipeline produces. + tokens = (WorkToken("General", Span(0, 7), role=Role.FAMILY), + WorkToken("Smith", Span(8, 13), role=Role.GIVEN)) + pn = _with_reports( + "General Smith", tokens, + PendingAmbiguity(AmbiguityKind.TITLE_OR_NAME, "read as ", (0, 1), + field_tail="")) + assert [a.detail for a in pn.ambiguities] == ["read as a family name"] diff --git a/tests/v2/pipeline/test_assign.py b/tests/v2/pipeline/test_assign.py index d96499f0..d2ed35ad 100644 --- a/tests/v2/pipeline/test_assign.py +++ b/tests/v2/pipeline/test_assign.py @@ -2,6 +2,8 @@ import pytest from nameparser._lexicon import Lexicon +from nameparser._pipeline import run +from nameparser._pipeline._assemble import assemble from nameparser._pipeline._assign import assign from nameparser._pipeline._classify import classify from nameparser._pipeline._extract import extract_delimited @@ -13,7 +15,7 @@ from nameparser._policy import ( FAMILY_FIRST, FAMILY_FIRST_GIVEN_LAST, Policy, Script, ) -from nameparser._types import AmbiguityKind, Role +from nameparser._types import Ambiguity, AmbiguityKind, Role #: The three read orders and the role a lone name word takes under #: each, shared by every parametrized case below that asks the same @@ -45,6 +47,17 @@ def _assigned(text: str, policy: Policy | None = None, extract_delimited(state)))))) +def _reported(text: str, policy: Policy | None = None, + lexicon: Lexicon | None = None) -> tuple[Ambiguity, ...]: + """The reports as the caller reads them. A report naming a field + is worded by assemble from the word's FINAL role (#626), so its + detail is complete only there, past the rules that still move a + word after assign.""" + return assemble(run(ParseState( + original=text, lexicon=lexicon or _LEX, + policy=policy or Policy()))).ambiguities + + def _by_role(state: ParseState, role: Role) -> str: return " ".join(t.text for t in state.tokens if t.role is role) @@ -145,8 +158,7 @@ def test_a_by_shape_pick_is_not_described_as_a_listed_member() -> None: for text, taken, declined in ( ("John Smith X.Y.Z.", "a suffix", "a name part"), ("Jack X.Y.Z.", "a family name", "a post-nominal")): - out = _assigned(text, lexicon=Lexicon.default()) - (amb,) = [a for a in out.ambiguities + (amb,) = [a for a in _reported(text, lexicon=Lexicon.default()) if a.kind is AmbiguityKind.SUFFIX_OR_NAME] assert amb.detail == ( f"'X.Y.Z.' is shaped like a post-nominal but listed in no " @@ -161,8 +173,8 @@ def test_jack_ma_s_two_detail_strings_are_verbatim() -> None: # strings verbatim: the credential lean also turns 'Jack' into the # only name word left, which is a SECOND fork (GIVEN_OR_FAMILY) # this one row now calls. - out = _assigned("Jack MA", lexicon=Lexicon.default()) - details = {a.kind.value: a.detail for a in out.ambiguities} + details = {a.kind.value: a.detail + for a in _reported("Jack MA", lexicon=Lexicon.default())} assert details == { "given-or-family": ( "'Jack' is the only name word and nothing else decides " @@ -221,16 +233,16 @@ def test_leading_ambiguous_particle_reads_as_given_with_ambiguity() -> None: @pytest.mark.parametrize("policy,role", _ORDERS) -def test_leading_particle_detail_names_the_role_it_took( +def test_leading_particle_detail_names_the_field_the_word_lands_in( policy: Policy | None, role: str) -> None: # The fork is the same under every order -- particle or name -- # but which role the head piece actually took is the assignment's - # answer, so the user-facing detail has to read it off the token - # rather than hardcode "given", exactly as SUFFIX_OR_NAME does. + # answer, so the user-facing detail names the field the word lands + # in rather than hardcoding "given", exactly as SUFFIX_OR_NAME does. # kind is public API and stays PARTICLE_OR_GIVEN throughout: the # fork really is "particle or given" even where the piece landed # in FAMILY. - (amb,) = _assigned("Van Johnson", policy).ambiguities + (amb,) = _reported("Van Johnson", policy) assert amb.kind is AmbiguityKind.PARTICLE_OR_GIVEN assert amb.detail == ( f"leading 'Van' may be a family-name particle; " @@ -250,12 +262,12 @@ def test_leading_particle_detail_names_the_role_it_took( "'John of Prince' is the only name unit and joins title " "vocabulary to a name word; read as a {role} name by convention"), ]) -def test_convention_details_name_the_role_the_assignment_took( +def test_convention_details_name_the_field_the_word_lands_in( text: str, kind: AmbiguityKind, detail: str, policy: Policy | None, role: str) -> None: # The three conventions this site reports all place a lone name - # word, and all three details have to READ the field back off the - # token rather than hardcode "given": the field follows the read + # word, and all three details name the field the word lands in + # rather than hardcoding "given": the field follows the read # order, which is why none of the three kinds names it. The # PARTICLE_OR_GIVEN test above is the precedent, and these are the # only other details at this site that name a field -- H4's PEEL @@ -263,7 +275,7 @@ def test_convention_details_name_the_role_the_assignment_took( # retags that word after assign under the default order. lex = _LEX.add(suffix_words={"rinpoche"}, conjunctions={"of"}, titles={"prince"}) - (amb,) = _assigned(text, policy, lexicon=lex).ambiguities + (amb,) = _reported(text, policy, lexicon=lex) assert amb.kind is kind assert amb.detail == detail.format(role=role) @@ -281,7 +293,7 @@ def test_leading_particle_detail_follows_the_effective_order() -> None: assert out.policy.script_orders[0][0] is Script.HAN assert _by_role(out, Role.FAMILY) == "毛" assert _by_role(out, Role.GIVEN) == "泽东" - (amb,) = out.ambiguities + (amb,) = _reported("毛 泽东", lexicon=han) assert amb.kind is AmbiguityKind.PARTICLE_OR_GIVEN assert amb.detail == ( "leading '毛' may be a family-name particle; " @@ -746,19 +758,25 @@ def test_the_comma_report_says_which_way_it_read_the_word() -> None: fork the parser considered and declined -- is the second one. """ detail = { - text: [a.detail for a in _assigned( - text, policy, lexicon=Lexicon.default()).ambiguities + text: [a.detail for a in _reported( + text, policy, lexicon=Lexicon.default()) if a.kind.value == "suffix-or-name"] - for text, policy in (("Smith, Ma", None), + for text, policy in (("Smith, Ma", None), (", Ma Dr.", None), ("Smith, A.B.", Policy(unlisted_dotted_suffixes=False))) } assert detail["Smith, Ma"] == [ "'Ma' after the comma is also an ordinary name word; read as " - "the given name"] + "a given name"] assert detail["Smith, A.B."] == [ "'A.B.' after the comma is also an ordinary name word; read as " - "the given name"] + "a given name"] + # #626's review: with nothing before the comma, H1 moves the word + # to the family behind the title after this report is made, so the + # field is worded from where it lands + assert detail[", Ma Dr."] == [ + "'Ma' after the comma is also an ordinary name word; read as " + "a family name"] def test_the_trailing_given_slot_takes_a_bare_class_member() -> None: diff --git a/tests/v2/test_cases.py b/tests/v2/test_cases.py index 1c2c1b51..300e0729 100644 --- a/tests/v2/test_cases.py +++ b/tests/v2/test_cases.py @@ -1,10 +1,13 @@ """Core runner over the shared case table. The facade runner (migration plan) consumes the same CASES.""" +import functools +import re from typing import Any import pytest -from nameparser import Parser, Policy, Role, locales, parser_for +from nameparser import ParsedName, Parser, Policy, Role, locales, parser_for +from nameparser._pipeline._state import FIELD_READING, NAME_ROLES from nameparser._policy import FAMILY_FIRST, FAMILY_FIRST_GIVEN_LAST from .cases import CASES, Case @@ -39,10 +42,33 @@ def test_case(case: Case) -> None: #: that declare one. _INVARIANT_ORDERS = (None, FAMILY_FIRST, FAMILY_FIRST_GIVEN_LAST) +_CASES_BY_ID = {c.id: c for c in CASES} -@pytest.mark.parametrize("order", _INVARIANT_ORDERS, - ids=lambda o: _ORDER_NAMES.get(o, "as-declared")) -@pytest.mark.parametrize("case", CASES, ids=lambda c: c.id) +#: The (row, order) pairs the order sweep checks: every row under each +#: order, a row with its own policy only as declared, and no locale row +#: (a pack's parser is not one of the three orders). Both invariants +#: below walk this one list. +_SWEPT_ROWS = [(case, order) for case in CASES for order in _INVARIANT_ORDERS + if case.locale is None + and not (order is not None and case.policy)] +_SWEPT_IDS = [f"{case.id}-" + f"{'as-declared' if order is None else _ORDER_NAMES[order]}" + for case, order in _SWEPT_ROWS] + + +@functools.cache +def _swept(case_id: str, + order: tuple[Role, Role, Role] | None) -> ParsedName: + """One parse per row and order, shared by every invariant the + sweep checks (AGENTS.md: a new invariant over an existing grid + joins that grid's walk).""" + case = _CASES_BY_ID[case_id] + parser = (_parser_for_case(case) if order is None + else Parser(policy=Policy(name_order=order))) + return parser.parse(case.text) + + +@pytest.mark.parametrize("case,order", _SWEPT_ROWS, ids=_SWEPT_IDS) def test_the_family_partitions_into_particles_and_base( case: Case, order: tuple[Role, Role, Role] | None) -> None: """rules.md#R2's invariant: a particle needs a base to attach to, @@ -58,11 +84,7 @@ def test_the_family_partitions_into_particles_and_base( views render particles first ("Vega, de la" is family 'Vega de la', particles 'de la', base 'Vega'). """ - if case.locale is not None or (order is not None and case.policy): - pytest.skip("row carries its own policy or locale") - parser = (_parser_for_case(case) if order is None - else Parser(policy=Policy(name_order=order))) - pn = parser.parse(case.text) + pn = _swept(case.id, order) if not pn.family: return assert pn.family_base, ( @@ -74,6 +96,90 @@ def test_the_family_partitions_into_particles_and_base( f"particles={pn.family_particles!r} + base={pn.family_base!r}") +_NAME_FIELDS = frozenset(NAME_ROLES) + +#: Every way a report's detail names the field its word was read into, +#: mapped to the fields that make it true: assemble's FIELD_READING, +#: which words every report a later rule could move (#626), plus the +#: phrasings worded where the field is settled -- a comma's credential +#: reading, the suffix run's and the numeral fork's ("reads as part of +#: the suffix run", "a generational suffix"), H4's "the name", P6's +#: "joins the family". A report that names a READING rather than a +#: field ("read as an initial") is out of scope, as is any phrasing +#: not listed here: the count below is what notices the list going +#: stale. +_FIELD_CLAIMS: dict[str, frozenset[Role]] = { + **{reading: frozenset({role}) for role, reading in FIELD_READING.items()}, + **dict.fromkeys(("a credential", "a generational suffix", + "part of the suffix run", "a post-nominal"), + frozenset({Role.SUFFIX})), + "a name": _NAME_FIELDS, "the name": _NAME_FIELDS, + "the family": frozenset({Role.FAMILY}), +} +_FIELD_CLAIM = re.compile( + r"\b(?:reads? as|joins) (" + + "|".join(map(re.escape, sorted(_FIELD_CLAIMS, key=len, reverse=True))) + + r")\b") + + +@pytest.mark.parametrize("case,order", _SWEPT_ROWS, ids=_SWEPT_IDS) +def test_a_report_names_the_field_its_word_lands_in( + case: Case, order: tuple[Role, Role, Role] | None) -> None: + """#626: a report's detail that names the field its word was read + into names the field the word holds in the RESULT. Rules in + post_rules move a word after assign reports it -- H1's move behind + a title, P1's family-first fold, P6's attachment, and under opt-in + policies the patronymic rotations and `middle_as_family`'s fold, + which rows carrying those policies reach as declared -- and six of + assign's reports were worded from the role at assign, so + 'Kim Min Do' under FAMILY_FIRST was told 'Do' was a middle name + when it landed in the family. Assemble words them now. + + Checked over `_SWEPT_ROWS`, which the count test below walks too. + + Negative control, measured 2026-10-08: this test over master's + parser (26cdb891) fails 14 parses, every one a report worded at + assign -- nine title-or-name join reports saying 'given' of a unit + H1 moved to the family behind a title (`Attorney General of + Minnesota`, `John of Prince Prof.`, `St St née`, `Freiherr von + Bishop X.Y.Z.` and five more rows, as declared), `Kim Min Do` + under FAMILY_FIRST ('middle', family, P6), `de Kim Ma` under + FAMILY_FIRST ('middle', given, P1), and the opt-in movers' rows as + declared: `Van Ivan Petrovich` and `Van Ali Veli oglu` ('given', + family, O1 and O2) and `Do Bishop Do` ('middle', family, O3). The + sixth emitter, the comma report on a word after an empty head, is + NOT in that count: master + worded it 'the given name', a phrasing the tree no longer emits + and so not one `_FIELD_CLAIMS` lists, and `, Ma Dr.` (H1 moving it + to the family) is pinned instead by test_assign.py's verbatim + detail. No row reached that shape until #626's review. + """ + pn = _swept(case.id, order) + for a in pn.ambiguities: + m = _FIELD_CLAIM.search(a.detail) + if m is None: + continue + landed = {t.role for t in a.tokens} + assert landed <= _FIELD_CLAIMS[m[1]], ( + f"{case.text!r}: {a.kind.value} says {m[0]!r} but its " + f"tokens landed in {sorted(r.value for r in landed)}: " + f"{a.detail!r}") + + +def test_the_field_sweep_sees_the_claims_it_checks() -> None: + """The sweep above passes vacuously if no detail matches + `_FIELD_CLAIM` -- an emitter reworded out of the list, or rows + dropped -- so the count of claims it checks is recorded here. + Move it deliberately when a row or a phrasing moves, never to 0.""" + claims = sum( + _FIELD_CLAIM.search(a.detail) is not None + for case, order in _SWEPT_ROWS + for a in _swept(case.id, order).ambiguities) + assert claims == 1155, ( + f"the field sweep checks {claims} claims, recorded as 1155 on " + f"2026-10-08") + + #: Case.__post_init__'s shape checks, each probed for the message that #: identifies it. A row here is a Case that must fail to construct, not #: one that ever joins CASES -- unlike test_case above, this exercises diff --git a/tests/v2/test_facade_cases.py b/tests/v2/test_facade_cases.py index e750017f..ed8bca42 100644 --- a/tests/v2/test_facade_cases.py +++ b/tests/v2/test_facade_cases.py @@ -105,6 +105,11 @@ # #606: the excluded salutation under the order Vietnamese names # use, which has no v1 spelling. "salutation_that_leads_a_surname_stays_a_name", + # #626's reviews: P1's family-first fold and O3's middle_as_family + # fold moving a reported word, under a declared order v1 has no + # spelling for + "a_pick_the_family_first_fold_moves_is_reported_where_it_lands", + "a_pick_the_middle_as_family_fold_moves_is_reported_in_the_family", # #518 review round: the `by_script` scope's positive controls. # An emptied script table has no v1 spelling -- v1 has no script # orders to empty -- so both rows are core-only, though the roles