Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions AGENTS.md

Large diffs are not rendered by default.

8 changes: 7 additions & 1 deletion docs/design/decisions.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion docs/design/mechanisms.md
Original file line number Diff line number Diff line change
Expand Up @@ -114,7 +114,7 @@ Problem shape. A retired default rides in on old pickles, but a user may have de

## AMBIGUITY-AT-THE-DECISION-SITE — emit where the branch is taken

Problem shape. An ambiguity report should fire exactly when the parse chose between live readings — no more, no less — and the choosing happens in code, not in vocabulary. Contract statement. Emit at the site that takes the branch, not where an ambiguous tag sits; when a fork's two branches are decided in different stages, EVERY deciding stage carries an emitter; and a branch that runs but changes nothing is not a decision and must not report. How it works. PARTICLE_OR_GIVEN fires from assignment for a lone leading particle and from grouping when a title shifts it off the front — for two years only the first site emitted. And keying on
Problem shape. An ambiguity report should fire exactly when the parse chose between live readings — no more, no less — and the choosing happens in code, not in vocabulary. Contract statement. Emit at the site that takes the branch, not where an ambiguous tag sits; when a fork's two branches are decided in different stages, EVERY deciding stage carries an emitter; and a branch that runs but changes nothing is not a decision and must not report. The decision is the site's, but a FIELD the report names is not: a later rule can still move the word, so the emitter leaves the field to assemble (`PendingAmbiguity.field_tail`), which words it from the final role (#626; a comma-structure report on a part past the second comma still names "suffix" for a part holding a title, #629). How it works. PARTICLE_OR_GIVEN fires from assignment for a lone leading particle and from grouping when a title shifts it off the front — for two years only the first site emitted. And keying on
"the code got here" instead of "the outcome differed" once reported
a fork for all ambiguous particles on "Dr. Van Jr.". Lives in. The AmbiguityKind emitters across _pipeline/ (rule A1 is the observable contract). Reach for it when. Adding any ambiguous vocabulary or any new fork — count the deciding sites, then count the emitters.

Expand Down
6 changes: 4 additions & 2 deletions docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -2429,7 +2429,9 @@ A1. Rationale: a caller can only act on doubt that is reported.
contradicts: a letter both a connective and an initial, read as
the generation it also spells, is neither of the two the
connective-or-initial fork offered, and only that fork's report
goes — the generation's own stands.
goes — the generation's own stands. And a report that names the
field its word was read into names the field the word holds in
the result, whatever rule moved it after the report was made.
"Van Johnson" → ambiguities=("particle-or-given",)
"JOHN QUINCY SMITH I" → ambiguities=("suffix-or-name",)
"Jane „JD Smith" → ambiguities=("unbalanced-delimiter",)
Expand All @@ -2439,7 +2441,7 @@ A1. Rationale: a caller can only act on doubt that is reported.
segmenter's own error, which propagates — a user-code error is
not a content error. (Needs the optional extra to demonstrate,
so no example line.)
history: decisions.md#A1 · interacts: P3, S2 · implemented: nameparser/_pipeline/_assemble.py, nameparser/_pipeline/_state.py
history: decisions.md#A1 · interacts: H1, O1, O2, O3, P1, P3, P6, S2 · implemented: nameparser/_pipeline/_assemble.py, nameparser/_pipeline/_state.py

A2. Rationale: an input with no name content names nobody, and
saying so beats inventing fields from punctuation.
Expand Down
2 changes: 2 additions & 0 deletions docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,8 @@ Release Log

- **Fix a comma part read wholly as suffixes reporting a particle in it as chained onto a name.** ``parse("John Smith, Jr., Freiherr von Richthofen").ambiguities`` names ``comma-structure`` alone, where 2.0 through 2.3 also named ``particle-or-given`` for ``von`` -- a word the same parse had put in the suffix, so the report described a reading it never made. This fix moves no field; the title change below (#603) moves ``Freiherr`` to the title. A part the parser consumes as suffixes, after a suffix comma or past the second comma, reports what the part is and nothing about its words as names, and the particle chain's credential-acronym report this release adds (above) keeps the same bound: ``John Smith, PhD Do Ma`` reports the comma's decision once and not again for ``Ma``. See the 2026-09-28 bullet of the ``C1`` entry in ``docs/design/decisions.md``

- **Fix an ambiguity report naming the field its word was first placed in rather than the one it ends up in.** ``parse("Attorney General of Minnesota").ambiguities`` says ``'General of Minnesota'`` was "read as a family name by convention", where 2.3.0 said "read as a given name" of a name whose family it was, and a leading ``Van`` that a patronymic rule rotates into the family name (``Van Ivan Petrovich`` under ``Policy(patronymic_rules=frozenset({PatronymicRule.EAST_SLAVIC}))``) is reported as read as a family name where 2.3.0 said a given one. A report that names the field one word was read into now names the field the word holds in the result, whatever rule moved it after the report was made. This fix moves no field and no report kind. See the #626 entry under ``A1`` in ``docs/design/decisions.md`` (closes #626)

- **A credential after the name core starts a suffix run to the end of its part.** ``HumanName("John Smith PhD Jones")`` gives suffix ``PhD Jones`` and reports ``suffix-or-name`` for the name word it took, where 2.3.0 gave middle ``Smith PhD``, last ``Jones``; ``Smith, John PhD Jones`` gives suffix ``PhD Jones`` where 2.3.0 gave middle ``Jones``, suffix ``PhD``. A title in the run is a title, so ``Eric H. Holder Jr. Attorney General`` gives title ``Attorney General``, suffix ``Jr.``, as the comma spelling ``Eric H. Holder, Jr. Attorney General`` already read, where 2.3.0 gave middle ``H. Holder Jr. Attorney``, last ``General``. Only an unambiguous credential or generational word starts the run, and only behind the name core -- two name words with no comma, or the given part after one. A credential that is also a name does not start one, nor does a surname particle or a single letter, so ``John Smith MA Jones`` and ``Mohamed Ali Abd Allah`` read as before, and a bare title word starts nothing: ``Mary Jane King Smith`` keeps middle ``Jane King``. See the ``S2`` entry of ``docs/design/decisions.md`` (closes #602)

- **Fix a comma part that opens with a credential being read as a given name.** ``HumanName("John Smith, PhD Jones")`` gives first ``John``, last ``Smith``, suffix ``PhD Jones`` and reports ``suffix-or-name`` for the name word it took, where 2.3.0 gave first ``PhD``, middle ``Jones``, last ``John Smith`` (and 1.4.0 title ``PhD``, first ``Jones``). With one word before the comma the whole part is suffixes: ``Doe, PhD Jones`` gives last ``Doe``, suffix ``PhD Jones``. A title word in the part stays a title. A word that is a title as well as a credential (``MD``, ``Ms``) opens nothing, so ``Smith, Ms Jane`` keeps title ``Ms``, first ``Jane``. An undeclared delimiter becomes one more word of the part: ``Steven Hardman, RN - CRNA`` gives first ``Steven``, last ``Hardman``, suffix ``RN - CRNA``, where 1.4.0 through 2.3.0 gave first ``RN``, middle ``-``, last ``Steven Hardman``. The split spelling now counts like the joined one, so ``Smith, John Ph. D. Jones`` gives suffix ``Ph. D. Jones``, where 2.3.0 gave middle ``Jones``, suffix ``Ph. D.``. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #603)
Expand Down
19 changes: 15 additions & 4 deletions nameparser/_pipeline/_assemble.py
Original file line number Diff line number Diff line change
@@ -1,6 +1,8 @@
"""Not a stage: converts the final ParseState into a public ParsedName.

Consumes: tokens (all roles set), dropped, ambiguities (by index).
Consumes: tokens (all roles set), dropped, ambiguities (by index;
one carrying ``field_tail`` is worded from its word's final role,
#626).
Produces: a validated ParsedName -- the constructor re-checks every
invariant (span order/bounds, ambiguity subset), so a pipeline bug
that would produce an invalid result dies HERE, not in a renderer
Expand All @@ -18,7 +20,7 @@
"""
from __future__ import annotations

from nameparser._pipeline._state import ParseState
from nameparser._pipeline._state import FIELD_READING, ParseState
from nameparser._types import (
Ambiguity, AmbiguityKind, ParsedName, Role, Token,
)
Expand Down Expand Up @@ -112,8 +114,17 @@ def assemble(state: ParseState) -> ParsedName:
and any(t.role in (Role.SUFFIX, Role.TITLE)
for t in materialized)):
continue
ambiguities.append(
Ambiguity(pending.kind, pending.detail, materialized))
detail = pending.detail
if pending.field_tail is not None:
# A report naming a field, in a name emptied above (no
# emitter of one reaches that today: every referent is a
# name word): no field is left to name, so it goes rather
# than ship the sentence's first half.
if not materialized:
continue
detail = (f"{detail}{FIELD_READING[materialized[0].role]}"
f"{pending.field_tail}")
ambiguities.append(Ambiguity(pending.kind, detail, materialized))
return ParsedName(original=state.original,
tokens=tuple(final.values()),
ambiguities=tuple(ambiguities))
73 changes: 36 additions & 37 deletions nameparser/_pipeline/_assign.py
Original file line number Diff line number Diff line change
Expand Up @@ -363,9 +363,10 @@ def _assign_main(seg_idx: int, state: ParseState,
#
# Every bare ambiguous acronym the FINAL peel had to resolve is one
# coin-flip each, in either direction, so the report collects
# rather than overwrites. Deferred to after assignment because the
# wording reads the role back, and which role "not peeled" means
# depends on name_order. (The roman-numeral fork needs no such
# rather than overwrites. Deferred to after assignment because what
# a pick is declined as depends on the role it took, and which role
# "not peeled" means depends on name_order; the field itself is
# worded by assemble (#626). (The roman-numeral fork needs no such
# deferral and is reported here.)
#
# Read once in group, after P3's connective joins and ahead of the
Expand Down Expand Up @@ -435,8 +436,6 @@ def _assign_main(seg_idx: int, state: ParseState,
# call budget on every parse otherwise (decisions.md#parse-cost).
if peeled.names <= 1:
head = pieces[name_pieces[0]]
token = tokens[head[0]]
assert token.role is not None
# Both conventions here turn on a lone name word, so both
# report at the site that places one
# (mechanisms.md#AMBIGUITY-AT-THE-DECISION-SITE), and one
Expand All @@ -458,16 +457,16 @@ def _assign_main(seg_idx: int, state: ParseState,
# `suffix-or-name`" (history: decisions.md#H4) -- only the
# word made into a name reports, and no name piece survived
# the peel, so the carve-out above made the first post-nominal
# the name. The role comes off the token for the reason stated
# at the particle emitter below.
# the name. Its field is worded by assemble from the final role
# (`PendingAmbiguity.field_tail`, #626).
if peeled.names == 0 and field_undecided:
text = " ".join(tokens[i].text for i in head)
ambiguities.append(PendingAmbiguity(
AmbiguityKind.SUFFIX_OR_NAME,
f"{text!r} is post-nominal vocabulary with no name word "
f"beside it; read as a {token.role.value} name rather than "
f"a post-nominal, nothing else being left to be the name",
tuple(head)))
f"beside it; read as ", tuple(head),
field_tail=" rather than a post-nominal, nothing else "
"being left to be the name"))
# One name piece off the peel: the convention placed a lone
# name word. A suffix beside it is not a decision -- 'Smith
# Jr.' and "'Smitty' Jones Jr." ARE this convention -- and the
Expand All @@ -489,16 +488,17 @@ def _assign_main(seg_idx: int, state: ParseState,
# Wales'): the leading-title peel took the whole name.
if any("vocab:title" in tokens[i].tags for i in head):
text = " ".join(tokens[i].text for i in head)
lone = len(head) == 1
ambiguities.append(PendingAmbiguity(
AmbiguityKind.TITLE_OR_NAME,
f"{text!r} is title vocabulary and the only name word "
f"the title peel left standing; read as the name by "
f"convention rather than as more title"
if len(head) == 1 else
if lone else
f"{text!r} is the only name unit and joins title "
f"vocabulary to a name word; read as a "
f"{token.role.value} name by convention",
tuple(head)))
f"vocabulary to a name word; read as ",
tuple(head),
field_tail=None if lone else " by convention"))
# rules.md#O5: "a name of one name word that nothing else has
# decided reads that word as the given name under the default
# given-first order, and as the family name under a declared
Expand All @@ -517,16 +517,17 @@ def _assign_main(seg_idx: int, state: ParseState,
ambiguities.append(PendingAmbiguity(
AmbiguityKind.GIVEN_OR_FAMILY,
f"{text!r} is the only name word and nothing else "
f"decides it; read as a {token.role.value} name by "
f"convention, which follows the read order",
tuple(head)))
f"decides it; read as ", tuple(head),
field_tail=" by convention, which follows the read "
"order"))
for piece in peeled.picks:
# every pick is in rest, so the loops above just gave it a role
# every pick is in rest, so the loops above just gave it a
# role; whether it was read as a suffix or a name settles what
# it was declined as, and which field assemble names (#626)
token = tokens[piece[0]]
assert token.role is not None
taken, declined = (
("a suffix", "a name part") if token.role is Role.SUFFIX
else (f"a {token.role.value} name", "a post-nominal"))
declined = ("a name part" if token.role is Role.SUFFIX
else "a post-nominal")
# A word in the class by SHAPE (rules.md#S2, #S3) is in no
# wordlist and may be written with its periods, so the listed
# member's wording would misdescribe it on both counts (#563)
Expand All @@ -537,27 +538,23 @@ def _assign_main(seg_idx: int, state: ParseState,
"and an ordinary name")
ambiguities.append(PendingAmbiguity(
AmbiguityKind.SUFFIX_OR_NAME,
f"{token.text!r} {what}; read as {taken} rather than "
f"{declined}",
piece))
f"{token.text!r} {what}; read as ", piece,
field_tail=f" rather than {declined}"))
# leading ambiguous particle read as a name (#121 surfaced)
if name_pieces:
head = pieces[name_pieces[0]]
if (len(head) == 1 and len(name_pieces) > 1
and "vocab:particle-ambiguous" in tokens[head[0]].tags):
# the loops above gave the head piece its role from
# `order`, which is _effective_order's answer and not
# necessarily name_order's -- a script_orders entry
# overrides it. So read the role off the token rather than
# assume given, or re-derive it here; same reason as
# SUFFIX_OR_NAME just above.
token = tokens[head[0]]
assert token.role is not None
# the field is assemble's to word from the final role
# (`PendingAmbiguity.field_tail`, #626), never assumed given:
# `order` above is _effective_order's answer, not
# necessarily name_order's, and a later rule may still move
# the word ('Van Ivan Petrovich' under the East Slavic rule
# rotates 'Van' into the family)
ambiguities.append(PendingAmbiguity(
AmbiguityKind.PARTICLE_OR_GIVEN,
f"leading {token.text!r} may be a family-name "
f"particle; read as a {token.role.value} name",
tuple(head)))
f"leading {tokens[head[0]].text!r} may be a family-name "
f"particle; read as ", tuple(head), field_tail=""))
return order


Expand Down Expand Up @@ -852,11 +849,13 @@ def assign(state: ParseState) -> ParseState:
if state.pieces[1] and len(state.pieces[1][0]) == 1:
i = state.pieces[1][0][0]
if not tokens[i].tags.isdisjoint(_AMBIGUOUS_CREDENTIAL_TAGS):
# the field is assemble's to word: with nothing before
# the comma and a title after the word, H1 moves it to
# the family (', Ma Dr.', #626's review)
ambiguities.append(PendingAmbiguity(
AmbiguityKind.SUFFIX_OR_NAME,
f"{tokens[i].text!r} after the comma is also an "
f"ordinary name word; read as the given name",
(i,)))
f"ordinary name word; read as ", (i,), field_tail=""))
# Segment 1 is read before segment 0's wholly-family pass. It
# consumes piece tags and text only -- nothing segment 0's
# read writes. (A title standing after the comma behind a whole
Expand Down
23 changes: 23 additions & 0 deletions nameparser/_pipeline/_state.py
Original file line number Diff line number Diff line change
Expand Up @@ -130,12 +130,35 @@ class PendingAmbiguity:
extract_delimited knows only a character offset, so it records that
and tokenize resolves it to the containing token's index. Stages
after tokenize set ``indices`` directly and leave ``origin`` None.

``field_tail`` is for a report that names the field its word was
read into. post_rules may still move the word (H1's move behind a
title, P1's family-first fold, P6's attachment, and under opt-in
policies the patronymic rotations and `middle_as_family`'s fold),
so the emitter
writes ``detail`` up to the field and the rest of its sentence
here -- ``""`` where the field ends the sentence, which is why
assemble tests ``is not None`` -- and assemble joins them around
the FIELD_READING of the first referent's final role, the one
place that knows it (#626). Until then ``detail`` is a fragment.
"""

kind: AmbiguityKind
detail: str
indices: tuple[int, ...] = ()
origin: int | None = None
field_tail: str | None = None


#: How a report names the field a word was read into, by final role.
FIELD_READING: Mapping[Role, str] = MappingProxyType({
Role.TITLE: "a title", Role.GIVEN: "a given name",
Role.MIDDLE: "a middle name", Role.FAMILY: "a family name",
Role.SUFFIX: "a suffix", Role.NICKNAME: "a nickname",
Role.MAIDEN: "a maiden name"})
# every role a referent can hold, or a new one is a KeyError at the
# first parse whose report lands in it
assert FIELD_READING.keys() == set(Role)


@dataclass(frozen=True, slots=True)
Expand Down
Loading
Loading