1"""Shared piece-level predicates for pipeline stages.
2
3How a PIECE reads -- its tokens, plus the tags classify wrote on them
4and the tags group derived for the piece -- where _vocab answers how a
5WORD reads from text alone. Both are consulted by more than one
6stage; the split is by what the question takes, not by
7which stage happens to ask (mechanisms.md#ONE-PREDICATE-PER-QUESTION).
8_vocab points here from its own side: "Text-level tests used by more
9than one stage; piece-level ones live in _pieces, the sibling layer
10over tokens-plus-tags." own_words (#289/#516) answers over the whole
11token STREAM rather than one piece -- pieces do not exist yet at the
12stages that call it -- but the question is still piece-shaped, not
13word-shaped: it reads token ROLE, which _vocab's text-level tests
14never take.
15
16Before this module those predicates lived in _group, not because
17grouping owned them but because assign imported group and could not be
18imported back, so group was the only place both stages could reach.
19They arrived there that way across three PRs -- #424 brought
20is_leading_title, leading_titles and trailing_start, #425 the peel
21(peel_walk, peel_trailing), #429 the no-name-segment test that
22#430 turned into segment_suffix_reading.
23is_title_piece and is_suffix_piece are older than any of that: they
24were group's from its first commit, and travel because the others
25call them.
26
27The import that forced all of it is the one #439 removed: assign no
28longer names _group at all. What still holds is the rule that replaced
29it, and tests/v2/test_layering.py is where it is written down -- a
30piece predicate may not depend on a stage, in either direction.
31
32The S2 trailing peel travels as the unit decisions.md describes --
33peel_walk, peel_trailing and trailing_start together -- though only
34the first two cross a stage boundary. trailing_titles joins them
35because it reads what that peel left: the two answer one question
36between them, where the tail of a name stops being the name -- and
37tail_reading is that one question, running them against each other to
38their fixed point for the two stages that must not disagree about the
39answer.
40
41Layering: imports _state and _vocab only; FOUR stages import it --
42_segment, _classify, _group and _assign, segment being the one the
43#289/#516 own-words span added -- and neither of the two it imports
44imports it back.
45
46Naming follows _vocab's: inside an already-private module the leading
47underscore marks module-PRIVATE, so the names other stages call are
48bare and only the internals keep it (_PERIOD_ABBREV here). Getting that
49backwards -- which this module did until the underscores came off --
50costs a reader the one cheap way to tell a shared predicate from a
51helper.
52"""
53from __future__ import annotations
54
55from collections.abc import Mapping, Sequence, Set
56from typing import NamedTuple
57
58from nameparser._pipeline._state import (
59 AMBIGUOUS_ACRONYM_TAG, SHAPE_ACRONYM_TAG, WorkToken,
60)
61from nameparser._pipeline._vocab import (
62 _PERIOD_ABBREV, Lean, ambiguous_lean, in_initialless_script,
63 is_trailing_numeral_suffix, tag_marker_runs,
64)
65
66
67# rules.md#P3: "both questions this rule asks of a name — how many
68# words it has, and whether it is written in one case — are asked of
69# the name's OWN words: a maiden marker taken as one, and the words it
70# takes (M2), are not among them, and neither is a delimited clause
71# (N1, M1)" (history: decisions.md#P3)
72def own_words(tokens: Sequence[WorkToken], comma_offsets: Sequence[int],
73 markers: frozenset[str],
74 marker_tags: Mapping[int, str] | None = None,
75 ) -> tuple[list[str], int]:
76 """The name's OWN word texts and the index the maiden clause
77 starts at -- one span for the two stages that ask about it
78 (#289/#516).
79
80 Own words are the role-less tokens before the clause: a delimited
81 clause's tokens arrive from extract with a role already set, and
82 everything from a maiden marker on is the clause. Appending a
83 clause to a name must not change how a word in the name reads.
84
85 `marker_tags` is the map `_vocab.tag_marker_runs` already built,
86 index -> "vocab:maiden-marker"/"...-cont"; a caller that has it
87 (classify) hands it over and pays no second walk. A caller that
88 runs BEFORE those tags exist (segment) omits it, and this
89 function calls `tag_marker_runs` itself to build the SAME map
90 classify would -- not an approximation of it, which is what a
91 from-scratch text walk (this module's earlier `first_marker_head`)
92 could disagree with on a name where a marker's head opens an entry
93 but no run completes ('z' of 'z domu'): measured, 'ANNA z Nowak,
94 MD' flipped the recorded one-case verdict under that approximation
95 (decisions.md#P3). Sharing the exact function instead makes the
96 two paths agree by construction, not by corpus luck.
97
98 Takes tokens/comma_offsets/markers rather than a whole ParseState:
99 both call sites have all three already, and passing them lets this
100 function sit beside the piece predicates rather than in _vocab
101 (mechanisms.md#ONE-PREDICATE-PER-QUESTION; the
102 _post_rules.suffix_entries precedent, AGENTS.md's named exception)
103 -- it answers with the SPAN, where `tag_marker_runs` answers only
104 which tokens open a run.
105
106 `marker_tags`' keys must arrive in index order for the walk below
107 to find the SMALLEST head in one pass: `tag_marker_runs` walks its
108 tokens left to right, so the first head it records is already the
109 smallest, and a caller building its own map must preserve that
110 order too.
111
112 A plain tuple, not a NamedTuple: measured 2026-09-17, wrapping
113 this in a `NamedTuple` (this module's `Peel` is one) cost the
114 reference name one more frame (413.00 vs the 412.00 band this
115 commit must hold) -- a NamedTuple's `__new__` is itself a call,
116 where a bare tuple literal is not. `Peel` can afford the frame
117 because assign builds one only where the trailing peel actually
118 ran; `own_words` returns on every parse.
119 """
120 if marker_tags is None:
121 marker_tags = tag_marker_runs(tokens, comma_offsets, markers)
122 clause_at = len(tokens)
123 for i, tag in marker_tags.items():
124 if tag == "vocab:maiden-marker" and tokens[i].role is None:
125 clause_at = i
126 break
127 return ([t.text for t in tokens[:clause_at] if t.role is None],
128 clause_at)
129
130
131# rules.md#H3: "successive title words at the start of the part
132# carrying the given name chain into one title; a title word
133# elsewhere in the name does not"
134def is_title_piece(piece: Sequence[int], ptags: Set[str],
135 tokens: Sequence[WorkToken]) -> bool:
136 if "title" in ptags:
137 return True
138 return len(piece) == 1 and "vocab:title" in tokens[piece[0]].tags
139
140
141# _PERIOD_ABBREV: imported from _vocab, not redefined here (#289/#516,
142# quality-review finding) -- _vocab.name_word_count needed the SAME
143# shape test is_leading_title asks (_vocab.is_title_shaped), and
144# layering only allows the move in that direction (_pieces may import
145# _vocab; _vocab may not import _pieces). tests/v2/test_regex_sync.py
146# still reaches it as `_pieces._PERIOD_ABBREV` -- an import binds the
147# same name here, so the sync test's target did not move. Out of
148# assign since #424 and in the piece layer since #439: the test is
149# assign's, and group's leading-particle scan and trailing-run walk
150# must start where assign starts.
151
152
153# rules.md#H2: "an abbreviation opening the part of the name that
154# carries the given name — the whole name, or the part after a
155# family comma — reads as a title even when unlisted"
156# (history: decisions.md#H2)
157def is_leading_title(piece: Sequence[int], ptags: Set[str],
158 tokens: Sequence[WorkToken]) -> bool:
159 if is_title_piece(piece, ptags, tokens):
160 return True
161 if len(piece) != 1:
162 return False
163 text = tokens[piece[0]].text
164 # INLINED rather than calling _vocab.is_title_shaped, which asks
165 # the exact same question (#289/#516, quality-review finding: the
166 # two must not drift, and did once -- name_word_count's own
167 # vocabulary-only title test read 'Xyz.' as a name word where this
168 # predicate reads it as a title, and the disagreement flipped a
169 # comma structure `Dr. Smith, Ed`'s LISTED spelling did not).
170 # THIS is the one home for the number, `is_title_shaped` pointing
171 # here rather than restating it: routing this hot path through the
172 # shared function costs one frame per call (`is_leading_title`
173 # runs on every leading piece of every parse, unlike
174 # name_word_count's comma-only path), moving the reference name
175 # from 412/449 to 417/454 -- five calls on `Dr. Juan de la Vega
176 # III`, recomputable with `uv run python
177 # tools/perf/call_count.py`. Kept as two spellings of ONE test
178 # instead -- if you touch one, touch both, and
179 # `test_is_title_shaped_and_is_leading_title_agree` (this module's
180 # own test file) checks it over the union of both predicates'
181 # example tables rather than leaving it to a sentence.
182 return (bool(_PERIOD_ABBREV.match(text))
183 and (text.isascii() or not in_initialless_script(text)))
184
185
186def leading_titles(pieces: Sequence[Sequence[int]],
187 ptags: Sequence[Set[str]],
188 tokens: Sequence[WorkToken]) -> int:
189 """How many leading pieces assign peels as titles: the first
190 non-title index. A title needs a following piece, unless the whole
191 segment is one title (v1 parity). And the run gives back its last
192 piece when that piece is a name candidate: where everything behind
193 the run is suffix pieces, the run gives back its last piece, when
194 that piece is one word and is not itself suffix vocabulary
195 (rules.md#H3, decisions.md#H3 -- the block at the floor below
196 carries the examples of each half, and the ordering its two inline
197 tag reads were measured on).
198 One definition, read by assign (which sets the roles) and by the
199 chain's trailing-run walk; the leading-particle scan shares the
200 predicate, is_leading_title, but stops at a title-and-particle
201 word (P4, #367, #424)."""
202 n = 0
203 while n < len(pieces):
204 if ((n + 1 < len(pieces) or len(pieces) == 1)
205 and is_leading_title(pieces[n], ptags[n], tokens)):
206 n += 1
207 continue
208 break
209 # rules.md#H3: "where everything behind the run is post-nominal,
210 # the run gives its last word back to the name, provided that word
211 # stands alone and is not itself suffix vocabulary"
212 #
213 # ONE WORD, because a joined unit led by a title is a title run and
214 # handing it back would lose the title: 'Prince of Wales Jr' reads
215 # title 'Prince of Wales', family 'Jr', not given 'Prince of Wales'
216 # with no title at all.
217 #
218 # Two residuals. A run whose last word IS suffix vocabulary is not
219 # given back, so 'Dr King MD PhD' still reads title 'Dr King MD',
220 # family 'PhD'. And the floor asks is_suffix_piece, which vetoes a
221 # bare initial-shaped numeral, so 'Dr King V' keeps the whole run
222 # as the title and reads the numeral as the name -- given 'V',
223 # 'king' being a given-name title and the run's last word, where
224 # 'Dr Smith V' reads suffix 'V'. That numeral fork is outside this
225 # floor (decisions.md#H3).
226 #
227 # The two inline tag reads are the cheapest NECESSARY condition for
228 # the piece behind the run to be a suffix piece at all --
229 # is_suffix_piece cannot answer yes without one of them -- so the
230 # ordinary titled name, whose next piece is no kind of suffix,
231 # leaves this branch without entering a frame. Measured on the
232 # plan's ordering rather than the shape below: asking the
233 # authoritative predicate first cost 8 calls per parse of the
234 # reference name (leading_titles runs four times), against a band
235 # with room for two (decisions.md#parse-cost). is_suffix_piece
236 # stays the predicate that ANSWERS, here and in the walk.
237 if (n and n < len(pieces)
238 and ("suffix" in ptags[n]
239 or "vocab:suffix" in tokens[pieces[n][0]].tags)
240 and len(pieces[n - 1]) == 1
241 and not is_suffix_piece(pieces[n - 1], ptags[n - 1],
242 tokens)):
243 for k in range(n, len(pieces)):
244 if not is_suffix_piece(pieces[k], ptags[k], tokens):
245 break
246 else:
247 # nothing behind the run but suffix pieces
248 n -= 1
249 return n
250
251
252def is_suffix_piece(piece: Sequence[int], ptags: Set[str],
253 tokens: Sequence[WorkToken]) -> bool:
254 if "suffix" in ptags:
255 return True
256 if len(piece) != 1:
257 return False
258 tags = tokens[piece[0]].tags
259 return "vocab:suffix" in tags and "initial" not in tags
260
261
262def _numeral_behind_the_initial_veto(piece: Sequence[int],
263 tokens: Sequence[WorkToken]) -> bool:
264 """Suffix vocabulary that is_suffix_piece refuses because it is
265 also initial-shaped: a ONE-CHARACTER entry, bare or with a period.
266
267 Named for the shape rather than enumerated, because the shape is
268 what the code tests and the enumeration goes stale -- in the
269 shipped lexicon it reaches i, v and 2, and NOT x or ix (roman, but
270 not suffix vocabulary) nor ii/iii/iv (suffix vocabulary, but two
271 characters, so never initial-shaped and never vetoed in the first
272 place). A caller adding a one-character suffix in a script that
273 has initials extends it.
274
275 The veto is right where such a word could be a middle initial, and
276 wrong where it is describing the suffix in front of it, which is
277 the only place this is asked from. The len(piece) != 1 guard is
278 defensive: a merged multi-token piece carries "suffix" in ptags, so
279 is_suffix_piece claims it one branch earlier and no reachable input
280 arrives here with one.
281 """
282 if len(piece) != 1:
283 return False
284 tags = tokens[piece[0]].tags
285 return "vocab:suffix" in tags and "initial" in tags
286
287
288def segment_suffix_reading(pieces: Sequence[Sequence[int]],
289 ptags: Sequence[Set[str]],
290 tokens: Sequence[WorkToken],
291 lenient: bool,
292 one_case: bool | None,
293 ) -> tuple[bool, ...] | None:
294 """How each piece of a no-name segment reads: True a suffix, False
295 a title. None when the segment holds a name word and so is not a
296 credential run at all.
297
298 `one_case` admits #289's credential lean: an ALL-CAPS member of
299 the ambiguous set inside a mixed-case name is a credential in this
300 slot even with one word before the comma, because the writing is
301 evidence the count does not have ('Smith, MA' -> family 'Smith',
302 suffix 'MA'). Only the LEAN reaches here: a token admitted to the
303 class by SHAPE takes the count instead, which is decided at the
304 comma and not in this walk ('Smith, A.B.' -> given 'A.B.').
305
306 ONE answer for two readers, both in _assign.py -- the no-name gate
307 and the router -- because they must agree piece for piece. #429
308 shipped the inverse of its own fix by deriving that agreement twice
309 (mechanisms.md#ONE-PREDICATE-PER-QUESTION). It answered for a third
310 until #436: group's one-entry join asked it too, and the render's
311 entry boundary is a rule over the written commas in post_rules now
312 (rules.md#R1), which asks this nothing.
313
314 rules.md#S2's initial veto keeps a roman numeral out of a suffix
315 reading, which is right after a NAME word: 'Smith, John V.' is a
316 middle initial (#432). After a SUFFIX word the numeral is
317 describing that suffix -- 'PSM I' is Professional Scrum Master
318 level I -- so the run continues through it, period included, an
319 initial there being no shape anyone writes (#430). A title resets
320 that: what follows a bare title is not continuing a credential.
321
322 None covers both ways a segment can fail to be a run: a name word
323 anywhere in it, and no pieces at all ('Doe,, Jr.', which holds no
324 title to read by).
325
326 `lenient` is Policy.lenient_comma_suffixes, and only the numeral
327 continuation consults it. C1: "by default a recognized suffix word
328 counts even written like an initial, while strict mode vetoes
329 initial-shaped words" -- so under strict the veto stands and the
330 run ends where it always did. Reading no policy here silently
331 overrode the one knob a caller sets to prevent exactly this.
332
333 The FAMILY_COMMA rule "segment 0 is wholly the family name" rests
334 on the writer having said where the family name ends. A comma
335 followed by no name word said no such thing -- 'John Smith, Dr.' is
336 'Dr. John Smith' with the honorific moved -- so the pre-comma name
337 keeps its positional read instead of being merged. Uses the same
338 is_leading_title predicate the peel does, period-abbreviation
339 inference included, so the two cannot disagree about what a title
340 is; a mixed run like 'Smith, Dr. Jr.' is a title and a postnominal,
341 each read where it stands, never a title run 'Dr. Jr.'.
342 """
343 if not pieces:
344 return None
345 out: list[bool] = []
346 for piece, tags in zip(pieces, ptags):
347 # the verdict just recorded IS "stands behind a suffix" -- keeping
348 # a separate flag meant maintaining that equality by hand at three
349 # sites, and a fourth branch that appended without assigning would
350 # have diverged silently
351 after_suffix = bool(out) and out[-1]
352 if is_suffix_piece(piece, tags, tokens):
353 out.append(True)
354 elif (len(piece) == 1
355 and AMBIGUOUS_ACRONYM_TAG in tokens[piece[0]].tags
356 and listed_lean(tokens[piece[0]], one_case)
357 == "credential"):
358 out.append(True)
359 elif (lenient and after_suffix
360 and _numeral_behind_the_initial_veto(piece, tokens)):
361 out.append(True)
362 elif is_leading_title(piece, tags, tokens):
363 out.append(False)
364 else:
365 return None
366 return tuple(out)
367
368
369class Peel(NamedTuple):
370 """What assign's trailing peel made of a walk. `names` is a count
371 of positions in the caller's `rest`: rest[:names] are the name
372 pieces and rest[names:] the suffixes. The other two are pieces --
373 token-index tuples, as PendingAmbiguity wants them -- and each is
374 one token long: `numeral` is the piece the roman-numeral fork
375 took (None when it did not fire; always the walk's last piece),
376 `picks` the bare ambiguous acronyms the peel had to resolve, in
377 peel order, either way (the last may sit at rest[names - 1])."""
378
379 names: int
380 numeral: tuple[int, ...] | None
381 picks: tuple[tuple[int, ...], ...]
382
383
384# rules.md#S2: "a trailing word of the suffix vocabulary reads as a
385# suffix — generational forms and credential acronyms alike, and an
386# ambiguous acronym written with its periods, one after each letter,
387# counts unambiguously; a single trailing period is the abbreviation
388# shape any word can wear and does not. A BARE ambiguous acronym is
389# consumed only when the name has words to spare"
390# (v1's are_suffixes tail rule, with the roman-numeral special)
391def peel_walk(start: int, ptags: Sequence[Set[str]],
392 skip: Set[int] = frozenset()) -> list[int]:
393 """The indices peel_trailing walks: `start` to the segment's end,
394 minus the group-flagged credential pieces (the Ph. D. merge),
395 which assign reads as suffixes at any position, and minus `skip`
396 -- a tail segment's delimiter cores, which are structure rather
397 than words (the maiden walk's case, #424). Built here and nowhere
398 else, so the walk's input cannot drift between assign and the
399 group sites that read it: the numeral fork is a last-piece
400 test that reads the piece before as rest[k - 2], which holds only
401 over this list."""
402 return [j for j in range(start, len(ptags))
403 if j not in skip and "suffix" not in ptags[j]]
404
405
406def trailing_start(start: int, pieces: Sequence[Sequence[int]],
407 ptags: Sequence[Set[str]], tokens: Sequence[WorkToken],
408 skip: Set[int] = frozenset(),
409 *, one_case: bool | None) -> int:
410 """Where assign's trailing suffix run begins, read over the pieces
411 as they stand from `start`: the index of the first piece the S2
412 peel takes, or len(pieces) when it takes none (#424). What P2's
413 chain and M2's walk stop before -- each had asked "is this a
414 suffix?" with the suffix-piece test, which vetoes a bare 'V' as
415 an initial (the #401 question), and so took a trailing numeral,
416 or a bare acronym with words to spare, into the family or the
417 maiden name.
418
419 Both forks, always. A caller that needs one of them alone -- the
420 maiden walk re-asking the numeral over the view its take would
421 leave, where the acronym fork's piece COUNT no longer describes
422 the name -- calls the `peel_walk` + `peel_trailing` pair this
423 wraps and reads the half it wants (#533). A `numeral_only` flag
424 lived here for that one caller and cost it a frame."""
425 rest = peel_walk(start, ptags, skip)
426 peeled = peel_trailing(rest, pieces, ptags, tokens, one_case)
427 return rest[peeled.names] if peeled.names < len(rest) else len(pieces)
428
429
430# #289/#516: the "listed member, not by-shape" test both
431# peel_trailing and segment_suffix_reading ask before reading the
432# lean -- shared here so the two cannot drift on what counts
433# (quality-review finding: it was spelled twice, once per site,
434# before this). Those two test "vocab:suffix-ambiguous" in tags
435# INLINE, before calling this, rather than leaving that cheap check to
436# this function's own body: measured, a caller whose `elif` reaches
437# this on every piece (segment_suffix_reading's does, one per
438# family-comma segment 1, member or not) pays one frame for the call
439# regardless of what is inside it, and the inline pre-check is what
440# keeps a non-member piece ("Smith, John"'s "John") from ever making
441# the call at all.
442#
443# A THIRD caller since #533 -- credential_at_the_given_slot just
444# below -- deliberately does NOT pre-check: it owns #531's reading
445# and leaves membership to its own callers (its docstring says so),
446# and the frame argument holds transitively because both of them ask
447# inline -- `AMBIGUOUS_ACRONYM_TAG in tok.tags` after a
448# `len(piece) == 1` at assign's given-part trailing slot, and the
449# same pair inside the `all(...)` of `_group.py`'s `_maiden_take`
450# view check. So no non-member piece reaches this function down that
451# route either.
452def listed_lean(token: WorkToken, one_case: bool | None) -> Lean | None:
453 """`ambiguous_lean` for a LISTED bare-ambiguous token, or None if
454 the token is not tagged a listed member, is admitted by SHAPE
455 instead (`SHAPE_ACRONYM_TAG`, a switch's doing, not the writing's),
456 or there is no case fact to ask at all."""
457 if (one_case is None or AMBIGUOUS_ACRONYM_TAG not in token.tags
458 or SHAPE_ACRONYM_TAG in token.tags):
459 return None
460 return ambiguous_lean(token.text, one_case)
461
462
463def credential_at_the_given_slot(token: WorkToken,
464 one_case: bool | None) -> bool:
465 """#531's reading of a class MEMBER ending the given part after a
466 family comma: the credential unless the writing says otherwise.
467 The caller decides membership and that the piece ends that part.
468
469 The words to spare are there by construction at that slot, so the
470 count says nothing and only the lean does; a member that is also
471 particle vocabulary reads as the credential on a POSITIVE lean
472 alone, P6's attachment keeping every other spelling.
473
474 Two callers since #533 -- assign's walk over the given part, and
475 the maiden walk's second check over the name the take would leave
476 (rules.md#M2) -- so the reading is a function rather than a
477 condition written twice
478 (mechanisms.md#ONE-PREDICATE-PER-QUESTION). It is a
479 text-and-tags question, which is what puts it in this module
480 rather than beside either caller.
481
482 The membership half of that contract is CHECKED rather than
483 trusted, because getting it wrong is silent: all three of
484 `listed_lean`'s None reasons fall through to the `particle` test
485 below, so a non-member handed in by mistake is answered True --
486 "read it as the credential" -- for a word the class never admitted.
487 An assert rather than a raise or a branch: it enters no Python
488 frame (measured -- 'Doe, John MA' stays at 311), it states the
489 contract where a reader of the function body meets it, and under
490 -O it is exactly the code that was here before.
491 """
492 assert AMBIGUOUS_ACRONYM_TAG in token.tags, (
493 f"credential_at_the_given_slot is #531's reading of a LISTED "
494 f"class member; {token.text!r} carries {sorted(token.tags)} "
495 f"and is not one. The caller decides membership -- test "
496 f"AMBIGUOUS_ACRONYM_TAG before calling")
497 lean = listed_lean(token, one_case)
498 return lean == "credential" or (lean is None
499 and "particle" not in token.tags)
500
501
502def peel_trailing(rest: Sequence[int], pieces: Sequence[Sequence[int]],
503 ptags: Sequence[Set[str]],
504 tokens: Sequence[WorkToken],
505 one_case: bool | None) -> Peel:
506 """The S2 trailing peel over `rest`, a peel_walk list. In the
507 piece layer rather than in assign because group's bound-given
508 reserve asks the same question of the view the join would leave
509 (#425): one walk, so the reserve and the assignment cannot drift. Pure -- the ambiguities are
510 returned for assign to report, in the order it always reported
511 them.
512
513 `one_case` is ParseState.one_case, the recorded fact: None means
514 nobody asked, which is every caller that has no state to ask with,
515 and reads as "no lean" -- rules.md#S2's count alone, the behavior
516 of every release before this one.
517 """
518 picks: list[tuple[int, ...]] = []
519 numeral: tuple[int, ...] | None = None
520 k = len(rest)
521 while k > 0:
522 piece = pieces[rest[k - 1]]
523 if is_suffix_piece(piece, ptags[rest[k - 1]], tokens):
524 k -= 1
525 continue
526 # a final single letter that is a roman numeral, after a piece
527 # that is not initial-shaped; the predicate's docstring carries
528 # the is_initial_shaped reasoning (#320)
529 if (k == len(rest) and k >= 2 and len(piece) == 1
530 and is_trailing_numeral_suffix(
531 tokens[piece[0]].text,
532 tokens[pieces[rest[k - 2]][0]].text)):
533 numeral = tuple(piece)
534 k -= 1
535 continue
536 # A bare ambiguous acronym ("MA", not "M.A.") is a credential
537 # only when peeling it still leaves a given AND a family name.
538 # With two pieces, "one of them is a credential" is the less
539 # likely reading, so it stays the family name -- "Jack MA" is a
540 # person, "John Smith MA" is a person with a degree. This is
541 # v1's reserve_last narrowed to the ambiguous set: 2.0
542 # deliberately peels UNambiguous suffixes even when nothing is
543 # left ("Smith PhD" -> suffix, a classified fix), because there
544 # the vocabulary is not in doubt.
545 bare_ambiguous = (len(piece) == 1
546 and AMBIGUOUS_ACRONYM_TAG in tokens[piece[0]].tags)
547 # #516, switch off: the writing still makes this token
548 # credential-SHAPED, and the parser is choosing the name
549 # reading over that one -- the fork the caller asked to be
550 # told about. Reported here, unconsumed, rather than folded
551 # into `bare_ambiguous` above: with the switch off the token
552 # never carries "vocab:suffix-ambiguous" (classify's own
553 # gate), so `bare_ambiguous` is already False and this is the
554 # ONLY place the report can be recorded. Switch ON, the token
555 # carries BOTH tags, so `not bare_ambiguous` is what stands
556 # this branch down and lets the consuming branch below take
557 # it; the two are not exclusive.
558 if (not bare_ambiguous and k >= 2 and len(piece) == 1
559 and SHAPE_ACRONYM_TAG in tokens[piece[0]].tags):
560 picks.append(tuple(piece))
561 break
562 # #289: written case is evidence the count does not have, and
563 # it overrides the count in BOTH directions -- an all-caps
564 # member of a mixed-case name is taken with nothing to spare
565 # ("Jack MA"), a Title-case one is declined with plenty
566 # ("John Smith Ma"). The lean is the LISTED set's alone: a
567 # token admitted to this class by SHAPE carries no writing
568 # convention to read, so it takes the count (decisions.md#S2).
569 #
570 # k < 2 means it is the only piece left, which is not the fork
571 # this reports -- and not a floor the lean moves: the walk
572 # starts after the leading title run, so one piece behind a
573 # title ("Mr MA") is exactly this case and must stay a name.
574 # The lean is computed only past this floor -- membership,
575 # then the floor, then the count-or-lean, in that order, with
576 # nothing computed a step earlier could discard.
577 if bare_ambiguous and k >= 2:
578 picks.append(tuple(piece))
579 lean = listed_lean(tokens[piece[0]], one_case)
580 # peeling still leaves given + family, or the writing says
581 # to peel anyway
582 if lean == "credential" or (lean is None and k >= 3):
583 k -= 1
584 continue
585 break
586 return Peel(k, numeral, tuple(picks))
587
588
589# rules.md#H5: "only a word the vocabulary knows as a title is one,
590# and a bare title word is a name word"
591# -- the trailing run's own predicate. NOT is_leading_title:
592# that predicate carries H2's unlisted-abbreviation inference, which is
593# the LEADING slot's shape rule and has no trailing counterpart, so
594# with it 'John Smith Xyz.' would lose its family name to a title
595# (decisions.md#H5). The vocabulary read is is_title_piece's, shared
596# with the leading run so the two cannot disagree about what a title
597# WORD is while disagreeing, deliberately, about what a title SHAPE is.
598def trailing_titles(rest: Sequence[int], pieces: Sequence[Sequence[int]],
599 ptags: Sequence[Set[str]],
600 tokens: Sequence[WorkToken]) -> int:
601 """How many pieces of `rest` the trailing title chain LEAVES
602 standing: `rest[:kept]` are the name pieces and `rest[kept:]` the
603 period-marked title words the chain took, in piece order. Counted
604 the way `peel_trailing` counts, so the two answers compose without
605 arithmetic at the call site. `rest` is the caller's NAME pieces:
606 on the no-comma path what the S2 peel left, after a family comma
607 the segment's pieces that the segment's own suffix reading does
608 not claim, and in `tail_reading` the leftovers of whichever peel
609 is current.
610 Floor: one name piece stands, so a name is never all title -- and
611 an empty `rest` returns 0, which is what leaves assign's
612 bare-suffix carve-out reached exactly as before.
613
614 ONE WORD per piece, the same gate the leading peel's give-back
615 uses: a joined unit is not the shape this reads, and the tokens of
616 one are not each a title word.
617
618 Every parse with a name word to place enters this frame -- assign
619 asks the question here rather than answering a cheaper version of
620 it inline (mechanisms.md#ONE-PREDICATE-PER-QUESTION) -- so what it
621 costs an ordinary name is one frame and one regex match. The
622 exceptions return before it: a segment that is all title, and a
623 comma part read wholly as a credential run, have no name piece to
624 hand this (52 of the 1289 corpus parses, measured 2026-09-09 --
625 'Coach', 'Lord of the Universe', 'Smith, Jr.', 'MD, PHD').
626 The shape test
627 runs BEFORE the vocabulary one to keep it at that: _PERIOD_ABBREV
628 is a compiled regex (a C call, no Python frame) where
629 is_title_piece is a call, and almost no name ends in a
630 period-marked word, so the ordinary parse pays the one match and
631 stops (decisions.md#parse-cost).
632 """
633 k = len(rest)
634 while k > 1:
635 idx = rest[k - 1]
636 piece = pieces[idx]
637 # no #323 veto on the shape here, unlike is_leading_title's:
638 # the shape is ANDed with is_title_piece, so the word is listed
639 # vocabulary, and a listed CJK title wearing a stop should read
640 # as a title.
641 if (len(piece) == 1
642 and _PERIOD_ABBREV.match(tokens[piece[0]].text)
643 and is_title_piece(piece, ptags[idx], tokens)):
644 k -= 1
645 continue
646 break
647 return k
648
649
650# rules.md#H5: "the title is TRANSPARENT to the suffix reading: where
651# two or more name words stand, what stands once the chain is taken
652# reads exactly as it would read written without the title, plus the
653# title"
654def tail_reading(rest: list[int], pieces: Sequence[Sequence[int]],
655 ptags: Sequence[Set[str]],
656 tokens: Sequence[WorkToken],
657 one_case: bool | None,
658 ) -> tuple[list[int], tuple[int, ...], Peel]:
659 """The S2 peel and the H5 chain read together to a FIXED POINT:
660 peel, chain, splice the chained pieces out, peel again over what
661 is left -- the name pieces the chain kept, then the pieces the
662 peel had taken, in original order -- until the chain takes
663 nothing. Where it takes nothing on the first pass, which is almost
664 every name, that first peel is the answer and the loop costs one
665 comparison.
666
667 Returns the walk with the chained pieces spliced out, the pieces
668 the chain took, and the FINAL peel -- whose numeral fork and
669 ambiguous picks are the ones assign reports. The walk is
670 partitioned by that peel exactly as a caller partitions its own:
671 `rest[:peel.names]` the name pieces, `rest[peel.names:]` the
672 suffixes. A bare tuple rather than a named one because a
673 NamedTuple's __new__ is a frame of its own on every parse
674 (decisions.md#parse-cost), and `_group_segment` returns its three
675 the same way.
676
677 Transparency is what the fixed point buys: 'X Prof. Y' reads
678 exactly as 'X Y' reads plus the title, however many titles are
679 written and wherever the peel then stops. Iterating ONCE reads a
680 second title only half way -- 'John Prof. MA Prof.' un-peeled the
681 acronym and re-exposed the first title, reading family 'Prof.'
682 with suffix 'MA' where 'John Prof. MA' reads family 'MA'.
683
684 One function for two readers -- assign's placement and group's
685 bound-given reserve (P5), which must count the name words assign
686 will leave. Deriving that agreement twice is what left the two
687 disagreeing at S2's bare-ambiguous reserve: 'abdul rahman MA'
688 declined the join and 'abdul rahman MA Prof.' took it
689 (mechanisms.md#ONE-PREDICATE-PER-QUESTION).
690
691 `rest` is a peel_walk list and is not mutated -- the splice
692 rebinds this local -- so a caller's own reference still names the
693 walk it built. Both callers read the one returned here instead,
694 which is the one the final peel partitions.
695 """
696 titled: list[int] = []
697 while True:
698 peeled = peel_trailing(rest, pieces, ptags, tokens, one_case)
699 kept = trailing_titles(rest[:peeled.names], pieces, ptags,
700 tokens)
701 if kept == peeled.names:
702 return rest, tuple(titled), peeled
703 # the chain's pieces reach this list back to front, so each
704 # run goes in FRONT of what the pass before it took
705 titled[:0] = rest[kept:peeled.names]
706 rest = rest[:kept] + rest[peeled.names:]
707
708
709# rules.md#H5: "the title is TRANSPARENT to the suffix reading: where
710# two or more name words stand, what stands once the chain is taken
711# reads exactly as it would read written without the title, plus the
712# title"
713def trailing_start_past_titles(start: int,
714 pieces: Sequence[Sequence[int]],
715 ptags: Sequence[Set[str]],
716 tokens: Sequence[WorkToken],
717 *, one_case: bool | None) -> int:
718 """`trailing_start` read through H5's chain: where assign's
719 trailing suffix run begins once a trailing TITLE has stopped
720 hiding it.
721
722 `trailing_start` reads the pieces as WRITTEN, so a title standing
723 behind the suffix run makes the peel take nothing and the answer
724 is `len(pieces)` -- the reading assign itself has not had since
725 H5, because assign runs the peel and the chain to their fixed
726 point instead (`tail_reading`). A caller using that answer as the
727 right bound of the NAME is told a credential is a name word:
728 'John Quincy Adams i MA Prof.' read family 'Adams i MA' where
729 'John Quincy Adams i MA' reads family 'Adams' and suffix 'i MA'
730 (#397 second review). Every caller that bounds the name wants
731 this one; `trailing_start` stays for the callers that count a
732 trailing run of the pieces as they stand.
733
734 The returned index bounds the name from the right, and the
735 trailing titles the chain spliced out are not under it: they end
736 the segment, so they stand at or past the first suffix piece
737 whenever there is one. Where the peel takes nothing even past the
738 chain this returns `len(pieces)` as `trailing_start` does, and a
739 trailing title is then inside the bound and refused by the title
740 test the callers already run beside it.
741 """
742 rest, _titled, peeled = tail_reading(peel_walk(start, ptags),
743 pieces, ptags, tokens, one_case)
744 return rest[peeled.names] if peeled.names < len(rest) else len(pieces)