Coverage for /pythoncovmergedfiles/medio/medio/usr/local/lib/python3.11/site-packages/nameparser/_pipeline/_pieces.py: 98%

Shortcuts on this page

r m x   toggle line displays

j k   next/prev highlighted chunk

0   (zero) top of page

1   (one) first highlighted chunk

132 statements  

1"""Shared piece-level predicates for pipeline stages. 

2 

3How a PIECE reads -- its tokens, plus the tags classify wrote on them 

4and the tags group derived for the piece -- where _vocab answers how a 

5WORD reads from text alone. Both are consulted by more than one 

6stage; the split is by what the question takes, not by 

7which stage happens to ask (mechanisms.md#ONE-PREDICATE-PER-QUESTION). 

8_vocab points here from its own side: "Text-level tests used by more 

9than one stage; piece-level ones live in _pieces, the sibling layer 

10over tokens-plus-tags." own_words (#289/#516) answers over the whole 

11token STREAM rather than one piece -- pieces do not exist yet at the 

12stages that call it -- but the question is still piece-shaped, not 

13word-shaped: it reads token ROLE, which _vocab's text-level tests 

14never take. 

15 

16Before this module those predicates lived in _group, not because 

17grouping owned them but because assign imported group and could not be 

18imported back, so group was the only place both stages could reach. 

19They arrived there that way across three PRs -- #424 brought 

20is_leading_title, leading_titles and trailing_start, #425 the peel 

21(peel_walk, peel_trailing), #429 the no-name-segment test that 

22#430 turned into segment_suffix_reading. 

23is_title_piece and is_suffix_piece are older than any of that: they 

24were group's from its first commit, and travel because the others 

25call them. 

26 

27The import that forced all of it is the one #439 removed: assign no 

28longer names _group at all. What still holds is the rule that replaced 

29it, and tests/v2/test_layering.py is where it is written down -- a 

30piece predicate may not depend on a stage, in either direction. 

31 

32The S2 trailing peel travels as the unit decisions.md describes -- 

33peel_walk, peel_trailing and trailing_start together -- though only 

34the first two cross a stage boundary. trailing_titles joins them 

35because it reads what that peel left: the two answer one question 

36between them, where the tail of a name stops being the name -- and 

37tail_reading is that one question, running them against each other to 

38their fixed point for the two stages that must not disagree about the 

39answer. 

40 

41Layering: imports _state and _vocab only; FOUR stages import it -- 

42_segment, _classify, _group and _assign, segment being the one the 

43#289/#516 own-words span added -- and neither of the two it imports 

44imports it back. 

45 

46Naming follows _vocab's: inside an already-private module the leading 

47underscore marks module-PRIVATE, so the names other stages call are 

48bare and only the internals keep it (_PERIOD_ABBREV here). Getting that 

49backwards -- which this module did until the underscores came off -- 

50costs a reader the one cheap way to tell a shared predicate from a 

51helper. 

52""" 

53from __future__ import annotations 

54 

55from collections.abc import Mapping, Sequence, Set 

56from typing import NamedTuple 

57 

58from nameparser._pipeline._state import ( 

59 AMBIGUOUS_ACRONYM_TAG, SHAPE_ACRONYM_TAG, WorkToken, 

60) 

61from nameparser._pipeline._vocab import ( 

62 _PERIOD_ABBREV, Lean, ambiguous_lean, in_initialless_script, 

63 is_trailing_numeral_suffix, tag_marker_runs, 

64) 

65 

66 

67# rules.md#P3: "both questions this rule asks of a name — how many 

68# words it has, and whether it is written in one case — are asked of 

69# the name's OWN words: a maiden marker taken as one, and the words it 

70# takes (M2), are not among them, and neither is a delimited clause 

71# (N1, M1)" (history: decisions.md#P3) 

72def own_words(tokens: Sequence[WorkToken], comma_offsets: Sequence[int], 

73 markers: frozenset[str], 

74 marker_tags: Mapping[int, str] | None = None, 

75 ) -> tuple[list[str], int]: 

76 """The name's OWN word texts and the index the maiden clause 

77 starts at -- one span for the two stages that ask about it 

78 (#289/#516). 

79 

80 Own words are the role-less tokens before the clause: a delimited 

81 clause's tokens arrive from extract with a role already set, and 

82 everything from a maiden marker on is the clause. Appending a 

83 clause to a name must not change how a word in the name reads. 

84 

85 `marker_tags` is the map `_vocab.tag_marker_runs` already built, 

86 index -> "vocab:maiden-marker"/"...-cont"; a caller that has it 

87 (classify) hands it over and pays no second walk. A caller that 

88 runs BEFORE those tags exist (segment) omits it, and this 

89 function calls `tag_marker_runs` itself to build the SAME map 

90 classify would -- not an approximation of it, which is what a 

91 from-scratch text walk (this module's earlier `first_marker_head`) 

92 could disagree with on a name where a marker's head opens an entry 

93 but no run completes ('z' of 'z domu'): measured, 'ANNA z Nowak, 

94 MD' flipped the recorded one-case verdict under that approximation 

95 (decisions.md#P3). Sharing the exact function instead makes the 

96 two paths agree by construction, not by corpus luck. 

97 

98 Takes tokens/comma_offsets/markers rather than a whole ParseState: 

99 both call sites have all three already, and passing them lets this 

100 function sit beside the piece predicates rather than in _vocab 

101 (mechanisms.md#ONE-PREDICATE-PER-QUESTION; the 

102 _post_rules.suffix_entries precedent, AGENTS.md's named exception) 

103 -- it answers with the SPAN, where `tag_marker_runs` answers only 

104 which tokens open a run. 

105 

106 `marker_tags`' keys must arrive in index order for the walk below 

107 to find the SMALLEST head in one pass: `tag_marker_runs` walks its 

108 tokens left to right, so the first head it records is already the 

109 smallest, and a caller building its own map must preserve that 

110 order too. 

111 

112 A plain tuple, not a NamedTuple: measured 2026-09-17, wrapping 

113 this in a `NamedTuple` (this module's `Peel` is one) cost the 

114 reference name one more frame (413.00 vs the 412.00 band this 

115 commit must hold) -- a NamedTuple's `__new__` is itself a call, 

116 where a bare tuple literal is not. `Peel` can afford the frame 

117 because assign builds one only where the trailing peel actually 

118 ran; `own_words` returns on every parse. 

119 """ 

120 if marker_tags is None: 

121 marker_tags = tag_marker_runs(tokens, comma_offsets, markers) 

122 clause_at = len(tokens) 

123 for i, tag in marker_tags.items(): 

124 if tag == "vocab:maiden-marker" and tokens[i].role is None: 

125 clause_at = i 

126 break 

127 return ([t.text for t in tokens[:clause_at] if t.role is None], 

128 clause_at) 

129 

130 

131# rules.md#H3: "successive title words at the start of the part 

132# carrying the given name chain into one title; a title word 

133# elsewhere in the name does not" 

134def is_title_piece(piece: Sequence[int], ptags: Set[str], 

135 tokens: Sequence[WorkToken]) -> bool: 

136 if "title" in ptags: 

137 return True 

138 return len(piece) == 1 and "vocab:title" in tokens[piece[0]].tags 

139 

140 

141# _PERIOD_ABBREV: imported from _vocab, not redefined here (#289/#516, 

142# quality-review finding) -- _vocab.name_word_count needed the SAME 

143# shape test is_leading_title asks (_vocab.is_title_shaped), and 

144# layering only allows the move in that direction (_pieces may import 

145# _vocab; _vocab may not import _pieces). tests/v2/test_regex_sync.py 

146# still reaches it as `_pieces._PERIOD_ABBREV` -- an import binds the 

147# same name here, so the sync test's target did not move. Out of 

148# assign since #424 and in the piece layer since #439: the test is 

149# assign's, and group's leading-particle scan and trailing-run walk 

150# must start where assign starts. 

151 

152 

153# rules.md#H2: "an abbreviation opening the part of the name that 

154# carries the given name — the whole name, or the part after a 

155# family comma — reads as a title even when unlisted" 

156# (history: decisions.md#H2) 

157def is_leading_title(piece: Sequence[int], ptags: Set[str], 

158 tokens: Sequence[WorkToken]) -> bool: 

159 if is_title_piece(piece, ptags, tokens): 

160 return True 

161 if len(piece) != 1: 

162 return False 

163 text = tokens[piece[0]].text 

164 # INLINED rather than calling _vocab.is_title_shaped, which asks 

165 # the exact same question (#289/#516, quality-review finding: the 

166 # two must not drift, and did once -- name_word_count's own 

167 # vocabulary-only title test read 'Xyz.' as a name word where this 

168 # predicate reads it as a title, and the disagreement flipped a 

169 # comma structure `Dr. Smith, Ed`'s LISTED spelling did not). 

170 # THIS is the one home for the number, `is_title_shaped` pointing 

171 # here rather than restating it: routing this hot path through the 

172 # shared function costs one frame per call (`is_leading_title` 

173 # runs on every leading piece of every parse, unlike 

174 # name_word_count's comma-only path), moving the reference name 

175 # from 412/449 to 417/454 -- five calls on `Dr. Juan de la Vega 

176 # III`, recomputable with `uv run python 

177 # tools/perf/call_count.py`. Kept as two spellings of ONE test 

178 # instead -- if you touch one, touch both, and 

179 # `test_is_title_shaped_and_is_leading_title_agree` (this module's 

180 # own test file) checks it over the union of both predicates' 

181 # example tables rather than leaving it to a sentence. 

182 return (bool(_PERIOD_ABBREV.match(text)) 

183 and (text.isascii() or not in_initialless_script(text))) 

184 

185 

186def leading_titles(pieces: Sequence[Sequence[int]], 

187 ptags: Sequence[Set[str]], 

188 tokens: Sequence[WorkToken]) -> int: 

189 """How many leading pieces assign peels as titles: the first 

190 non-title index. A title needs a following piece, unless the whole 

191 segment is one title (v1 parity). And the run gives back its last 

192 piece when that piece is a name candidate: where everything behind 

193 the run is suffix pieces, the run gives back its last piece, when 

194 that piece is one word and is not itself suffix vocabulary 

195 (rules.md#H3, decisions.md#H3 -- the block at the floor below 

196 carries the examples of each half, and the ordering its two inline 

197 tag reads were measured on). 

198 One definition, read by assign (which sets the roles) and by the 

199 chain's trailing-run walk; the leading-particle scan shares the 

200 predicate, is_leading_title, but stops at a title-and-particle 

201 word (P4, #367, #424).""" 

202 n = 0 

203 while n < len(pieces): 

204 if ((n + 1 < len(pieces) or len(pieces) == 1) 

205 and is_leading_title(pieces[n], ptags[n], tokens)): 

206 n += 1 

207 continue 

208 break 

209 # rules.md#H3: "where everything behind the run is post-nominal, 

210 # the run gives its last word back to the name, provided that word 

211 # stands alone and is not itself suffix vocabulary" 

212 # 

213 # ONE WORD, because a joined unit led by a title is a title run and 

214 # handing it back would lose the title: 'Prince of Wales Jr' reads 

215 # title 'Prince of Wales', family 'Jr', not given 'Prince of Wales' 

216 # with no title at all. 

217 # 

218 # Two residuals. A run whose last word IS suffix vocabulary is not 

219 # given back, so 'Dr King MD PhD' still reads title 'Dr King MD', 

220 # family 'PhD'. And the floor asks is_suffix_piece, which vetoes a 

221 # bare initial-shaped numeral, so 'Dr King V' keeps the whole run 

222 # as the title and reads the numeral as the name -- given 'V', 

223 # 'king' being a given-name title and the run's last word, where 

224 # 'Dr Smith V' reads suffix 'V'. That numeral fork is outside this 

225 # floor (decisions.md#H3). 

226 # 

227 # The two inline tag reads are the cheapest NECESSARY condition for 

228 # the piece behind the run to be a suffix piece at all -- 

229 # is_suffix_piece cannot answer yes without one of them -- so the 

230 # ordinary titled name, whose next piece is no kind of suffix, 

231 # leaves this branch without entering a frame. Measured on the 

232 # plan's ordering rather than the shape below: asking the 

233 # authoritative predicate first cost 8 calls per parse of the 

234 # reference name (leading_titles runs four times), against a band 

235 # with room for two (decisions.md#parse-cost). is_suffix_piece 

236 # stays the predicate that ANSWERS, here and in the walk. 

237 if (n and n < len(pieces) 

238 and ("suffix" in ptags[n] 

239 or "vocab:suffix" in tokens[pieces[n][0]].tags) 

240 and len(pieces[n - 1]) == 1 

241 and not is_suffix_piece(pieces[n - 1], ptags[n - 1], 

242 tokens)): 

243 for k in range(n, len(pieces)): 

244 if not is_suffix_piece(pieces[k], ptags[k], tokens): 

245 break 

246 else: 

247 # nothing behind the run but suffix pieces 

248 n -= 1 

249 return n 

250 

251 

252def is_suffix_piece(piece: Sequence[int], ptags: Set[str], 

253 tokens: Sequence[WorkToken]) -> bool: 

254 if "suffix" in ptags: 

255 return True 

256 if len(piece) != 1: 

257 return False 

258 tags = tokens[piece[0]].tags 

259 return "vocab:suffix" in tags and "initial" not in tags 

260 

261 

262def _numeral_behind_the_initial_veto(piece: Sequence[int], 

263 tokens: Sequence[WorkToken]) -> bool: 

264 """Suffix vocabulary that is_suffix_piece refuses because it is 

265 also initial-shaped: a ONE-CHARACTER entry, bare or with a period. 

266 

267 Named for the shape rather than enumerated, because the shape is 

268 what the code tests and the enumeration goes stale -- in the 

269 shipped lexicon it reaches i, v and 2, and NOT x or ix (roman, but 

270 not suffix vocabulary) nor ii/iii/iv (suffix vocabulary, but two 

271 characters, so never initial-shaped and never vetoed in the first 

272 place). A caller adding a one-character suffix in a script that 

273 has initials extends it. 

274 

275 The veto is right where such a word could be a middle initial, and 

276 wrong where it is describing the suffix in front of it, which is 

277 the only place this is asked from. The len(piece) != 1 guard is 

278 defensive: a merged multi-token piece carries "suffix" in ptags, so 

279 is_suffix_piece claims it one branch earlier and no reachable input 

280 arrives here with one. 

281 """ 

282 if len(piece) != 1: 

283 return False 

284 tags = tokens[piece[0]].tags 

285 return "vocab:suffix" in tags and "initial" in tags 

286 

287 

288def segment_suffix_reading(pieces: Sequence[Sequence[int]], 

289 ptags: Sequence[Set[str]], 

290 tokens: Sequence[WorkToken], 

291 lenient: bool, 

292 one_case: bool | None, 

293 ) -> tuple[bool, ...] | None: 

294 """How each piece of a no-name segment reads: True a suffix, False 

295 a title. None when the segment holds a name word and so is not a 

296 credential run at all. 

297 

298 `one_case` admits #289's credential lean: an ALL-CAPS member of 

299 the ambiguous set inside a mixed-case name is a credential in this 

300 slot even with one word before the comma, because the writing is 

301 evidence the count does not have ('Smith, MA' -> family 'Smith', 

302 suffix 'MA'). Only the LEAN reaches here: a token admitted to the 

303 class by SHAPE takes the count instead, which is decided at the 

304 comma and not in this walk ('Smith, A.B.' -> given 'A.B.'). 

305 

306 ONE answer for two readers, both in _assign.py -- the no-name gate 

307 and the router -- because they must agree piece for piece. #429 

308 shipped the inverse of its own fix by deriving that agreement twice 

309 (mechanisms.md#ONE-PREDICATE-PER-QUESTION). It answered for a third 

310 until #436: group's one-entry join asked it too, and the render's 

311 entry boundary is a rule over the written commas in post_rules now 

312 (rules.md#R1), which asks this nothing. 

313 

314 rules.md#S2's initial veto keeps a roman numeral out of a suffix 

315 reading, which is right after a NAME word: 'Smith, John V.' is a 

316 middle initial (#432). After a SUFFIX word the numeral is 

317 describing that suffix -- 'PSM I' is Professional Scrum Master 

318 level I -- so the run continues through it, period included, an 

319 initial there being no shape anyone writes (#430). A title resets 

320 that: what follows a bare title is not continuing a credential. 

321 

322 None covers both ways a segment can fail to be a run: a name word 

323 anywhere in it, and no pieces at all ('Doe,, Jr.', which holds no 

324 title to read by). 

325 

326 `lenient` is Policy.lenient_comma_suffixes, and only the numeral 

327 continuation consults it. C1: "by default a recognized suffix word 

328 counts even written like an initial, while strict mode vetoes 

329 initial-shaped words" -- so under strict the veto stands and the 

330 run ends where it always did. Reading no policy here silently 

331 overrode the one knob a caller sets to prevent exactly this. 

332 

333 The FAMILY_COMMA rule "segment 0 is wholly the family name" rests 

334 on the writer having said where the family name ends. A comma 

335 followed by no name word said no such thing -- 'John Smith, Dr.' is 

336 'Dr. John Smith' with the honorific moved -- so the pre-comma name 

337 keeps its positional read instead of being merged. Uses the same 

338 is_leading_title predicate the peel does, period-abbreviation 

339 inference included, so the two cannot disagree about what a title 

340 is; a mixed run like 'Smith, Dr. Jr.' is a title and a postnominal, 

341 each read where it stands, never a title run 'Dr. Jr.'. 

342 """ 

343 if not pieces: 

344 return None 

345 out: list[bool] = [] 

346 for piece, tags in zip(pieces, ptags): 

347 # the verdict just recorded IS "stands behind a suffix" -- keeping 

348 # a separate flag meant maintaining that equality by hand at three 

349 # sites, and a fourth branch that appended without assigning would 

350 # have diverged silently 

351 after_suffix = bool(out) and out[-1] 

352 if is_suffix_piece(piece, tags, tokens): 

353 out.append(True) 

354 elif (len(piece) == 1 

355 and AMBIGUOUS_ACRONYM_TAG in tokens[piece[0]].tags 

356 and listed_lean(tokens[piece[0]], one_case) 

357 == "credential"): 

358 out.append(True) 

359 elif (lenient and after_suffix 

360 and _numeral_behind_the_initial_veto(piece, tokens)): 

361 out.append(True) 

362 elif is_leading_title(piece, tags, tokens): 

363 out.append(False) 

364 else: 

365 return None 

366 return tuple(out) 

367 

368 

369class Peel(NamedTuple): 

370 """What assign's trailing peel made of a walk. `names` is a count 

371 of positions in the caller's `rest`: rest[:names] are the name 

372 pieces and rest[names:] the suffixes. The other two are pieces -- 

373 token-index tuples, as PendingAmbiguity wants them -- and each is 

374 one token long: `numeral` is the piece the roman-numeral fork 

375 took (None when it did not fire; always the walk's last piece), 

376 `picks` the bare ambiguous acronyms the peel had to resolve, in 

377 peel order, either way (the last may sit at rest[names - 1]).""" 

378 

379 names: int 

380 numeral: tuple[int, ...] | None 

381 picks: tuple[tuple[int, ...], ...] 

382 

383 

384# rules.md#S2: "a trailing word of the suffix vocabulary reads as a 

385# suffix — generational forms and credential acronyms alike, and an 

386# ambiguous acronym written with its periods, one after each letter, 

387# counts unambiguously; a single trailing period is the abbreviation 

388# shape any word can wear and does not. A BARE ambiguous acronym is 

389# consumed only when the name has words to spare" 

390# (v1's are_suffixes tail rule, with the roman-numeral special) 

391def peel_walk(start: int, ptags: Sequence[Set[str]], 

392 skip: Set[int] = frozenset()) -> list[int]: 

393 """The indices peel_trailing walks: `start` to the segment's end, 

394 minus the group-flagged credential pieces (the Ph. D. merge), 

395 which assign reads as suffixes at any position, and minus `skip` 

396 -- a tail segment's delimiter cores, which are structure rather 

397 than words (the maiden walk's case, #424). Built here and nowhere 

398 else, so the walk's input cannot drift between assign and the 

399 group sites that read it: the numeral fork is a last-piece 

400 test that reads the piece before as rest[k - 2], which holds only 

401 over this list.""" 

402 return [j for j in range(start, len(ptags)) 

403 if j not in skip and "suffix" not in ptags[j]] 

404 

405 

406def trailing_start(start: int, pieces: Sequence[Sequence[int]], 

407 ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], 

408 skip: Set[int] = frozenset(), 

409 *, one_case: bool | None) -> int: 

410 """Where assign's trailing suffix run begins, read over the pieces 

411 as they stand from `start`: the index of the first piece the S2 

412 peel takes, or len(pieces) when it takes none (#424). What P2's 

413 chain and M2's walk stop before -- each had asked "is this a 

414 suffix?" with the suffix-piece test, which vetoes a bare 'V' as 

415 an initial (the #401 question), and so took a trailing numeral, 

416 or a bare acronym with words to spare, into the family or the 

417 maiden name. 

418 

419 Both forks, always. A caller that needs one of them alone -- the 

420 maiden walk re-asking the numeral over the view its take would 

421 leave, where the acronym fork's piece COUNT no longer describes 

422 the name -- calls the `peel_walk` + `peel_trailing` pair this 

423 wraps and reads the half it wants (#533). A `numeral_only` flag 

424 lived here for that one caller and cost it a frame.""" 

425 rest = peel_walk(start, ptags, skip) 

426 peeled = peel_trailing(rest, pieces, ptags, tokens, one_case) 

427 return rest[peeled.names] if peeled.names < len(rest) else len(pieces) 

428 

429 

430# #289/#516: the "listed member, not by-shape" test both 

431# peel_trailing and segment_suffix_reading ask before reading the 

432# lean -- shared here so the two cannot drift on what counts 

433# (quality-review finding: it was spelled twice, once per site, 

434# before this). Those two test "vocab:suffix-ambiguous" in tags 

435# INLINE, before calling this, rather than leaving that cheap check to 

436# this function's own body: measured, a caller whose `elif` reaches 

437# this on every piece (segment_suffix_reading's does, one per 

438# family-comma segment 1, member or not) pays one frame for the call 

439# regardless of what is inside it, and the inline pre-check is what 

440# keeps a non-member piece ("Smith, John"'s "John") from ever making 

441# the call at all. 

442# 

443# A THIRD caller since #533 -- credential_at_the_given_slot just 

444# below -- deliberately does NOT pre-check: it owns #531's reading 

445# and leaves membership to its own callers (its docstring says so), 

446# and the frame argument holds transitively because both of them ask 

447# inline -- `AMBIGUOUS_ACRONYM_TAG in tok.tags` after a 

448# `len(piece) == 1` at assign's given-part trailing slot, and the 

449# same pair inside the `all(...)` of `_group.py`'s `_maiden_take` 

450# view check. So no non-member piece reaches this function down that 

451# route either. 

452def listed_lean(token: WorkToken, one_case: bool | None) -> Lean | None: 

453 """`ambiguous_lean` for a LISTED bare-ambiguous token, or None if 

454 the token is not tagged a listed member, is admitted by SHAPE 

455 instead (`SHAPE_ACRONYM_TAG`, a switch's doing, not the writing's), 

456 or there is no case fact to ask at all.""" 

457 if (one_case is None or AMBIGUOUS_ACRONYM_TAG not in token.tags 

458 or SHAPE_ACRONYM_TAG in token.tags): 

459 return None 

460 return ambiguous_lean(token.text, one_case) 

461 

462 

463def credential_at_the_given_slot(token: WorkToken, 

464 one_case: bool | None) -> bool: 

465 """#531's reading of a class MEMBER ending the given part after a 

466 family comma: the credential unless the writing says otherwise. 

467 The caller decides membership and that the piece ends that part. 

468 

469 The words to spare are there by construction at that slot, so the 

470 count says nothing and only the lean does; a member that is also 

471 particle vocabulary reads as the credential on a POSITIVE lean 

472 alone, P6's attachment keeping every other spelling. 

473 

474 Two callers since #533 -- assign's walk over the given part, and 

475 the maiden walk's second check over the name the take would leave 

476 (rules.md#M2) -- so the reading is a function rather than a 

477 condition written twice 

478 (mechanisms.md#ONE-PREDICATE-PER-QUESTION). It is a 

479 text-and-tags question, which is what puts it in this module 

480 rather than beside either caller. 

481 

482 The membership half of that contract is CHECKED rather than 

483 trusted, because getting it wrong is silent: all three of 

484 `listed_lean`'s None reasons fall through to the `particle` test 

485 below, so a non-member handed in by mistake is answered True -- 

486 "read it as the credential" -- for a word the class never admitted. 

487 An assert rather than a raise or a branch: it enters no Python 

488 frame (measured -- 'Doe, John MA' stays at 311), it states the 

489 contract where a reader of the function body meets it, and under 

490 -O it is exactly the code that was here before. 

491 """ 

492 assert AMBIGUOUS_ACRONYM_TAG in token.tags, ( 

493 f"credential_at_the_given_slot is #531's reading of a LISTED " 

494 f"class member; {token.text!r} carries {sorted(token.tags)} " 

495 f"and is not one. The caller decides membership -- test " 

496 f"AMBIGUOUS_ACRONYM_TAG before calling") 

497 lean = listed_lean(token, one_case) 

498 return lean == "credential" or (lean is None 

499 and "particle" not in token.tags) 

500 

501 

502def peel_trailing(rest: Sequence[int], pieces: Sequence[Sequence[int]], 

503 ptags: Sequence[Set[str]], 

504 tokens: Sequence[WorkToken], 

505 one_case: bool | None) -> Peel: 

506 """The S2 trailing peel over `rest`, a peel_walk list. In the 

507 piece layer rather than in assign because group's bound-given 

508 reserve asks the same question of the view the join would leave 

509 (#425): one walk, so the reserve and the assignment cannot drift. Pure -- the ambiguities are 

510 returned for assign to report, in the order it always reported 

511 them. 

512 

513 `one_case` is ParseState.one_case, the recorded fact: None means 

514 nobody asked, which is every caller that has no state to ask with, 

515 and reads as "no lean" -- rules.md#S2's count alone, the behavior 

516 of every release before this one. 

517 """ 

518 picks: list[tuple[int, ...]] = [] 

519 numeral: tuple[int, ...] | None = None 

520 k = len(rest) 

521 while k > 0: 

522 piece = pieces[rest[k - 1]] 

523 if is_suffix_piece(piece, ptags[rest[k - 1]], tokens): 

524 k -= 1 

525 continue 

526 # a final single letter that is a roman numeral, after a piece 

527 # that is not initial-shaped; the predicate's docstring carries 

528 # the is_initial_shaped reasoning (#320) 

529 if (k == len(rest) and k >= 2 and len(piece) == 1 

530 and is_trailing_numeral_suffix( 

531 tokens[piece[0]].text, 

532 tokens[pieces[rest[k - 2]][0]].text)): 

533 numeral = tuple(piece) 

534 k -= 1 

535 continue 

536 # A bare ambiguous acronym ("MA", not "M.A.") is a credential 

537 # only when peeling it still leaves a given AND a family name. 

538 # With two pieces, "one of them is a credential" is the less 

539 # likely reading, so it stays the family name -- "Jack MA" is a 

540 # person, "John Smith MA" is a person with a degree. This is 

541 # v1's reserve_last narrowed to the ambiguous set: 2.0 

542 # deliberately peels UNambiguous suffixes even when nothing is 

543 # left ("Smith PhD" -> suffix, a classified fix), because there 

544 # the vocabulary is not in doubt. 

545 bare_ambiguous = (len(piece) == 1 

546 and AMBIGUOUS_ACRONYM_TAG in tokens[piece[0]].tags) 

547 # #516, switch off: the writing still makes this token 

548 # credential-SHAPED, and the parser is choosing the name 

549 # reading over that one -- the fork the caller asked to be 

550 # told about. Reported here, unconsumed, rather than folded 

551 # into `bare_ambiguous` above: with the switch off the token 

552 # never carries "vocab:suffix-ambiguous" (classify's own 

553 # gate), so `bare_ambiguous` is already False and this is the 

554 # ONLY place the report can be recorded. Switch ON, the token 

555 # carries BOTH tags, so `not bare_ambiguous` is what stands 

556 # this branch down and lets the consuming branch below take 

557 # it; the two are not exclusive. 

558 if (not bare_ambiguous and k >= 2 and len(piece) == 1 

559 and SHAPE_ACRONYM_TAG in tokens[piece[0]].tags): 

560 picks.append(tuple(piece)) 

561 break 

562 # #289: written case is evidence the count does not have, and 

563 # it overrides the count in BOTH directions -- an all-caps 

564 # member of a mixed-case name is taken with nothing to spare 

565 # ("Jack MA"), a Title-case one is declined with plenty 

566 # ("John Smith Ma"). The lean is the LISTED set's alone: a 

567 # token admitted to this class by SHAPE carries no writing 

568 # convention to read, so it takes the count (decisions.md#S2). 

569 # 

570 # k < 2 means it is the only piece left, which is not the fork 

571 # this reports -- and not a floor the lean moves: the walk 

572 # starts after the leading title run, so one piece behind a 

573 # title ("Mr MA") is exactly this case and must stay a name. 

574 # The lean is computed only past this floor -- membership, 

575 # then the floor, then the count-or-lean, in that order, with 

576 # nothing computed a step earlier could discard. 

577 if bare_ambiguous and k >= 2: 

578 picks.append(tuple(piece)) 

579 lean = listed_lean(tokens[piece[0]], one_case) 

580 # peeling still leaves given + family, or the writing says 

581 # to peel anyway 

582 if lean == "credential" or (lean is None and k >= 3): 

583 k -= 1 

584 continue 

585 break 

586 return Peel(k, numeral, tuple(picks)) 

587 

588 

589# rules.md#H5: "only a word the vocabulary knows as a title is one, 

590# and a bare title word is a name word" 

591# -- the trailing run's own predicate. NOT is_leading_title: 

592# that predicate carries H2's unlisted-abbreviation inference, which is 

593# the LEADING slot's shape rule and has no trailing counterpart, so 

594# with it 'John Smith Xyz.' would lose its family name to a title 

595# (decisions.md#H5). The vocabulary read is is_title_piece's, shared 

596# with the leading run so the two cannot disagree about what a title 

597# WORD is while disagreeing, deliberately, about what a title SHAPE is. 

598def trailing_titles(rest: Sequence[int], pieces: Sequence[Sequence[int]], 

599 ptags: Sequence[Set[str]], 

600 tokens: Sequence[WorkToken]) -> int: 

601 """How many pieces of `rest` the trailing title chain LEAVES 

602 standing: `rest[:kept]` are the name pieces and `rest[kept:]` the 

603 period-marked title words the chain took, in piece order. Counted 

604 the way `peel_trailing` counts, so the two answers compose without 

605 arithmetic at the call site. `rest` is the caller's NAME pieces: 

606 on the no-comma path what the S2 peel left, after a family comma 

607 the segment's pieces that the segment's own suffix reading does 

608 not claim, and in `tail_reading` the leftovers of whichever peel 

609 is current. 

610 Floor: one name piece stands, so a name is never all title -- and 

611 an empty `rest` returns 0, which is what leaves assign's 

612 bare-suffix carve-out reached exactly as before. 

613 

614 ONE WORD per piece, the same gate the leading peel's give-back 

615 uses: a joined unit is not the shape this reads, and the tokens of 

616 one are not each a title word. 

617 

618 Every parse with a name word to place enters this frame -- assign 

619 asks the question here rather than answering a cheaper version of 

620 it inline (mechanisms.md#ONE-PREDICATE-PER-QUESTION) -- so what it 

621 costs an ordinary name is one frame and one regex match. The 

622 exceptions return before it: a segment that is all title, and a 

623 comma part read wholly as a credential run, have no name piece to 

624 hand this (52 of the 1289 corpus parses, measured 2026-09-09 -- 

625 'Coach', 'Lord of the Universe', 'Smith, Jr.', 'MD, PHD'). 

626 The shape test 

627 runs BEFORE the vocabulary one to keep it at that: _PERIOD_ABBREV 

628 is a compiled regex (a C call, no Python frame) where 

629 is_title_piece is a call, and almost no name ends in a 

630 period-marked word, so the ordinary parse pays the one match and 

631 stops (decisions.md#parse-cost). 

632 """ 

633 k = len(rest) 

634 while k > 1: 

635 idx = rest[k - 1] 

636 piece = pieces[idx] 

637 # no #323 veto on the shape here, unlike is_leading_title's: 

638 # the shape is ANDed with is_title_piece, so the word is listed 

639 # vocabulary, and a listed CJK title wearing a stop should read 

640 # as a title. 

641 if (len(piece) == 1 

642 and _PERIOD_ABBREV.match(tokens[piece[0]].text) 

643 and is_title_piece(piece, ptags[idx], tokens)): 

644 k -= 1 

645 continue 

646 break 

647 return k 

648 

649 

650# rules.md#H5: "the title is TRANSPARENT to the suffix reading: where 

651# two or more name words stand, what stands once the chain is taken 

652# reads exactly as it would read written without the title, plus the 

653# title" 

654def tail_reading(rest: list[int], pieces: Sequence[Sequence[int]], 

655 ptags: Sequence[Set[str]], 

656 tokens: Sequence[WorkToken], 

657 one_case: bool | None, 

658 ) -> tuple[list[int], tuple[int, ...], Peel]: 

659 """The S2 peel and the H5 chain read together to a FIXED POINT: 

660 peel, chain, splice the chained pieces out, peel again over what 

661 is left -- the name pieces the chain kept, then the pieces the 

662 peel had taken, in original order -- until the chain takes 

663 nothing. Where it takes nothing on the first pass, which is almost 

664 every name, that first peel is the answer and the loop costs one 

665 comparison. 

666 

667 Returns the walk with the chained pieces spliced out, the pieces 

668 the chain took, and the FINAL peel -- whose numeral fork and 

669 ambiguous picks are the ones assign reports. The walk is 

670 partitioned by that peel exactly as a caller partitions its own: 

671 `rest[:peel.names]` the name pieces, `rest[peel.names:]` the 

672 suffixes. A bare tuple rather than a named one because a 

673 NamedTuple's __new__ is a frame of its own on every parse 

674 (decisions.md#parse-cost), and `_group_segment` returns its three 

675 the same way. 

676 

677 Transparency is what the fixed point buys: 'X Prof. Y' reads 

678 exactly as 'X Y' reads plus the title, however many titles are 

679 written and wherever the peel then stops. Iterating ONCE reads a 

680 second title only half way -- 'John Prof. MA Prof.' un-peeled the 

681 acronym and re-exposed the first title, reading family 'Prof.' 

682 with suffix 'MA' where 'John Prof. MA' reads family 'MA'. 

683 

684 One function for two readers -- assign's placement and group's 

685 bound-given reserve (P5), which must count the name words assign 

686 will leave. Deriving that agreement twice is what left the two 

687 disagreeing at S2's bare-ambiguous reserve: 'abdul rahman MA' 

688 declined the join and 'abdul rahman MA Prof.' took it 

689 (mechanisms.md#ONE-PREDICATE-PER-QUESTION). 

690 

691 `rest` is a peel_walk list and is not mutated -- the splice 

692 rebinds this local -- so a caller's own reference still names the 

693 walk it built. Both callers read the one returned here instead, 

694 which is the one the final peel partitions. 

695 """ 

696 titled: list[int] = [] 

697 while True: 

698 peeled = peel_trailing(rest, pieces, ptags, tokens, one_case) 

699 kept = trailing_titles(rest[:peeled.names], pieces, ptags, 

700 tokens) 

701 if kept == peeled.names: 

702 return rest, tuple(titled), peeled 

703 # the chain's pieces reach this list back to front, so each 

704 # run goes in FRONT of what the pass before it took 

705 titled[:0] = rest[kept:peeled.names] 

706 rest = rest[:kept] + rest[peeled.names:] 

707 

708 

709# rules.md#H5: "the title is TRANSPARENT to the suffix reading: where 

710# two or more name words stand, what stands once the chain is taken 

711# reads exactly as it would read written without the title, plus the 

712# title" 

713def trailing_start_past_titles(start: int, 

714 pieces: Sequence[Sequence[int]], 

715 ptags: Sequence[Set[str]], 

716 tokens: Sequence[WorkToken], 

717 *, one_case: bool | None) -> int: 

718 """`trailing_start` read through H5's chain: where assign's 

719 trailing suffix run begins once a trailing TITLE has stopped 

720 hiding it. 

721 

722 `trailing_start` reads the pieces as WRITTEN, so a title standing 

723 behind the suffix run makes the peel take nothing and the answer 

724 is `len(pieces)` -- the reading assign itself has not had since 

725 H5, because assign runs the peel and the chain to their fixed 

726 point instead (`tail_reading`). A caller using that answer as the 

727 right bound of the NAME is told a credential is a name word: 

728 'John Quincy Adams i MA Prof.' read family 'Adams i MA' where 

729 'John Quincy Adams i MA' reads family 'Adams' and suffix 'i MA' 

730 (#397 second review). Every caller that bounds the name wants 

731 this one; `trailing_start` stays for the callers that count a 

732 trailing run of the pieces as they stand. 

733 

734 The returned index bounds the name from the right, and the 

735 trailing titles the chain spliced out are not under it: they end 

736 the segment, so they stand at or past the first suffix piece 

737 whenever there is one. Where the peel takes nothing even past the 

738 chain this returns `len(pieces)` as `trailing_start` does, and a 

739 trailing title is then inside the bound and refused by the title 

740 test the callers already run beside it. 

741 """ 

742 rest, _titled, peeled = tail_reading(peel_walk(start, ptags), 

743 pieces, ptags, tokens, one_case) 

744 return rest[peeled.names] if peeled.names < len(rest) else len(pieces)