Coverage for /pythoncovmergedfiles/medio/medio/usr/local/lib/python3.11/site-packages/nameparser/_pipeline/_script_segment.py: 83%

Shortcuts on this page

r m x   toggle line displays

j k   next/prev highlighted chunk

0   (zero) top of page

1   (one) first highlighted chunk

122 statements  

1"""Stage: script_segment (#271, #272, #308, #312). 

2 

3Consumes: tokens, segments, structure, interpunct_offsets, segmenter, 

4one_case (where segment recorded it; None everywhere else, which is 

5what this stage's suffix-run predicate read for every name before 

62.4). 

7Produces: tokens, by two independent splits into sub-slices -- a 

8listed honorific peeled off the END of the name's last 

9non-post-nominal token, in whichever of the name's runs that falls 

10(the tail token also carries this module's _PEELED_TAG), and the 

11first activated-script token of the name segment split into n+1 -- 

12segments (index runs remapped past the insertions), ambiguities 

13(indices likewise remapped, plus a SEGMENTATION report when more than 

14one split was vocabulary-supported, or when a segmenter's answer 

15scored under the confidence floor). 

16Reads: Policy.segment_scripts, Lexicon.surnames, 

17Lexicon.honorific_tails, ParseState.segmenter, and Lexicon suffix 

18vocabulary through TWO predicates, which are NOT one another's 

19singular and plural. _vocab.is_suffix_strict asks whether a single 

20token is a post-nominal, initial veto included (the peel's scan-back 

21and the surname site). _vocab.is_wholly_suffix asks segment's own 

22suffix-comma question of a whole RUN, through the POLICY-selected 

23token test plus period_joined_vocab, delimiter handling and the 

24Ph./D. merge -- so the run predicate says yes both to tokens the 

25token predicate VETOES ("V.", "V", "I") and to tokens it never sees 

26as suffixes at all ("Msc.Ed.", "J.씨", which reach it through 

27period_joined_vocab). The initial-shaped words are the class #319 was 

28ABOUT, not the whole of the disagreement, and the extra routes listed 

29above are where the rest of it comes from; 

30test_is_wholly_suffix_is_not_the_plural_of_is_post_nominal pins the 

31"V." half. The run predicate is how the peel declines a run that is 

32not name text (#319), but only where the name's own run offers a 

33site, since a glued honorific is itself part of what makes a run read 

34as suffix-shaped; it owns two further Policy fields 

35(lenient_comma_suffixes picks the strict or lenient token test -- it 

36flips "田中さん, V." -- and extra_suffix_delimiters both counts a bare 

37delimiter-core token as a suffix and splits a token on a core, the 

38split being the half that flips "田中さん, Jr./V." under {"/"} and the 

39bare core the half that flips "田中さん, /"). 

40 

41Implements rules W1 (the vocabulary/segmenter division), W2 (the 

42glued-honorific peel) and W3 (the writer's divisions are respected) 

43of docs/design/rules.md, cited at their code below; the decision 

44chain (#308, #312, #319, the vetting bars, the measured 

45spaced-honorific trade) is decisions.md#W1, #W2 and #W3. Both splits 

46make sub-slices of one token, rewriting nothing -- spans still index 

47the original exactly, so the anti-#100 invariant holds by 

48construction. Both also match on the token's CORE, its 

49text.rstrip(FULL_STOPS) (#323) -- rstrip and never strip, because 

50both offsets are measured from the text's START, so the core has to 

51stay a prefix of the text for the offset the match yields to land 

52where the match did. The peel runs FIRST, so suffix classification 

53can claim the tail and the surname match or segmenter consult sees 

54the name rather than name-plus-honorific; its 

55ASCII bail sits above everything here, so a caller-added LATIN tail 

56fires only on a name carrying at least one non-ASCII character (see 

57the bail's own comment, and honorific_tails' field note). 

58""" 

59from __future__ import annotations 

60 

61import dataclasses 

62import functools 

63from collections.abc import Sequence 

64 

65from nameparser._lexicon import FULL_STOPS 

66from nameparser._pipeline._state import ( 

67 ParseState, PendingAmbiguity, Structure, WorkToken, 

68) 

69from nameparser._pipeline._vocab import ( 

70 effective_script, is_suffix_strict, is_wholly_suffix, 

71) 

72from nameparser._types import AmbiguityKind, Segmentation, Span 

73 

74#: Marks the tail token the peel below MANUFACTURED, so the segmenter's 

75#: neighbour test can tell it from a token somebody wrote. Namespaced, 

76#: therefore unstable provenance rather than API (_types.STABLE_TAGS is 

77#: the whole stable set, and FOLDED_TAG is the precedent for a 

78#: structural marker carrying this prefix); it is vocabulary-derived 

79#: besides, since honorific_tails is what licensed the split. Emitter 

80#: and reader are both in this module, so unlike FOLDED_TAG it needs no 

81#: home in _types. 

82_PEELED_TAG = "vocab:peeled-honorific" 

83 

84#: Segmenter answers scoring below this attach a SEGMENTATION report 

85#: (amendment 2026-07-29 section 3). Kept at the drafted 0.9 after 

86#: measuring namedivider 0.4.1 over 112 names (#272 Task 5), which 

87#: found a distribution the amendment did not anticipate: the scores 

88#: are BIMODAL, not clustered near 1. A rule-based division (the kana 

89#: boundary in 高橋みなみ, or a two-character name) scores exactly 1.0; 

90#: a kanji-statistics division scores a softmax over the candidate cut 

91#: positions, observed in 0.23-0.68 and driven more by name LENGTH 

92#: (median 0.61 at three characters, 0.37 at five -- the softmax is 

93#: over len-1 candidates) than by correctness: the wrong answers in 

94#: the sample scored 0.32/0.60/0.60/0.61, straddling the correct 

95#: median of 0.52. So no floor INSIDE that band separates error from 

96#: success, and the only cut the data supports is between a stated 

97#: certainty and a statistical guess -- which is what section 3's 

98#: epistemic argument asks for anyway. 0.9 sits in the empty gap 

99#: (0.68 to 1.0), far from both modes, and reads as "confident" for a 

100#: third-party segmenter with a calibrated score too. Consequence, 

101#: pinned by tests and stated so it can be checked: every 

102#: STATISTICALLY divided name carries a SEGMENTATION report, and every 

103#: RULE-divided one -- the kana boundary, a two-character name, 

104#: namedivider's specific-name rules -- carries none. 

105#: namedivider scores its own two-character rule 1.0, so this floor 

106#: keeps that division silent -- see locales/ja.py for why the 

107#: presumption is accepted as stated. 

108#: Not configurable in 2.x (YAGNI). 

109_SEGMENTER_CONFIDENCE_FLOOR = 0.9 

110 

111 

112def _remap(run: tuple[int, ...], split_at: int, 

113 added: int) -> tuple[int, ...]: 

114 """One segment run after token `split_at` split into `added` + 1 

115 pieces: the extra pieces join the run immediately after it, and 

116 every later index shifts by as many.""" 

117 out: list[int] = [] 

118 for j in run: 

119 if j == split_at: 

120 out.extend(range(split_at, split_at + added + 1)) 

121 else: 

122 out.append(j + added if j > split_at else j) 

123 return tuple(out) 

124 

125 

126def _pieces(text: str, splits: tuple[int, ...]) -> tuple[str, ...]: 

127 """`text` cut at every offset in `splits`: n offsets, n+1 pieces.""" 

128 cuts = (0, *splits, len(text)) 

129 return tuple(text[a:b] for a, b in zip(cuts, cuts[1:])) 

130 

131 

132def _split(state: ParseState, i: int, splits: tuple[int, ...], 

133 detail: str | None, tail_tag: str | None = None) -> ParseState: 

134 """Cut token `i` at every offset in `splits`, recording `detail` as 

135 a SEGMENTATION report when there is one, and adding `tail_tag` to 

136 the LAST piece when there is one of those. 

137 

138 The offsets arrive non-empty, ascending and interior whatever chose 

139 them. From a segmenter: non-empty is the caller's own check, 

140 strictly ascending with each >= 1 is Segmentation.__post_init__'s, 

141 and the last offset is < len(text) -- checked by the caller. From 

142 the vocabulary: the single offset is >= 1 and < len(text) by the 

143 range(cap, 0, -1) construction and its len-1 cap. 

144 

145 The ONE split path: the vocabulary hit is the single-offset case 

146 and the segmenter's answer the general one, so neither can drift 

147 from the other's index arithmetic. `tail_tag` is part of keeping it 

148 one path: the piece is tagged HERE, where it is built, rather than 

149 by the caller, which would have to re-derive where its own tail 

150 landed after handing the index arithmetic off. Only the peel passes 

151 one, and the piece it wants is the last by construction -- the 

152 honorific it cut off the end; the surname path passes nothing and 

153 the default leaves it exactly as it was.""" 

154 token = state.tokens[i] 

155 base = token.span.start 

156 parts: list[WorkToken] = [] 

157 start = 0 

158 for piece in _pieces(token.text, splits): 

159 end = start + len(piece) 

160 parts.append(dataclasses.replace( 

161 token, text=piece, span=Span(base + start, base + end))) 

162 start = end 

163 if tail_tag is not None: 

164 parts[-1] = dataclasses.replace( 

165 parts[-1], tags=parts[-1].tags | {tail_tag}) 

166 added = len(splits) 

167 tokens = state.tokens[:i] + tuple(parts) + state.tokens[i + 1:] 

168 # Every index the earlier stages recorded is now stale past the 

169 # split point: the segment runs group's own iteration rests on, and 

170 # the ambiguities from extract_delimited (resolved to indices by 

171 # tokenize) and segment. An ambiguity ON the split token keeps 

172 # pointing at the head. 

173 segments = tuple(_remap(run, i, added) for run in state.segments) 

174 ambiguities = tuple( 

175 dataclasses.replace(a, indices=tuple( 

176 j + added if j > i else j for j in a.indices)) 

177 for a in state.ambiguities) 

178 if detail is not None: 

179 ambiguities += (PendingAmbiguity( 

180 AmbiguityKind.SEGMENTATION, detail, 

181 tuple(range(i, i + added + 1))),) 

182 return dataclasses.replace(state, tokens=tokens, segments=segments, 

183 ambiguities=ambiguities) 

184 

185 

186@functools.lru_cache(maxsize=16) 

187def _longest_entry(entries: frozenset[str]) -> int: 

188 """The longest entry in a vocabulary, cached per-vocabulary rather 

189 than recomputed per parse (Lexicon is frozen and slotted, so it 

190 cannot carry a cached_property of its own). The frozenset is 

191 hashable and a process holds only a handful of distinct 

192 vocabularies -- the default one, plus one per constructed pack 

193 parser -- so maxsize=16 bounds pathological many-lexicon churn 

194 without ever evicting in normal use. Keyed by the frozenset VALUE, 

195 not by (lexicon, field): the two callers pass surnames and 

196 honorific_tails, but every lexicon that leaves KOREAN_SURNAMES 

197 alone shares one entry, so the count is distinct SETS rather than 

198 lexicons times fields. 

199 

200 Callers must pass a NON-EMPTY vocabulary: max() of an empty set 

201 raises, and both call sites sit under a match guard that cannot 

202 reach it with one.""" 

203 return max(map(len, entries)) 

204 

205 

206def _is_post_nominal(state: ParseState, i: int) -> bool: 

207 """Whether token `i` is post-nominal VOCABULARY. Two of this 

208 stage's decisions ask it and neither is about script: an honorific 

209 is neither a surname SITE nor the token a glued honorific hangs off 

210 (#308). 

211 

212 It answers a VOCABULARY question, not a positional one, and the two 

213 callers do not spend the answer the same way. In TRAILING position 

214 vocabulary and position agree, so the peel site steps over a True 

215 and goes on scanning. In LEADING position they disagree -- 양 is 

216 the family name there, whatever the suffix set says -- so the 

217 surname site reads a True as an ANSWER and declines, rather than as 

218 a token to step past. 

219 

220 STRICT, not lenient: the initial veto applies, so "V." is not a 

221 post-nominal here though bare "v" is a suffix word. That agrees 

222 with what classify does with the same token downstream -- "V." is 

223 a middle initial -- and the difference is reachable under the 

224 default lexicon: "田中さん V." stops its scan-back at the initial 

225 and does not peel, while "田中さん II" steps over II and does. That 

226 pair is the case table's ja_honorific_glued_before_an_initial and 

227 ja_honorific_glued_before_a_roman_suffix, added because the clause 

228 could NAME the discriminating input while swapping in 

229 is_suffix_lenient still failed no test -- the shape a "unify the 

230 two suffix predicates" refactor would have walked straight past. 

231 

232 Suffixes only, deliberately. Should a CJK entry ever join titles, 

233 the surname site would want it excluded while the peel site must 

234 not: a title never trails, so widening the peel's scan-back would 

235 move it off the real last name token. Split the predicate then, not 

236 before.""" 

237 return is_suffix_strict(state.tokens[i].text, state.lexicon) 

238 

239 

240# rules.md#W2: "A part that is not name text — a post-nominal word 

241# standing on its own — is never the name's end: the split-off steps 

242# past it to the name word behind, and never dissects it." (history: 

243# decisions.md#W2) 

244# The scan-back below is that clause, and it needs no punctuation to 

245# fire. '김민준 박사님' steps past 박사님 to 김민준, finds no listed 

246# tail there and returns None -- so 박사님 is left whole for suffix 

247# classification rather than cut into 박사 + 님, which is what the 

248# same input gives when the step is removed. '선생님' is post-nominal 

249# entire with no name word behind it, so the scan yields no site at 

250# all and the token stays whole. Both are contract-tier corpus names 

251# and rules.md#W2 example lines; decisions.md#cjk-comma-demotion 

252# carries the forced-predicate measurement behind them. 

253def _peel_site(state: ParseState, flat: Sequence[int], 

254 tails: frozenset[str]) -> tuple[int, int] | None: 

255 """Where a peel would land in the token run `flat`: the index of the 

256 token to cut and the OFFSET to cut it at (the listed tail, and any 

257 full stops riding behind it, are what comes off the end), or None 

258 where that run offers no peel. 

259 

260 Two callers, one answer, which is the point of naming it. The peel 

261 itself asks it once, of the runs it decided to scan. The gate above 

262 that decision asks it of segments[0] ALONE, to find out whether 

263 declining the second run would cost the only site. Sharing the scan 

264 with the peel rather than approximating it there is what makes the 

265 gate's answer mean what it says: a gate that merely looked for a 

266 token ENDING in a listed tail would count a token that is a tail 

267 entire, which the scan-back skips as a site -- "선생님, J.씨" would 

268 decline on a site the peel then cannot use, and lose the peel in 

269 exactly the way the gate exists to prevent. (Only the lone-token 

270 shape of that divergence is reachable from the gate, and the gate's 

271 own other conjunct is what makes that so: SUFFIX_COMMA wants BOTH 

272 a suffix-shaped second run and more than one word before the 

273 comma, and the gate is asked only where the first of those is 

274 already true -- so a run this call ever sees under FAMILY_COMMA 

275 failed on the second, and segments[0] holds at most one token. The 

276 word-count alone would not say that: "Dr 김민준, 지훈" has two 

277 words before the comma and is FAMILY_COMMA.) 

278 

279 Callers must pass a NON-EMPTY `tails` -- _longest_entry's 

280 precondition, which the stage's own early return supplies.""" 

281 i = next((j for j in reversed(flat) 

282 if not _is_post_nominal(state, j)), None) 

283 if i is None: 

284 # nothing but post-nominals, or no tokens at all 

285 return None 

286 text = state.tokens[i].text 

287 # The tail is matched on the token's core (#323): a stop glued after 

288 # the honorific ('김민준씨.', '田中さん.') stands between the listed 

289 # tail and the token's end and used to defeat the match. What the 

290 # text read instead depended on the stop: the ASCII spelling went to 

291 # a title downstream (H2's shape is ASCII-period-only), while a 

292 # fullwidth or ideographic stop left the whole text a lone name word 

293 # ('田中さん。' read given). The cut lands BEFORE the tail, so the 

294 # stops ride with the honorific piece and the token text is never 

295 # rewritten. TRAILING only: a leading stop is not between the name 

296 # and its honorific, and the offset returned below is the core's 

297 # length less the tail's, one subtraction. 

298 core = text.rstrip(FULL_STOPS) 

299 # range/cap construction identical to the surname match below, and 

300 # for the same two reasons: longest-first, and a len-1 cap that 

301 # makes the offset interior by construction (_split's contract). An 

302 # empty or one-character core caps below 1 and the loop is empty. 

303 cap = min(_longest_entry(tails), len(core) - 1) 

304 for length in range(cap, 0, -1): 

305 if core[-length:] in tails: 

306 return i, len(core) - length 

307 return None 

308 

309 

310# rules.md#W2: "a listed honorific glued to the end of the name's 

311# last name word splits off once and reads as a suffix." (history: 

312# decisions.md#W2) 

313# That the peel also reaches ACROSS a family comma is stated at 

314# rules.md#W3 instead, which is a tolerated rule since the 

315# 2026-09-01 comma demotion -- the crossing is what the parser does 

316# today, not something W2 promises. W3 also carries the reading of a 

317# period a listing leaves behind (decisions.md#cjk-comma-demotion, 

318# amended by decisions.md#cjk-full-stops): a period on a SEPARATE 

319# post-nominal word rides into the suffix and moves nothing ('様.' 

320# is post-nominal-strict and the scan steps past it as it steps past 

321# '様'), and since #323 a period glued to the honorific's OWN token 

322# rides with the honorific too -- _peel_site matches the tail through 

323# the trailing stop and cuts before it, so '田中さん.' divides where 

324# '田中さん' does, and '김민준씨.' where '김민준씨' does (both are 

325# case rows now; before 2026-09-10 they read as a title, measured 

326# 2026-09-05 and pinned by nothing, and for one commit of the #323 

327# branch the surname site read '김민준씨.' as 김 + 민준씨.). Neither 

328# reading is a promise; the step past the post-nominal word itself 

329# is W2's, above, and is. 

330def _peel_honorific_tail(state: ParseState) -> ParseState: 

331 """#308: split a listed honorific off the END of the name's last 

332 NON-POST-NOMINAL token -- 田中さん -> 田中 + さん -- and let 

333 the existing machinery do the rest. Suffix classification claims 

334 the tail downstream (every honorific_tails entry is a suffix word 

335 too, enforced by Lexicon), and the segmentation half below then 

336 sees the remainder rather than the glued whole, so 김민준씨 splits 

337 김 + 민준 and a configured segmenter is handed 山田太郎 rather 

338 than 山田太郎様. 

339 

340 Scanning back over post-nominals rather than taking the last token 

341 outright does three things at once. An unrelated trailing suffix 

342 cannot hide the peel site, so "김민준씨 Jr." answers as the 

343 comma-written "Dr 김민준씨, Jr." does -- one name, two spellings, 

344 one parse. A token that IS a tail (씨, さん, and the nested 선생님, 

345 which the cap alone would peel to 선생 + 님) is skipped as a site 

346 and stays whole: every tail is a suffix word by the Lexicon 

347 invariant, which is what makes that guard hold, and the cap below 

348 only keeps the offset interior for _split. And the scan answers 

349 None rather than indexing anything, which two reachable inputs 

350 need: a name that is nothing but post-nominals ("씨"), and one 

351 whose name runs are empty (", , 씨" scopes to two empty ones). 

352 Nothing here rests on a structural gate landing first, then, which 

353 is what let #312 move both of the stage's gates below it. 

354 

355 WHICH tokens are scanned is a separate decision, and the one a 

356 reader is likeliest to undo: the NAME's segment runs, flattened, 

357 and never state.tokens or every segment. The alternatives agree 

358 except where extract_delimited has already claimed a token or a 

359 comma has closed the name, both of which this stage can still see 

360 -- so "김민준씨 (Jimmy)" and "Dr 김민준씨, V." are the inputs that 

361 tell the three apart. See the note at the scan itself. 

362 

363 That last token is the last of the NAME, which reaches into a 

364 maiden clause: maiden tokens are still main-stream here 

365 (extract_delimited has masked only bracketed content), so 

366 "김민준 née 박씨" peels 씨 off the MAIDEN name 박씨 and hands it to 

367 the person's suffix list -- "née Ms. Park". Intended rather than 

368 incidental in that direction: the honorific is the reader's 

369 regardless of which of her names it was glued to, and a 

370 name-final honorific is exactly what this peels. Since #312 that 

371 reach extends to FAMILY_COMMA along with the rest of the crossing 

372 -- "김, 민준 née 박씨" now routes 씨 to suffix, where before #312 

373 the stage returned at the family-comma gate above the peel and 

374 left it in maiden "박씨". 

375 

376 The other direction is a LIMIT, stated rather than fixed: a maiden 

377 clause pushes the site off the person's own name, so "김민준씨 née 

378 박" does NOT peel and gives given "민준씨" -- the original bug, 

379 intact behind a marker. "김민준씨 née 박씨" shows both at once, two 

380 identical honorifics of which only the maiden's is routed. Chasing 

381 the marker into the site scan is scope creep for an uncommon input, 

382 and it could not be done here anyway: classify has not run, so the 

383 marker tokens carry no tag this stage could read. 

384 

385 Longest-first, and ONE peel: a remainder that itself ends in a 

386 listed tail is accepted rather than chased. No SHIPPED input 

387 witnesses that any more -- 박사님 was the last one, and adding it 

388 is what removed the case: 김민준박사님 now gives up 박사님 entire, 

389 since longest-first reaches the whole honorific. The pin needs a 

390 tail whose remainder ends in a DIFFERENT tail, the two not 

391 themselves a listed entry, and no pair in the shipped vocabulary 

392 has that shape; the stage test test_one_peel_never_a_stack carries 

393 it on a synthetic lexicon, which is now its only witness. No 

394 script precondition on the remainder either, since the tail alone 

395 is the license: Andersonさん peels. 

396 

397 Emits no ambiguity, unlike the surname fork below, though 

398 longest-first does CHOOSE here too -- 김선생님 gives 선생님 where 

399 님 also matches. The difference is what the runner-up is: a second 

400 matching surname is a competing READING of the name, which a 

401 caller may prefer, while a shorter tail leaves a remainder that is 

402 not a name at all (김선생), so there is nothing to adjudicate.""" 

403 tails = state.lexicon.honorific_tails 

404 if not tails: 

405 return state 

406 # The NAME's runs, which is more than segments[0] but never every 

407 # segment. A comma is how a writer says which runs are the name: a 

408 # FAMILY comma splits the name itself across two of them ("김, 

409 # 민준씨" is (0,) and (1,)), and the honorific is as often glued to 

410 # the given name as to the family, so the peel has to cross that 

411 # boundary (#312). Every OTHER structure keeps the whole name in 

412 # segments[0], so no boundary has to be crossed to find the site 

413 # there -- which is why the asymmetry is the point rather than an 

414 # accident of two cases. Note that "the rest is post-nominals" is 

415 # true of the SUFFIX comma only: under NO_COMMA segment returns 

416 # exactly one run and any trailing post-nominal is INSIDE it 

417 # ("김민준씨 Jr." is one run of two tokens), which is why the 

418 # scan-back above steps over such a token rather than simply never 

419 # reaching it. 

420 # The second run is only NAME text when segment read it as one, 

421 # which the structure alone does not say: SUFFIX_COMMA also wants 

422 # more than one word before the comma, so a one-word part turns a 

423 # wholly suffix-shaped remainder into FAMILY_COMMA anyway ("田中さん, 

424 # V." is that input). So ask segment's own predicate instead of 

425 # inferring the answer from the structure it produced (#319). 

426 # is_wholly_suffix, NOT the plural of _is_post_nominal: the two 

427 # disagree on the initial-shaped suffix words ("V.", "V", "I"), 

428 # which is the class #319 was reported about, and on everything 

429 # the run predicate's extra routes reach and the token predicate 

430 # does not ("Msc.Ed." and "J.씨" by period_joined_vocab, both of 

431 # which this change also moves). Reaching into such a 

432 # run put the site on "V.", which ends in no listed tail, so the 

433 # peel silently abandoned and さん stayed glued to the family -- 

434 # while "田中さん, PhD" peeled all along, because "PhD" satisfies 

435 # the strict test and the scan-back stepped over it. One credential, 

436 # two spellings, two answers FROM THE PEEL -- and the peel's answer 

437 # is the only one that moved: the peeled remainder is 田中 and lands 

438 # in family under every spelling, while where the CREDENTIAL lands 

439 # is assign's question and still differs ("PhD" a title, "V." a 

440 # given, "Ph. D." a suffix beside さん). 

441 # A junk tail is the worse shape of the same reach: in 

442 # "김민준씨, J.씨" the site lands on the junk "J.씨", so master 

443 # peeled THAT 씨 and left the person's own glued inside family 

444 # "김민준씨". Reachable only where the run is genuinely 

445 # suffix-shaped, which "J.씨" is by period_joined_vocab; a run of 

446 # ordinary name text is scanned on purpose, and a junk tail further 

447 # out than the second run is held off by the scope rule instead 

448 # ("Dr 김민준씨, Jr., 박씨" is SUFFIX_COMMA with 박씨 in a third 

449 # run, so it never reaches here at all). 

450 # An EMPTY second run stays in scope and contributes nothing: 

451 # is_wholly_suffix is False on it by its own contract (v1 read 

452 # "Doe,, Jr." as a family comma), which is the reading this line 

453 # wants anyway -- flattening an empty run adds no site. 

454 # Policy(lenient_comma_suffixes=False) keeps the old answer for the 

455 # INITIAL-shaped suffixes specifically, which is where the 

456 # strict/lenient gap lives: the knob drops this call to the strict 

457 # predicate too, so is_wholly_suffix(["V."]) is False, the run reads 

458 # as name text, it IS scanned, and the peel is abandoned on "V." as 

459 # before -- family "田中さん", given "V.". It is not a blanket 

460 # freeze of the old behavior, and "田中さん, Ph. D." is the input 

461 # that shows the difference: the Ph./D. merge folds that pair to a 

462 # form is_suffix_strict accepts, so the run is declined and the peel 

463 # fires under the strict knob as well. 

464 # Flattening the SEGMENTS rather than state.tokens is load-bearing 

465 # too: extracted nickname and maiden content is in tokens but in NO 

466 # segment, and scanning tokens would put the peel site on a 

467 # nickname ("김민준씨 (Jimmy)" -> the site becomes Jimmy and 

468 # nothing peels). ko_honorific_glued_given_nickname pins that. 

469 # And declining takes a SECOND condition: segments[0] must hold a 

470 # peel site of its own. is_wholly_suffix reaches period_joined_vocab, 

471 # which calls a run suffix-shaped when ANY period-chunk is suffix 

472 # VOCABULARY -- and every honorific tail is a suffix word by the 

473 # Lexicon invariant, so a glued honorific is itself the evidence. 

474 # The predicate is circular at THIS call site alone -- not because 

475 # segment asks it any earlier (nothing has peeled at either call) 

476 # but because segment SPENDS the answer differently: it reads the 

477 # run's shape and stops, the answer being the structure, while the 

478 # peel reads the same shape and then decides whether to go strip 

479 # the very honorific that produced it. "이, J.씨" reads as wholly suffix 

480 # only because of the 씨 the peel exists to remove, and declining a 

481 # run that holds the only site does not fall back to some other 

482 # site -- it loses the peel outright, and with it the given name, 

483 # which lands in suffix as "J.씨". Asking for a site in segments[0] 

484 # keeps the #319 answer wherever the peel has somewhere else to go 

485 # ("田中さん, V." still declines, さん is right there) and gives the 

486 # circular case back to master's reading. The two-honorific input 

487 # "김민준씨, J.씨" is where the choice is visible and deliberate: 

488 # both runs offer a site, so the decline stands and the person's own 

489 # 씨 is peeled rather than the junk one behind the comma. 

490 runs = state.segments[:1] 

491 if state.structure is Structure.FAMILY_COMMA: 

492 second = [state.tokens[j].text for j in state.segments[1]] 

493 # `one_case` is passed for the reason the field exists: the 

494 # credential lean is part of what "wholly suffix" MEANS since 

495 # #289, and a stage that asks the question without it gets a 

496 # different answer from `segment`, which asked it with the fact 

497 # in hand one stage earlier. Left out, 'Kim김민준씨, MA' read 

498 # 'MA' as name material, declined nothing, and scanned the 

499 # second run -- where the site is 'MA' itself, which carries no 

500 # listed tail -- so the person's own 씨 went unpeeled (family 

501 # 'Kim김민준씨') while 'Kim김민준씨, PhD' peeled it, one name in 

502 # two spellings parsed two ways (review round, #289/#516). 

503 # `segment` records the fact only where a comma form could turn 

504 # ON it, so this read is None for every other name and the 

505 # predicate then behaves exactly as it did before 2.4. 

506 if not (is_wholly_suffix(second, state.lexicon, state.policy, 

507 one_case=state.one_case) 

508 and _peel_site(state, state.segments[0], tails)): 

509 runs = state.segments[:2] 

510 site = _peel_site(state, [j for seg in runs for j in seg], tails) 

511 if site is None: 

512 return state 

513 i, offset = site 

514 # The tail carries a tag because this stage MANUFACTURED it. The 

515 # segmenter's neighbour test below needs to tell it from a token 

516 # somebody wrote, and no vocabulary question can: the two spellings 

517 # put the same word in the same place, and only the provenance 

518 # differs. 

519 return _split(state, i, (offset,), None, tail_tag=_PEELED_TAG) 

520 

521 

522def _split_surname_site(state: ParseState) -> ParseState: 

523 """The stage's other split: the first activated-script token of the 

524 name part is matched longest-first against Lexicon.surnames, and a 

525 hit splits it in two; where the vocabulary declines, an optional 

526 Parser(segmenter=...) is consulted instead. 

527 

528 A sibling of _peel_honorific_tail rather than a continuation of it. 

529 The two answer different questions -- this one asks where a name 

530 divides into surname and given, the peel asks whether a token ends 

531 in a word that can never end a name -- which is why they carry 

532 different gates: the FAMILY comma and the 间隔号 gate this half 

533 alone (#312), and segment_scripts below gates it alone too. That 

534 is also why the stage entry below interleaves gates with its two 

535 calls rather than running one cascade.""" 

536 scripts = state.policy.segment_scripts 

537 # an empty VOCABULARY deliberately does not bail here -- see below 

538 if not scripts: 

539 return state 

540 # segments[0] is the NAME part under both remaining structures 

541 # (everything, under NO_COMMA); later segments are suffixes. Its 

542 # members are main-stream token indices by construction, so 

543 # extracted nickname/maiden content is unreachable from here. For 

544 # this site that is merely tidy -- the first script-written token 

545 # is the same either way. It is the PEEL above that reads this run 

546 # load-bearingly, and in the two structures that reach here it 

547 # reads exactly this one, only backwards; the argument lives at 

548 # that scan. 

549 i = next((i for i in state.segments[0] 

550 if effective_script(state.tokens[i].text) in scripts), None) 

551 # A post-nominal in the surname's own position is not a site to 

552 # skip past but an answer: a surname LEADS, so if the leading 

553 # script-written token is an honorific there is no surname here to 

554 # find. Scanning ON would reach the given name, which is exactly 

555 # what the first-token rule exists to prevent -- 지 is a listed 

556 # surname, so "양 지훈" (양 is a surname AND a shipped honorific) 

557 # would have its own given name split in half. Declining also 

558 # covers the token the peel above manufactures, which is the first 

559 # and only script-written one in "Anderson선생님". 

560 if i is None or _is_post_nominal(state, i): 

561 return state 

562 token = state.tokens[i] 

563 text = token.text 

564 surnames = state.lexicon.surnames 

565 # The vocabulary match reads the token's CORE (#323). Without it 

566 # the stop WAS the remainder: '김.' matched its own head and became 

567 # 김 + '.'; the stop rides with the remainder instead, so '김민준.' 

568 # divides as 김 + '민준.'. A leading stop never reaches THIS site -- 

569 # the classification fold in front rstrips too 

570 # (_vocab._normalized_for_script) and hides such a token before it 

571 # arrives -- while the peel above has no such gate in front of it 

572 # and its own rstrip is what does the work there. 

573 core = text.rstrip(FULL_STOPS) 

574 # A token that IS a surname never splits: a bare "남궁" must not 

575 # become 남 + 궁 just because the single-syllable surname also 

576 # matches -- there is nothing to split off, and a lone token's 

577 # role is the order resolution's call, stop or no stop. 

578 if core in surnames: 

579 return state 

580 # Longest-first (compound-before-single falls out of it), capped 

581 # so the remainder is never empty. Direct membership, no 

582 # _normalize: the script gate admits only CJK text, and a CJK 

583 # entry's stored form is its NFC composition with no case or 

584 # edge-stop change (#322) -- so an entry AUTHORED in NFD is 

585 # composed on the way in, raw NFD input matches nothing here, and 

586 # the name goes unsplit, which is the no-split rules.md#W1 already 

587 # accepts for NFD text. An empty vocabulary skips the match rather 

588 # than bailing the stage (_longest_entry's max() has nothing to 

589 # take): a surname-less lexicon declines every token, which is 

590 # exactly the condition the segmenter is consulted on, so an early 

591 # bail would make a configured segmenter silently inert under 

592 # Lexicon.empty() -- the JA pack's own shape. 

593 matches: list[int] = [] 

594 if surnames: 

595 cap = min(_longest_entry(surnames), len(core) - 1) 

596 matches = [length for length in range(cap, 0, -1) 

597 if core[:length] in surnames] 

598 if matches: 

599 take = matches[0] 

600 detail = None 

601 if len(matches) > 1: 

602 # more than one vocabulary-supported split: longest-first 

603 # DECIDED a fork, and the deciding stage records it. A 

604 # single-match split chose nothing, so it stays silent -- 

605 # a dictionary certainty and the statistical guess below 

606 # are different epistemic states, and these two emission 

607 # rules say so. 

608 chosen = " + ".join(map(repr, _pieces(text, (take,)))) 

609 other = " + ".join(map(repr, _pieces(text, (matches[1],)))) 

610 detail = (f"{text!r} splits as {chosen} on the longest " 

611 f"surname; {other} also reads") 

612 return _split(state, i, (take,), detail) 

613 if state.segmenter is None: 

614 return state # the first activated-script token decides 

615 # The segmenter's precondition, which the vocabulary has no twin of: 

616 # it is asked where an UNDIVIDED name divides, so it may only be 

617 # shown a token that is the whole name. Where the name part carries 

618 # a second script-written token the writer already drew this 

619 # stage's missing boundary -- "山田 太郎" is divided, and dividing 

620 # its family again yields 山 + 田 + 太郎 (namedivider answers for 

621 # any string, and scores a two-character one 1.0 by rule, so no 

622 # confidence check would catch it). The neighbour counts whatever 

623 # script it is written in, ACTIVATED or not -- effective_script is 

624 # merely non-None -- because a katakana or hangul neighbour is a 

625 # boundary its writer drew just as deliberately as a Han one: under 

626 # the JA pack "山田太郎 マイケル" declines, though katakana is in no 

627 # activation set. A Latin title or suffix is NOT such a boundary: 

628 # it says nothing about where the CJK name splits, so "Dr 阿明日, 

629 # Jr." still reaches the segmenter. Vocabulary keeps its own rule 

630 # -- a listed surname is a certainty about that exact string, 

631 # whoever else stands beside it. 

632 # The one neighbour that does not count is the one this stage 

633 # MANUFACTURED (#308): a glued 山田太郎様 has no writer-drawn 

634 # boundary anywhere, so the 様 the peel just cut off cannot be read 

635 # as one -- the precondition must see what the writer wrote, which 

636 # was a single undivided token. A SPACED 様 does count, and the 

637 # reason is weaker than "its writer drew that boundary and chose 

638 # to write 山田太郎 as a unit": in "山田太郎 様" the unit the writer 

639 # drew is the WHOLE NAME, honorific and all, so that story is false 

640 # for the very input it describes. What the code relies on is that 

641 # by POSITION a spaced honorific is indistinguishable from a spaced 

642 # name element, so the test conservatively counts it. Measured, the 

643 # trade is worth it: counting them keeps 佐藤 氏, 田中 様, 鈴木 先生 

644 # and 中村 教授 whole -- all four divide bare under the JA pack (佐 

645 # + 藤, 田 + 中, 鈴 + 木, 中 + 村) -- and costs the single division 

646 # 山田太郎 様. Four real surnames against one. 

647 # Provenance, not vocabulary: the two spellings put the same word 

648 # in the same place, and asking the suffix set instead cannot 

649 # separate them. 

650 if any(j != i and effective_script(state.tokens[j].text) is not None 

651 and _PEELED_TAG not in state.tokens[j].tags 

652 for j in state.segments[0]): 

653 return state 

654 # No try/except around the call: rules.md#A1's Accepted clause 

655 # ("a user-supplied segmenter's own error propagates"). The two 

656 # checks below are that same doctrine, curated, 

657 # and they are where the line this module draws is easiest to state: 

658 # a PROTOCOL VIOLATION BY THE SEGMENTER AUTHOR RAISES, while an 

659 # ADAPTER'S DEFENSE AGAINST ITS LIBRARY DECLINES. Both checks here 

660 # are the first kind -- a wrong answer TYPE and an answer indexing 

661 # past the token it was handed are stage-detectable bugs in 

662 # user-supplied code, inside the declared totality exception, so 

663 # they get the same treatment the callable's own exceptions get. 

664 # locales/ja.py is the second kind: its repertoire, length, 

665 # reconstruction and score guards all return None, because what 

666 # they defend against is namedivider answering a question nobody 

667 # asked it, which is a fact about the CONTENT, not a broken 

668 # protocol. Bounded like every message here: the type's NAME, never 

669 # its contents. 

670 # The CORE, not the raw token (#323): handed the raw token, a 

671 # segmenter answering len-1 would make the stop a piece of its own 

672 # ('田中太郎' + '。'), where against the core the trailing stops 

673 # ride with the last piece -- '山田太郎.' answered at 2 divides as 

674 # 山田 + '太郎.'. 

675 answer = state.segmenter(core) 

676 if answer is not None and not isinstance(answer, Segmentation): 

677 # a duck-typed answer carrying a .splits of its own would 

678 # otherwise wander into the split path and surface as a 

679 # ValueError naming Token, pointing the reader at nameparser's 

680 # insides instead of at their segmenter 

681 raise TypeError( 

682 f"segmenter must return Segmentation or None, got " 

683 f"{type(answer).__name__}") 

684 if answer is None or not answer.splits: 

685 return state # declined, or confidently one token 

686 # splits[-1] is the max -- Segmentation enforced ascending. The 

687 # upper bound is the half Segmentation cannot check, since it never 

688 # sees the text; an offset at or past the end would make an empty 

689 # piece. Declining silently here (as this did before the review) 

690 # made an off-by-one segmenter undebuggable: every answer it gave 

691 # vanished, and the parse merely looked unsegmented. 

692 # The bound is the CORE's length, the string the segmenter was 

693 # actually handed: an offset it could not have derived from its own 

694 # input is the same author bug whether or not `text` is longer. 

695 if answer.splits[-1] >= len(core): 

696 raise ValueError( 

697 f"segmenter returned splits beyond the token: last offset " 

698 f"{answer.splits[-1]}, the segmenter was given " 

699 f"{len(core)} characters") 

700 conf = answer.confidence 

701 detail = None 

702 if conf is not None and conf < _SEGMENTER_CONFIDENCE_FLOOR: 

703 reading = " + ".join(map(repr, _pieces(text, answer.splits))) 

704 detail = (f"{text!r} splits as {reading} on a segmenter answer " 

705 f"scoring {conf:.2f}, under the " 

706 f"{_SEGMENTER_CONFIDENCE_FLOOR} confidence floor") 

707 return _split(state, i, answer.splits, detail) 

708 

709 

710# rules.md#W1: "an undivided word in the family position of a name 

711# written in an activated script divides after a recognized surname, 

712# the longest recognized surname first; where the vocabulary 

713# recognizes nothing, an optional segmenter may divide instead" 

714# rules.md#W3: "under a family comma the pre-comma text is the 

715# family by declaration and never divides, and the post-comma side 

716# is given text with no family to find" (history: decisions.md#W3) 

717def script_segment(state: ParseState) -> ParseState: 

718 if state.original.isascii(): 

719 # spans index the original exactly (the anti-#100 invariant), 

720 # so an ASCII original has only ASCII tokens: nothing here is 

721 # in any script's ranges. It also short-circuits the PEEL, 

722 # which has no script gate of its own -- so a caller-configured 

723 # ASCII tail never fires, and one non-ASCII character anywhere 

724 # in the name switches it on. Correct for the CJK vocabulary 

725 # that ships, stated because honorific_tails is public: see 

726 # that field's own note. Latin orthography SPACES its 

727 # post-nominals ("John Smith PhD"), so the glued position this 

728 # peel exists for is a CJK one to begin with. What Latin does 

729 # glue is PREnominal and period-joined (Mr.Smith, Lt.Gov.), 

730 # which _vocab.period_joined_vocab classifies as one token 

731 # rather than splitting -- a different mechanism, and no peel 

732 # site either way. A glued Latin POST-nominal is spelled the 

733 # same way, so it reaches that same mechanism rather than 

734 # nothing: period_joined_vocab reads "Smith.Jr." as a title 

735 # ('jr' is title vocabulary as well as suffix vocabulary), and 

736 # where position allows, that wins -- "Smith.Jr. Anderson" 

737 # gives title "Smith.Jr.", family "Anderson". Whatever it 

738 # decides, this bail is what settles the question here: it 

739 # returns above all of it, so no ASCII input reaches the peel. 

740 return state 

741 if not state.segments: 

742 return state 

743 # #312: the peel runs in front of both gates below, because both 

744 # answer where a name divides into surname and given, and the peel 

745 # does not ask that. It asks whether a token ends in a word that 

746 # can never end a name, and a comma or a dot elsewhere in the 

747 # string does not change the answer. Placing it above the block 

748 # also keeps a future segmentation gate from silently capturing 

749 # it: a new gate lands with its siblings, below. 

750 state = _peel_honorific_tail(state) 

751 if state.structure is Structure.FAMILY_COMMA: 

752 return state # the comma already drew the SURNAME boundary 

753 if state.interpunct_offsets: 

754 # #298: a 间隔号-divided name is a transcription -- its pieces 

755 # are syllable groups, not surname+given, so neither the 

756 # vocabulary nor the segmenter applies (codepoint-scoped: the 

757 # nakaguro records nothing and gates nothing, spec decision 5). 

758 # State-global like the FAMILY_COMMA gate above: a marker 

759 # anywhere in the name reads the WHOLE name as a transcription 

760 # listing, so even an un-dotted hangul token beside a dotted 

761 # one stays whole. Scoped to the surname split since #312: an 

762 # honorific glued to a transcription is still an honorific. 

763 return state 

764 return _split_surname_site(state)