1"""Stage: script_segment (#271, #272, #308, #312).
2
3Consumes: tokens, segments, structure, interpunct_offsets, segmenter,
4one_case (where segment recorded it; None everywhere else, which is
5what this stage's suffix-run predicate read for every name before
62.4).
7Produces: tokens, by two independent splits into sub-slices -- a
8listed honorific peeled off the END of the name's last
9non-post-nominal token, in whichever of the name's runs that falls
10(the tail token also carries this module's _PEELED_TAG), and the
11first activated-script token of the name segment split into n+1 --
12segments (index runs remapped past the insertions), ambiguities
13(indices likewise remapped, plus a SEGMENTATION report when more than
14one split was vocabulary-supported, or when a segmenter's answer
15scored under the confidence floor).
16Reads: Policy.segment_scripts, Lexicon.surnames,
17Lexicon.honorific_tails, ParseState.segmenter, and Lexicon suffix
18vocabulary through TWO predicates, which are NOT one another's
19singular and plural. _vocab.is_suffix_strict asks whether a single
20token is a post-nominal, initial veto included (the peel's scan-back
21and the surname site). _vocab.is_wholly_suffix asks segment's own
22suffix-comma question of a whole RUN, through the POLICY-selected
23token test plus period_joined_vocab, delimiter handling and the
24Ph./D. merge -- so the run predicate says yes both to tokens the
25token predicate VETOES ("V.", "V", "I") and to tokens it never sees
26as suffixes at all ("Msc.Ed.", "J.씨", which reach it through
27period_joined_vocab). The initial-shaped words are the class #319 was
28ABOUT, not the whole of the disagreement, and the extra routes listed
29above are where the rest of it comes from;
30test_is_wholly_suffix_is_not_the_plural_of_is_post_nominal pins the
31"V." half. The run predicate is how the peel declines a run that is
32not name text (#319), but only where the name's own run offers a
33site, since a glued honorific is itself part of what makes a run read
34as suffix-shaped; it owns two further Policy fields
35(lenient_comma_suffixes picks the strict or lenient token test -- it
36flips "田中さん, V." -- and extra_suffix_delimiters both counts a bare
37delimiter-core token as a suffix and splits a token on a core, the
38split being the half that flips "田中さん, Jr./V." under {"/"} and the
39bare core the half that flips "田中さん, /").
40
41Implements rules W1 (the vocabulary/segmenter division), W2 (the
42glued-honorific peel) and W3 (the writer's divisions are respected)
43of docs/design/rules.md, cited at their code below; the decision
44chain (#308, #312, #319, the vetting bars, the measured
45spaced-honorific trade) is decisions.md#W1, #W2 and #W3. Both splits
46make sub-slices of one token, rewriting nothing -- spans still index
47the original exactly, so the anti-#100 invariant holds by
48construction. Both also match on the token's CORE, its
49text.rstrip(FULL_STOPS) (#323) -- rstrip and never strip, because
50both offsets are measured from the text's START, so the core has to
51stay a prefix of the text for the offset the match yields to land
52where the match did. The peel runs FIRST, so suffix classification
53can claim the tail and the surname match or segmenter consult sees
54the name rather than name-plus-honorific; its
55ASCII bail sits above everything here, so a caller-added LATIN tail
56fires only on a name carrying at least one non-ASCII character (see
57the bail's own comment, and honorific_tails' field note).
58"""
59from __future__ import annotations
60
61import dataclasses
62import functools
63from collections.abc import Sequence
64
65from nameparser._lexicon import FULL_STOPS
66from nameparser._pipeline._state import (
67 ParseState, PendingAmbiguity, Structure, WorkToken,
68)
69from nameparser._pipeline._vocab import (
70 effective_script, is_suffix_strict, is_wholly_suffix,
71)
72from nameparser._types import AmbiguityKind, Segmentation, Span
73
74#: Marks the tail token the peel below MANUFACTURED, so the segmenter's
75#: neighbour test can tell it from a token somebody wrote. Namespaced,
76#: therefore unstable provenance rather than API (_types.STABLE_TAGS is
77#: the whole stable set, and FOLDED_TAG is the precedent for a
78#: structural marker carrying this prefix); it is vocabulary-derived
79#: besides, since honorific_tails is what licensed the split. Emitter
80#: and reader are both in this module, so unlike FOLDED_TAG it needs no
81#: home in _types.
82_PEELED_TAG = "vocab:peeled-honorific"
83
84#: Segmenter answers scoring below this attach a SEGMENTATION report
85#: (amendment 2026-07-29 section 3). Kept at the drafted 0.9 after
86#: measuring namedivider 0.4.1 over 112 names (#272 Task 5), which
87#: found a distribution the amendment did not anticipate: the scores
88#: are BIMODAL, not clustered near 1. A rule-based division (the kana
89#: boundary in 高橋みなみ, or a two-character name) scores exactly 1.0;
90#: a kanji-statistics division scores a softmax over the candidate cut
91#: positions, observed in 0.23-0.68 and driven more by name LENGTH
92#: (median 0.61 at three characters, 0.37 at five -- the softmax is
93#: over len-1 candidates) than by correctness: the wrong answers in
94#: the sample scored 0.32/0.60/0.60/0.61, straddling the correct
95#: median of 0.52. So no floor INSIDE that band separates error from
96#: success, and the only cut the data supports is between a stated
97#: certainty and a statistical guess -- which is what section 3's
98#: epistemic argument asks for anyway. 0.9 sits in the empty gap
99#: (0.68 to 1.0), far from both modes, and reads as "confident" for a
100#: third-party segmenter with a calibrated score too. Consequence,
101#: pinned by tests and stated so it can be checked: every
102#: STATISTICALLY divided name carries a SEGMENTATION report, and every
103#: RULE-divided one -- the kana boundary, a two-character name,
104#: namedivider's specific-name rules -- carries none.
105#: namedivider scores its own two-character rule 1.0, so this floor
106#: keeps that division silent -- see locales/ja.py for why the
107#: presumption is accepted as stated.
108#: Not configurable in 2.x (YAGNI).
109_SEGMENTER_CONFIDENCE_FLOOR = 0.9
110
111
112def _remap(run: tuple[int, ...], split_at: int,
113 added: int) -> tuple[int, ...]:
114 """One segment run after token `split_at` split into `added` + 1
115 pieces: the extra pieces join the run immediately after it, and
116 every later index shifts by as many."""
117 out: list[int] = []
118 for j in run:
119 if j == split_at:
120 out.extend(range(split_at, split_at + added + 1))
121 else:
122 out.append(j + added if j > split_at else j)
123 return tuple(out)
124
125
126def _pieces(text: str, splits: tuple[int, ...]) -> tuple[str, ...]:
127 """`text` cut at every offset in `splits`: n offsets, n+1 pieces."""
128 cuts = (0, *splits, len(text))
129 return tuple(text[a:b] for a, b in zip(cuts, cuts[1:]))
130
131
132def _split(state: ParseState, i: int, splits: tuple[int, ...],
133 detail: str | None, tail_tag: str | None = None) -> ParseState:
134 """Cut token `i` at every offset in `splits`, recording `detail` as
135 a SEGMENTATION report when there is one, and adding `tail_tag` to
136 the LAST piece when there is one of those.
137
138 The offsets arrive non-empty, ascending and interior whatever chose
139 them. From a segmenter: non-empty is the caller's own check,
140 strictly ascending with each >= 1 is Segmentation.__post_init__'s,
141 and the last offset is < len(text) -- checked by the caller. From
142 the vocabulary: the single offset is >= 1 and < len(text) by the
143 range(cap, 0, -1) construction and its len-1 cap.
144
145 The ONE split path: the vocabulary hit is the single-offset case
146 and the segmenter's answer the general one, so neither can drift
147 from the other's index arithmetic. `tail_tag` is part of keeping it
148 one path: the piece is tagged HERE, where it is built, rather than
149 by the caller, which would have to re-derive where its own tail
150 landed after handing the index arithmetic off. Only the peel passes
151 one, and the piece it wants is the last by construction -- the
152 honorific it cut off the end; the surname path passes nothing and
153 the default leaves it exactly as it was."""
154 token = state.tokens[i]
155 base = token.span.start
156 parts: list[WorkToken] = []
157 start = 0
158 for piece in _pieces(token.text, splits):
159 end = start + len(piece)
160 parts.append(dataclasses.replace(
161 token, text=piece, span=Span(base + start, base + end)))
162 start = end
163 if tail_tag is not None:
164 parts[-1] = dataclasses.replace(
165 parts[-1], tags=parts[-1].tags | {tail_tag})
166 added = len(splits)
167 tokens = state.tokens[:i] + tuple(parts) + state.tokens[i + 1:]
168 # Every index the earlier stages recorded is now stale past the
169 # split point: the segment runs group's own iteration rests on, and
170 # the ambiguities from extract_delimited (resolved to indices by
171 # tokenize) and segment. An ambiguity ON the split token keeps
172 # pointing at the head.
173 segments = tuple(_remap(run, i, added) for run in state.segments)
174 ambiguities = tuple(
175 dataclasses.replace(a, indices=tuple(
176 j + added if j > i else j for j in a.indices))
177 for a in state.ambiguities)
178 if detail is not None:
179 ambiguities += (PendingAmbiguity(
180 AmbiguityKind.SEGMENTATION, detail,
181 tuple(range(i, i + added + 1))),)
182 return dataclasses.replace(state, tokens=tokens, segments=segments,
183 ambiguities=ambiguities)
184
185
186@functools.lru_cache(maxsize=16)
187def _longest_entry(entries: frozenset[str]) -> int:
188 """The longest entry in a vocabulary, cached per-vocabulary rather
189 than recomputed per parse (Lexicon is frozen and slotted, so it
190 cannot carry a cached_property of its own). The frozenset is
191 hashable and a process holds only a handful of distinct
192 vocabularies -- the default one, plus one per constructed pack
193 parser -- so maxsize=16 bounds pathological many-lexicon churn
194 without ever evicting in normal use. Keyed by the frozenset VALUE,
195 not by (lexicon, field): the two callers pass surnames and
196 honorific_tails, but every lexicon that leaves KOREAN_SURNAMES
197 alone shares one entry, so the count is distinct SETS rather than
198 lexicons times fields.
199
200 Callers must pass a NON-EMPTY vocabulary: max() of an empty set
201 raises, and both call sites sit under a match guard that cannot
202 reach it with one."""
203 return max(map(len, entries))
204
205
206def _is_post_nominal(state: ParseState, i: int) -> bool:
207 """Whether token `i` is post-nominal VOCABULARY. Two of this
208 stage's decisions ask it and neither is about script: an honorific
209 is neither a surname SITE nor the token a glued honorific hangs off
210 (#308).
211
212 It answers a VOCABULARY question, not a positional one, and the two
213 callers do not spend the answer the same way. In TRAILING position
214 vocabulary and position agree, so the peel site steps over a True
215 and goes on scanning. In LEADING position they disagree -- 양 is
216 the family name there, whatever the suffix set says -- so the
217 surname site reads a True as an ANSWER and declines, rather than as
218 a token to step past.
219
220 STRICT, not lenient: the initial veto applies, so "V." is not a
221 post-nominal here though bare "v" is a suffix word. That agrees
222 with what classify does with the same token downstream -- "V." is
223 a middle initial -- and the difference is reachable under the
224 default lexicon: "田中さん V." stops its scan-back at the initial
225 and does not peel, while "田中さん II" steps over II and does. That
226 pair is the case table's ja_honorific_glued_before_an_initial and
227 ja_honorific_glued_before_a_roman_suffix, added because the clause
228 could NAME the discriminating input while swapping in
229 is_suffix_lenient still failed no test -- the shape a "unify the
230 two suffix predicates" refactor would have walked straight past.
231
232 Suffixes only, deliberately. Should a CJK entry ever join titles,
233 the surname site would want it excluded while the peel site must
234 not: a title never trails, so widening the peel's scan-back would
235 move it off the real last name token. Split the predicate then, not
236 before."""
237 return is_suffix_strict(state.tokens[i].text, state.lexicon)
238
239
240# rules.md#W2: "A part that is not name text — a post-nominal word
241# standing on its own — is never the name's end: the split-off steps
242# past it to the name word behind, and never dissects it." (history:
243# decisions.md#W2)
244# The scan-back below is that clause, and it needs no punctuation to
245# fire. '김민준 박사님' steps past 박사님 to 김민준, finds no listed
246# tail there and returns None -- so 박사님 is left whole for suffix
247# classification rather than cut into 박사 + 님, which is what the
248# same input gives when the step is removed. '선생님' is post-nominal
249# entire with no name word behind it, so the scan yields no site at
250# all and the token stays whole. Both are contract-tier corpus names
251# and rules.md#W2 example lines; decisions.md#cjk-comma-demotion
252# carries the forced-predicate measurement behind them.
253def _peel_site(state: ParseState, flat: Sequence[int],
254 tails: frozenset[str]) -> tuple[int, int] | None:
255 """Where a peel would land in the token run `flat`: the index of the
256 token to cut and the OFFSET to cut it at (the listed tail, and any
257 full stops riding behind it, are what comes off the end), or None
258 where that run offers no peel.
259
260 Two callers, one answer, which is the point of naming it. The peel
261 itself asks it once, of the runs it decided to scan. The gate above
262 that decision asks it of segments[0] ALONE, to find out whether
263 declining the second run would cost the only site. Sharing the scan
264 with the peel rather than approximating it there is what makes the
265 gate's answer mean what it says: a gate that merely looked for a
266 token ENDING in a listed tail would count a token that is a tail
267 entire, which the scan-back skips as a site -- "선생님, J.씨" would
268 decline on a site the peel then cannot use, and lose the peel in
269 exactly the way the gate exists to prevent. (Only the lone-token
270 shape of that divergence is reachable from the gate, and the gate's
271 own other conjunct is what makes that so: SUFFIX_COMMA wants BOTH
272 a suffix-shaped second run and more than one word before the
273 comma, and the gate is asked only where the first of those is
274 already true -- so a run this call ever sees under FAMILY_COMMA
275 failed on the second, and segments[0] holds at most one token. The
276 word-count alone would not say that: "Dr 김민준, 지훈" has two
277 words before the comma and is FAMILY_COMMA.)
278
279 Callers must pass a NON-EMPTY `tails` -- _longest_entry's
280 precondition, which the stage's own early return supplies."""
281 i = next((j for j in reversed(flat)
282 if not _is_post_nominal(state, j)), None)
283 if i is None:
284 # nothing but post-nominals, or no tokens at all
285 return None
286 text = state.tokens[i].text
287 # The tail is matched on the token's core (#323): a stop glued after
288 # the honorific ('김민준씨.', '田中さん.') stands between the listed
289 # tail and the token's end and used to defeat the match. What the
290 # text read instead depended on the stop: the ASCII spelling went to
291 # a title downstream (H2's shape is ASCII-period-only), while a
292 # fullwidth or ideographic stop left the whole text a lone name word
293 # ('田中さん。' read given). The cut lands BEFORE the tail, so the
294 # stops ride with the honorific piece and the token text is never
295 # rewritten. TRAILING only: a leading stop is not between the name
296 # and its honorific, and the offset returned below is the core's
297 # length less the tail's, one subtraction.
298 core = text.rstrip(FULL_STOPS)
299 # range/cap construction identical to the surname match below, and
300 # for the same two reasons: longest-first, and a len-1 cap that
301 # makes the offset interior by construction (_split's contract). An
302 # empty or one-character core caps below 1 and the loop is empty.
303 cap = min(_longest_entry(tails), len(core) - 1)
304 for length in range(cap, 0, -1):
305 if core[-length:] in tails:
306 return i, len(core) - length
307 return None
308
309
310# rules.md#W2: "a listed honorific glued to the end of the name's
311# last name word splits off once and reads as a suffix." (history:
312# decisions.md#W2)
313# That the peel also reaches ACROSS a family comma is stated at
314# rules.md#W3 instead, which is a tolerated rule since the
315# 2026-09-01 comma demotion -- the crossing is what the parser does
316# today, not something W2 promises. W3 also carries the reading of a
317# period a listing leaves behind (decisions.md#cjk-comma-demotion,
318# amended by decisions.md#cjk-full-stops): a period on a SEPARATE
319# post-nominal word rides into the suffix and moves nothing ('様.'
320# is post-nominal-strict and the scan steps past it as it steps past
321# '様'), and since #323 a period glued to the honorific's OWN token
322# rides with the honorific too -- _peel_site matches the tail through
323# the trailing stop and cuts before it, so '田中さん.' divides where
324# '田中さん' does, and '김민준씨.' where '김민준씨' does (both are
325# case rows now; before 2026-09-10 they read as a title, measured
326# 2026-09-05 and pinned by nothing, and for one commit of the #323
327# branch the surname site read '김민준씨.' as 김 + 민준씨.). Neither
328# reading is a promise; the step past the post-nominal word itself
329# is W2's, above, and is.
330def _peel_honorific_tail(state: ParseState) -> ParseState:
331 """#308: split a listed honorific off the END of the name's last
332 NON-POST-NOMINAL token -- 田中さん -> 田中 + さん -- and let
333 the existing machinery do the rest. Suffix classification claims
334 the tail downstream (every honorific_tails entry is a suffix word
335 too, enforced by Lexicon), and the segmentation half below then
336 sees the remainder rather than the glued whole, so 김민준씨 splits
337 김 + 민준 and a configured segmenter is handed 山田太郎 rather
338 than 山田太郎様.
339
340 Scanning back over post-nominals rather than taking the last token
341 outright does three things at once. An unrelated trailing suffix
342 cannot hide the peel site, so "김민준씨 Jr." answers as the
343 comma-written "Dr 김민준씨, Jr." does -- one name, two spellings,
344 one parse. A token that IS a tail (씨, さん, and the nested 선생님,
345 which the cap alone would peel to 선생 + 님) is skipped as a site
346 and stays whole: every tail is a suffix word by the Lexicon
347 invariant, which is what makes that guard hold, and the cap below
348 only keeps the offset interior for _split. And the scan answers
349 None rather than indexing anything, which two reachable inputs
350 need: a name that is nothing but post-nominals ("씨"), and one
351 whose name runs are empty (", , 씨" scopes to two empty ones).
352 Nothing here rests on a structural gate landing first, then, which
353 is what let #312 move both of the stage's gates below it.
354
355 WHICH tokens are scanned is a separate decision, and the one a
356 reader is likeliest to undo: the NAME's segment runs, flattened,
357 and never state.tokens or every segment. The alternatives agree
358 except where extract_delimited has already claimed a token or a
359 comma has closed the name, both of which this stage can still see
360 -- so "김민준씨 (Jimmy)" and "Dr 김민준씨, V." are the inputs that
361 tell the three apart. See the note at the scan itself.
362
363 That last token is the last of the NAME, which reaches into a
364 maiden clause: maiden tokens are still main-stream here
365 (extract_delimited has masked only bracketed content), so
366 "김민준 née 박씨" peels 씨 off the MAIDEN name 박씨 and hands it to
367 the person's suffix list -- "née Ms. Park". Intended rather than
368 incidental in that direction: the honorific is the reader's
369 regardless of which of her names it was glued to, and a
370 name-final honorific is exactly what this peels. Since #312 that
371 reach extends to FAMILY_COMMA along with the rest of the crossing
372 -- "김, 민준 née 박씨" now routes 씨 to suffix, where before #312
373 the stage returned at the family-comma gate above the peel and
374 left it in maiden "박씨".
375
376 The other direction is a LIMIT, stated rather than fixed: a maiden
377 clause pushes the site off the person's own name, so "김민준씨 née
378 박" does NOT peel and gives given "민준씨" -- the original bug,
379 intact behind a marker. "김민준씨 née 박씨" shows both at once, two
380 identical honorifics of which only the maiden's is routed. Chasing
381 the marker into the site scan is scope creep for an uncommon input,
382 and it could not be done here anyway: classify has not run, so the
383 marker tokens carry no tag this stage could read.
384
385 Longest-first, and ONE peel: a remainder that itself ends in a
386 listed tail is accepted rather than chased. No SHIPPED input
387 witnesses that any more -- 박사님 was the last one, and adding it
388 is what removed the case: 김민준박사님 now gives up 박사님 entire,
389 since longest-first reaches the whole honorific. The pin needs a
390 tail whose remainder ends in a DIFFERENT tail, the two not
391 themselves a listed entry, and no pair in the shipped vocabulary
392 has that shape; the stage test test_one_peel_never_a_stack carries
393 it on a synthetic lexicon, which is now its only witness. No
394 script precondition on the remainder either, since the tail alone
395 is the license: Andersonさん peels.
396
397 Emits no ambiguity, unlike the surname fork below, though
398 longest-first does CHOOSE here too -- 김선생님 gives 선생님 where
399 님 also matches. The difference is what the runner-up is: a second
400 matching surname is a competing READING of the name, which a
401 caller may prefer, while a shorter tail leaves a remainder that is
402 not a name at all (김선생), so there is nothing to adjudicate."""
403 tails = state.lexicon.honorific_tails
404 if not tails:
405 return state
406 # The NAME's runs, which is more than segments[0] but never every
407 # segment. A comma is how a writer says which runs are the name: a
408 # FAMILY comma splits the name itself across two of them ("김,
409 # 민준씨" is (0,) and (1,)), and the honorific is as often glued to
410 # the given name as to the family, so the peel has to cross that
411 # boundary (#312). Every OTHER structure keeps the whole name in
412 # segments[0], so no boundary has to be crossed to find the site
413 # there -- which is why the asymmetry is the point rather than an
414 # accident of two cases. Note that "the rest is post-nominals" is
415 # true of the SUFFIX comma only: under NO_COMMA segment returns
416 # exactly one run and any trailing post-nominal is INSIDE it
417 # ("김민준씨 Jr." is one run of two tokens), which is why the
418 # scan-back above steps over such a token rather than simply never
419 # reaching it.
420 # The second run is only NAME text when segment read it as one,
421 # which the structure alone does not say: SUFFIX_COMMA also wants
422 # more than one word before the comma, so a one-word part turns a
423 # wholly suffix-shaped remainder into FAMILY_COMMA anyway ("田中さん,
424 # V." is that input). So ask segment's own predicate instead of
425 # inferring the answer from the structure it produced (#319).
426 # is_wholly_suffix, NOT the plural of _is_post_nominal: the two
427 # disagree on the initial-shaped suffix words ("V.", "V", "I"),
428 # which is the class #319 was reported about, and on everything
429 # the run predicate's extra routes reach and the token predicate
430 # does not ("Msc.Ed." and "J.씨" by period_joined_vocab, both of
431 # which this change also moves). Reaching into such a
432 # run put the site on "V.", which ends in no listed tail, so the
433 # peel silently abandoned and さん stayed glued to the family --
434 # while "田中さん, PhD" peeled all along, because "PhD" satisfies
435 # the strict test and the scan-back stepped over it. One credential,
436 # two spellings, two answers FROM THE PEEL -- and the peel's answer
437 # is the only one that moved: the peeled remainder is 田中 and lands
438 # in family under every spelling, while where the CREDENTIAL lands
439 # is assign's question and still differs ("PhD" a title, "V." a
440 # given, "Ph. D." a suffix beside さん).
441 # A junk tail is the worse shape of the same reach: in
442 # "김민준씨, J.씨" the site lands on the junk "J.씨", so master
443 # peeled THAT 씨 and left the person's own glued inside family
444 # "김민준씨". Reachable only where the run is genuinely
445 # suffix-shaped, which "J.씨" is by period_joined_vocab; a run of
446 # ordinary name text is scanned on purpose, and a junk tail further
447 # out than the second run is held off by the scope rule instead
448 # ("Dr 김민준씨, Jr., 박씨" is SUFFIX_COMMA with 박씨 in a third
449 # run, so it never reaches here at all).
450 # An EMPTY second run stays in scope and contributes nothing:
451 # is_wholly_suffix is False on it by its own contract (v1 read
452 # "Doe,, Jr." as a family comma), which is the reading this line
453 # wants anyway -- flattening an empty run adds no site.
454 # Policy(lenient_comma_suffixes=False) keeps the old answer for the
455 # INITIAL-shaped suffixes specifically, which is where the
456 # strict/lenient gap lives: the knob drops this call to the strict
457 # predicate too, so is_wholly_suffix(["V."]) is False, the run reads
458 # as name text, it IS scanned, and the peel is abandoned on "V." as
459 # before -- family "田中さん", given "V.". It is not a blanket
460 # freeze of the old behavior, and "田中さん, Ph. D." is the input
461 # that shows the difference: the Ph./D. merge folds that pair to a
462 # form is_suffix_strict accepts, so the run is declined and the peel
463 # fires under the strict knob as well.
464 # Flattening the SEGMENTS rather than state.tokens is load-bearing
465 # too: extracted nickname and maiden content is in tokens but in NO
466 # segment, and scanning tokens would put the peel site on a
467 # nickname ("김민준씨 (Jimmy)" -> the site becomes Jimmy and
468 # nothing peels). ko_honorific_glued_given_nickname pins that.
469 # And declining takes a SECOND condition: segments[0] must hold a
470 # peel site of its own. is_wholly_suffix reaches period_joined_vocab,
471 # which calls a run suffix-shaped when ANY period-chunk is suffix
472 # VOCABULARY -- and every honorific tail is a suffix word by the
473 # Lexicon invariant, so a glued honorific is itself the evidence.
474 # The predicate is circular at THIS call site alone -- not because
475 # segment asks it any earlier (nothing has peeled at either call)
476 # but because segment SPENDS the answer differently: it reads the
477 # run's shape and stops, the answer being the structure, while the
478 # peel reads the same shape and then decides whether to go strip
479 # the very honorific that produced it. "이, J.씨" reads as wholly suffix
480 # only because of the 씨 the peel exists to remove, and declining a
481 # run that holds the only site does not fall back to some other
482 # site -- it loses the peel outright, and with it the given name,
483 # which lands in suffix as "J.씨". Asking for a site in segments[0]
484 # keeps the #319 answer wherever the peel has somewhere else to go
485 # ("田中さん, V." still declines, さん is right there) and gives the
486 # circular case back to master's reading. The two-honorific input
487 # "김민준씨, J.씨" is where the choice is visible and deliberate:
488 # both runs offer a site, so the decline stands and the person's own
489 # 씨 is peeled rather than the junk one behind the comma.
490 runs = state.segments[:1]
491 if state.structure is Structure.FAMILY_COMMA:
492 second = [state.tokens[j].text for j in state.segments[1]]
493 # `one_case` is passed for the reason the field exists: the
494 # credential lean is part of what "wholly suffix" MEANS since
495 # #289, and a stage that asks the question without it gets a
496 # different answer from `segment`, which asked it with the fact
497 # in hand one stage earlier. Left out, 'Kim김민준씨, MA' read
498 # 'MA' as name material, declined nothing, and scanned the
499 # second run -- where the site is 'MA' itself, which carries no
500 # listed tail -- so the person's own 씨 went unpeeled (family
501 # 'Kim김민준씨') while 'Kim김민준씨, PhD' peeled it, one name in
502 # two spellings parsed two ways (review round, #289/#516).
503 # `segment` records the fact only where a comma form could turn
504 # ON it, so this read is None for every other name and the
505 # predicate then behaves exactly as it did before 2.4.
506 if not (is_wholly_suffix(second, state.lexicon, state.policy,
507 one_case=state.one_case)
508 and _peel_site(state, state.segments[0], tails)):
509 runs = state.segments[:2]
510 site = _peel_site(state, [j for seg in runs for j in seg], tails)
511 if site is None:
512 return state
513 i, offset = site
514 # The tail carries a tag because this stage MANUFACTURED it. The
515 # segmenter's neighbour test below needs to tell it from a token
516 # somebody wrote, and no vocabulary question can: the two spellings
517 # put the same word in the same place, and only the provenance
518 # differs.
519 return _split(state, i, (offset,), None, tail_tag=_PEELED_TAG)
520
521
522def _split_surname_site(state: ParseState) -> ParseState:
523 """The stage's other split: the first activated-script token of the
524 name part is matched longest-first against Lexicon.surnames, and a
525 hit splits it in two; where the vocabulary declines, an optional
526 Parser(segmenter=...) is consulted instead.
527
528 A sibling of _peel_honorific_tail rather than a continuation of it.
529 The two answer different questions -- this one asks where a name
530 divides into surname and given, the peel asks whether a token ends
531 in a word that can never end a name -- which is why they carry
532 different gates: the FAMILY comma and the 间隔号 gate this half
533 alone (#312), and segment_scripts below gates it alone too. That
534 is also why the stage entry below interleaves gates with its two
535 calls rather than running one cascade."""
536 scripts = state.policy.segment_scripts
537 # an empty VOCABULARY deliberately does not bail here -- see below
538 if not scripts:
539 return state
540 # segments[0] is the NAME part under both remaining structures
541 # (everything, under NO_COMMA); later segments are suffixes. Its
542 # members are main-stream token indices by construction, so
543 # extracted nickname/maiden content is unreachable from here. For
544 # this site that is merely tidy -- the first script-written token
545 # is the same either way. It is the PEEL above that reads this run
546 # load-bearingly, and in the two structures that reach here it
547 # reads exactly this one, only backwards; the argument lives at
548 # that scan.
549 i = next((i for i in state.segments[0]
550 if effective_script(state.tokens[i].text) in scripts), None)
551 # A post-nominal in the surname's own position is not a site to
552 # skip past but an answer: a surname LEADS, so if the leading
553 # script-written token is an honorific there is no surname here to
554 # find. Scanning ON would reach the given name, which is exactly
555 # what the first-token rule exists to prevent -- 지 is a listed
556 # surname, so "양 지훈" (양 is a surname AND a shipped honorific)
557 # would have its own given name split in half. Declining also
558 # covers the token the peel above manufactures, which is the first
559 # and only script-written one in "Anderson선생님".
560 if i is None or _is_post_nominal(state, i):
561 return state
562 token = state.tokens[i]
563 text = token.text
564 surnames = state.lexicon.surnames
565 # The vocabulary match reads the token's CORE (#323). Without it
566 # the stop WAS the remainder: '김.' matched its own head and became
567 # 김 + '.'; the stop rides with the remainder instead, so '김민준.'
568 # divides as 김 + '민준.'. A leading stop never reaches THIS site --
569 # the classification fold in front rstrips too
570 # (_vocab._normalized_for_script) and hides such a token before it
571 # arrives -- while the peel above has no such gate in front of it
572 # and its own rstrip is what does the work there.
573 core = text.rstrip(FULL_STOPS)
574 # A token that IS a surname never splits: a bare "남궁" must not
575 # become 남 + 궁 just because the single-syllable surname also
576 # matches -- there is nothing to split off, and a lone token's
577 # role is the order resolution's call, stop or no stop.
578 if core in surnames:
579 return state
580 # Longest-first (compound-before-single falls out of it), capped
581 # so the remainder is never empty. Direct membership, no
582 # _normalize: the script gate admits only CJK text, and a CJK
583 # entry's stored form is its NFC composition with no case or
584 # edge-stop change (#322) -- so an entry AUTHORED in NFD is
585 # composed on the way in, raw NFD input matches nothing here, and
586 # the name goes unsplit, which is the no-split rules.md#W1 already
587 # accepts for NFD text. An empty vocabulary skips the match rather
588 # than bailing the stage (_longest_entry's max() has nothing to
589 # take): a surname-less lexicon declines every token, which is
590 # exactly the condition the segmenter is consulted on, so an early
591 # bail would make a configured segmenter silently inert under
592 # Lexicon.empty() -- the JA pack's own shape.
593 matches: list[int] = []
594 if surnames:
595 cap = min(_longest_entry(surnames), len(core) - 1)
596 matches = [length for length in range(cap, 0, -1)
597 if core[:length] in surnames]
598 if matches:
599 take = matches[0]
600 detail = None
601 if len(matches) > 1:
602 # more than one vocabulary-supported split: longest-first
603 # DECIDED a fork, and the deciding stage records it. A
604 # single-match split chose nothing, so it stays silent --
605 # a dictionary certainty and the statistical guess below
606 # are different epistemic states, and these two emission
607 # rules say so.
608 chosen = " + ".join(map(repr, _pieces(text, (take,))))
609 other = " + ".join(map(repr, _pieces(text, (matches[1],))))
610 detail = (f"{text!r} splits as {chosen} on the longest "
611 f"surname; {other} also reads")
612 return _split(state, i, (take,), detail)
613 if state.segmenter is None:
614 return state # the first activated-script token decides
615 # The segmenter's precondition, which the vocabulary has no twin of:
616 # it is asked where an UNDIVIDED name divides, so it may only be
617 # shown a token that is the whole name. Where the name part carries
618 # a second script-written token the writer already drew this
619 # stage's missing boundary -- "山田 太郎" is divided, and dividing
620 # its family again yields 山 + 田 + 太郎 (namedivider answers for
621 # any string, and scores a two-character one 1.0 by rule, so no
622 # confidence check would catch it). The neighbour counts whatever
623 # script it is written in, ACTIVATED or not -- effective_script is
624 # merely non-None -- because a katakana or hangul neighbour is a
625 # boundary its writer drew just as deliberately as a Han one: under
626 # the JA pack "山田太郎 マイケル" declines, though katakana is in no
627 # activation set. A Latin title or suffix is NOT such a boundary:
628 # it says nothing about where the CJK name splits, so "Dr 阿明日,
629 # Jr." still reaches the segmenter. Vocabulary keeps its own rule
630 # -- a listed surname is a certainty about that exact string,
631 # whoever else stands beside it.
632 # The one neighbour that does not count is the one this stage
633 # MANUFACTURED (#308): a glued 山田太郎様 has no writer-drawn
634 # boundary anywhere, so the 様 the peel just cut off cannot be read
635 # as one -- the precondition must see what the writer wrote, which
636 # was a single undivided token. A SPACED 様 does count, and the
637 # reason is weaker than "its writer drew that boundary and chose
638 # to write 山田太郎 as a unit": in "山田太郎 様" the unit the writer
639 # drew is the WHOLE NAME, honorific and all, so that story is false
640 # for the very input it describes. What the code relies on is that
641 # by POSITION a spaced honorific is indistinguishable from a spaced
642 # name element, so the test conservatively counts it. Measured, the
643 # trade is worth it: counting them keeps 佐藤 氏, 田中 様, 鈴木 先生
644 # and 中村 教授 whole -- all four divide bare under the JA pack (佐
645 # + 藤, 田 + 中, 鈴 + 木, 中 + 村) -- and costs the single division
646 # 山田太郎 様. Four real surnames against one.
647 # Provenance, not vocabulary: the two spellings put the same word
648 # in the same place, and asking the suffix set instead cannot
649 # separate them.
650 if any(j != i and effective_script(state.tokens[j].text) is not None
651 and _PEELED_TAG not in state.tokens[j].tags
652 for j in state.segments[0]):
653 return state
654 # No try/except around the call: rules.md#A1's Accepted clause
655 # ("a user-supplied segmenter's own error propagates"). The two
656 # checks below are that same doctrine, curated,
657 # and they are where the line this module draws is easiest to state:
658 # a PROTOCOL VIOLATION BY THE SEGMENTER AUTHOR RAISES, while an
659 # ADAPTER'S DEFENSE AGAINST ITS LIBRARY DECLINES. Both checks here
660 # are the first kind -- a wrong answer TYPE and an answer indexing
661 # past the token it was handed are stage-detectable bugs in
662 # user-supplied code, inside the declared totality exception, so
663 # they get the same treatment the callable's own exceptions get.
664 # locales/ja.py is the second kind: its repertoire, length,
665 # reconstruction and score guards all return None, because what
666 # they defend against is namedivider answering a question nobody
667 # asked it, which is a fact about the CONTENT, not a broken
668 # protocol. Bounded like every message here: the type's NAME, never
669 # its contents.
670 # The CORE, not the raw token (#323): handed the raw token, a
671 # segmenter answering len-1 would make the stop a piece of its own
672 # ('田中太郎' + '。'), where against the core the trailing stops
673 # ride with the last piece -- '山田太郎.' answered at 2 divides as
674 # 山田 + '太郎.'.
675 answer = state.segmenter(core)
676 if answer is not None and not isinstance(answer, Segmentation):
677 # a duck-typed answer carrying a .splits of its own would
678 # otherwise wander into the split path and surface as a
679 # ValueError naming Token, pointing the reader at nameparser's
680 # insides instead of at their segmenter
681 raise TypeError(
682 f"segmenter must return Segmentation or None, got "
683 f"{type(answer).__name__}")
684 if answer is None or not answer.splits:
685 return state # declined, or confidently one token
686 # splits[-1] is the max -- Segmentation enforced ascending. The
687 # upper bound is the half Segmentation cannot check, since it never
688 # sees the text; an offset at or past the end would make an empty
689 # piece. Declining silently here (as this did before the review)
690 # made an off-by-one segmenter undebuggable: every answer it gave
691 # vanished, and the parse merely looked unsegmented.
692 # The bound is the CORE's length, the string the segmenter was
693 # actually handed: an offset it could not have derived from its own
694 # input is the same author bug whether or not `text` is longer.
695 if answer.splits[-1] >= len(core):
696 raise ValueError(
697 f"segmenter returned splits beyond the token: last offset "
698 f"{answer.splits[-1]}, the segmenter was given "
699 f"{len(core)} characters")
700 conf = answer.confidence
701 detail = None
702 if conf is not None and conf < _SEGMENTER_CONFIDENCE_FLOOR:
703 reading = " + ".join(map(repr, _pieces(text, answer.splits)))
704 detail = (f"{text!r} splits as {reading} on a segmenter answer "
705 f"scoring {conf:.2f}, under the "
706 f"{_SEGMENTER_CONFIDENCE_FLOOR} confidence floor")
707 return _split(state, i, answer.splits, detail)
708
709
710# rules.md#W1: "an undivided word in the family position of a name
711# written in an activated script divides after a recognized surname,
712# the longest recognized surname first; where the vocabulary
713# recognizes nothing, an optional segmenter may divide instead"
714# rules.md#W3: "under a family comma the pre-comma text is the
715# family by declaration and never divides, and the post-comma side
716# is given text with no family to find" (history: decisions.md#W3)
717def script_segment(state: ParseState) -> ParseState:
718 if state.original.isascii():
719 # spans index the original exactly (the anti-#100 invariant),
720 # so an ASCII original has only ASCII tokens: nothing here is
721 # in any script's ranges. It also short-circuits the PEEL,
722 # which has no script gate of its own -- so a caller-configured
723 # ASCII tail never fires, and one non-ASCII character anywhere
724 # in the name switches it on. Correct for the CJK vocabulary
725 # that ships, stated because honorific_tails is public: see
726 # that field's own note. Latin orthography SPACES its
727 # post-nominals ("John Smith PhD"), so the glued position this
728 # peel exists for is a CJK one to begin with. What Latin does
729 # glue is PREnominal and period-joined (Mr.Smith, Lt.Gov.),
730 # which _vocab.period_joined_vocab classifies as one token
731 # rather than splitting -- a different mechanism, and no peel
732 # site either way. A glued Latin POST-nominal is spelled the
733 # same way, so it reaches that same mechanism rather than
734 # nothing: period_joined_vocab reads "Smith.Jr." as a title
735 # ('jr' is title vocabulary as well as suffix vocabulary), and
736 # where position allows, that wins -- "Smith.Jr. Anderson"
737 # gives title "Smith.Jr.", family "Anderson". Whatever it
738 # decides, this bail is what settles the question here: it
739 # returns above all of it, so no ASCII input reaches the peel.
740 return state
741 if not state.segments:
742 return state
743 # #312: the peel runs in front of both gates below, because both
744 # answer where a name divides into surname and given, and the peel
745 # does not ask that. It asks whether a token ends in a word that
746 # can never end a name, and a comma or a dot elsewhere in the
747 # string does not change the answer. Placing it above the block
748 # also keeps a future segmentation gate from silently capturing
749 # it: a new gate lands with its siblings, below.
750 state = _peel_honorific_tail(state)
751 if state.structure is Structure.FAMILY_COMMA:
752 return state # the comma already drew the SURNAME boundary
753 if state.interpunct_offsets:
754 # #298: a 间隔号-divided name is a transcription -- its pieces
755 # are syllable groups, not surname+given, so neither the
756 # vocabulary nor the segmenter applies (codepoint-scoped: the
757 # nakaguro records nothing and gates nothing, spec decision 5).
758 # State-global like the FAMILY_COMMA gate above: a marker
759 # anywhere in the name reads the WHOLE name as a transcription
760 # listing, so even an un-dotted hangul token beside a dotted
761 # one stays whole. Scoped to the surname split since #312: an
762 # honorific glued to a transcription is still an honorific.
763 return state
764 return _split_surname_site(state)