Repository navigation
Conversation
br0 and prevLine follow LB9, so combining marks attach to the previous character, but the rules compared them against the raw prev and prevPrev runes and the raw previous general category. Track the runes prevLine and prevPrevLine were computed for and compare against those. U+6F22 U+201D U+0308 U+6F22 now breaks before the last ideograph, and U+6F22 U+0308 U+201C U+6F22 breaks before the quote. LineBreakTest.txt has no EastAsian/CM/QU triples, so both become regression tests.
Ignore attached combining marks and ZWJ when rules inspect the next base character. The previous-context fix still prohibited breaks before opening quotes with attached marks in East Asian text. Resolve SA marks with LB1 and share the lookahead across the affected rules, including numeric prefixes and decimal marks.
egonelbre
requested review from
andydotxyz,
benoitkugler and
whereswaldon
as code owners
September 30, 2026 09:28
Contributor
Author
|
For reference this is how other libraries are breaking things
|
Contributor
Author
|
With regards to line breaking issues, if ICU78 is the standard, it could be helpful to add a fuzzer to find all the discrepancies. However, I'm not sure how often ICU78 itself has bug -- or are there intentional diversions from the spec. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
LB9 says a combining mark or ZWJ attaches to the character before it, and the
segmenter already applies that when it computes the class of the previous
character. The rules that compare runes rather than classes did not get the
same treatment. LB19, LB19a and LB30 looked at the raw previous rune and its
general category, so a mark sitting between an ideograph and a quote was
treated as its own character on both sides.
The first commit tracks which runes the collapsed previous classes were
computed for and compares against those. Two cases change:
LineBreakTest.txt has no EastAsian, CM, QU triples, so neither case was
covered. Both are added as regression tests.
The second commit does the same for lookahead. The rules that inspect the
next base character, including the numeric prefix and decimal mark rules,
now skip attached marks and ZWJ and resolve SA marks through LB1. Without
this the first commit still prohibited a break before an opening quote with
an attached mark in East Asian text.
I am less sure of these two than of the ones in #290. The reasoning is that
the rules should see the same collapsed characters on both sides, but I do
not have a reference implementation that agrees or disagrees on these
inputs, so a second look at the affected rules would help.