Skip to content

[segmenter] apply LB9 on both sides of the line break rules - #295

Open
egonelbre wants to merge 2 commits into
go-text:mainfrom
egonelbre:fix/segmenter-lb9
Open

egonelbre wants to merge 2 commits into
go-text:mainfrom
egonelbre:fix/segmenter-lb9

Conversation

@egonelbre

Copy link
Copy Markdown
Contributor

LB9 says a combining mark or ZWJ attaches to the character before it, and the
segmenter already applies that when it computes the class of the previous
character. The rules that compare runes rather than classes did not get the
same treatment. LB19, LB19a and LB30 looked at the raw previous rune and its
general category, so a mark sitting between an ideograph and a quote was
treated as its own character on both sides.

The first commit tracks which runes the collapsed previous classes were
computed for and compares against those. Two cases change:

  • U+6F22 U+201D U+0308 U+6F22 now breaks before the last ideograph.
  • U+6F22 U+0308 U+201C U+6F22 now breaks before the quote.

LineBreakTest.txt has no EastAsian, CM, QU triples, so neither case was
covered. Both are added as regression tests.

The second commit does the same for lookahead. The rules that inspect the
next base character, including the numeric prefix and decimal mark rules,
now skip attached marks and ZWJ and resolve SA marks through LB1. Without
this the first commit still prohibited a break before an opening quote with
an attached mark in East Asian text.

I am less sure of these two than of the ones in #290. The reasoning is that
the rules should see the same collapsed characters on both sides, but I do
not have a reference implementation that agrees or disagrees on these
inputs, so a second look at the affected rules would help.

br0 and prevLine follow LB9, so combining marks attach to the previous
character, but the rules compared them against the raw prev and
prevPrev runes and the raw previous general category. Track the runes
prevLine and prevPrevLine were computed for and compare against those.
U+6F22 U+201D U+0308 U+6F22 now breaks before the last ideograph, and
U+6F22 U+0308 U+201C U+6F22 breaks before the quote. LineBreakTest.txt
has no EastAsian/CM/QU triples, so both become regression tests.
Ignore attached combining marks and ZWJ when rules inspect the next base
character. The previous-context fix still prohibited breaks before
opening quotes with attached marks in East Asian text. Resolve SA marks
with LB1 and share the lookahead across the affected rules, including
numeric prefixes and decimal marks.
@egonelbre

Copy link
Copy Markdown
Contributor Author

For reference this is how other libraries are breaking things

Input ICU 78 libunibreak 8.0 this branch main
U+6F22 U+201D U+0308 U+6F22 3 3 3 none
U+6F22 U+0308 U+201C U+6F22 2, 4 2 2, 4 none
U+6F22 U+0308 U+201C U+0308 U+6F22 2 none 2 none
1)2% none none 3 3

@egonelbre

Copy link
Copy Markdown
Contributor Author

With regards to line breaking issues, if ICU78 is the standard, it could be helpful to add a fuzzer to find all the discrepancies. However, I'm not sure how often ICU78 itself has bug -- or are there intentional diversions from the spec.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant