PeachDrawing.Text

PeachDrawing.Text.Unicode

LineBreaker Class

The Unicode Line Breaking Algorithm (UAX #14): where a line of text may, and where it must, end.

public static class LineBreaker

Inheritance System.Object → LineBreaker

Remarks

The algorithm answers for every position between two characters. FindOpportunities(ReadOnlySpan<char>, LineBreakOptions) returns one answer for each UTF-16 index of the text and one for its end, so opportunities[i] says what happens to a line that would end just before text[i]. Inside a surrogate pair and between a character and the combining marks that follow it the answer is Prohibited, except after a space, a hard break or a zero width space, where a combining mark stands alone.

The result is the algorithm’s own view of the text. A host that lays text out still decides what to do with it: whether the space at a break stays on the line, how a word too long for a line is split, and where hyphenation adds breaks. Thai and Khmer, which write no spaces between words, are broken at the words a word list finds (see ComplexContext); the other Complex_Context scripts are broken as their letters, so a line of them has no opportunities where the script writes no spaces.

The rules are checked against Unicode’s own LineBreakTest.txt conformance file, with Strict, which is the algorithm’s own default, and GeneralCategory, which resolves the Complex_Context class as rule LB1 does.

Methods  
FindOpportunities(ReadOnlySpan<char>, LineBreakOptions) Finds where a line of text may or must break.