PeachDrawing.Text
PeachDrawing.Text.Unicode
Segmenter Class
The default boundaries of Unicode Text Segmentation (UAX #29): where one grapheme cluster, word or sentence ends and the next begins.
public static class Segmenter
Inheritance System.Object → Segmenter
Remarks
Every method returns the boundaries as UTF-16 indices into the text, in increasing order, including the start of the text (0) and its end (the length). Text with nothing in it has no boundaries. A boundary never falls between the two halves of a surrogate pair. Pieces are the text between neighbouring boundaries.
These are the default rules of the Unicode Character Database that ships with the library; they are not tailored to a
language. Each algorithm is checked against Unicode’s own conformance file for it (GraphemeBreakTest.txt,
WordBreakTest.txt, SentenceBreakTest.txt).
| Methods | |
|---|---|
| FindGraphemeBoundaries(ReadOnlySpan<char>) | Finds the boundaries between extended grapheme clusters: the units a reader sees as one character, such as a letter with its accents, a Hangul syllable, an emoji sequence or a flag. |
| FindSentenceBoundaries(ReadOnlySpan<char>) | Finds the boundaries between sentences. A sentence keeps its closing punctuation, the spaces after it and its paragraph separator. |
| FindWordBoundaries(ReadOnlySpan<char>) | Finds the boundaries between words, as a search or a double click would treat them: runs of letters and numbers, with the punctuation inside them, are one word, a run of spaces is one piece, and every other character is a piece of its own. |