PeachDrawing.Text

PeachDrawing.Text.Unicode

Segmenter Class

The default boundaries of Unicode Text Segmentation (UAX #29): where one grapheme cluster, word or sentence ends and the next begins.

public static class Segmenter

Inheritance System.Object → Segmenter

Remarks

Every method returns the boundaries as UTF-16 indices into the text, in increasing order, including the start of the text (0) and its end (the length). Text with nothing in it has no boundaries. A boundary never falls between the two halves of a surrogate pair. Pieces are the text between neighbouring boundaries.

These are the default rules of the Unicode Character Database that ships with the library; they are not tailored to a language. Each algorithm is checked against Unicode’s own conformance file for it (GraphemeBreakTest.txt, WordBreakTest.txt, SentenceBreakTest.txt).

Methods  
FindGraphemeBoundaries(ReadOnlySpan<char>) Finds the boundaries between extended grapheme clusters: the units a reader sees as one character, such as a letter with its accents, a Hangul syllable, an emoji sequence or a flag.
FindSentenceBoundaries(ReadOnlySpan<char>) Finds the boundaries between sentences. A sentence keeps its closing punctuation, the spaces after it and its paragraph separator.
FindWordBoundaries(ReadOnlySpan<char>) Finds the boundaries between words, as a search or a double click would treat them: runs of letters and numbers, with the punctuation inside them, are one word, a run of spaces is one piece, and every other character is a piece of its own.