PeachDrawing.Text

PeachDrawing.Text.Unicode

UseCategory Enum

The category a character has in the Universal Shaping Engine (USE), the model that shapes Indic scripts: the alphabet its syllable grammar is written over and its reordering acts on.

public enum UseCategory : System.Byte

Fields

O 0

A character that takes no part in syllable structure: punctuation, another script’s letters, stray marks. It forms a syllable on its own and is never reordered.

B 1

A syllable’s base: an ordinary consonant, an independent vowel, a digit or an avagraha.

CGJ 2

A combining grapheme joiner or zero-width joiner, which is dropped from the syllable grammar altogether.

H 3

A halant (virama), which suppresses the inherent vowel of the consonant before it.

ZWNJ 4

A zero-width non-joiner, which stops conjuncts and ligatures forming across it.

R 5

A repha, the raised form of a leading consonant. No character has this category statically: it is given to a leading glyph that the font’s rphf feature actually substituted.

CMAbv 6

A consonant modifier that sits above the base. None occur in Devanagari itself.

CMBlw 7

A consonant modifier that sits below the base, such as the nukta.

VPre 8

A dependent vowel sign (matra) that is written before the base: the one category that reordering moves, to just after the nearest halant or the start of the syllable.

VAbv 9

A dependent vowel sign above the base.

VBlw 10

A dependent vowel sign below the base.

VPst 11

A dependent vowel sign written after the base.

VMPre 12

A vowel modifier (bindu, visarga, tone mark) written before the base.

VMAbv 13

A vowel modifier above the base, such as the anusvara and candrabindu.

VMBlw 14

A vowel modifier below the base, such as the anudatta.

VMPst 15

A vowel modifier written after the base, such as the visarga.

GB 16

A consonant placeholder that stands where a base consonant would, which is Bengali anji. It leads a syllable as a base does and is reordered as one.

FMAbv 17

A final syllable modifier above the base, which is the Bengali sandhi mark.

Remarks

USE categories are not a Unicode property. They are worked out from the Indic syllabic and positional categories and the general category of each character, by Classify(int).

Only the categories that Devanagari, Bengali, Gujarati and Tamil can produce are here. A script that needs more (medial consonants, or a static reordering killer) needs new members.