PeachDrawing.Text

PeachDrawing.Text is the font and text engine PeachPDF renders HTML with, published as its own NuGet package so other .NET applications can use it without PeachPDF. It has no third-party dependencies (only its own sibling data package, PeachDrawing.Text.Data, described below), is trimmable and Native AOT compatible, and is versioned in lockstep with PeachPDF: the same version number for every release, and PeachPDF depends on it.

dotnet add package PeachDrawing.Text

Status: pre-1.0. The library is being opened up area by area. Today the public surface is font loading and matching (FontSet and the types around it), what a Typeface says about itself (metrics, glyph mapping and advances), shaping, glyph outlines and colour glyphs, the MATH table, and the PeachDrawing.Text.Unicode namespace, all described below. Font subsetting for embedding, and paragraph layout (PeachDrawing.Text.Layout), described below, are public too. Until 1.0, the public API may change between releases.

What the engine does

Fonts: FontSet, families and matching

A FontSet is the fonts a piece of text can be set in: the fonts installed on the machine, plus the ones you add to it. A font you add under the name of an installed family joins that family for that set only, taking the place of the face with the same weight, slant, width and code point ranges. Two sets never see each other’s fonts, so two callers can register different data under one family name. A set is not safe for concurrent use: give each thread its own. Add the fonts before you match: an answer a set has already given is remembered and does not change when a font is added afterwards, so that measuring and drawing one piece of text cannot end up in different faces.

using PeachDrawing.Text;

var fonts = new FontSet();

// TrueType, OpenType (glyf and CFF), WOFF and WOFF2 are recognised by their content.
TypefaceFamily brand = fonts.AddFile("Brand-Regular.otf", new AddOptions { FamilyName = "Brand" });
fonts.AddFile("Brand-Bold.otf", new AddOptions { FamilyName = "Brand", Weight = 700 });

// Ask a family for the face that fits a query. A query has a weight, a width class, a slant and, optionally, a
// character the face has to be able to draw.
if (brand.TryMatch(new TypefaceQuery(Weight: 600, IsItalic: true), out TypefaceMatch match))
{
    Typeface face = match.Typeface;              // the face that matched; it has no size
    SyntheticStyle toFake = match.Synthesis;     // what is still missing: Bold, Italic, both, or None
}

TryMatch follows CSS Fonts 4 face matching: the width first, then the slant (upright, italic or oblique), then the weight, taking the nearest face when none is exact, so a request for condensed italic text gets the condensed face of a family whose condensed face is upright and whose italic face is of normal width, and the lean is faked. Among faces that declare an oblique range, the query’s ObliqueAngle chooses the one that holds the angle or else the nearest; a face declared italic beats an oblique range for an italic request with no angle. For an explicit oblique <angle> of 0 degrees or more, the oblique ranges leaning the same way as the angle are tried first, then a declared italic face, and only then the ranges leaning the other way. A request for oblique 0deg is upright’s equivalent on this scale, and a genuinely upright face is preferred to any oblique range - one that merely includes 0 as much as one that excludes it. The weight is a number, not a whole number: 350.5 is a weight, and a face whose range holds it is preferred to one that only holds 350. A face is taken to cover the characters of its unicode-range if it has one, and the ones its cmap maps otherwise. Synthesis says what the caller has to fake because the face falls short: bold when 600 or more was asked for and the face is lighter, italic when italic was asked for and the face is upright.

A typeface has no size. Text size belongs to whoever draws the text, and the same Typeface serves every size.

What a typeface says about itself

Everything a Typeface reports is in design units: whole numbers on the grid the font was drawn on, UnitsPerEm of them to the em. To get a length at a size, multiply by the size and divide by UnitsPerEm.

TypefaceMetrics metrics = face.Metrics;
double size = 16;
double ascent = size * metrics.CellAscent / metrics.UnitsPerEm;

if (face.TryMapRune(new Rune('A'), out ushort glyph))
{
    double advance = size * face.GetAdvance(glyph) / metrics.UnitsPerEm;
}

AddOptions is the counterpart of the descriptors of a CSS @font-face rule: the family name, weight, italic, width class and unicode-range to register a font under in place of what the file itself declares.

Other things a FontSet does:

Shaping: PeachDrawing.Text.Shaping

Shaper.Shape turns text into glyphs in one face: the cmap mapping, the substitutions of the font’s GSUB table and the positioning of its GPOS table.

using PeachDrawing.Text.Shaping;

GlyphRun run = Shaper.Shape(face, "office", new ShapeSettings(Caps: CapsMode.SmallCaps));

foreach (PlacedGlyph glyph in run.Glyphs)
{
    // glyph.GlyphIndex is the glyph, glyph.ClusterStart and ClusterLength say which characters it stands for, and the
    // deltas and offsets are the GPOS adjustments, in design units.
    double advance = face.GetAdvance((ushort)glyph.GlyphIndex) + glyph.XAdvanceDelta;
}

Outlines and colour glyphs: PeachDrawing.Text.Outlines

Typeface.TryGetOutline reads the shape of a glyph as data: a GlyphOutline of closed contours, each a start point and a list of segments that are straight lines or cubic curves, in design units with the y axis up. A glyph is filled by the nonzero winding rule, which is how it gets its counters. TrueType quadratic curves are raised to cubic ones, so a consumer has two kinds of segment to draw. This form is never grid-fitted; for text drawn into pixels, see Grid fitting (hinting).

using PeachDrawing.Text.Outlines;

if (face.TryMapRune(new Rune('g'), out ushort glyph) && face.TryGetOutline(glyph, out GlyphOutline outline))
{
    foreach (OutlineContour contour in outline.Contours)
    {
        // move to contour.Start, then for each OutlineSegment draw a line to End, or a cubic through Control1 and Control2
    }

    // Where the ink lies across a band between two heights: what text-decoration-skip-ink needs.
    foreach (var (start, end) in outline.Crossings(bandLow: -200, bandHigh: -100)) { /* ... */ }
}

Grid fitting (hinting)

A font can carry hints that move the points of a glyph, at one size, so that stems, x-heights and baselines land on whole pixels: a TrueType font as small programs, a font with CFF (PostScript) outlines as stem hints and blue zones in its charstrings. That makes small text drawn into a pixel raster sharper, and means nothing for vector output. An OutlineRequest asks for it: a size in pixels per em and a GridFitting.

var request = new OutlineRequest { PixelsPerEm = 11, GridFitting = GridFitting.Standard };

if (face.TryGetOutline(glyph, request, out GlyphOutline fitted))
{
    // fitted.IsGridFitted: the font's hints were applied. Coordinates are in pixels at 11 ppem, y up, origin at (0, 0).
    // fitted.GridFittedAdvance: the advance after fitting, rounded to a whole number of pixels as the font's hinting leaves it
    // (in Monochrome mode from the font's hdmx table where it has one for the size).
}
var request = new OutlineRequest { PixelsPerEm = 9, GridFitting = GridFitting.Standard, StemDarkening = true };

The instruction interpreter and the CFF engine are ports of FreeType’s (the CFF engine is the one Adobe contributed to FreeType), which is why the package carries the FreeType Project License notices and Adobe’s (see Licences). They give the same fitted points as FreeType 2.14.3, in 26.6 fixed point, for every font the test suite checks them with, including fonts made of random programs and charstrings, and variable fonts at many locations of their design spaces, that FreeType is compared with point for point.

Variable fonts

A variable font is one file that holds a whole design space: axes such as weight and width, and the outlines and metrics at every point in between. Typeface.IsVariable says whether a face is one, Typeface.Axes lists its axes (a VariationAxis with a tag, a range and a default; the tags the specification registers are in AxisTags), and Typeface.NamedVariations lists the named locations the font declares. Typeface.WithAxes returns the typeface at a location.

if (face.IsVariable)
{
    Typeface bold = face.WithAxes([new AxisSetting(AxisTags.Weight, 700)]);
    Typeface condensedBold = bold.WithAxes([new AxisSetting(AxisTags.Width, 80)]);   // builds on what bold has

    ushort glyph = ...;
    int advance = condensedBold.GetAdvance(glyph);                  // design units at that location
    bool hasOutline = condensedBold.TryGetOutline(glyph, out GlyphOutline outline);
}

Mathematics: PeachDrawing.Text.OpenType

A face made for setting mathematics has a MATH table, and Typeface.HasMathData says so. Typeface.MathData returns it as a MathTable with three parts. Constants holds the values a math layout algorithm positions fractions, radicals, scripts, stacks and limits with, in design units apart from the percentages. GlyphInfo answers per glyph: the italics correction, the horizontal position an accent attaches at, and whether the glyph is an extended shape. Variants gives the glyphs that stretch (fences, radicals, accents, arrows) their pre-sized variants and, for a size beyond the largest, the parts to assemble them from.

using PeachDrawing.Text.OpenType;

if (face.MathData is MathTable math && face.TryMapRune(new Rune('('), out ushort paren))
{
    double axis = math.Constants.AxisHeight;                      // design units above the baseline
    MathGlyphConstruction? tall = math.Variants.GetVerticalConstruction(paren);

    foreach (MathGlyphVariant variant in tall?.Variants ?? [])    // smallest to largest
    {
        // the first variant whose AdvanceMeasurement reaches the size you need is the one to draw
    }

    if (tall?.Assembly is MathGlyphAssembly assembly)
    {
        // bottom to top: repeat the parts that IsExtender until the target height is reached,
        // overlapping neighbours by at most their connector lengths and at least Variants.MinConnectorOverlap
    }
}

The per-glyph corner kerning of MathKernInfo and the device tables that adjust a value at particular sizes are not read.

Embedding: PeachDrawing.Text.Export

A document that embeds a font wants only the glyphs it uses. TypefaceExporter.ExportSubset cuts a typeface down to the glyph indices you give it and returns the bytes of a font file, an ExportedFont.

using PeachDrawing.Text.Export;

ExportedFont subset = TypefaceExporter.ExportSubset(face, usedGlyphs, keepCharacterMap: false);
byte[] fontFile = subset.Data.ToArray();
// subset.HasCffOutlines says which kind of font stream to write; subset.IsSubset says whether it was cut down.

What a font descriptor records about a face comes from the members you already have: Typeface.Metrics (with IsSymbolic, IsFixedPitch, HasSerifs, IsItalicStyle and FirstCharIndex for the descriptor flags), Typeface.GetAdvance for widths, Typeface.FullName for a base font name, and Typeface.ContentHash (a 128-bit hash of the font data, so two different fonts never share one) to key a cache of what you made from a face.

Laying out text: PeachDrawing.Text.Layout

ParagraphBuilder collects text and styles, and Build() prepares a Paragraph: the text’s direction (UAX #9), scripts, joining and line break opportunities (UAX #14) are worked out once, and the paragraph can then be laid out at any width. A Paragraph is immutable and can be laid out from several threads.

var builder = new ParagraphBuilder(new RunStyle(typeface, 16))
    .SetStyle(new ParagraphStyle { Align = TextAlign.Start, OverflowWrap = OverflowWrap.BreakWord });
builder.AddText("Some ").PushRun(new RunStyle(bold, 24)).AddText("large").PopRun().AddText(" text.");
Paragraph paragraph = builder.Build();

ParagraphLayout layout = paragraph.Layout(availableWidth: 300);
foreach (LineBox line in layout.Lines)
{
    foreach (PlacedRun run in line.Runs)          // left to right, in the order they are drawn
    {
        // run.Glyphs.Glyphs are in drawing order; run.X is the left edge and run.Baseline the baseline, in layout units.
    }
}

Layout units are the units of RunStyle.Size; coordinates run right and down from the top left of the paragraph.

Drawing a layout

PeachDrawing.Text stops at positions; painting is the caller’s. With a PeachDrawing.Core canvas (the RasterCanvas from the PeachDrawing package, or any other Canvas), one call paints a laid-out paragraph:

var layout = new ParagraphBuilder(new RunStyle(typeface, 16))
    .AddText("Hello, ").PushRun(new RunStyle(bold, 16)).AddText("world").PopRun()
    .Build().Layout(availableWidth: 300);

canvas.DrawParagraph(layout, new PaintPoint(10, 10), PaintColor.FromArgb(255, 0, 0, 0));

Glyphs are drawn where the shaper put them (nothing is reshaped), from the face each run was shaped in, so fallback faces, kerning, ligatures, bitmap glyphs and COLR/CPAL colour glyphs come out exactly as laid out. The overload taking a Func<PlacedRun, ParagraphPaint> chooses a colour and TextDecorations (underline, overline, line-through, drawn from the face’s own metrics) per run, and an optional callback receives each inline box’s bounds. canvas.DrawGlyphRun paints a single GlyphRun from Shaper.Shape. Layout units are the canvas’s user units.

The PeachDrawing.Text.Unicode namespace

Each entry point is a static class named for the algorithm or property it implements, and takes plain strings, runes and arrays.

Line breaking and text segmentation

LineBreaker.FindOpportunities implements the Unicode Line Breaking Algorithm (UAX #14). It answers for every UTF-16 index of a paragraph, and one past its end: Prohibited, Allowed or Mandatory for a line that would end just before that character. Where you break is still yours to decide: whether the space at a break stays on the line, how a word too long for a line is split, and where hyphenation adds breaks.

using PeachDrawing.Text.Unicode;

string text = "Wrap this line, please.\nNext line.";
LineBreakOpportunity[] opportunities = LineBreaker.FindOpportunities(text);

for (int i = 1; i < text.Length; i++)
{
    if (opportunities[i] == LineBreakOpportunity.Allowed) { /* a line may end before text[i] */ }
    if (opportunities[i] == LineBreakOpportunity.Mandatory) { /* it must */ }
}

LineBreakOptions applies the tailorings of CSS Text: WordBreak (a WordBreakMode: Normal, BreakAll, KeepAll) is word-break, and Strictness (Auto, Loose, Normal, Strict, Anywhere) is line-break. Strict is the algorithm’s own default, in which a small kana or a wave dash may not start a line; Auto (which is Normal) also lets a wave dash and the katakana double hyphen start one, and Loose lets a line start with a small kana, an iteration mark, and with a hyphen after an ideograph; between two ellipses it may break, but not before one. Language (a BCP 47 tag, or null) is what the rest depends on: the wave dash and the katakana double hyphen may start a line in Normal and Loose only where the language is Chinese or Japanese, and there Loose also lets a line start with a middle dot, the colon and semicolon of CJK text and a fullwidth or double exclamation or question mark, end before a suffix and after a prefix of East Asian width (%, ℃, ¥), which the number rules would otherwise keep with their digits. Anywhere allows a break after every grapheme cluster, whatever the character rules say, and keeps only hard line breaks.

Thai, Lao, Khmer and Burmese write no spaces between words, so no rule of the algorithm can find where a line may end (UAX #14 leaves those characters, its SA or Complex_Context class, to a dictionary). The library carries a word list for each, taken from ICU’s break-iterator dictionaries, and by default LineBreaker allows a break between the words it finds, as browsers do. It chooses the words by looking a few words ahead for the choice that covers the text best, preferring the longer word when two choices cover it alike; a stretch that no word matches stays whole, cut off from the words around it; and it never breaks inside a syllable (no break before a dependent vowel, tone mark or other sign, after a leading vowel, inside a Khmer or Burmese subscript/stacked consonant, or before a Burmese asat that closes the syllable before it). The script decides, not Language, and WordBreak, Strictness and overflow wrapping apply on top of it. A word list is read the first time text of its script is analysed (about 0.44 MB of embedded data in all, in the PeachDrawing.Text.Data package this one depends on, Brotli-compressed like the rest of its Unicode data), and is kept for the life of the process. A compound that the list has as one word stays whole even where a browser splits it. Set LineBreakOptions.ComplexContext to ComplexContextBreaking.GeneralCategory to have no opportunity inside a run of these scripts, which is what rule LB1 itself falls back to (a caller with its own dictionary wants that); the other Complex_Context scripts (Tai Tham, Cham and the rest) have no word list and always get it. A host with no Brotli decoder of its own (WebAssembly, at the time of writing) gets no word list either, and every script falls back the same way, unless it registers one with PeachDrawing.Text.Compression.BrotliDecompression.SetDecompressor - see The Unicode/hyphenation/dictionary data, and its Brotli decoder seam below.

Segmenter finds the boundaries of UAX #29: FindGraphemeBoundaries (extended grapheme clusters: a letter with its accents, a Hangul syllable, an emoji sequence, a flag), FindWordBoundaries and FindSentenceBoundaries. Each returns increasing UTF-16 indices, including the start and the end of the text, and none for empty text; the pieces are the text between neighbouring boundaries.

int[] clusters = Segmenter.FindGraphemeBoundaries("e\u0301\U0001F1FA\U0001F1F8");   // [0, 2, 6]

Bidirectional text

UAX #9 works in two steps, and so does the API. Bidi.Analyze resolves an embedding level for every UTF-16 code unit of a paragraph. Once your layout has decided where lines break, Bidi.ReorderLine puts one line’s runs in the order they are drawn.

using PeachDrawing.Text.Unicode;

string text = "abc אבג";
BidiAnalysis analysis = Bidi.Analyze(text, BaseDirection.Ltr);

foreach (BidiRun run in Bidi.ReorderLine(analysis.Levels, 0, text.Length))
{
    string piece = text.Substring(run.Start, run.Length);
    // A run at an odd level reads right to left: Mirror reverses it and swaps mirrored characters such as brackets.
    Console.WriteLine(run.IsRtl ? Bidi.Mirror(piece, run.Level) : piece);
}

BaseDirection.Auto detects the direction from the first strong character. A host with its own markup, such as CSS unicode-bidi or an SVG direction attribute, passes EmbeddingSpan values to describe embeddings the text itself does not spell out. Bidi.ClassOf returns a character’s Bidi_Class, and Bidi.TryGetMirror finds a bracket’s counterpart.

The implementation is checked against Unicode’s BidiCharacterTest.txt conformance file. Rule L1’s last clause, which resets the whitespace at the end of each line, is left to the caller because it depends on whether the caller lays out in characters, words or glyphs.

Scripts and OpenType tags

Scripts.Of returns the Unicode Script of a character (Latin, Arabic, Han, and the shared Common and Inherited). Scripts.Resolve gives every character of a text the script it is to be treated as, so that a comma between two Arabic words counts as Arabic (UAX #24 section 5.1). OpenTypeTags.ForScript and OpenTypeTags.ForLanguage turn a script name or a BCP 47 language tag into the four-letter tag an OpenType font’s layout tables are keyed by. Both answer null for a script or language the built-in table does not cover.

Vertical text, invisible characters, hyphenation and emoji

The Unicode/hyphenation/dictionary data, and its Brotli decoder seam

The tables above (Bidi, Script, vertical orientation, Arabic joining, the Indic Use tables), the hyphenation patterns and the Thai/Lao/Khmer/Burmese word lists ship Brotli-compressed, in a separate package, PeachDrawing.Text.Data, that PeachDrawing.Text depends on (see Fonts above for what “no third-party dependencies” means alongside this). A host whose Brotli decoder does not work - WebAssembly in a browser, at the time of writing, where System.IO.Compression.BrotliStream throws PlatformNotSupportedException - gets an empty table or an unhyphenated line instead of a failed render, exactly as before this data moved packages. PeachDrawing.Text.Compression.BrotliDecompression.SetDecompressor lets a host register a managed Brotli decoder of its own instead, to recover that data there; call it once, before using any feature backed by this data, since each table is read once and cached for the life of the process.

PeachDrawing.Text.Brotli is a ready-made decoder for that seam: a pure-managed port of google/brotli’s own C# decoder, with no dependency beyond PeachDrawing.Text itself, kept as a separate opt-in project rather than folded into PeachDrawing.Text so a host that never needs it never pays for it. Call PeachDrawing.Text.Brotli.ManagedBrotliDecompressor.Register() once at startup:

using PeachDrawing.Text.Brotli;

if (OperatingSystem.IsBrowser())
{
    ManagedBrotliDecompressor.Register();
}

PeachPDF.Demo.BlazorWasm’s Program.cs does exactly this, which is how its own WOFF2 fonts, hyphens: auto and Thai/Lao/Khmer/ Burmese dictionary line breaking all work in the browser.

Licences

The engine (PeachDrawing.Text) is BSD 3-Clause. It carries its third-party notices with it, in THIRD-PARTY-LICENSES.md: the font readers derive from PDFsharp (MIT), several shaping algorithms are ports of HarfBuzz code, and the TrueType instruction interpreter and Adobe’s CFF engine that do the grid fitting are ports of FreeType’s (under the FreeType Project License, whose text ships in the package as FTL.TXT, with Adobe’s patent licence grant for the CFF engine; an application that redistributes the package has to credit the FreeType Team in its documentation). The data tables - the Unicode Character Database, the hyph-utf8 pattern collection and ICU’s Thai, Lao, Khmer and Burmese word lists - live in PeachDrawing.Text.Data and carry their notices in that package’s own THIRD-PARTY-LICENSES.md. See License for the whole list.