Documentation
¶
Overview ¶
Package sanitize filters text someone else wrote before a renderer shows it, in a terminal or on a web page, by one ordered set of character rules and one pinned Unicode property table, so that every implementation shows the same characters for the same value.
Filter applies four rules, in order, each to what the one before it returned:
- a tab becomes a space, and U+2028 and U+2029 become LF;
- every control character but LF (C0, DEL and C1) goes, every escape sequence goes whole, the bidirectional controls (U+061C, U+200E, U+200F, U+202A to U+202E, U+2066 to U+2069), the zero-width and invisible characters (U+00AD, U+034F, U+115F, U+1160, U+180E, U+200B, U+200C, U+2060 to U+2064, U+3164, U+FEFF, U+FFA0) and the supplementary variation selectors (U+E0100 to U+E01EF) go, and every tag character (U+E0000 to U+E007F) goes unless its run, with the character kept before it, is an emoji tag sequence the table lists (the flags of England, Scotland and Wales), which stays whole;
- U+FE00 to U+FE0D go, and U+FE0E or U+FE0F stays only where, with the character kept before it, it forms an emoji variation sequence the table lists, so at most one selector follows a character;
- U+200D stays only where the nearest character kept before it, passing over U+FE0F and emoji modifiers (U+1F3FB to U+1F3FF), and the character after it are both Extended_Pictographic.
"The character before" is always the last one the rule has kept, so a character a rule removes never stands between two it keeps. An escape sequence is what termsafe recognises as one: a CSI (ESC [ or U+009B) to its final byte, an OSC, DCS, SOS, PM or APC string to BEL, U+009C or ESC \, ESC with intermediates to its final byte, and ESC with any other one character; a sequence the value ends inside runs to the end.
Invalid UTF-8 first becomes U+FFFD, one for each maximal ill-formed subsequence, as the WHATWG decoder (TextDecoder) replaces it, so a value read from bytes filters the same in both languages.
The table is generated from Unicode 15.1's emoji data (emoji-data.txt, emoji-variation-sequences.txt and emoji-sequences.txt, under third_party/unicode) by tools/vectors, and is the same bytes the TypeScript package reads; neither uses its runtime's own Unicode properties. The filter neither bounds a value's length nor renders it: a terminal passes what it returns through termsafe, and a web page inserts it as text.
Index ¶
Examples ¶
Constants ¶
const UnicodeVersion = "15.1"
UnicodeVersion is the version of the Unicode data the table holds.
Variables ¶
This section is empty.
Functions ¶
func Filter ¶
Filter returns s with the four rules applied. It never returns a control character but LF, nor any character rule 2 lists.
Example ¶
A value from a record somebody else signed, filtered before a renderer shows it: the escape sequence, the bidirectional override and the hidden tag characters go, the tab becomes a space, and the emoji keep their selectors and joiner, and the flag its tags.
package main
import (
"fmt"
"github.com/lightwebinc/bcommon/sanitize"
)
func main() {
hostile := "ship it\t\x1b]0;pwned\x07\u202e!\U000E0069\U000E0067\U000E006E \u2764\ufe0f\u200d\U0001F525 " +
"\U0001F3F4\U000E0067\U000E0062\U000E0077\U000E006C\U000E0073\U000E007F"
fmt.Printf("%+q\n", sanitize.Filter(hostile))
fmt.Printf("%+q\n", sanitize.Filter("a\u200db\ufe0f\ufe0f"))
}
Output: "ship it ! \u2764\ufe0f\u200d\U0001f525 \U0001f3f4\U000e0067\U000e0062\U000e0077\U000e006c\U000e0073\U000e007f" "ab"
Types ¶
This section is empty.