chomp

package module
v0.4.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 14, 2024 License: MIT Imports: 3 Imported by: 2

README

= Chomp

A parser combinator library for chomping strings (_a rune at a time_) in Go. A more intuitive way to parse text without having to write a single regex. Happy to chomp both ASCII and Unicode (_it all tastes the same_).

Inspired by https://github.com/rust-bakery/nom[nom] 💜.

== Design

At the heart of `chomp` is a combinator. A higher-order function capable of parsing text under a defined condition and returning a tuple `(1,2,3)`:

- `1`: the remaining unparsed (_or unchomped_) text.
- `2`: the parsed (_or chomped_) text.
- `3`: an error if the combinator failed to parse.

Here's a sneak peek at its definition:

[source,go]
----
type Result interface {
	string | []string
}

type Combinator[T Result] func(string) (string, T, error)
----

A combinator in its simplest form would look like this:

[source,go]
----
func Tag(str string) chomp.Combinator[string] {
	return func(s string) (string, string, error) {
		if strings.HasPrefix(s, str) {
			// Return a tuple containing:
			// 1. the remaining string after the prefix
			// 2. the matched prefix
			// 3. no error
			return s[len(str):], str, nil
		}

		return s, "", chomp.CombinatorParseError{
			Input: str,
			Text: s,
			Type: "tag",
		}
	}
}
----

The true power of `chomp` comes from the ability to build parsers by chaining (_or combining_) combinators together.

== Writing a Parser Combinator

Take a look at one of the examples of how to write a parser combinator.

. https://github.com/purpleclay/chomp/blob/main/examples/gpg/main.go[GPG Private Key parser]
. https://github.com/purpleclay/chomp/blob/main/examples/git-diff/main.go[Git Diff parser]

A full glossary of combinators can be be viewed xref:docs/combinators.adoc[here].

== Why use Chomp?

- Combinators are very easy to write and combine into more complex parsers.
- Code written with chomp looks like natural grammar and is easy to understand, maintain and extend.
- It is incredibly easy to unit test.

== Badges

image:https://img.shields.io/github/actions/workflow/status/purpleclay/chomp/ci.yml?style=flat-square&logo=go["Build Status", link=https://github.com/purpleclay/chomp/actions?workflow=ci]
image:https://img.shields.io/badge/license-MIT-blue.svg?style=flat-square["License MIT", link=/LICENSE]
image:https://goreportcard.com/badge/github.com/purpleclay/chomp?style=flat-square["Go Report Card", link=https://goreportcard.com/report/github.com/purpleclay/chomp]
image:https://img.shields.io/github/go-mod/go-version/purpleclay/chomp.svg?style=flat-square["Go Version", link=go.mod]
image:https://app.deepsource.com/gh/purpleclay/chomp.svg/?label=active+issues&show_trend=false&token=DFB8RRar8iHJrVaNF7e9JaVm["DeepSource", link=https://app.deepsource.com/gh/purpleclay/chomp/]

Documentation

Overview

Package chomp provides a parser combinator library for chomping strings (a byte at a time) in Go. A more intuitive way to parse text without having to write a single regex

Index

Constants

This section is empty.

Variables

View Source
var (
	// IsDigit determines whether a rune is a decimal digit. A rune is classed
	// as a digit if it is between the ASCII range of '0' or '9', or if it belongs
	// within the Unicode [Nd] category.
	//
	// [Nd]: https://www.fileformat.info/info/unicode/category/Nd/list.htm
	IsDigit = isDigit{}

	// IsLetter determines if a rune is a letter. A rune is classed as a letter
	// if it is between the ASCII range of 'a' and 'z' (including its uppercase
	// equivalents), or it belongs within any of the Unicode letter categories:
	// [Lu] [LI] [Lt] [Lm] [Lo].
	//
	// [Lu]: https://www.fileformat.info/info/unicode/category/Lu/list.htm
	// [LI]: https://www.fileformat.info/info/unicode/category/Ll/list.htm
	// [Lt]: https://www.fileformat.info/info/unicode/category/Lt/list.htm
	// [Lm]: https://www.fileformat.info/info/unicode/category/Lm/list.htm
	// [Lo]: https://www.fileformat.info/info/unicode/category/Lo/list.htm
	IsLetter = isLetter{}

	// IsAlphanumeric determines whether a rune is a decimal digit or a letter.
	// This convenience method wraps the existing [IsDigit] and [IsLetter]
	// predicates.
	IsAlphanumeric = isAlphanumeric{}

	// IsLineEnding determines whether a rune is one of the following ASCII
	// line ending characters '\r' or '\n'.
	IsLineEnding = isLineEnding{}
)

Functions

This section is empty.

Types

type Combinator

type Combinator[T Result] func(string) (string, T, error)

Combinator is a higher-order function capable of parsing text under a defined condition. Combinators can be combined to form more complex parsers. Upon success, a combinator will return both the unparsed and parsed text. All combinators are strict and must parse its input. Any failure to do so should raise a CombinatorParseError.

func All

func All[T Result](c ...Combinator[T]) Combinator[[]string]

All will match the input text against a series of [Combinator]s. All combinators must match in the order provided.

chomp.All(
	chomp.Tag("Hello"),
	chomp.Until("W"),
	chomp.Tag("World!"))("Hello, World!")
// ("", []string{"Hello", ", ", "World!"}, nil)

func Any

func Any(str string) Combinator[string]

Any must match at least one character from the provided sequence at the beginning of the input text. Parsing stops upon the first unmatched character.

chomp.Any("eH")("Hello, World!")
// ("llo, World!", "He", nil)

func BracketAngled

func BracketAngled() Combinator[string]

BracketAngled will match any text delimited (or surrounded) by a pair of <angled brackets>.

chomp.BracketAngled()("<Hello, World!>")
// ("", "Hello, World!", nil)

func BracketSquare

func BracketSquare() Combinator[string]

BracketSquare will match any text delimited (or surrounded) by a pair of [square brackets].

chomp.BracketSquare()("[Hello, World!]")
// ("", "Hello, World!", nil)

func Crlf

func Crlf() Combinator[string]

Crlf must match either a CR '\r' or CRLF '\r\n' line ending.

chomp.Crlf()("\r\nHello")
// ("Hello", "\r\n", nil)

func Delimited

func Delimited[T, U, V Result](left Combinator[T], str Combinator[U], right Combinator[V]) Combinator[U]

Delimited will match a series of combinators against the input text. All must match, with the delimiters being discarded.

chomp.Delimited(
	chomp.Tag("'"),
	chomp.Tag("Hello, World!"),
	chomp.Tag("'"))("'Hello, World!'")
// ("", "Hello, World!", nil)

func Eol added in v0.3.0

func Eol() Combinator[string]

Eol will scan and return any text before any ASCII line ending characters. Line endings are discarded.

chomp.Eol()(`Hello, World!\nIt's a great day!`)
// ("It's a great day!", "Hello, World!", nil)

func First

func First[T Result](c ...Combinator[T]) Combinator[T]

First will match the input text against a series of [Combinator]s. Matching stops as soon as the first combinator succeeds. One Combinator must match. For better performance, try and order the combinators from most to least likely to match.

chomp.First(
	chomp.Tag("Good Morning"),
	chomp.Tag("Hello"))("Good Morning, World!")
// (" ,World!", "Good Morning", nil)

func Flatten added in v0.4.0

func Flatten(c Combinator[[]string]) Combinator[string]

Flatten the output from a Combinator by joining all extracted values into a string.

chomp.Flatten(
	chomp.Many(chomp.Parentheses()),
)("(H)(el)(lo), World!")
// (", World!", "Hello", nil)

func I

I extracts and returns a single string from the result of the inner Combinator. Combinators of differing return types can be successfully chained together while using this conversion combinator.

chomp.I(chomp.SepPair(
	chomp.Tag("Hello"),
	chomp.Tag(", "),
	chomp.Tag("World")), 1)("Hello, World!")
// ("!", "World", nil)

func Many added in v0.2.0

func Many[T Result](c Combinator[T]) Combinator[[]string]

Many will scan the input text, and it must match the Combinator at least once. This Combinator is greedy and will continuously execute until the first failed match. It is the equivalent of calling ManyN with an argument of 1.

chomp.Many(one.Of("Ho"))("Hello, World!")
// ("ello, World!", []string{"H"}, nil)

func ManyN added in v0.2.0

func ManyN[T Result](c Combinator[T], n uint) Combinator[[]string]

ManyN will scan the input text and match the Combinator a minimum number of times. This Combinator is greedy and will continuously execute until the first failed match.

chomp.ManyN(chomp.OneOf("W"), 0)("Hello, World!")
// ("Hello, World!", nil, nil)

func NoneOf

func NoneOf(str string) Combinator[string]

NoneOf must not match a single character at the beginning of the text from the provided sequence.

chomp.NoneOf("loWrd!e")("Hello, World!")
// ("ello, World!", "H", nil)

func Not

func Not(str string) Combinator[string]

Not must not match at least one character at the beginning of the input text from the provided sequence. Parsing stops upon the first matched character.

chomp.Not("ol")("Hello, World!")
// ("llo, World!", "He", nil)

func OneOf

func OneOf(str string) Combinator[string]

OneOf must match a single character at the beginning of the text from the provided sequence.

chomp.OneOf("!,eH")("Hello, World!")
// ("ello, World!", "H", nil)

func Opt

func Opt[T Result](c Combinator[T]) Combinator[T]

Opt allows a Combinator to be optional by discarding its returned error and not modifying the input text upon failure.

chomp.Opt(chomp.Tag("Hey"))("Hello, World!")
// ("Hello, World!", "", nil)

func Pair

func Pair[T, U Result](c1 Combinator[T], c2 Combinator[U]) Combinator[[]string]

Pair will scan the input text and match each Combinator in turn. Both combinators must match.

chomp.Pair(chomp.Tag("Hello,"), chomp.Tag(" World"))("Hello, World!")
// ("!", []string{"Hello,", " World"}, nil)

func Parentheses

func Parentheses() Combinator[string]

Parentheses will match any text delimited (or surrounded) by a pair of (parentheses).

chomp.Parentheses()("(Hello, World!)")
// ("", "Hello, World!", nil)

func Peek added in v0.3.0

func Peek[T Result](c Combinator[T]) Combinator[T]

Peek will scan the text and apply the Combinator without consuming any input. Useful if you need to look ahead.

chomp.Peek(chomp.Tag("Hello"))("Hello, World!")
// ("Hello, World!", "Hello", nil)

chomp.Peek(
	chomp.Many(chomp.Suffixed(chomp.Tag(" "), chomp.Until(" "))),
)("Hello and Good Morning!")
// ("Hello and Good Morning!", []string{"Hello", "and", "Good"}, nil)

func Prefixed added in v0.2.0

func Prefixed(c, pre Combinator[string]) Combinator[string]

Prefixed will scan the input text for a defined prefix and discard it before matching the remaining text against the Combinator. Both combinators must match.

chomp.Prefixed(
	chomp.Tag("Hello"),
	chomp.Tag(`"`))(`"Hello, World!"`)
// (`, World!"`, "Hello", nil)

func QuoteDouble

func QuoteDouble() Combinator[string]

QuoteDouble will match any text delimited (or surrounded) by a pair of "double quotes".

chomp.DoubleQuote()(`"Hello, World!"`)
// ("", "Hello, World!", nil)

func QuoteSingle

func QuoteSingle() Combinator[string]

QuoteSingle will match any text delimited (or surrounded) by a pair of 'single quotes'.

chomp.QuoteSingle()("'Hello, World!'")
// ("", "Hello, World!", nil)

func Repeat

func Repeat[T Result](c Combinator[T], n uint) Combinator[[]string]

Repeat will scan the input text and match the combinator the defined number of times. Every execution must match.

chomp.Repeat(chomp.Parentheses(), 2)("(Hello)(World)(!)")
// ("(!)", []string{"(Hello)", "(World)"}, nil)

func RepeatRange added in v0.2.0

func RepeatRange[T Result](c Combinator[T], n, m uint) Combinator[[]string]

RepeatRange will scan the input text and match the Combinator between a minimum and maximum number of times. It must match the expected minimum number of times.

chomp.RepeatRange(chomp.OneOf("Hleo"), 1, 8)("Hello, World!")
// (", World!", []string{"H", "e", "l", "l", "o"}, nil)

func S

S wraps the result of the inner Combinator within a string slice. Combinators of differing return types can be successfully chained together while using this conversion combinator.

chomp.S(chomp.Until(","))("Hello, World!")
// (", World!", []string{"Hello"}, nil)

func SepPair

func SepPair[T, U, V Result](c1 Combinator[T], sep Combinator[U], c2 Combinator[V]) Combinator[[]string]

SepPair will scan the input text and match each Combinator, discarding the separator's output. All combinators must match.

chomp.SepPair(
	chomp.Tag("Hello"),
	chomp.Tag(", "),
	chomp.Tag("World"))("Hello, World!")
// ("!", []string{"Hello", "World"}, nil)

func Suffixed added in v0.2.0

func Suffixed(c, suf Combinator[string]) Combinator[string]

Suffixed will scan the input text against the Combinator before matching a suffix and discarding it. Both combinators must match.

chomp.Suffixed(
	chomp.Tag("Hello"),
	chomp.Tag(", "))("Hello, World!")
// ("World!", "Hello", nil)

func Tag

func Tag(str string) Combinator[string]

Tag must match a series of characters at the beginning of the input text in the exact order and case provided.

chomp.Tag("Hello")("Hello, World!")
// (", World!", "Hello", nil)

func Until

func Until(str string) Combinator[string]

Until will scan the input text for the first occurrence of the provided series of characters. Everything until that point in the text will be matched.

chomp.Until("World")("Hello, World!")
// ("World!", "Hello, ", nil)

func While

func While(p Predicate) Combinator[string]

While will scan the input text, testing each character against the provided Predicate. The Predicate must match at least one character.

chomp.While(chomp.IsLetter)("Hello, World!")
// (", World!", "Hello", nil)

func WhileN added in v0.3.0

func WhileN(p Predicate, n uint) Combinator[string]

WhileN will scan the input text, testing each character against the provided Predicate. The Predicate must match at least n characters. If n is zero, this becomes an optional combinator.

chomp.WhileN(chomp.IsLetter, 1)("Hello, World!")
// (", World!", "Hello", nil)

chomp.WhileN(chomp.IsDigit, 0)("Hello, World!")
// ("Hello, World!", "", nil)

func WhileNM added in v0.3.0

func WhileNM(p Predicate, n, m uint) Combinator[string]

WhileNM will scan the input text, testing each character against the provided Predicate. The Predicate must match a minimum of n and upto a maximum of m characters. If n is zero, this becomes an optional combinator.

chomp.WhileNM(chomp.IsLetter, 1, 8)("Hello, World!")
// (", World!", "Hello", nil)

func WhileNot

func WhileNot(p Predicate) Combinator[string]

WhileNot will scan the input text, testing each character against the provided Predicate. The Predicate must not match at least one character. It has the inverse behavior of While.

chomp.WhileNot(chomp.IsDigit)("Hello, World!")
// ("", "Hello, World!", nil)

func WhileNotN added in v0.3.0

func WhileNotN(p Predicate, n uint) Combinator[string]

WhileNotN will scan the input text, testing each character against the provided Predicate. The Predicate must not match at least n characters. If n is zero, this becomes an optional combinator. It has the inverse behavior of WhileN.

chomp.WhileNotN(chomp.IsDigit, 1)("Hello, World!")
// ("", "Hello, World!", nil)

chomp.WhileNotN(chomp.IsLetter, 0)("Hello, World!")
// ("Hello, World!", "", nil)

func WhileNotNM added in v0.3.0

func WhileNotNM(p Predicate, n, m uint) Combinator[string]

WhileNotNM will scan the input text, testing each character against the provided Predicate. The Predicate must not match a minimum of n and upto a maximum of m characters. If n is zero, this becomes an optional combinator. It has the inverse behavior of WhileNM.

chomp.WhileNotNM(chomp.IsLetter, 1, 8)("20240709 was a great day")
// (" was a great day", "20240709", nil)

type CombinatorParseError

type CombinatorParseError struct {
	// Input to the [Combinator]. This can be empty, as a combinator may
	// not require any input to parse the text.
	Input string

	// Text that was being parsed by the [Combinator]. This will be truncated
	// in the error message.
	Text string

	// Type of [Combinator] that failed.
	Type string
}

CombinatorParseError defines an error that is raised when a combinator fails to parse the input text under its expected condition.

func (CombinatorParseError) Error

func (e CombinatorParseError) Error() string

Error returns a friendly string representation of the current error.

type MappedCombinator added in v0.3.0

type MappedCombinator[S any, T Result] func(string) (string, S, error)

MappedCombinator is a function capable of converting the output from a Combinator into any given type. Upon success, it will return the unparsed text, along with the mapped value. All combinators are strict and must parse its input. Any failure to do so should raise a CombinatorParseError. It is designed for exclusive use by the Map function

func Map added in v0.3.0

func Map[S any, T Result](c Combinator[T], mapper func(in T) S) MappedCombinator[S, T]

Map the result of a Combinator to any other type

chomp.Map(
	chomp.While(chomp.IsDigit),
	func (in string) int { return len(in) })("123456")
// ("", 6, nil)

type ParserError

type ParserError struct {
	// Err contains the [CombinatorParseError] that caused the parser to fail.
	Err error

	// Type of [Parser] that failed.
	Type string
}

ParserError defines an error that is raised when a parser fails to parse the input text due to a failed Combinator.

func (ParserError) Error

func (e ParserError) Error() string

Error returns a friendly string representation of the current error.

func (ParserError) Unwrap

func (e ParserError) Unwrap() error

Unwrap returns the inner CombinatorParseError.

type Predicate

type Predicate interface {
	// Match a rune against a defined expression, returning true
	// if the condition is met
	Match(r rune) bool

	// Returns the name of the predicate for error handling
	fmt.Stringer
}

Predicate defines an expression that will return either true or false

type RangedParserError added in v0.2.0

type RangedParserError struct {
	// Err contains the [CombinatorParseError] that caused the parser to fail.
	Err error

	// Range contains the execution details of the ranged parser.
	Exec RangedParserExec

	// Type of [Parser] that failed.
	Type string
}

RangedParserError defines an error that is raised when a ranged parser fails to parse the input text due to a failed Combinator within the expected execution range.

func (RangedParserError) Error added in v0.2.0

func (e RangedParserError) Error() string

Error returns a friendly string representation of the current error.

func (RangedParserError) Unwrap added in v0.2.0

func (e RangedParserError) Unwrap() error

Unwrap returns the inner CombinatorParseError.

type RangedParserExec added in v0.2.0

type RangedParserExec struct {
	// Min is the minimum number of expected executions.
	Min uint

	// Max is the maximum number of possible executions.
	Max uint

	// Count contains the number of executions.
	Count uint
}

RangedParserExec details how a ranged Combinator was exeucted.

func RangeExecution added in v0.2.0

func RangeExecution(i ...uint) RangedParserExec

RangeExecution is a utility method for setting a RangedParserExec.

  • With one argument, the [RangeParserExec.Count] is set.
  • With two arguments, the [RangeParserExec.Count] and [RangeParserExec.Min] are set.
  • With three arguments, the [RangeParserExec.Count]], [RangeParserExec.Min] and [RangeParserExec.Max] are set.
  • If four or more arguments are provided, a default RangedParserExec will be returned.

func (RangedParserExec) String added in v0.2.0

func (e RangedParserExec) String() string

String returns a string representation of a RangedParserExec.

type Result

type Result interface {
	string | []string
}

Result is the expected output from a Combinator.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL