core

package
v1.2.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 14, 2026 License: MIT Imports: 3 Imported by: 0

Documentation

Overview

Package core provides universal types and mechanics for transliteration. This package is script-agnostic - it knows nothing about Brahmic, Arabic, etc.

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

This section is empty.

Types

type Category

type Category int

Category classifies symbols for parsing. Base categories are defined here; scripts can extend with their own.

const (
	CatUnknown   Category = iota // Unknown/unrecognized
	CatVowel                     // Independent vowels
	CatConsonant                 // Consonants
	CatNumber                    // Numerals
	CatSymbol                    // Punctuation, other symbols

)

func (Category) String

func (c Category) String() string

type DebugInfo

type DebugInfo struct {
	Input  string      // Original input
	Output string      // Final output
	Units  []UnitDebug // Parsed units
	Traces []RuleTrace // Rule applications
}

DebugInfo contains debugging information from transliteration.

type Engine

type Engine struct {
	// contains filtered or unexported fields
}

Engine orchestrates the transliteration pipeline.

func NewEngine

func NewEngine(lang Language, scheme Scheme, opts ...EngineOption) *Engine

NewEngine creates an engine for a language + scheme combination. Panics if the language returns a nil Script.

func (*Engine) Language

func (e *Engine) Language() Language

Language returns the engine's language.

func (*Engine) RuleEngine

func (e *Engine) RuleEngine() *RuleEngine

RuleEngine returns the engine's rule engine for direct manipulation. Use this to enable/disable rules after engine creation.

func (*Engine) Scheme

func (e *Engine) Scheme() Scheme

Scheme returns the engine's scheme.

func (*Engine) Transliterate

func (e *Engine) Transliterate(input string) string

Transliterate converts script text to romanized form using default options.

func (*Engine) TransliterateDebug

func (e *Engine) TransliterateDebug(input string, opts Options) (string, *DebugInfo)

TransliterateDebug converts script text and returns debug information.

func (*Engine) TransliterateWithOptions

func (e *Engine) TransliterateWithOptions(input string, opts Options) string

TransliterateWithOptions converts script text to romanized form with custom options.

type EngineOption

type EngineOption func(*engineConfig)

EngineOption configures engine creation.

func WithDisabledRules

func WithDisabledRules(patterns ...string) EngineOption

WithDisabledRules disables rules matching the given patterns at engine creation. Patterns can be exact names or glob patterns (e.g., "schwa.*", "vowel.long-aa.*").

func WithEnabledRules

func WithEnabledRules(patterns ...string) EngineOption

WithEnabledRules enables rules matching the given patterns at engine creation. Useful for enabling rules that are disabled by default. Patterns can be exact names or glob patterns (e.g., "vowel.long-aa.all").

type Language

type Language interface {
	// Name returns the language identifier (e.g., "hindi", "marathi").
	Name() string

	// Script returns the script used by this language.
	Script() Script

	// Symbols returns the symbol map for this language.
	Symbols() SymbolMap

	// ScriptConfig returns script-specific configuration.
	// For Brahmic: brahmic.Config{Halant: "्", Nukta: "़"}
	ScriptConfig() interface{}

	// Rules returns the complete rule catalog for this language.
	// This includes ALL possible rules; schemes select from this catalog.
	Rules() RuleCatalog
}

Language defines language-specific behavior within a script.

type LexiconProvider

type LexiconProvider interface {
	LexiconLookup(word string) (string, bool)
}

LexiconProvider is an optional interface a Language may implement to supply a word-level romanization lexicon. When Options.Lexicon is set, the engine consults it before the rule pipeline; a hit short-circuits transliteration.

type Options

type Options struct {
	// LongVowels outputs "aa" for all ā (aa-matra) positions.
	LongVowels bool
	// SimpleNasals uses simplified nasal endings (करें→karen instead of karein).
	SimpleNasals bool
	// KeepMedialSchwa disables CCV schwa deletion, retaining schwa in more positions.
	// This produces output closer to some datasets (जनता→janata instead of janta).
	KeepMedialSchwa bool
	// SchwaModel uses the learned decision-tree schwa classifier for inherent-schwa
	// decisions instead of the hand-written heuristic schwa rules. See
	// lang/hindi/schwa_model.go and docs/reviews for the held-out evaluation.
	SchwaModel bool
	// Lexicon consults the language's high-confidence romanization lexicon first;
	// known words return the attested human spelling, unknown words fall through
	// to the rule engine. Requires the Language to implement LexiconProvider.
	Lexicon bool
	// Rerank generates candidate romanizations under several rule configurations
	// and picks the one a character language model finds most natural. Requires
	// the Language to implement Reranker. A lexicon hit (with Lexicon) wins first.
	Rerank bool
	// Debug enables debug output showing rule applications.
	Debug bool
}

Options configures transliteration behavior.

func DefaultOptions

func DefaultOptions() Options

DefaultOptions returns the default transliteration options.

type Parser

type Parser interface {
	// Parse converts input text into a Word structure.
	Parse(input string, symbols SymbolMap) *Word
}

Parser converts script input into Units.

type Position

type Position struct {
	Offset int // Byte offset in original string
	Rune   int // Rune (character) index
}

Position tracks location in source text for debugging.

type Renderer

type Renderer interface {
	// Render converts a Word into a romanized string.
	Render(word *Word) string
}

Renderer converts Units into romanized output.

type Reranker

type Reranker interface {
	RerankRomans(candidates []string) string
}

Reranker is an optional interface a Language may implement to choose the most natural romanization among candidates produced under different rule configurations. Used when Options.Rerank is set.

type Rule

type Rule struct {
	Name            string
	Phase           RulePhase
	Scope           RuleScope
	Priority        int // 0-99 within scope
	Mode            RuleMode
	DisabledDefault bool   // If true, rule is disabled by default (must be explicitly enabled)
	Conditional     string // If non-empty, rule depends on this runtime option (e.g., "LongVowels")
	Condition       func(*Unit, *Word) bool
	Action          func(*Unit, *Word)
}

Rule represents a single transliteration rule.

func AppendIfFound

func AppendIfFound(rules []Rule, catalog []Rule, name string) []Rule

AppendIfFound appends a rule to the slice if found by name. Returns the (possibly extended) slice.

func FindInSlice

func FindInSlice(rules []Rule, name string) *Rule

FindInSlice finds a rule by name in a slice of rules. Returns nil if not found.

func (*Rule) EffectivePriority

func (r *Rule) EffectivePriority() int

EffectivePriority calculates the overall priority across scopes. Higher values run first.

type RuleCatalog

type RuleCatalog struct {
	Schwa     []Rule // All schwa-related rules
	Consonant []Rule // All consonant rules
	Vowel     []Rule // All vowel rules
	Render    []Rule // All output transform rules
}

RuleCatalog organizes rules by category. Languages provide complete catalogs; schemes select from them.

func (RuleCatalog) AllRules

func (c RuleCatalog) AllRules() []Rule

AllRules returns all rules in the catalog as a flat slice.

func (RuleCatalog) FindByName

func (c RuleCatalog) FindByName(name string) *Rule

FindByName finds a rule by name across all categories. Returns nil if not found.

type RuleEngine

type RuleEngine struct {
	// contains filtered or unexported fields
}

RuleEngine applies rules to parsed words.

func NewRuleEngine

func NewRuleEngine(rules []Rule) *RuleEngine

NewRuleEngine creates a new rule engine with the given rules. Panics if any rule has nil Condition or Action functions.

func (*RuleEngine) AddRule

func (e *RuleEngine) AddRule(r Rule) error

AddRule adds a rule to the engine. Returns error if there's a priority conflict or invalid rule.

func (*RuleEngine) Apply

func (e *RuleEngine) Apply(word *Word)

Apply executes all rules on the word in phase order.

func (*RuleEngine) DisableRule

func (e *RuleEngine) DisableRule(pattern string) int

DisableRule disables rules matching the given pattern. Pattern can be an exact name or a glob pattern using '*' as wildcard. Examples: "schwa.delete.ccv", "schwa.*", "schwa.delete.*" Returns the count of rules matched and disabled.

func (*RuleEngine) EnableDebug

func (e *RuleEngine) EnableDebug(enabled bool)

EnableDebug enables debug trace collection.

func (*RuleEngine) EnableRule

func (e *RuleEngine) EnableRule(pattern string) int

EnableRule enables rules matching the given pattern. Pattern can be an exact name or a glob pattern using '*' as wildcard. Examples: "vowel.long-aa.all", "vowel.*", "vowel.long-aa.*" Returns the count of rules matched and enabled.

func (*RuleEngine) IsDisabled

func (e *RuleEngine) IsDisabled(name string) bool

IsDisabled returns true if the rule with the given name is currently disabled.

func (*RuleEngine) ListRules

func (e *RuleEngine) ListRules(pattern string) []RuleStatus

ListRules returns all rules with their enabled/disabled status. If pattern is non-empty, only rules matching the pattern are returned.

func (*RuleEngine) Rules

func (e *RuleEngine) Rules() []Rule

Rules returns a copy of all registered rules.

func (*RuleEngine) RulesForPhase

func (e *RuleEngine) RulesForPhase(phase RulePhase) []Rule

RulesForPhase returns a copy of rules for a specific phase, sorted by priority.

func (*RuleEngine) SetDebugMetaExtractor

func (e *RuleEngine) SetDebugMetaExtractor(fn func(*Unit) string)

SetDebugMetaExtractor sets a function to extract script-specific metadata.

func (*RuleEngine) Traces

func (e *RuleEngine) Traces() []RuleTrace

Traces returns collected debug traces (call after Apply).

type RuleMode

type RuleMode int

RuleMode determines execution behavior.

const (
	ModeExclusive RuleMode = iota // First match wins for this unit
	ModeAlways                    // Always run if condition matches
	ModeFallback                  // Only if no other rule acted on unit
)

func (RuleMode) String

func (m RuleMode) String() string

type RulePhase

type RulePhase int

RulePhase determines when a rule executes in the pipeline.

const (
	PhaseSchwa     RulePhase = iota // Schwa deletion/retention decisions
	PhaseConsonant                  // Consonant modifications (e.g., व→w)
	PhaseVowel                      // Vowel modifications (e.g., aa→ā)
	PhaseRender                     // Final output adjustments
)

func (RulePhase) String

func (p RulePhase) String() string

type RuleScope

type RuleScope int

RuleScope determines the priority tier for a rule. Higher scopes have higher effective priority.

const (
	ScopeUniversal RuleScope = iota // Base 0: applies to all languages
	ScopeScript                     // Base 100: script-specific (Brahmic)
	ScopeLanguage                   // Base 200: language-specific (Hindi)
	ScopeScheme                     // Base 300: scheme-specific (IAST)
)

func (RuleScope) String

func (s RuleScope) String() string

type RuleStatus

type RuleStatus struct {
	Name        string
	Phase       RulePhase
	Scope       RuleScope
	Priority    int
	Enabled     bool
	Conditional string // If non-empty, rule depends on this runtime option
}

RuleStatus represents a rule and its current enabled/disabled state.

type RuleTrace

type RuleTrace struct {
	Phase    string // Phase name (Schwa, Consonant, Vowel, Render)
	Rule     string // Rule name
	Unit     string // Unit characters (e.g., "क")
	UnitIdx  int    // Unit index in word
	Before   string // BaseRom before rule
	After    string // BaseRom after rule (empty if unchanged)
	Metadata string // Additional info (e.g., schwa state)
}

RuleTrace records a single rule application for debugging.

type Scheme

type Scheme interface {
	// Name returns the scheme identifier (e.g., "colloquial", "iast").
	Name() string

	// SelectRules selects rules from the language's catalog.
	// Returns the rules to apply for this scheme.
	SelectRules(catalog RuleCatalog) []Rule
}

Scheme selects which rules to use from a language's catalog. Schemes are language-agnostic - they work with any language.

type Script

type Script interface {
	// Name returns the script identifier (e.g., "brahmic", "arabic").
	Name() string

	// NewParser creates a parser for this script.
	// The config parameter is script-specific (e.g., brahmic.Config).
	NewParser(config interface{}) Parser

	// NewRenderer creates a renderer for this script.
	NewRenderer() Renderer

	// PrepareWord performs script-specific processing after parsing.
	// For example, Brahmic scripts identify consonant runs here.
	PrepareWord(word *Word)

	// Categories returns script-specific categories.
	Categories() []Category

	// DebugMetaExtractor returns a function that extracts script-specific
	// metadata for debugging (e.g., schwa state for Brahmic).
	// Returns nil if no metadata is available.
	DebugMetaExtractor() func(*Unit) string
}

Script defines a script family (Brahmic, Arabic, Latin, etc.). Scripts provide parsing and rendering capabilities.

type SymbolInfo

type SymbolInfo struct {
	Category Category
	BaseRom  string
}

SymbolInfo holds romanization info for a symbol.

type SymbolMap

type SymbolMap map[string]SymbolInfo

SymbolMap maps script characters to their romanization info.

type Unit

type Unit struct {
	// Source tracking
	Runes []rune   // Original characters (1-3 for conjuncts)
	Start Position // Start position in source
	End   Position // End position in source

	// Classification
	Type    UnitType
	BaseRom string // Base romanization (modifiable by rules)

	// Navigation (bidirectional links)
	Prev *Unit
	Next *Unit

	// Script-specific extension point
	// Scripts store their own data here (e.g., SchwaState for Brahmic)
	ScriptData interface{}
}

Unit represents a single phonetic unit in the parsed word.

func (*Unit) IsWordFinal

func (u *Unit) IsWordFinal() bool

IsWordFinal returns true if this is the last unit in the word.

func (*Unit) IsWordInitial

func (u *Unit) IsWordInitial() bool

IsWordInitial returns true if this is the first unit in the word.

func (*Unit) String

func (u *Unit) String() string

String returns the original characters as a string.

type UnitDebug

type UnitDebug struct {
	Index    int    // Position in word
	Chars    string // Original characters
	Type     string // Unit type
	BaseRom  string // Base romanization
	RunePos  int    // Rune position in original string
	Metadata string // Script-specific info
}

UnitDebug contains debug info for a single unit.

type UnitType

type UnitType int

UnitType classifies parsed phonetic units.

const (
	UnitVowel     UnitType = iota // Vowels and matras that replace inherent schwa
	UnitModifier                  // Modifiers that follow the vowel (anusvara, visarga, chandrabindu)
	UnitConsonant                 // Single consonant
	UnitConjunct                  // Multi-character conjunct (e.g., ज्ञ)
	UnitNumber                    // Numerals
	UnitSymbol                    // Other symbols (punctuation, etc.)
)

func (UnitType) String

func (t UnitType) String() string

type Word

type Word struct {
	Units    []*Unit // All parsed units in order
	Original string  // Original input string
	Options  Options // Transliteration options
}

Word is the complete parsed representation of an input word.

func NewWord

func NewWord(original string) *Word

NewWord creates a new empty Word.

func (*Word) AddUnit

func (w *Word) AddUnit(u *Unit)

AddUnit appends a unit and maintains bidirectional links.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL