Documentation
¶
Overview ¶
Package core provides universal types and mechanics for transliteration. This package is script-agnostic - it knows nothing about Brahmic, Arabic, etc.
Index ¶
- type Category
- type DebugInfo
- type Engine
- func (e *Engine) Language() Language
- func (e *Engine) RuleEngine() *RuleEngine
- func (e *Engine) Scheme() Scheme
- func (e *Engine) Transliterate(input string) string
- func (e *Engine) TransliterateDebug(input string, opts Options) (string, *DebugInfo)
- func (e *Engine) TransliterateWithOptions(input string, opts Options) string
- type EngineOption
- type Language
- type LexiconProvider
- type Options
- type Parser
- type Position
- type Renderer
- type Reranker
- type Rule
- type RuleCatalog
- type RuleEngine
- func (e *RuleEngine) AddRule(r Rule) error
- func (e *RuleEngine) Apply(word *Word)
- func (e *RuleEngine) DisableRule(pattern string) int
- func (e *RuleEngine) EnableDebug(enabled bool)
- func (e *RuleEngine) EnableRule(pattern string) int
- func (e *RuleEngine) IsDisabled(name string) bool
- func (e *RuleEngine) ListRules(pattern string) []RuleStatus
- func (e *RuleEngine) Rules() []Rule
- func (e *RuleEngine) RulesForPhase(phase RulePhase) []Rule
- func (e *RuleEngine) SetDebugMetaExtractor(fn func(*Unit) string)
- func (e *RuleEngine) Traces() []RuleTrace
- type RuleMode
- type RulePhase
- type RuleScope
- type RuleStatus
- type RuleTrace
- type Scheme
- type Script
- type SymbolInfo
- type SymbolMap
- type Unit
- type UnitDebug
- type UnitType
- type Word
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Category ¶
type Category int
Category classifies symbols for parsing. Base categories are defined here; scripts can extend with their own.
type DebugInfo ¶
type DebugInfo struct {
Input string // Original input
Output string // Final output
Units []UnitDebug // Parsed units
Traces []RuleTrace // Rule applications
}
DebugInfo contains debugging information from transliteration.
type Engine ¶
type Engine struct {
// contains filtered or unexported fields
}
Engine orchestrates the transliteration pipeline.
func NewEngine ¶
func NewEngine(lang Language, scheme Scheme, opts ...EngineOption) *Engine
NewEngine creates an engine for a language + scheme combination. Panics if the language returns a nil Script.
func (*Engine) RuleEngine ¶
func (e *Engine) RuleEngine() *RuleEngine
RuleEngine returns the engine's rule engine for direct manipulation. Use this to enable/disable rules after engine creation.
func (*Engine) Transliterate ¶
Transliterate converts script text to romanized form using default options.
func (*Engine) TransliterateDebug ¶
TransliterateDebug converts script text and returns debug information.
type EngineOption ¶
type EngineOption func(*engineConfig)
EngineOption configures engine creation.
func WithDisabledRules ¶
func WithDisabledRules(patterns ...string) EngineOption
WithDisabledRules disables rules matching the given patterns at engine creation. Patterns can be exact names or glob patterns (e.g., "schwa.*", "vowel.long-aa.*").
func WithEnabledRules ¶
func WithEnabledRules(patterns ...string) EngineOption
WithEnabledRules enables rules matching the given patterns at engine creation. Useful for enabling rules that are disabled by default. Patterns can be exact names or glob patterns (e.g., "vowel.long-aa.all").
type Language ¶
type Language interface {
// Name returns the language identifier (e.g., "hindi", "marathi").
Name() string
// Script returns the script used by this language.
Script() Script
// Symbols returns the symbol map for this language.
Symbols() SymbolMap
// ScriptConfig returns script-specific configuration.
// For Brahmic: brahmic.Config{Halant: "्", Nukta: "़"}
ScriptConfig() interface{}
// Rules returns the complete rule catalog for this language.
// This includes ALL possible rules; schemes select from this catalog.
Rules() RuleCatalog
}
Language defines language-specific behavior within a script.
type LexiconProvider ¶
LexiconProvider is an optional interface a Language may implement to supply a word-level romanization lexicon. When Options.Lexicon is set, the engine consults it before the rule pipeline; a hit short-circuits transliteration.
type Options ¶
type Options struct {
// LongVowels outputs "aa" for all ā (aa-matra) positions.
LongVowels bool
// SimpleNasals uses simplified nasal endings (करें→karen instead of karein).
SimpleNasals bool
// KeepMedialSchwa disables CCV schwa deletion, retaining schwa in more positions.
// This produces output closer to some datasets (जनता→janata instead of janta).
KeepMedialSchwa bool
// SchwaModel uses the learned decision-tree schwa classifier for inherent-schwa
// decisions instead of the hand-written heuristic schwa rules. See
// lang/hindi/schwa_model.go and docs/reviews for the held-out evaluation.
SchwaModel bool
// Lexicon consults the language's high-confidence romanization lexicon first;
// known words return the attested human spelling, unknown words fall through
// to the rule engine. Requires the Language to implement LexiconProvider.
Lexicon bool
// Rerank generates candidate romanizations under several rule configurations
// and picks the one a character language model finds most natural. Requires
// the Language to implement Reranker. A lexicon hit (with Lexicon) wins first.
Rerank bool
// Debug enables debug output showing rule applications.
Debug bool
}
Options configures transliteration behavior.
func DefaultOptions ¶
func DefaultOptions() Options
DefaultOptions returns the default transliteration options.
type Parser ¶
type Parser interface {
// Parse converts input text into a Word structure.
Parse(input string, symbols SymbolMap) *Word
}
Parser converts script input into Units.
type Position ¶
type Position struct {
Offset int // Byte offset in original string
Rune int // Rune (character) index
}
Position tracks location in source text for debugging.
type Renderer ¶
type Renderer interface {
// Render converts a Word into a romanized string.
Render(word *Word) string
}
Renderer converts Units into romanized output.
type Reranker ¶
Reranker is an optional interface a Language may implement to choose the most natural romanization among candidates produced under different rule configurations. Used when Options.Rerank is set.
type Rule ¶
type Rule struct {
Name string
Phase RulePhase
Scope RuleScope
Priority int // 0-99 within scope
Mode RuleMode
DisabledDefault bool // If true, rule is disabled by default (must be explicitly enabled)
Conditional string // If non-empty, rule depends on this runtime option (e.g., "LongVowels")
Condition func(*Unit, *Word) bool
Action func(*Unit, *Word)
}
Rule represents a single transliteration rule.
func AppendIfFound ¶
AppendIfFound appends a rule to the slice if found by name. Returns the (possibly extended) slice.
func FindInSlice ¶
FindInSlice finds a rule by name in a slice of rules. Returns nil if not found.
func (*Rule) EffectivePriority ¶
EffectivePriority calculates the overall priority across scopes. Higher values run first.
type RuleCatalog ¶
type RuleCatalog struct {
Schwa []Rule // All schwa-related rules
Consonant []Rule // All consonant rules
Vowel []Rule // All vowel rules
Render []Rule // All output transform rules
}
RuleCatalog organizes rules by category. Languages provide complete catalogs; schemes select from them.
func (RuleCatalog) AllRules ¶
func (c RuleCatalog) AllRules() []Rule
AllRules returns all rules in the catalog as a flat slice.
func (RuleCatalog) FindByName ¶
func (c RuleCatalog) FindByName(name string) *Rule
FindByName finds a rule by name across all categories. Returns nil if not found.
type RuleEngine ¶
type RuleEngine struct {
// contains filtered or unexported fields
}
RuleEngine applies rules to parsed words.
func NewRuleEngine ¶
func NewRuleEngine(rules []Rule) *RuleEngine
NewRuleEngine creates a new rule engine with the given rules. Panics if any rule has nil Condition or Action functions.
func (*RuleEngine) AddRule ¶
func (e *RuleEngine) AddRule(r Rule) error
AddRule adds a rule to the engine. Returns error if there's a priority conflict or invalid rule.
func (*RuleEngine) Apply ¶
func (e *RuleEngine) Apply(word *Word)
Apply executes all rules on the word in phase order.
func (*RuleEngine) DisableRule ¶
func (e *RuleEngine) DisableRule(pattern string) int
DisableRule disables rules matching the given pattern. Pattern can be an exact name or a glob pattern using '*' as wildcard. Examples: "schwa.delete.ccv", "schwa.*", "schwa.delete.*" Returns the count of rules matched and disabled.
func (*RuleEngine) EnableDebug ¶
func (e *RuleEngine) EnableDebug(enabled bool)
EnableDebug enables debug trace collection.
func (*RuleEngine) EnableRule ¶
func (e *RuleEngine) EnableRule(pattern string) int
EnableRule enables rules matching the given pattern. Pattern can be an exact name or a glob pattern using '*' as wildcard. Examples: "vowel.long-aa.all", "vowel.*", "vowel.long-aa.*" Returns the count of rules matched and enabled.
func (*RuleEngine) IsDisabled ¶
func (e *RuleEngine) IsDisabled(name string) bool
IsDisabled returns true if the rule with the given name is currently disabled.
func (*RuleEngine) ListRules ¶
func (e *RuleEngine) ListRules(pattern string) []RuleStatus
ListRules returns all rules with their enabled/disabled status. If pattern is non-empty, only rules matching the pattern are returned.
func (*RuleEngine) Rules ¶
func (e *RuleEngine) Rules() []Rule
Rules returns a copy of all registered rules.
func (*RuleEngine) RulesForPhase ¶
func (e *RuleEngine) RulesForPhase(phase RulePhase) []Rule
RulesForPhase returns a copy of rules for a specific phase, sorted by priority.
func (*RuleEngine) SetDebugMetaExtractor ¶
func (e *RuleEngine) SetDebugMetaExtractor(fn func(*Unit) string)
SetDebugMetaExtractor sets a function to extract script-specific metadata.
func (*RuleEngine) Traces ¶
func (e *RuleEngine) Traces() []RuleTrace
Traces returns collected debug traces (call after Apply).
type RuleScope ¶
type RuleScope int
RuleScope determines the priority tier for a rule. Higher scopes have higher effective priority.
type RuleStatus ¶
type RuleStatus struct {
Name string
Phase RulePhase
Scope RuleScope
Priority int
Enabled bool
Conditional string // If non-empty, rule depends on this runtime option
}
RuleStatus represents a rule and its current enabled/disabled state.
type RuleTrace ¶
type RuleTrace struct {
Phase string // Phase name (Schwa, Consonant, Vowel, Render)
Rule string // Rule name
Unit string // Unit characters (e.g., "क")
UnitIdx int // Unit index in word
Before string // BaseRom before rule
After string // BaseRom after rule (empty if unchanged)
Metadata string // Additional info (e.g., schwa state)
}
RuleTrace records a single rule application for debugging.
type Scheme ¶
type Scheme interface {
// Name returns the scheme identifier (e.g., "colloquial", "iast").
Name() string
// SelectRules selects rules from the language's catalog.
// Returns the rules to apply for this scheme.
SelectRules(catalog RuleCatalog) []Rule
}
Scheme selects which rules to use from a language's catalog. Schemes are language-agnostic - they work with any language.
type Script ¶
type Script interface {
// Name returns the script identifier (e.g., "brahmic", "arabic").
Name() string
// NewParser creates a parser for this script.
// The config parameter is script-specific (e.g., brahmic.Config).
NewParser(config interface{}) Parser
// NewRenderer creates a renderer for this script.
NewRenderer() Renderer
// PrepareWord performs script-specific processing after parsing.
// For example, Brahmic scripts identify consonant runs here.
PrepareWord(word *Word)
// Categories returns script-specific categories.
Categories() []Category
// DebugMetaExtractor returns a function that extracts script-specific
// metadata for debugging (e.g., schwa state for Brahmic).
// Returns nil if no metadata is available.
DebugMetaExtractor() func(*Unit) string
}
Script defines a script family (Brahmic, Arabic, Latin, etc.). Scripts provide parsing and rendering capabilities.
type SymbolInfo ¶
SymbolInfo holds romanization info for a symbol.
type SymbolMap ¶
type SymbolMap map[string]SymbolInfo
SymbolMap maps script characters to their romanization info.
type Unit ¶
type Unit struct {
// Source tracking
Runes []rune // Original characters (1-3 for conjuncts)
Start Position // Start position in source
End Position // End position in source
// Classification
Type UnitType
BaseRom string // Base romanization (modifiable by rules)
// Navigation (bidirectional links)
Prev *Unit
Next *Unit
// Script-specific extension point
// Scripts store their own data here (e.g., SchwaState for Brahmic)
ScriptData interface{}
}
Unit represents a single phonetic unit in the parsed word.
func (*Unit) IsWordFinal ¶
IsWordFinal returns true if this is the last unit in the word.
func (*Unit) IsWordInitial ¶
IsWordInitial returns true if this is the first unit in the word.
type UnitDebug ¶
type UnitDebug struct {
Index int // Position in word
Chars string // Original characters
Type string // Unit type
BaseRom string // Base romanization
RunePos int // Rune position in original string
Metadata string // Script-specific info
}
UnitDebug contains debug info for a single unit.
type UnitType ¶
type UnitType int
UnitType classifies parsed phonetic units.
const ( UnitVowel UnitType = iota // Vowels and matras that replace inherent schwa UnitModifier // Modifiers that follow the vowel (anusvara, visarga, chandrabindu) UnitConsonant // Single consonant UnitConjunct // Multi-character conjunct (e.g., ज्ञ) UnitNumber // Numerals UnitSymbol // Other symbols (punctuation, etc.) )