Documentation
¶
Index ¶
- type Build
- type Context
- func (c *Context) Close() error
- func (c *Context) Embed(text string, normalize bool) ([]float32, error)
- func (c *Context) EmbedTokens(tokens []int32, normalize bool) ([]float32, error)
- func (c *Context) Eval(text string, addSpecial, parseSpecial bool) (EvalResult, error)
- func (c *Context) Generate(prompt string, p Params) (Result, error)
- func (c *Context) GenerateWithDraft(draft *Context, prompt string, p Params, nDraft int) (Result, error)
- func (c *Context) Interrupt() error
- func (c *Context) LoadState(state []byte) error
- func (c *Context) Reset() error
- func (c *Context) SaveState() ([]byte, error)
- func (c *Context) Score(text string) (ScoreResult, error)
- func (c *Context) ScoreChoices(choices []string) ([]ScoreResult, error)
- func (c *Context) SetLoRA(adapters []LoRAWeight) error
- func (c *Context) Stream(prompt string, p Params, onPiece func(string)) (Result, error)
- type ContextParams
- type EvalResult
- type FS
- type Instance
- type Llama
- type LoRA
- type LoRAWeight
- type Message
- type Model
- func (m *Model) ApplyChatTemplate(messages []Message, templateOverride string, addAssistant bool) (string, error)
- func (m *Model) Close() error
- func (m *Model) Detokenize(tokens []int32, renderSpecial bool) (string, error)
- func (m *Model) Info() (ModelInfo, error)
- func (m *Model) LoadLoRA(path string) (*LoRA, error)
- func (m *Model) NewContext(p ContextParams) (*Context, error)
- func (m *Model) TokenToPiece(token int32, renderSpecial bool) (string, error)
- func (m *Model) Tokenize(text string, addSpecial, parseSpecial bool) ([]int32, error)
- type ModelInfo
- type Option
- type Params
- type Result
- type ScoreResult
- type Snapshot
- type SnapshotBuilder
- type StopReason
- type Timings
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Build ¶
type Build struct {
SIMD bool `json:"simd"`
Threads bool `json:"threads"`
Exceptions bool `json:"exceptions"`
MaxDevices int `json:"max_devices"`
}
Build reports how the embedded engine was compiled. Diagnostics only; see BuildInfo.
type Context ¶
type Context struct {
// contains filtered or unexported fields
}
Context is one inference context over a Model: its own KV cache and its own sampling state. Independent conversations get independent contexts over the same weights.
func (*Context) Embed ¶
Embed returns the embedding of text. The context must have been created with ContextParams.Embeddings set. normalize applies L2 normalisation.
func (*Context) EmbedTokens ¶
EmbedTokens is Embed from token ids instead of text.
func (*Context) Eval ¶
func (c *Context) Eval(text string, addSpecial, parseSpecial bool) (EvalResult, error)
Eval decodes text into the context's KV cache without sampling — prompt prefill for a later Generate on the same context, or plain evaluation. Positions continue from the cache's current end; Reset starts over. addSpecial and parseSpecial mirror Tokenize.
func (*Context) GenerateWithDraft ¶
func (c *Context) GenerateWithDraft(draft *Context, prompt string, p Params, nDraft int) (Result, error)
GenerateWithDraft is Generate accelerated by speculative decoding: a draft context over a smaller model with the SAME vocabulary proposes up to nDraft tokens per round (0 picks a default) and this context verifies them in one batch. Every emitted token is sampled by THIS context's sampler chain, so the output distribution is exactly Generate's — the draft only trades its cheap decodes for larger verification batches.
Both contexts' KV caches restart from the prompt. Result.NDrafted and Result.NAccepted report how well the draft anticipated the target.
func (*Context) Interrupt ¶
Interrupt stops a running generation at its next token WITHOUT executing any engine code: it writes the interrupt flag straight into linear memory. Calling in instead would need the lock the generation holds, and would share the C stack with it. Generate then returns what it has, with Reason == StopInterrupted.
Safe to call from any goroutine, including while a generation runs. A call when nothing is running is a no-op: the flag is cleared when generation starts.
func (*Context) LoadState ¶
LoadState restores a state produced by SaveState. The context must be over the same model with compatible parameters, or the engine rejects the payload.
func (*Context) SaveState ¶
SaveState serializes the context's state — KV cache, sampling RNG and the prompt-prefix history — into a byte slice that LoadState can restore, also in a later process. Save a long system prompt's state once and every future context skips re-decoding it: restore the blob and either continue positionally (Eval-prefill style) or generate with Params.CachePrompt, which picks up the saved prefix immediately.
func (*Context) Score ¶
func (c *Context) Score(text string) (ScoreResult, error)
Score computes the teacher-forced negative log-likelihood of text under the model, the quantity llama.cpp's perplexity tool averages. It decodes text into the context's KV cache; call Reset before generating afterwards.
func (*Context) ScoreChoices ¶ added in v0.4.0
func (c *Context) ScoreChoices(choices []string) ([]ScoreResult, error)
ScoreChoices scores candidate continuations of the context's CURRENT cache state: for each choice it returns the negative log-likelihood of the choice's tokens as the continuation of what the context has already decoded. Call it right after decoding the shared stem (Eval, or a Generate whose prompt just ran) — the first token of every choice is scored from the live next-position logits, the rest teacher-forced. Each choice is rolled back out of the KV cache before the next, so the choices never see each other and the context ends exactly where it started. Choices must be non-empty and must not contain newlines (the wire format separates them with '\n').
func (*Context) SetLoRA ¶
func (c *Context) SetLoRA(adapters []LoRAWeight) error
SetLoRA replaces the context's ENTIRE adapter configuration — the set semantics of llama.cpp's own API. An empty (or nil) slice removes every adapter, so SetLoRA(nil) is the spelling of "clear".
func (*Context) Stream ¶
Stream generates from prompt like Generate and calls onPiece with each piece of text as it is decoded, then returns the same complete Result.
onPiece runs inside the generation, on the goroutine that called Stream, so it must not call back into this context or into any other part of the engine. Keep it short: the generation is stopped while it runs.
The pieces concatenate to Result.Text, with one exception: a stop string is delivered as it is decoded and only then trimmed from Result.Text, so with Params.Stop set the stream can run a few characters past the returned text.
A nil onPiece makes this exactly Generate.
type ContextParams ¶
type ContextParams struct {
// NCtx is the context window in tokens. Zero means the model's training
// context length. The KV cache scales with it.
NCtx uint32
// NBatch and NUBatch are the logical and physical batch sizes. Zero means
// llama.cpp's defaults.
NBatch uint32
NUBatch uint32
// NThreads is how many threads ggml may use. It has an effect only in a
// threads-enabled build (BuildInfo.Threads); the single-threaded wasm
// clamps to 1.
NThreads uint32
// NSeqMax is the number of sequences the context can hold at once. A
// value > 1 turns on the unified KV cache: NCtx stays the TOTAL cell
// budget shared by every sequence, and ScoreChoices batches its
// teacher-forced candidates into one decode (one sequence per candidate,
// all sharing the stem). Zero or 1 keeps the single-sequence default,
// where ScoreChoices decodes candidates one at a time.
NSeqMax uint32
// Embeddings puts the context in embedding mode, which Context.Embed
// requires and which disables generation.
Embeddings bool
// RopeFreqBase and RopeFreqScale override the model's RoPE
// configuration, the knobs behind context-length extension schemes.
// Zero keeps the model's own values.
RopeFreqBase float32
RopeFreqScale float32
}
ContextParams configures a Context. The zero value asks for the model's own training context length and llama.cpp's batch defaults.
type EvalResult ¶
EvalResult reports what Eval put into the KV cache: NTokens from this call, NPast the total sequence length now cached.
type FS ¶
FS is the filesystem the guest sees when Config.FS is set. base.NewMemFS returns a ready-made in-memory implementation, so a model can be served out of memory rather than off disk.
type Instance ¶ added in v0.3.0
type Instance struct {
// contains filtered or unexported fields
}
Instance is one engine instance forked from a Snapshot.
type Llama ¶
type Llama struct {
// contains filtered or unexported fields
}
Llama is one engine instance: llama.cpp compiled to wasm, with its own view of linear memory and its own C heap. Load models into it and open contexts over them. Instances are independent — several run concurrently — while within one instance calls are serialised.
Memory an instance never writes is physically shared with the other instances in the process: engines are built over copy-on-write maps of a process-wide image, and instances that load a model an earlier instance already loaded (same options, same path) share the weights the same way instead of loading them again. Set GO_LLAMA_NO_SHARED_IMAGE to force every instance onto a private allocation.
func New ¶
New brings up an engine instance. The zero-option instance gives the guest the host filesystem, an empty environment and discarded stdio — the filesystem access a Go program reading a model file already has.
func (*Llama) Close ¶
Close releases the instance, including the copy-on-write memory mapping when the engine is backed by one. Models and contexts opened from it must be closed first; using any of them afterward fails — the engine's memory is gone.
func (*Llama) LoadModel ¶
LoadModel loads a GGUF model from path into this instance.
path is a guest path. Without WithPreopenDir / WithFS the guest sees the host filesystem, so a relative path is resolved against the working directory first; with a scoped filesystem it is used as given, relative to that root.
type LoRA ¶
type LoRA struct {
// contains filtered or unexported fields
}
LoRA is a loaded LoRA adapter. It adapts contexts over the model that loaded it (Context.SetLoRA) and must not outlive that model.
type LoRAWeight ¶
LoRAWeight pairs an adapter with the scale to apply it at.
type Model ¶
type Model struct {
// contains filtered or unexported fields
}
Model is a loaded GGUF model. Several contexts can share one, and several models can be loaded into one Llama instance at once.
func (*Model) ApplyChatTemplate ¶
func (m *Model) ApplyChatTemplate(messages []Message, templateOverride string, addAssistant bool) (string, error)
ApplyChatTemplate renders messages into a prompt with the model's chat template. templateOverride replaces it, and is required when the GGUF carries none; addAssistant appends the generation prefix.
func (*Model) Detokenize ¶
Detokenize renders tokens back to text.
func (*Model) LoadLoRA ¶
LoadLoRA loads a LoRA adapter GGUF for this model. path resolves like LoadModel's.
func (*Model) NewContext ¶
func (m *Model) NewContext(p ContextParams) (*Context, error)
NewContext creates an inference context over the model.
func (*Model) TokenToPiece ¶
TokenToPiece renders one token. A byte-level token can render to invalid UTF-8 on its own; accumulate pieces before treating the result as text.
type ModelInfo ¶
type ModelInfo struct {
// Desc is llama.cpp's own one-line description, e.g. "llama 7B Q4_K - Medium".
Desc string `json:"desc"`
// NParams is the parameter count; SizeBytes the on-disk tensor size.
NParams uint64 `json:"n_params"`
SizeBytes uint64 `json:"size_bytes"`
// NCtxTrain is the context length the model was trained with — the
// largest context worth asking for.
NCtxTrain int `json:"n_ctx_train"`
NEmbd int `json:"n_embd"`
NLayer int `json:"n_layer"`
NHead int `json:"n_head"`
NHeadKV int `json:"n_head_kv"`
NVocab int `json:"n_vocab"`
// HasEncoder / HasDecoder describe the architecture; a plain causal LM
// has only a decoder.
HasEncoder bool `json:"has_encoder"`
HasDecoder bool `json:"has_decoder"`
// BOSToken / EOSToken are the vocabulary's sentence delimiters, -1
// when the model defines none. AddBOS is whether the tokenizer
// prepends BOSToken by convention (what Tokenize's addSpecial obeys).
BOSToken int32 `json:"bos_token"`
EOSToken int32 `json:"eos_token"`
AddBOS bool `json:"add_bos"`
// ChatTemplate is the Jinja-ish template string the GGUF carries, or ""
// when it has none. Chat needs a template from somewhere: either this or
// SamplingParams-independent override passed to ApplyChatTemplate.
ChatTemplate string `json:"chat_template"`
}
ModelInfo describes a loaded model. Everything here comes from the GGUF metadata, so a field is zero when the file does not carry it.
type Option ¶
type Option func(*config)
An Option configures a Llama at New time.
func WithEnv ¶
WithEnv sets the environment the guest sees. Without it the guest gets an empty environment: the host process environment is never leaked.
func WithFS ¶
WithFS replaces the filesystem backend entirely — feed a model from memory or a virtual tree. It takes precedence over WithPreopenDir. base.NewMemFS returns a ready-made in-memory implementation.
func WithMaxMemory ¶
WithMaxMemory caps linear-memory growth, so a model bigger than expected fails inside the guest instead of growing the host process. Zero means the engine's default ceiling: for an instance backed by a copy-on-write mapping that is the 64 GiB of address space the mapping reserves (untouched pages cost nothing), and for a private allocation there is no cap.
func WithMemoryReserve ¶
WithMemoryReserve pre-reserves linear memory. Loading a multi-gigabyte model otherwise grows memory in steps, copying it forward each time; reserving up front avoids that. Zero uses the engine's default.
func WithPreopenDir ¶
WithPreopenDir scopes the guest filesystem to a host directory. Guest paths then resolve inside it, so a model is named relative to it. Without it the guest sees the whole host filesystem, where absolute paths work as written.
func WithStderr ¶
func WithStdout ¶
WithStdout and WithStderr capture what the guest writes to fd 1 and fd 2. Unset discards, the default, because llama.cpp's own logging is silenced inside the bridge anyway.
type Params ¶
type Params struct {
// NPredict caps the number of tokens generated. Zero or negative means
// "as many as the context allows".
NPredict int
// Temperature <= 0 selects greedy decoding, which makes generation
// deterministic and skips the truncation samplers below.
Temperature float32
// TopK, TopP, MinP and TypicalP truncate the candidate set before
// sampling. Zero disables each of them, as does 1.0 for the two
// probability-mass ones.
TopK int
TopP float32
MinP float32
TypicalP float32
// RepeatPenalty, PresencePenalty and FrequencyPenalty discourage
// repetition over the last RepeatLastN tokens. Zero disables all three;
// so does RepeatPenalty 1.0, which is llama.cpp's own spelling of "off".
RepeatPenalty float32
RepeatLastN int
PresencePenalty float32
FrequencyPenalty float32
// Seed makes sampling reproducible. Zero is a fixed seed, so a
// temperature above zero still replays identically; set a varying seed to
// get varying output.
Seed uint32
// Mirostat selects the mirostat sampling algorithm: 0 off, 1 v1, 2 v2.
// When on it replaces the TopK/TopP/MinP/TypicalP truncation samplers,
// as in llama.cpp. MirostatTau is the target entropy (0 means llama.cpp's
// 5.0) and MirostatEta the learning rate (0 means 0.1).
Mirostat int
MirostatTau float32
MirostatEta float32
// IgnoreEOS keeps generating past end-of-generation tokens by excluding
// them from sampling, like llama.cpp's --ignore-eos. Generation then runs
// to NPredict, a stop string, or the context limit.
IgnoreEOS bool
// LogitBias adds a bias to specific tokens' logits before sampling.
// math.Inf(-1) (or any very negative value) forbids a token outright.
LogitBias map[int32]float32
// Grammar is a GBNF grammar constraining the output. Empty means none.
Grammar string
// Stop ends generation when the output first contains one of these
// strings — even mid-token — and the text is cut at the match.
Stop []string
// CachePrompt treats the prompt as the WHOLE intended context: the
// longest prefix already in the context's KV cache is kept, whatever the
// cache holds beyond it is dropped, and only the rest is decoded —
// llama.cpp server's cache_prompt. Requests that share a long constant
// preamble (a system prompt, a routing configuration) then pay only for
// the part that changed; Result.NCached reports the reuse. Off, the
// prompt appends at the cache's current end, which is what an
// Eval-prefill continuation expects.
CachePrompt bool
}
Params configures one generation.
The zero value is greedy decoding with no truncation samplers, no penalties and no token limit: the most predictable thing the model can do, and reproducible run to run. Every field is honoured exactly as written — a zero is a decision ("no top-k", "temperature 0 means greedy"), not "use some default" — so what a caller leaves out cannot be overridden by a default buried in the engine.
type Result ¶
type Result struct {
// Text is the generated text, with a matched stop string removed.
Text string `json:"text"`
// Tokens are the tokens behind Text.
Tokens []int32 `json:"tokens"`
// NPrompt is how many tokens the prompt occupied; NDecoded how many were
// generated.
NPrompt int `json:"n_prompt"`
NDecoded int `json:"n_decoded"`
// NCached is how many leading prompt tokens a Params.CachePrompt
// generation reused from the KV cache instead of re-decoding. Zero
// without CachePrompt.
NCached int `json:"n_cached"`
// NDrafted / NAccepted are GenerateWithDraft's speculation counters —
// how many tokens the draft proposed and how many the target accepted.
// Zero on a plain Generate.
NDrafted int `json:"n_drafted"`
NAccepted int `json:"n_accepted"`
// Reason says why generation stopped; Interrupted is a shorthand for
// Reason == StopInterrupted.
Reason StopReason `json:"stop_reason"`
Interrupted bool `json:"interrupted"`
// Timings splits the generation's wall time between the prompt pass
// and the decode loop.
Timings Timings `json:"timings"`
}
Result is what Generate returns.
type ScoreResult ¶
type ScoreResult struct {
NTokens int32 `json:"n_tokens"`
NScored int32 `json:"n_scored"`
NLL float64 `json:"nll"`
}
ScoreResult reports the teacher-forced negative log-likelihood of a text: the model decodes the tokenized text once and NLL sums -log softmax(logits_i)[token_{i+1}] over the NScored predicting positions. Perplexity is math.Exp(NLL / NScored).
type Snapshot ¶ added in v0.3.0
type Snapshot struct {
// contains filtered or unexported fields
}
Snapshot is the copy-on-write image of a prepared instance. It stays valid for the life of the process; forks reference it and share its pages.
func NewSnapshot ¶ added in v0.3.0
func NewSnapshot(build func(*SnapshotBuilder) error, opts ...Option) (*Snapshot, error)
NewSnapshot boots an engine on a snapshot image and hands it to build to prepare: load the model, create contexts, decode their prompt prefixes (Params.CachePrompt), and Register the ones forks will use. When build returns, every kept context's threadpool is detached (threads are host constructs and cannot be captured) and the image is sealed.
The builder instance itself is consumed by the build; do not retain the *Llama, models or contexts it produced beyond the callback.
type SnapshotBuilder ¶ added in v0.3.0
type SnapshotBuilder struct {
*Llama
// contains filtered or unexported fields
}
SnapshotBuilder is the instance handed to a NewSnapshot callback: a normal *Llama plus Register, which records the contexts forks will use.
type StopReason ¶
type StopReason string
StopReason says why generation ended.
const ( // StopEOS: the model emitted an end-of-generation token. StopEOS StopReason = "eos" // StopLength: the token budget ran out (Params.NPredict, or the context // filled up). StopLength StopReason = "length" // StopString: one of Params.Stop matched the tail of the output, which // was trimmed off. StopString StopReason = "stop" // StopInterrupted: Context.Interrupt was called from another goroutine. StopInterrupted StopReason = "interrupted" )
Directories
¶
| Path | Synopsis |
|---|---|
|
cmd/perfgate
command
Command perfgate is the CI entry point of internal/perfgate.
|
Command perfgate is the CI entry point of internal/perfgate. |
|
perfgate
Package perfgate compares go-llama's throughput against a native llama.cpp build measured on the same machine, and gates CI on the ratio: a change that makes go-llama more than a threshold slower RELATIVE TO NATIVE than the recorded baseline fails.
|
Package perfgate compares go-llama's throughput against a native llama.cpp build measured on the same machine, and gates CI on the ratio: a change that makes go-llama more than a threshold slower RELATIVE TO NATIVE than the recorded baseline fails. |