Documentation
¶
Index ¶
- Variables
- type Build
- type Context
- func (c *Context) Close() error
- func (c *Context) Embed(text string, normalize bool) ([]float32, error)
- func (c *Context) EmbedTokens(tokens []int32, normalize bool) ([]float32, error)
- func (c *Context) Eval(text string, addSpecial, parseSpecial bool) (EvalResult, error)
- func (c *Context) Generate(prompt string, p Params) (Result, error)
- func (c *Context) GenerateWithDraft(draft *Context, prompt string, p Params, nDraft int) (Result, error)
- func (c *Context) Interrupt() error
- func (c *Context) LoadState(state []byte) error
- func (c *Context) Model() *Model
- func (c *Context) Reset() error
- func (c *Context) SaveState() ([]byte, error)
- func (c *Context) Score(text string) (ScoreResult, error)
- func (c *Context) ScoreChoices(choices []string) ([]ScoreResult, error)
- func (c *Context) SetLoRA(adapters []LoRAWeight) error
- func (c *Context) Slots() (*Slots, error)
- func (c *Context) Stream(prompt string, p Params, onPiece func(string)) (Result, error)
- type ContextParams
- type EvalResult
- type FS
- type Instance
- type Llama
- type LoRA
- type LoRAWeight
- type Message
- type Model
- func (m *Model) ApplyChatTemplate(messages []Message, templateOverride string, addAssistant bool) (string, error)
- func (m *Model) Close() error
- func (m *Model) Detokenize(tokens []int32, renderSpecial bool) (string, error)
- func (m *Model) Info() (ModelInfo, error)
- func (m *Model) LoadLoRA(path string) (*LoRA, error)
- func (m *Model) NewContext(p ContextParams) (*Context, error)
- func (m *Model) Tensors() ([]TensorInfo, error)
- func (m *Model) TokenToPiece(token int32, renderSpecial bool) (string, error)
- func (m *Model) Tokenize(text string, addSpecial, parseSpecial bool) ([]int32, error)
- type ModelInfo
- type Option
- type Params
- type Result
- type ScoreResult
- type Slots
- type SlotsStatus
- type Snapshot
- type SnapshotBuilder
- type StopReason
- type Task
- type TaskCompletion
- type TaskOutput
- type TensorInfo
- type Timings
Constants ¶
This section is empty.
Variables ¶
var ErrInstanceClosed = bridge.ErrEngineClosed
ErrInstanceClosed is in the error of every call on a closed instance, and of a Slots task the instance's Close ended.
var ErrMemoryLeaked = errors.New("llama: close: memory kept mapped, worker threads may still run in it")
ErrMemoryLeaked is returned by Close when the instance's memory was deliberately kept mapped, because a trap left it unknown whether the instance's worker threads had been joined (see Llama.Close).
var ErrSlotsClosed = errors.New("llama: slots are closed")
ErrSlotsClosed is the error of every task a Slots dropped when it was closed, and of a Post after Close.
var ErrSlotsRunning = errors.New("llama: slots: the context already has a scheduler; close it first")
ErrSlotsRunning is returned by Context.Slots while an earlier Slots of the context is still open.
var ErrTaskRunning = errors.New("llama: task is still running")
ErrTaskRunning is returned by TaskCompletion.Result while the task has not finished.
Functions ¶
This section is empty.
Types ¶
type Build ¶
type Build struct {
SIMD bool `json:"simd"`
Threads bool `json:"threads"`
Exceptions bool `json:"exceptions"`
MaxDevices int `json:"max_devices"`
}
Build reports how the embedded engine was compiled. Diagnostics only; see BuildInfo.
type Context ¶
type Context struct {
// contains filtered or unexported fields
}
Context is one inference context over a Model: its own KV cache and its own sampling state. Independent conversations get independent contexts over the same weights.
func (*Context) Close ¶
Close frees the context, stopping and joining its threadpool workers. A generation in flight on the context is interrupted (it returns at its next token, with StopInterrupted) and any other call waited for. A Slots still scheduling on the context is closed too; its tasks end with ErrSlotsClosed. The Model outlives it.
func (*Context) Embed ¶
Embed returns the embedding of text. The context must have been created with ContextParams.Embeddings set. normalize applies L2 normalisation.
func (*Context) EmbedTokens ¶
EmbedTokens is Embed from token ids instead of text.
func (*Context) Eval ¶
func (c *Context) Eval(text string, addSpecial, parseSpecial bool) (EvalResult, error)
Eval decodes text into the context's KV cache without sampling — prompt prefill for a later Generate on the same context, or plain evaluation. Positions continue from the cache's current end; Reset starts over. addSpecial and parseSpecial mirror Tokenize.
func (*Context) GenerateWithDraft ¶
func (c *Context) GenerateWithDraft(draft *Context, prompt string, p Params, nDraft int) (Result, error)
GenerateWithDraft is Generate accelerated by speculative decoding: a draft context over a smaller model with the SAME vocabulary proposes up to nDraft tokens per round (0 picks a default) and this context verifies them in one batch. Every emitted token is sampled by THIS context's sampler chain, so the output distribution is exactly Generate's — the draft only trades its cheap decodes for larger verification batches.
Both contexts' KV caches restart from the prompt. Result.NDrafted and Result.NAccepted report how well the draft anticipated the target.
func (*Context) Interrupt ¶
Interrupt stops a running generation at its next token WITHOUT executing any engine code: it writes the interrupt flag straight into linear memory. Calling in instead would need the lock the generation holds, and would share the C stack with it. Generate then returns what it has, with Reason == StopInterrupted.
Safe to call from any goroutine, including while a generation runs, and while a Close is waiting for one (Close interrupts it itself). A call when nothing is running is a no-op: the flag is cleared when generation starts.
func (*Context) LoadState ¶
LoadState restores a state produced by SaveState. The context must be over the same model with compatible parameters, or the engine rejects the payload.
func (*Context) SaveState ¶
SaveState serializes the context's state — KV cache, sampling RNG and the prompt-prefix history — into a byte slice that LoadState can restore, also in a later process. Save a long system prompt's state once and every future context skips re-decoding it: restore the blob and either continue positionally (Eval-prefill style) or generate with Params.CachePrompt, which picks up the saved prefix immediately.
func (*Context) Score ¶
func (c *Context) Score(text string) (ScoreResult, error)
Score computes the teacher-forced negative log-likelihood of text under the model, the quantity llama.cpp's perplexity tool averages. It decodes text into the context's KV cache; call Reset before generating afterwards.
func (*Context) ScoreChoices ¶ added in v0.4.0
func (c *Context) ScoreChoices(choices []string) ([]ScoreResult, error)
ScoreChoices scores candidate continuations of the context's CURRENT cache state: for each choice it returns the negative log-likelihood of the choice's tokens as the continuation of what the context has already decoded. Call it right after decoding the shared stem (Eval, or a Generate whose prompt just ran) — the first token of every choice is scored from the live next-position logits, the rest teacher-forced. Each choice is rolled back out of the KV cache before the next, so the choices never see each other and the context ends exactly where it started. Choices must be non-empty and must not contain newlines (the wire format separates them with '\n').
func (*Context) SetLoRA ¶
func (c *Context) SetLoRA(adapters []LoRAWeight) error
SetLoRA replaces the context's ENTIRE adapter configuration — the set semantics of llama.cpp's own API. An empty (or nil) slice removes every adapter, so SetLoRA(nil) is the spelling of "clear".
func (*Context) Slots ¶ added in v0.5.0
Slots starts a scheduler over the context's slots. The context must have been created with NSeqMax set to the number of slots wanted (1 still works: tasks then run one at a time, each on its own sequence). A context runs one Slots at a time: a second call while the first is open returns ErrSlotsRunning. Closing the context closes its Slots.
func (*Context) Stream ¶
Stream generates from prompt like Generate and calls onPiece with each piece of text as it is decoded, then returns the same complete Result.
onPiece runs inside the generation, on the goroutine that called Stream, so it must not call back into this context or into any other part of the engine. Keep it short: the generation is stopped while it runs.
The pieces concatenate to Result.Text, with one exception: a stop string is delivered as it is decoded and only then trimmed from Result.Text, so with Params.Stop set the stream can run a few characters past the returned text.
A nil onPiece makes this exactly Generate.
type ContextParams ¶
type ContextParams struct {
// NCtx is the context window in tokens. Zero means the model's training
// context length. The KV cache scales with it.
NCtx uint32
// NBatch and NUBatch are the logical and physical batch sizes. Zero means
// llama.cpp's defaults.
NBatch uint32
NUBatch uint32
// NThreads is how many threads ggml may use. It has an effect only in a
// threads-enabled build (BuildInfo.Threads); the single-threaded wasm
// clamps to 1.
NThreads uint32
// NSeqMax is the number of sequences the context can hold at once: the
// slots of a Slots, or the candidates ScoreChoices decodes in one batch
// (one sequence per candidate, all sharing the stem). Zero or 1 is the
// single-sequence default, where ScoreChoices decodes candidates one at
// a time.
NSeqMax uint32
// KVUnified chooses how the NSeqMax sequences share the KV cache, as
// llama.cpp's kv_unified does. False, the default, gives each sequence
// its own stream of NCtx / NSeqMax cells: attention over a sequence then
// costs only its own cells, which is what independent tasks on a Slots
// want. True shares one buffer of NCtx cells across the sequences, where
// copying a sequence is metadata rather than a copy: what ScoreChoices'
// batched path leans on (its candidates all share the stem), and what a
// prompt prefix shared across slots needs. The measured difference for
// 64 independent tasks was 903 vs 500 tok/s in favour of streams; for
// ScoreChoices with NSeqMax > 1, set it.
KVUnified bool
// Embeddings puts the context in embedding mode, which Context.Embed
// requires and which disables generation.
Embeddings bool
// RopeFreqBase and RopeFreqScale override the model's RoPE
// configuration, the knobs behind context-length extension schemes.
// Zero keeps the model's own values.
RopeFreqBase float32
RopeFreqScale float32
}
ContextParams configures a Context. The zero value asks for the model's own training context length and llama.cpp's batch defaults.
type EvalResult ¶
EvalResult reports what Eval put into the KV cache: NTokens from this call, NPast the total sequence length now cached.
type FS ¶
FS is the filesystem the guest sees when Config.FS is set. base.NewMemFS returns a ready-made in-memory implementation, so a model can be served out of memory rather than off disk.
type Instance ¶ added in v0.3.0
type Instance struct {
// contains filtered or unexported fields
}
Instance is one engine instance forked from a Snapshot.
func (*Instance) Close ¶ added in v0.3.0
Close releases the fork: its kept contexts are freed, which joins the threadpool workers Fork attached, and then its copy-on-write mapping is unmapped, returning every private page. The kept contexts need no Close of their own; one that was closed already is simply skipped.
If freeing a context trapped, whether its workers were joined is unknown, and the mapping is kept rather than unmapped under threads that may still run in it: Close then returns ErrMemoryLeaked along with the context's error. Any error of a kept context's Close is returned.
type Llama ¶
type Llama struct {
// contains filtered or unexported fields
}
Llama is one engine instance: llama.cpp compiled to wasm, with its own view of linear memory and its own C heap. Load models into it and open contexts over them. Instances are independent — several run concurrently — while within one instance calls are serialised.
Memory an instance never writes is physically shared with the other instances in the process: engines are built over copy-on-write maps of a process-wide image, and instances that load a model an earlier instance already loaded (same options, same path) share the weights the same way instead of loading them again. Set GO_LLAMA_NO_SHARED_IMAGE to force every instance onto a private allocation.
func New ¶
New brings up an engine instance. The zero-option instance gives the guest the host filesystem, an empty environment and discarded stdio — the filesystem access a Go program reading a model file already has.
func (*Llama) Close ¶
Close releases the instance, including the copy-on-write memory mapping when the engine is backed by one. Contexts still open are closed first — a generation running on one is interrupted and waited for, and the free joins the context's threadpool workers, which must not outlive the memory they run in — and their Close errors are returned; models need no Close of their own. Using anything of the instance afterward fails with ErrInstanceClosed.
If freeing a context trapped, whether its workers were joined is unknown, and unmapping the memory under threads that may still run in it would crash the process: Close then leaves the memory mapped for the life of the process and reports ErrMemoryLeaked. The instance is closed either way.
func (*Llama) LoadModel ¶
LoadModel loads a GGUF model from path into this instance.
path is a guest path. Without WithPreopenDir / WithFS the guest sees the host filesystem, so a relative path is resolved against the working directory first; with a scoped filesystem it is used as given, relative to that root.
type LoRA ¶
type LoRA struct {
// contains filtered or unexported fields
}
LoRA is a loaded LoRA adapter. It adapts contexts over the model that loaded it (Context.SetLoRA) and must not outlive that model.
type LoRAWeight ¶
LoRAWeight pairs an adapter with the scale to apply it at.
type Model ¶
type Model struct {
// contains filtered or unexported fields
}
Model is a loaded GGUF model. Several contexts can share one, and several models can be loaded into one Llama instance at once.
func (*Model) ApplyChatTemplate ¶
func (m *Model) ApplyChatTemplate(messages []Message, templateOverride string, addAssistant bool) (string, error)
ApplyChatTemplate renders messages into a prompt with the model's chat template. templateOverride replaces it, and is required when the GGUF carries none; addAssistant appends the generation prefix.
func (*Model) Close ¶
Close frees the model, closing its contexts still open first: a context must not outlive its model. Their Close errors are returned. Nothing to do once the instance's Close has begun: that frees every context and then the memory the model is in.
func (*Model) Detokenize ¶
Detokenize renders tokens back to text.
func (*Model) LoadLoRA ¶
LoadLoRA loads a LoRA adapter GGUF for this model. path resolves like LoadModel's.
func (*Model) NewContext ¶
func (m *Model) NewContext(p ContextParams) (*Context, error)
NewContext creates an inference context over the model.
func (*Model) Tensors ¶ added in v0.5.0
func (m *Model) Tensors() ([]TensorInfo, error)
Tensors lists the model's weight tensors with the buffer each landed in, which decides the kernel path (repacked GEMV/GEMM versus per-row dot).
func (*Model) TokenToPiece ¶
TokenToPiece renders one token. A byte-level token can render to invalid UTF-8 on its own; accumulate pieces before treating the result as text.
type ModelInfo ¶
type ModelInfo struct {
// Desc is llama.cpp's own one-line description, e.g. "llama 7B Q4_K - Medium".
Desc string `json:"desc"`
// NParams is the parameter count; SizeBytes the on-disk tensor size.
NParams uint64 `json:"n_params"`
SizeBytes uint64 `json:"size_bytes"`
// NCtxTrain is the context length the model was trained with — the
// largest context worth asking for.
NCtxTrain int `json:"n_ctx_train"`
NEmbd int `json:"n_embd"`
NLayer int `json:"n_layer"`
NHead int `json:"n_head"`
NHeadKV int `json:"n_head_kv"`
NVocab int `json:"n_vocab"`
// HasEncoder / HasDecoder describe the architecture; a plain causal LM
// has only a decoder.
HasEncoder bool `json:"has_encoder"`
HasDecoder bool `json:"has_decoder"`
// BOSToken / EOSToken are the vocabulary's sentence delimiters, -1
// when the model defines none. AddBOS is whether the tokenizer
// prepends BOSToken by convention (what Tokenize's addSpecial obeys).
BOSToken int32 `json:"bos_token"`
EOSToken int32 `json:"eos_token"`
AddBOS bool `json:"add_bos"`
// ChatTemplate is the Jinja-ish template string the GGUF carries, or ""
// when it has none. Chat needs a template from somewhere: either this or
// SamplingParams-independent override passed to ApplyChatTemplate.
ChatTemplate string `json:"chat_template"`
}
ModelInfo describes a loaded model. Everything here comes from the GGUF metadata, so a field is zero when the file does not carry it.
type Option ¶
type Option func(*config)
An Option configures a Llama at New time.
func WithEnv ¶
WithEnv sets the environment the guest sees. Without it the guest gets an empty environment: the host process environment is never leaked.
func WithFS ¶
WithFS replaces the filesystem backend entirely — feed a model from memory or a virtual tree. It takes precedence over WithPreopenDir. base.NewMemFS returns a ready-made in-memory implementation.
func WithMaxMemory ¶
WithMaxMemory caps linear-memory growth, so a model bigger than expected fails inside the guest instead of growing the host process. Zero means the engine's default ceiling: for an instance backed by a copy-on-write mapping that is the 64 GiB of address space the mapping reserves (untouched pages cost nothing), and for a private allocation there is no cap.
func WithMemoryReserve ¶
WithMemoryReserve pre-reserves linear memory. Loading a multi-gigabyte model otherwise grows memory in steps, copying it forward each time; reserving up front avoids that. Zero uses the engine's default.
func WithPreopenDir ¶
WithPreopenDir scopes the guest filesystem to a host directory. Guest paths then resolve inside it, so a model is named relative to it. Without it the guest sees the whole host filesystem, where absolute paths work as written.
func WithStderr ¶
func WithStdout ¶
WithStdout and WithStderr capture what the guest writes to fd 1 and fd 2. Unset discards, the default, because llama.cpp's own logging is silenced inside the bridge anyway.
type Params ¶
type Params struct {
// NPredict caps the number of tokens generated. Zero or negative means
// "as many as the context allows".
NPredict int
// Temperature <= 0 selects greedy decoding, which makes generation
// deterministic and skips the truncation samplers below.
Temperature float32
// TopK, TopP, MinP and TypicalP truncate the candidate set before
// sampling. Zero disables each of them, as does 1.0 for the two
// probability-mass ones.
TopK int
TopP float32
MinP float32
TypicalP float32
// RepeatPenalty, PresencePenalty and FrequencyPenalty discourage
// repetition over the last RepeatLastN tokens. Zero disables all three;
// so does RepeatPenalty 1.0, which is llama.cpp's own spelling of "off".
RepeatPenalty float32
RepeatLastN int
PresencePenalty float32
FrequencyPenalty float32
// Seed makes sampling reproducible. Zero is a fixed seed, so a
// temperature above zero still replays identically; set a varying seed to
// get varying output.
Seed uint32
// Mirostat selects the mirostat sampling algorithm: 0 off, 1 v1, 2 v2.
// When on it replaces the TopK/TopP/MinP/TypicalP truncation samplers,
// as in llama.cpp. MirostatTau is the target entropy (0 means llama.cpp's
// 5.0) and MirostatEta the learning rate (0 means 0.1).
Mirostat int
MirostatTau float32
MirostatEta float32
// IgnoreEOS keeps generating past end-of-generation tokens by excluding
// them from sampling, like llama.cpp's --ignore-eos. Generation then runs
// to NPredict, a stop string, or the context limit.
IgnoreEOS bool
// LogitBias adds a bias to specific tokens' logits before sampling.
// math.Inf(-1) (or any very negative value) forbids a token outright.
LogitBias map[int32]float32
// Grammar is a GBNF grammar constraining the output. Empty means none.
Grammar string
// Stop ends generation when the output first contains one of these
// strings — even mid-token — and the text is cut at the match.
Stop []string
// CachePrompt treats the prompt as the WHOLE intended context: the
// longest prefix already in the context's KV cache is kept, whatever the
// cache holds beyond it is dropped, and only the rest is decoded —
// llama.cpp server's cache_prompt. Requests that share a long constant
// preamble (a system prompt, a routing configuration) then pay only for
// the part that changed; Result.NCached reports the reuse. Off, the
// prompt appends at the cache's current end, which is what an
// Eval-prefill continuation expects.
CachePrompt bool
}
Params configures one generation.
The zero value is greedy decoding with no truncation samplers, no penalties and no token limit: the most predictable thing the model can do, and reproducible run to run. Every field is honoured exactly as written — a zero is a decision ("no top-k", "temperature 0 means greedy"), not "use some default" — so what a caller leaves out cannot be overridden by a default buried in the engine.
type Result ¶
type Result struct {
// Text is the generated text, with a matched stop string removed.
Text string `json:"text"`
// Tokens are the tokens behind Text.
Tokens []int32 `json:"tokens"`
// NPrompt is how many tokens the prompt occupied; NDecoded how many were
// generated.
NPrompt int `json:"n_prompt"`
NDecoded int `json:"n_decoded"`
// NCached is how many leading prompt tokens a Params.CachePrompt
// generation reused from the KV cache instead of re-decoding. Zero
// without CachePrompt.
NCached int `json:"n_cached"`
// NDrafted / NAccepted are GenerateWithDraft's speculation counters —
// how many tokens the draft proposed and how many the target accepted.
// Zero on a plain Generate.
NDrafted int `json:"n_drafted"`
NAccepted int `json:"n_accepted"`
// Reason says why generation stopped; Interrupted is a shorthand for
// Reason == StopInterrupted.
Reason StopReason `json:"stop_reason"`
Interrupted bool `json:"interrupted"`
// Timings splits the generation's wall time between the prompt pass
// and the decode loop.
Timings Timings `json:"timings"`
}
Result is what Generate returns.
type ScoreResult ¶
type ScoreResult struct {
NTokens int32 `json:"n_tokens"`
NScored int32 `json:"n_scored"`
NLL float64 `json:"nll"`
}
ScoreResult reports the teacher-forced negative log-likelihood of a text: the model decodes the tokenized text once and NLL sums -log softmax(logits_i)[token_{i+1}] over the NScored predicting positions. Perplexity is math.Exp(NLL / NScored).
type Slots ¶ added in v0.5.0
type Slots struct {
// contains filtered or unexported fields
}
Slots runs tasks concurrently on one Context — continuous batching, after llama.cpp's server. The context holds as many slots as its ContextParams.NSeqMax, each running one task on its own KV sequence; one scheduling step decodes a single batch drawn from every busy slot, so the slots share the per-step cost instead of each paying it.
A posted task waits for a free slot in FIFO order, and for enough free cells in the shared cache: its prompt plus its NPredict must fit next to what the running tasks may still write (an NPredict of zero means "until the context ends", which reserves the whole remainder).
While a Slots holds tasks, the single-sequence methods of its Context (Generate, Stream, Score, ScoreChoices, Eval, Embed, SaveState, LoadState) refuse to run; Reset drops every task. Close cancels every task and stops the scheduler; the Context stays usable.
func (*Slots) Close ¶ added in v0.5.0
Close cancels every task (their completions finish with ErrSlotsClosed) and stops the scheduler. It returns once the scheduler has returned, which is after every task has ended; a concurrent second Close returns then too.
func (*Slots) Post ¶ added in v0.5.0
Post queues task and returns its completion. The task is cancelled when ctx is done: its slot is released, and its completion finishes with the text produced so far, StopInterrupted, and ctx.Err() as the error.
func (*Slots) SetSystemPrompt ¶ added in v0.5.0
SetSystemPrompt decodes text once and shares its KV cells with every task whose prompt starts with it, so such a task decodes only the rest (its Result.NCached counts the reused tokens). This is the system prompt of llama.cpp's server of old: it occupies one of the context's sequences, leaving NSeqMax-1 slots for tasks, and with KVUnified the sharing is free (a copy is metadata); without it each task copies the cells, still far cheaper than decoding them. The text is tokenized on its own, so a task's prompt should be exactly this text followed by the rest; a prompt that tokenizes differently at the boundary simply decodes in full. Refused while tasks are held; an empty text clears it.
func (*Slots) Status ¶ added in v0.5.0
func (s *Slots) Status() (SlotsStatus, error)
Status reports the slots, the tasks and the cache cells in use.
type SlotsStatus ¶ added in v0.5.0
type SlotsStatus struct {
NSlots int `json:"n_slots"`
Active int `json:"active"`
Queued int `json:"queued"`
NCtx int `json:"n_ctx"`
Used int `json:"used"`
}
SlotsStatus is a snapshot of a Slots: how many slots there are and how many hold a task, how many tasks wait for one, and how much of the context's KV cache (NCtx cells) the running tasks hold.
type Snapshot ¶ added in v0.3.0
type Snapshot struct {
// contains filtered or unexported fields
}
Snapshot is the copy-on-write image of a prepared instance. It stays valid for the life of the process; forks reference it and share its pages.
func NewSnapshot ¶ added in v0.3.0
func NewSnapshot(build func(*SnapshotBuilder) error, opts ...Option) (*Snapshot, error)
NewSnapshot boots an engine on a snapshot image and hands it to build to prepare: load the model, create contexts, decode their prompt prefixes (Params.CachePrompt), and Register the ones forks will use. When build returns, every kept context's threadpool is freed — its workers joined, since threads are host constructs and cannot be captured — and the image is sealed.
The builder instance itself is consumed by the build; do not retain the *Llama, models or contexts it produced beyond the callback.
type SnapshotBuilder ¶ added in v0.3.0
type SnapshotBuilder struct {
*Llama
// contains filtered or unexported fields
}
SnapshotBuilder is the instance handed to a NewSnapshot callback: a normal *Llama plus Register, which records the contexts forks will use.
func (*SnapshotBuilder) Register ¶ added in v0.3.0
func (b *SnapshotBuilder) Register(name string, c *Context) error
Register records a prepared context under name, making it available from every fork of the snapshot via Instance.Context(name). A context is registered once: every fork frees each of its kept contexts on Close, so the same context under two names would be freed twice.
type StopReason ¶
type StopReason string
StopReason says why generation ended.
const ( // StopEOS: the model emitted an end-of-generation token. StopEOS StopReason = "eos" // StopLength: the token budget ran out (Params.NPredict, or the context // filled up). StopLength StopReason = "length" // StopString: one of Params.Stop matched the tail of the output, which // was trimmed off. StopString StopReason = "stop" // StopInterrupted: Context.Interrupt was called from another goroutine. StopInterrupted StopReason = "interrupted" )
type Task ¶ added in v0.5.0
Task is what a Slots runs: a prompt with its sampling parameters. It is a plain value — Post does not modify it, and the same Task can be posted any number of times. Params.CachePrompt is not available to a task: every task decodes its own sequence.
type TaskCompletion ¶ added in v0.5.0
type TaskCompletion struct {
// contains filtered or unexported fields
}
TaskCompletion is a posted task: its outputs as they are produced, and its result once it has finished.
func (*TaskCompletion) ID ¶ added in v0.5.0
func (t *TaskCompletion) ID() int
ID is the task's id in the engine; ids start at 1 for each Context.
func (*TaskCompletion) Outputs ¶ added in v0.5.0
func (t *TaskCompletion) Outputs() iter.Seq[TaskOutput]
Outputs yields the task's outputs as they are produced and returns when the task has finished. It never blocks the scheduler: outputs are kept until read, so a slow reader only delays itself. Leaving the loop early stops nothing; cancel the Post's ctx to stop the task.
func (*TaskCompletion) Result ¶ added in v0.5.0
func (t *TaskCompletion) Result() (Result, error)
Result returns the finished task's result without waiting; ErrTaskRunning while it has not finished.
func (*TaskCompletion) Wait ¶ added in v0.5.0
func (t *TaskCompletion) Wait() (Result, error)
Wait blocks until the task has finished and returns its result: the same Result a Generate of the task would return. A cancelled task returns the text produced so far with StopInterrupted and the cancellation error.
type TaskOutput ¶ added in v0.5.0
TaskOutput is one produced token: its id and the text it renders to. The texts of a task's outputs concatenate to the Text of its Result, except that a matched stop string is delivered as it is decoded and only afterwards trimmed from the result.
type TensorInfo ¶ added in v0.5.0
type TensorInfo struct {
Name string `json:"name"`
// Type is the ggml type name ("q4_K", "q8_0", "f16", ...).
Type string `json:"type"`
// Shape is ne[0..]: the row length first.
Shape []int64 `json:"ne"`
// Buffer names the backend buffer holding the tensor: "CPU_REPACK" for
// tensors the CPU backend repacked for its GEMV/GEMM kernels (rows a
// multiple of the repack width, a repackable type), "CPU" otherwise.
Buffer string `json:"buffer"`
// VecDotType is the activation type the tensor's dot product quantizes
// to ("q8_K" for K-quants, "q8_0" / "q8_1" for the legacy types).
VecDotType string `json:"vec_dot_type"`
}
TensorInfo describes one weight tensor of a loaded model and the kernel path the CPU backend gave it.
func (TensorInfo) Repacked ¶ added in v0.5.0
func (t TensorInfo) Repacked() bool
Repacked reports whether the CPU backend repacked the tensor for its GEMV/GEMM kernels; tensors that are not repacked run the per-row dot.
Directories
¶
| Path | Synopsis |
|---|---|
|
cmd/perfgate
command
Command perfgate is the CI entry point of internal/perfgate.
|
Command perfgate is the CI entry point of internal/perfgate. |
|
perfgate
Package perfgate compares go-llama's throughput against a native llama.cpp build measured on the same machine, and gates CI on the ratio: a change that makes go-llama more than a threshold slower RELATIVE TO NATIVE than the recorded baseline fails.
|
Package perfgate compares go-llama's throughput against a native llama.cpp build measured on the same machine, and gates CI on the ratio: a change that makes go-llama more than a threshold slower RELATIVE TO NATIVE than the recorded baseline fails. |