llama

package module
v0.5.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 15, 2026 License: MIT Imports: 19 Imported by: 0

README

go-llama

Go Reference CI

llama.cpp in pure Go — run GGUF models anywhere Go runs. No cgo, no shared library, one static binary.

The inference engine is llama.cpp compiled to WebAssembly and then translated to Go — no wasm runtime is involved at run time — built on the llamawasm2go module.

package main

import (
	"fmt"

	llama "github.com/goccy/go-llama"
)

func main() {
	inst, err := llama.New()
	if err != nil {
		panic(err)
	}
	defer inst.Close()

	model, err := inst.LoadModel("model.gguf")
	if err != nil {
		panic(err)
	}
	defer model.Close()

	ctx, err := model.NewContext(llama.ContextParams{NCtx: 2048})
	if err != nil {
		panic(err)
	}
	defer ctx.Close()

	res, err := ctx.Generate("Once upon a time", llama.Params{
		NPredict:    128,
		Temperature: 0.8,
		TopP:        0.95,
	})
	if err != nil {
		panic(err)
	}
	fmt.Println(res.Text)
}

Features

  • Pure Go: works anywhere Go compiles; CGO_ENABLED=0 friendly.
  • Independent instances: llama.New builds an engine with its own linear memory; create several and run them concurrently and in isolation.
  • One model, many contexts: contexts share the weights and keep their own KV cache, which is how to serve independent conversations from one model. One instance can also hold several models — what speculative decoding needs.
  • Sampling: temperature, top-k, top-p, min-p, typical-p, repetition / presence / frequency penalties, seeds, and GBNF grammars.
  • Streaming: Context.Stream calls you back with each piece of text as it is decoded.
  • Interruptible: Context.Interrupt stops a running generation from another goroutine.
  • Speculative decoding, LoRA adapters, chat templates, embeddings, scoring, and state save / load.
  • Configurable sandbox: options on New scope the guest to one directory, hand it an in-memory filesystem, cap its memory, and capture its stdio.
inst, err := llama.New(
	llama.WithPreopenDir("/srv/models"), // the only directory the guest can see
	llama.WithMaxMemory(6<<30),          // fail inside the guest, not in the host
)

Instances, models and contexts

llama.New returns a *Llama — one engine instance, with its own linear memory and C heap. It is fully independent of any other instance, so several can run concurrently on separate goroutines.

Within an instance you load one or more models with LoadModel; each model spawns contexts with NewContext that share its weights and keep their own KV cache. Because a whole instance carries the engine's memory, the common shape is one instance with as many models and contexts as you need — reach for a second instance when you want hard isolation between them.

inst, _ := llama.New()
defer inst.Close()

target, _ := inst.LoadModel("qwen2.5-3b.gguf")
draft, _ := inst.LoadModel("qwen2.5-0.5b.gguf") // same instance: speculative decoding

Close contexts, then models, then the instance — using any handle after its owner is closed is a use-after-free in the engine, and closed handles are refused.

Streaming and interruption

The engine is a single translated module with one C stack, so the goroutine running Generate is the only one that can be inside it. Streaming and interruption reach a running generation from opposite directions around that constraint:

res, err := ctx.Stream("Once upon a time", llama.Params{NPredict: 512},
	func(piece string) { fmt.Print(piece) })

Stream calls onPiece once per decoded token, on the generating goroutine itself — so there is no concurrency and nothing to drop, but the callback must be short and must not call back into the engine. It returns the same complete Result as Generate (the pieces concatenate to Result.Text; a Params.Stop string is delivered as decoded and only then trimmed, so the stream can run a few characters past the returned text). A nil onPiece makes Stream exactly Generate.

Interrupt goes the other way: it writes one aligned word straight into linear memory (never calling into the engine), which the generation loop reads once per token. It is safe to call from any goroutine while a generation runs; Generate then returns what it has with Reason == StopInterrupted.

Slots: many requests on one context

A Context created with NSeqMax slots can run that many tasks at once — continuous batching, in llama.cpp's server vocabulary (slot, task, post, system prompt). One scheduling step decodes a single batch drawn from every busy slot, so the slots share the per-step cost instead of each paying it.

ctx, _ := model.NewContext(llama.ContextParams{NCtx: 8192, NSeqMax: 32})
slots, _ := ctx.Slots()
defer slots.Close()

// Optional: an instruction every request starts with is decoded once and
// shared; a task that starts with it decodes only the rest.
slots.SetSystemPrompt(instruction)

task := llama.Task{Prompt: instruction + userText, Params: llama.Params{NPredict: 16}}
cmpl, err := slots.Post(ctx, task)        // ctx is a context.Context: cancel it to drop the task
for out := range cmpl.Outputs() {         // each token as it is produced (optional)
	fmt.Print(out.Text)
}
res, err := cmpl.Wait()                    // the same Result Generate returns

Task is a plain value; Post never modifies it and returns the task's TaskCompletion. Outputs is pull-based — outputs are kept until read, so a slow reader never stalls the batch — and Result returns the finished result without waiting (ErrTaskRunning before). A task waits for a free slot in FIFO order and for enough free cells; its prompt plus NPredict must fit a sequence's window. A greedy task produces exactly what Generate produces for the same prompt when the prompt is decoded in the same chunks; a long prompt split differently across steps can land on the other side of a near-tie, as Generate itself does with a different NBatch.

ContextParams.KVUnified follows llama.cpp's kv_unified: off (the default) gives each sequence its own stream of NCtx / NSeqMax cells, the right layout for independent tasks; on shares one buffer, which makes the system prompt's sharing free and is what ScoreChoices with NSeqMax > 1 relies on. While a Slots holds tasks the context's single-sequence methods refuse.

Measured on an M-series laptop with 8 threads and a 0.5B Q8_0 model, 100 requests arriving at once (a 45-token shared instruction plus a short user part, 16 tokens each): eight forked instances completed them at p50 4.3 s / p99 7.8 s (205 tok/s); 32 slots with the shared system prompt at p50 1.5 s / p99 2.4 s (655 tok/s), first token at p50 0.85 s.

Memory

wasm32 caps linear memory at 4 GiB, and the model weights plus every context's KV cache live inside it. Target quantized models comfortably under that — roughly 3B parameters at Q4 — and size ContextParams.NCtx accordingly. WithMaxMemory caps growth so an oversized model fails in the guest instead of growing the host process, and WithMemoryReserve reserves up front so a large load does not repeatedly grow and copy.

Performance

The generated Go is compiled by the Go compiler, and on amd64/arm64 most of it ships as assembly derived from that compilation. The SIMD kernels ggml relies on are native NEON on arm64 and SSE on amd64 (the latter under GOAMD64=v2 or higher — set it, or the vector helpers fall back to scalar Go). arm64 is the flagship target, where the dot-product kernels lower to SDOT/SMMLA.

Supply-chain verification

internal/llama.go in this repository is a release artifact of llama-wasm, not hand-written code. It is refreshed with:

make llama LLAMA_WASM_VERSION=v0.1.0

which downloads it and verifies its build-provenance attestation against llama-wasm's release workflow. CI re-runs that verification (make verify) on every push.

Testing

make test        # fetches a tiny GGUF into testdata/ and runs the suite

Point the suite at your own model with GO_LLAMA_TEST_MODEL=/path/to.gguf.

License

MIT (see LICENSE). llama.cpp is MIT; the embedded engine is a derivative work of it.

Documentation

Index

Constants

This section is empty.

Variables

View Source
var ErrInstanceClosed = bridge.ErrEngineClosed

ErrInstanceClosed is in the error of every call on a closed instance, and of a Slots task the instance's Close ended.

View Source
var ErrMemoryLeaked = errors.New("llama: close: memory kept mapped, worker threads may still run in it")

ErrMemoryLeaked is returned by Close when the instance's memory was deliberately kept mapped, because a trap left it unknown whether the instance's worker threads had been joined (see Llama.Close).

View Source
var ErrSlotsClosed = errors.New("llama: slots are closed")

ErrSlotsClosed is the error of every task a Slots dropped when it was closed, and of a Post after Close.

View Source
var ErrSlotsRunning = errors.New("llama: slots: the context already has a scheduler; close it first")

ErrSlotsRunning is returned by Context.Slots while an earlier Slots of the context is still open.

View Source
var ErrTaskRunning = errors.New("llama: task is still running")

ErrTaskRunning is returned by TaskCompletion.Result while the task has not finished.

Functions

This section is empty.

Types

type Build

type Build struct {
	SIMD       bool `json:"simd"`
	Threads    bool `json:"threads"`
	Exceptions bool `json:"exceptions"`
	MaxDevices int  `json:"max_devices"`
}

Build reports how the embedded engine was compiled. Diagnostics only; see BuildInfo.

type Context

type Context struct {
	// contains filtered or unexported fields
}

Context is one inference context over a Model: its own KV cache and its own sampling state. Independent conversations get independent contexts over the same weights.

func (*Context) Close

func (c *Context) Close() error

Close frees the context, stopping and joining its threadpool workers. A generation in flight on the context is interrupted (it returns at its next token, with StopInterrupted) and any other call waited for. A Slots still scheduling on the context is closed too; its tasks end with ErrSlotsClosed. The Model outlives it.

func (*Context) Embed

func (c *Context) Embed(text string, normalize bool) ([]float32, error)

Embed returns the embedding of text. The context must have been created with ContextParams.Embeddings set. normalize applies L2 normalisation.

func (*Context) EmbedTokens

func (c *Context) EmbedTokens(tokens []int32, normalize bool) ([]float32, error)

EmbedTokens is Embed from token ids instead of text.

func (*Context) Eval

func (c *Context) Eval(text string, addSpecial, parseSpecial bool) (EvalResult, error)

Eval decodes text into the context's KV cache without sampling — prompt prefill for a later Generate on the same context, or plain evaluation. Positions continue from the cache's current end; Reset starts over. addSpecial and parseSpecial mirror Tokenize.

func (*Context) Generate

func (c *Context) Generate(prompt string, p Params) (Result, error)

Generate runs generation from prompt and returns the whole result.

func (*Context) GenerateWithDraft

func (c *Context) GenerateWithDraft(draft *Context, prompt string, p Params, nDraft int) (Result, error)

GenerateWithDraft is Generate accelerated by speculative decoding: a draft context over a smaller model with the SAME vocabulary proposes up to nDraft tokens per round (0 picks a default) and this context verifies them in one batch. Every emitted token is sampled by THIS context's sampler chain, so the output distribution is exactly Generate's — the draft only trades its cheap decodes for larger verification batches.

Both contexts' KV caches restart from the prompt. Result.NDrafted and Result.NAccepted report how well the draft anticipated the target.

func (*Context) Interrupt

func (c *Context) Interrupt() error

Interrupt stops a running generation at its next token WITHOUT executing any engine code: it writes the interrupt flag straight into linear memory. Calling in instead would need the lock the generation holds, and would share the C stack with it. Generate then returns what it has, with Reason == StopInterrupted.

Safe to call from any goroutine, including while a generation runs, and while a Close is waiting for one (Close interrupts it itself). A call when nothing is running is a no-op: the flag is cleared when generation starts.

func (*Context) LoadState

func (c *Context) LoadState(state []byte) error

LoadState restores a state produced by SaveState. The context must be over the same model with compatible parameters, or the engine rejects the payload.

func (*Context) Model added in v0.5.1

func (c *Context) Model() *Model

Model returns the model the context is over.

func (*Context) Reset

func (c *Context) Reset() error

Reset drops the KV cache so the next generation starts from nothing.

func (*Context) SaveState

func (c *Context) SaveState() ([]byte, error)

SaveState serializes the context's state — KV cache, sampling RNG and the prompt-prefix history — into a byte slice that LoadState can restore, also in a later process. Save a long system prompt's state once and every future context skips re-decoding it: restore the blob and either continue positionally (Eval-prefill style) or generate with Params.CachePrompt, which picks up the saved prefix immediately.

func (*Context) Score

func (c *Context) Score(text string) (ScoreResult, error)

Score computes the teacher-forced negative log-likelihood of text under the model, the quantity llama.cpp's perplexity tool averages. It decodes text into the context's KV cache; call Reset before generating afterwards.

func (*Context) ScoreChoices added in v0.4.0

func (c *Context) ScoreChoices(choices []string) ([]ScoreResult, error)

ScoreChoices scores candidate continuations of the context's CURRENT cache state: for each choice it returns the negative log-likelihood of the choice's tokens as the continuation of what the context has already decoded. Call it right after decoding the shared stem (Eval, or a Generate whose prompt just ran) — the first token of every choice is scored from the live next-position logits, the rest teacher-forced. Each choice is rolled back out of the KV cache before the next, so the choices never see each other and the context ends exactly where it started. Choices must be non-empty and must not contain newlines (the wire format separates them with '\n').

func (*Context) SetLoRA

func (c *Context) SetLoRA(adapters []LoRAWeight) error

SetLoRA replaces the context's ENTIRE adapter configuration — the set semantics of llama.cpp's own API. An empty (or nil) slice removes every adapter, so SetLoRA(nil) is the spelling of "clear".

func (*Context) Slots added in v0.5.0

func (c *Context) Slots() (*Slots, error)

Slots starts a scheduler over the context's slots. The context must have been created with NSeqMax set to the number of slots wanted (1 still works: tasks then run one at a time, each on its own sequence). A context runs one Slots at a time: a second call while the first is open returns ErrSlotsRunning. Closing the context closes its Slots.

func (*Context) Stream

func (c *Context) Stream(prompt string, p Params, onPiece func(string)) (Result, error)

Stream generates from prompt like Generate and calls onPiece with each piece of text as it is decoded, then returns the same complete Result.

onPiece runs inside the generation, on the goroutine that called Stream, so it must not call back into this context or into any other part of the engine. Keep it short: the generation is stopped while it runs.

The pieces concatenate to Result.Text, with one exception: a stop string is delivered as it is decoded and only then trimmed from Result.Text, so with Params.Stop set the stream can run a few characters past the returned text.

A nil onPiece makes this exactly Generate.

type ContextParams

type ContextParams struct {
	// NCtx is the context window in tokens. Zero means the model's training
	// context length. The KV cache scales with it.
	NCtx uint32
	// NBatch and NUBatch are the logical and physical batch sizes. Zero means
	// llama.cpp's defaults.
	NBatch  uint32
	NUBatch uint32
	// NThreads is how many threads ggml may use. It has an effect only in a
	// threads-enabled build (BuildInfo.Threads); the single-threaded wasm
	// clamps to 1.
	NThreads uint32
	// NSeqMax is the number of sequences the context can hold at once: the
	// slots of a Slots, or the candidates ScoreChoices decodes in one batch
	// (one sequence per candidate, all sharing the stem). Zero or 1 is the
	// single-sequence default, where ScoreChoices decodes candidates one at
	// a time.
	NSeqMax uint32
	// KVUnified chooses how the NSeqMax sequences share the KV cache, as
	// llama.cpp's kv_unified does. False, the default, gives each sequence
	// its own stream of NCtx / NSeqMax cells: attention over a sequence then
	// costs only its own cells, which is what independent tasks on a Slots
	// want. True shares one buffer of NCtx cells across the sequences, where
	// copying a sequence is metadata rather than a copy: what ScoreChoices'
	// batched path leans on (its candidates all share the stem), and what a
	// prompt prefix shared across slots needs. The measured difference for
	// 64 independent tasks was 903 vs 500 tok/s in favour of streams; for
	// ScoreChoices with NSeqMax > 1, set it.
	KVUnified bool
	// Embeddings puts the context in embedding mode, which Context.Embed
	// requires and which disables generation.
	Embeddings bool
	// RopeFreqBase and RopeFreqScale override the model's RoPE
	// configuration, the knobs behind context-length extension schemes.
	// Zero keeps the model's own values.
	RopeFreqBase  float32
	RopeFreqScale float32
}

ContextParams configures a Context. The zero value asks for the model's own training context length and llama.cpp's batch defaults.

type EvalResult

type EvalResult struct {
	NTokens int32 `json:"n_tokens"`
	NPast   int32 `json:"n_past"`
}

EvalResult reports what Eval put into the KV cache: NTokens from this call, NPast the total sequence length now cached.

type FS

type FS = base.FS

FS is the filesystem the guest sees when Config.FS is set. base.NewMemFS returns a ready-made in-memory implementation, so a model can be served out of memory rather than off disk.

type Instance added in v0.3.0

type Instance struct {
	// contains filtered or unexported fields
}

Instance is one engine instance forked from a Snapshot.

func (*Instance) Close added in v0.3.0

func (f *Instance) Close() error

Close releases the fork: its kept contexts are freed, which joins the threadpool workers Fork attached, and then its copy-on-write mapping is unmapped, returning every private page. The kept contexts need no Close of their own; one that was closed already is simply skipped.

If freeing a context trapped, whether its workers were joined is unknown, and the mapping is kept rather than unmapped under threads that may still run in it: Close then returns ErrMemoryLeaked along with the context's error. Any error of a kept context's Close is returned.

func (*Instance) Context added in v0.3.0

func (f *Instance) Context(name string) *Context

Context returns the kept context by name, or nil.

type Llama

type Llama struct {
	// contains filtered or unexported fields
}

Llama is one engine instance: llama.cpp compiled to wasm, with its own view of linear memory and its own C heap. Load models into it and open contexts over them. Instances are independent — several run concurrently — while within one instance calls are serialised.

Memory an instance never writes is physically shared with the other instances in the process: engines are built over copy-on-write maps of a process-wide image, and instances that load a model an earlier instance already loaded (same options, same path) share the weights the same way instead of loading them again. Set GO_LLAMA_NO_SHARED_IMAGE to force every instance onto a private allocation.

func New

func New(opts ...Option) (*Llama, error)

New brings up an engine instance. The zero-option instance gives the guest the host filesystem, an empty environment and discarded stdio — the filesystem access a Go program reading a model file already has.

func (*Llama) BuildInfo

func (l *Llama) BuildInfo() (Build, error)

BuildInfo reports how the embedded engine was compiled.

func (*Llama) Close

func (l *Llama) Close() error

Close releases the instance, including the copy-on-write memory mapping when the engine is backed by one. Contexts still open are closed first — a generation running on one is interrupted and waited for, and the free joins the context's threadpool workers, which must not outlive the memory they run in — and their Close errors are returned; models need no Close of their own. Using anything of the instance afterward fails with ErrInstanceClosed.

If freeing a context trapped, whether its workers were joined is unknown, and unmapping the memory under threads that may still run in it would crash the process: Close then leaves the memory mapped for the life of the process and reports ErrMemoryLeaked. The instance is closed either way.

func (*Llama) LoadModel

func (l *Llama) LoadModel(path string) (*Model, error)

LoadModel loads a GGUF model from path into this instance.

path is a guest path. Without WithPreopenDir / WithFS the guest sees the host filesystem, so a relative path is resolved against the working directory first; with a scoped filesystem it is used as given, relative to that root.

type LoRA

type LoRA struct {
	// contains filtered or unexported fields
}

LoRA is a loaded LoRA adapter. It adapts contexts over the model that loaded it (Context.SetLoRA) and must not outlive that model.

func (*LoRA) Close

func (l *LoRA) Close() error

Close frees the adapter. Contexts still using it must clear it first.

type LoRAWeight

type LoRAWeight struct {
	Adapter *LoRA
	Scale   float32
}

LoRAWeight pairs an adapter with the scale to apply it at.

type Message

type Message struct {
	Role    string `json:"role"`
	Content string `json:"content"`
}

Message is one turn of a chat, as ApplyChatTemplate consumes it.

type Model

type Model struct {
	// contains filtered or unexported fields
}

Model is a loaded GGUF model. Several contexts can share one, and several models can be loaded into one Llama instance at once.

func (*Model) ApplyChatTemplate

func (m *Model) ApplyChatTemplate(messages []Message, templateOverride string, addAssistant bool) (string, error)

ApplyChatTemplate renders messages into a prompt with the model's chat template. templateOverride replaces it, and is required when the GGUF carries none; addAssistant appends the generation prefix.

func (*Model) Close

func (m *Model) Close() error

Close frees the model, closing its contexts still open first: a context must not outlive its model. Their Close errors are returned. Nothing to do once the instance's Close has begun: that frees every context and then the memory the model is in.

func (*Model) Detokenize

func (m *Model) Detokenize(tokens []int32, renderSpecial bool) (string, error)

Detokenize renders tokens back to text.

func (*Model) Info

func (m *Model) Info() (ModelInfo, error)

Info returns the model's metadata.

func (*Model) LoadLoRA

func (m *Model) LoadLoRA(path string) (*LoRA, error)

LoadLoRA loads a LoRA adapter GGUF for this model. path resolves like LoadModel's.

func (*Model) NewContext

func (m *Model) NewContext(p ContextParams) (*Context, error)

NewContext creates an inference context over the model.

func (*Model) Tensors added in v0.5.0

func (m *Model) Tensors() ([]TensorInfo, error)

Tensors lists the model's weight tensors with the buffer each landed in, which decides the kernel path (repacked GEMV/GEMM versus per-row dot).

func (*Model) TokenToPiece

func (m *Model) TokenToPiece(token int32, renderSpecial bool) (string, error)

TokenToPiece renders one token. A byte-level token can render to invalid UTF-8 on its own; accumulate pieces before treating the result as text.

func (*Model) Tokenize

func (m *Model) Tokenize(text string, addSpecial, parseSpecial bool) ([]int32, error)

Tokenize splits text into tokens. addSpecial adds the model's BOS/EOS convention; parseSpecial lets special-token text ("<|im_start|>") tokenize as that token rather than as its characters.

type ModelInfo

type ModelInfo struct {
	// Desc is llama.cpp's own one-line description, e.g. "llama 7B Q4_K - Medium".
	Desc string `json:"desc"`
	// NParams is the parameter count; SizeBytes the on-disk tensor size.
	NParams   uint64 `json:"n_params"`
	SizeBytes uint64 `json:"size_bytes"`
	// NCtxTrain is the context length the model was trained with — the
	// largest context worth asking for.
	NCtxTrain int `json:"n_ctx_train"`
	NEmbd     int `json:"n_embd"`
	NLayer    int `json:"n_layer"`
	NHead     int `json:"n_head"`
	NHeadKV   int `json:"n_head_kv"`
	NVocab    int `json:"n_vocab"`
	// HasEncoder / HasDecoder describe the architecture; a plain causal LM
	// has only a decoder.
	HasEncoder bool `json:"has_encoder"`
	HasDecoder bool `json:"has_decoder"`
	// BOSToken / EOSToken are the vocabulary's sentence delimiters, -1
	// when the model defines none. AddBOS is whether the tokenizer
	// prepends BOSToken by convention (what Tokenize's addSpecial obeys).
	BOSToken int32 `json:"bos_token"`
	EOSToken int32 `json:"eos_token"`
	AddBOS   bool  `json:"add_bos"`
	// ChatTemplate is the Jinja-ish template string the GGUF carries, or ""
	// when it has none. Chat needs a template from somewhere: either this or
	// SamplingParams-independent override passed to ApplyChatTemplate.
	ChatTemplate string `json:"chat_template"`
}

ModelInfo describes a loaded model. Everything here comes from the GGUF metadata, so a field is zero when the file does not carry it.

type Option

type Option func(*config)

An Option configures a Llama at New time.

func WithEnv

func WithEnv(env []string) Option

WithEnv sets the environment the guest sees. Without it the guest gets an empty environment: the host process environment is never leaked.

func WithFS

func WithFS(fs FS) Option

WithFS replaces the filesystem backend entirely — feed a model from memory or a virtual tree. It takes precedence over WithPreopenDir. base.NewMemFS returns a ready-made in-memory implementation.

func WithMaxMemory

func WithMaxMemory(bytes uint64) Option

WithMaxMemory caps linear-memory growth, so a model bigger than expected fails inside the guest instead of growing the host process. Zero means the engine's default ceiling: for an instance backed by a copy-on-write mapping that is the 64 GiB of address space the mapping reserves (untouched pages cost nothing), and for a private allocation there is no cap.

func WithMemoryReserve

func WithMemoryReserve(bytes int) Option

WithMemoryReserve pre-reserves linear memory. Loading a multi-gigabyte model otherwise grows memory in steps, copying it forward each time; reserving up front avoids that. Zero uses the engine's default.

func WithPreopenDir

func WithPreopenDir(dir string) Option

WithPreopenDir scopes the guest filesystem to a host directory. Guest paths then resolve inside it, so a model is named relative to it. Without it the guest sees the whole host filesystem, where absolute paths work as written.

func WithStderr

func WithStderr(w io.Writer) Option

func WithStdout

func WithStdout(w io.Writer) Option

WithStdout and WithStderr capture what the guest writes to fd 1 and fd 2. Unset discards, the default, because llama.cpp's own logging is silenced inside the bridge anyway.

type Params

type Params struct {
	// NPredict caps the number of tokens generated. Zero or negative means
	// "as many as the context allows".
	NPredict int
	// Temperature <= 0 selects greedy decoding, which makes generation
	// deterministic and skips the truncation samplers below.
	Temperature float32
	// TopK, TopP, MinP and TypicalP truncate the candidate set before
	// sampling. Zero disables each of them, as does 1.0 for the two
	// probability-mass ones.
	TopK     int
	TopP     float32
	MinP     float32
	TypicalP float32
	// RepeatPenalty, PresencePenalty and FrequencyPenalty discourage
	// repetition over the last RepeatLastN tokens. Zero disables all three;
	// so does RepeatPenalty 1.0, which is llama.cpp's own spelling of "off".
	RepeatPenalty    float32
	RepeatLastN      int
	PresencePenalty  float32
	FrequencyPenalty float32
	// Seed makes sampling reproducible. Zero is a fixed seed, so a
	// temperature above zero still replays identically; set a varying seed to
	// get varying output.
	Seed uint32
	// Mirostat selects the mirostat sampling algorithm: 0 off, 1 v1, 2 v2.
	// When on it replaces the TopK/TopP/MinP/TypicalP truncation samplers,
	// as in llama.cpp. MirostatTau is the target entropy (0 means llama.cpp's
	// 5.0) and MirostatEta the learning rate (0 means 0.1).
	Mirostat    int
	MirostatTau float32
	MirostatEta float32
	// IgnoreEOS keeps generating past end-of-generation tokens by excluding
	// them from sampling, like llama.cpp's --ignore-eos. Generation then runs
	// to NPredict, a stop string, or the context limit.
	IgnoreEOS bool
	// LogitBias adds a bias to specific tokens' logits before sampling.
	// math.Inf(-1) (or any very negative value) forbids a token outright.
	LogitBias map[int32]float32
	// Grammar is a GBNF grammar constraining the output. Empty means none.
	Grammar string
	// Stop ends generation when the output first contains one of these
	// strings — even mid-token — and the text is cut at the match.
	Stop []string
	// CachePrompt treats the prompt as the WHOLE intended context: the
	// longest prefix already in the context's KV cache is kept, whatever the
	// cache holds beyond it is dropped, and only the rest is decoded —
	// llama.cpp server's cache_prompt. Requests that share a long constant
	// preamble (a system prompt, a routing configuration) then pay only for
	// the part that changed; Result.NCached reports the reuse. Off, the
	// prompt appends at the cache's current end, which is what an
	// Eval-prefill continuation expects.
	CachePrompt bool
}

Params configures one generation.

The zero value is greedy decoding with no truncation samplers, no penalties and no token limit: the most predictable thing the model can do, and reproducible run to run. Every field is honoured exactly as written — a zero is a decision ("no top-k", "temperature 0 means greedy"), not "use some default" — so what a caller leaves out cannot be overridden by a default buried in the engine.

type Result

type Result struct {
	// Text is the generated text, with a matched stop string removed.
	Text string `json:"text"`
	// Tokens are the tokens behind Text.
	Tokens []int32 `json:"tokens"`
	// NPrompt is how many tokens the prompt occupied; NDecoded how many were
	// generated.
	NPrompt  int `json:"n_prompt"`
	NDecoded int `json:"n_decoded"`
	// NCached is how many leading prompt tokens a Params.CachePrompt
	// generation reused from the KV cache instead of re-decoding. Zero
	// without CachePrompt.
	NCached int `json:"n_cached"`
	// NDrafted / NAccepted are GenerateWithDraft's speculation counters —
	// how many tokens the draft proposed and how many the target accepted.
	// Zero on a plain Generate.
	NDrafted  int `json:"n_drafted"`
	NAccepted int `json:"n_accepted"`
	// Reason says why generation stopped; Interrupted is a shorthand for
	// Reason == StopInterrupted.
	Reason      StopReason `json:"stop_reason"`
	Interrupted bool       `json:"interrupted"`
	// Timings splits the generation's wall time between the prompt pass
	// and the decode loop.
	Timings Timings `json:"timings"`
}

Result is what Generate returns.

type ScoreResult

type ScoreResult struct {
	NTokens int32   `json:"n_tokens"`
	NScored int32   `json:"n_scored"`
	NLL     float64 `json:"nll"`
}

ScoreResult reports the teacher-forced negative log-likelihood of a text: the model decodes the tokenized text once and NLL sums -log softmax(logits_i)[token_{i+1}] over the NScored predicting positions. Perplexity is math.Exp(NLL / NScored).

type Slots added in v0.5.0

type Slots struct {
	// contains filtered or unexported fields
}

Slots runs tasks concurrently on one Context — continuous batching, after llama.cpp's server. The context holds as many slots as its ContextParams.NSeqMax, each running one task on its own KV sequence; one scheduling step decodes a single batch drawn from every busy slot, so the slots share the per-step cost instead of each paying it.

A posted task waits for a free slot in FIFO order, and for enough free cells in the shared cache: its prompt plus its NPredict must fit next to what the running tasks may still write (an NPredict of zero means "until the context ends", which reserves the whole remainder).

While a Slots holds tasks, the single-sequence methods of its Context (Generate, Stream, Score, ScoreChoices, Eval, Embed, SaveState, LoadState) refuse to run; Reset drops every task. Close cancels every task and stops the scheduler; the Context stays usable.

func (*Slots) Close added in v0.5.0

func (s *Slots) Close() error

Close cancels every task (their completions finish with ErrSlotsClosed) and stops the scheduler. It returns once the scheduler has returned, which is after every task has ended; a concurrent second Close returns then too.

func (*Slots) Post added in v0.5.0

func (s *Slots) Post(ctx context.Context, task Task) (*TaskCompletion, error)

Post queues task and returns its completion. The task is cancelled when ctx is done: its slot is released, and its completion finishes with the text produced so far, StopInterrupted, and ctx.Err() as the error.

func (*Slots) SetSystemPrompt added in v0.5.0

func (s *Slots) SetSystemPrompt(text string) error

SetSystemPrompt decodes text once and shares its KV cells with every task whose prompt starts with it, so such a task decodes only the rest (its Result.NCached counts the reused tokens). This is the system prompt of llama.cpp's server of old: it occupies one of the context's sequences, leaving NSeqMax-1 slots for tasks, and with KVUnified the sharing is free (a copy is metadata); without it each task copies the cells, still far cheaper than decoding them. The text is tokenized on its own, so a task's prompt should be exactly this text followed by the rest; a prompt that tokenizes differently at the boundary simply decodes in full. Refused while tasks are held; an empty text clears it.

func (*Slots) Status added in v0.5.0

func (s *Slots) Status() (SlotsStatus, error)

Status reports the slots, the tasks and the cache cells in use.

type SlotsStatus added in v0.5.0

type SlotsStatus struct {
	NSlots int `json:"n_slots"`
	Active int `json:"active"`
	Queued int `json:"queued"`
	NCtx   int `json:"n_ctx"`
	Used   int `json:"used"`
}

SlotsStatus is a snapshot of a Slots: how many slots there are and how many hold a task, how many tasks wait for one, and how much of the context's KV cache (NCtx cells) the running tasks hold.

type Snapshot added in v0.3.0

type Snapshot struct {
	// contains filtered or unexported fields
}

Snapshot is the copy-on-write image of a prepared instance. It stays valid for the life of the process; forks reference it and share its pages.

func NewSnapshot added in v0.3.0

func NewSnapshot(build func(*SnapshotBuilder) error, opts ...Option) (*Snapshot, error)

NewSnapshot boots an engine on a snapshot image and hands it to build to prepare: load the model, create contexts, decode their prompt prefixes (Params.CachePrompt), and Register the ones forks will use. When build returns, every kept context's threadpool is freed — its workers joined, since threads are host constructs and cannot be captured — and the image is sealed.

The builder instance itself is consumed by the build; do not retain the *Llama, models or contexts it produced beyond the callback.

func (*Snapshot) Fork added in v0.3.0

func (s *Snapshot) Fork(nThreads uint32) (*Instance, error)

Fork brings up an instance from the snapshot: the prepared contexts are live immediately, with fresh threadpools of nThreads workers each (nThreads <= 1 keeps them single-threaded). Close the fork when done with it — its private pages are returned then.

type SnapshotBuilder added in v0.3.0

type SnapshotBuilder struct {
	*Llama
	// contains filtered or unexported fields
}

SnapshotBuilder is the instance handed to a NewSnapshot callback: a normal *Llama plus Register, which records the contexts forks will use.

func (*SnapshotBuilder) Register added in v0.3.0

func (b *SnapshotBuilder) Register(name string, c *Context) error

Register records a prepared context under name, making it available from every fork of the snapshot via Instance.Context(name). A context is registered once: every fork frees each of its kept contexts on Close, so the same context under two names would be freed twice.

type StopReason

type StopReason string

StopReason says why generation ended.

const (
	// StopEOS: the model emitted an end-of-generation token.
	StopEOS StopReason = "eos"
	// StopLength: the token budget ran out (Params.NPredict, or the context
	// filled up).
	StopLength StopReason = "length"
	// StopString: one of Params.Stop matched the tail of the output, which
	// was trimmed off.
	StopString StopReason = "stop"
	// StopInterrupted: Context.Interrupt was called from another goroutine.
	StopInterrupted StopReason = "interrupted"
)

type Task added in v0.5.0

type Task struct {
	Prompt string
	Params Params
}

Task is what a Slots runs: a prompt with its sampling parameters. It is a plain value — Post does not modify it, and the same Task can be posted any number of times. Params.CachePrompt is not available to a task: every task decodes its own sequence.

type TaskCompletion added in v0.5.0

type TaskCompletion struct {
	// contains filtered or unexported fields
}

TaskCompletion is a posted task: its outputs as they are produced, and its result once it has finished.

func (*TaskCompletion) ID added in v0.5.0

func (t *TaskCompletion) ID() int

ID is the task's id in the engine; ids start at 1 for each Context.

func (*TaskCompletion) Outputs added in v0.5.0

func (t *TaskCompletion) Outputs() iter.Seq[TaskOutput]

Outputs yields the task's outputs as they are produced and returns when the task has finished. It never blocks the scheduler: outputs are kept until read, so a slow reader only delays itself. Leaving the loop early stops nothing; cancel the Post's ctx to stop the task.

func (*TaskCompletion) Result added in v0.5.0

func (t *TaskCompletion) Result() (Result, error)

Result returns the finished task's result without waiting; ErrTaskRunning while it has not finished.

func (*TaskCompletion) Wait added in v0.5.0

func (t *TaskCompletion) Wait() (Result, error)

Wait blocks until the task has finished and returns its result: the same Result a Generate of the task would return. A cancelled task returns the text produced so far with StopInterrupted and the cancellation error.

type TaskOutput added in v0.5.0

type TaskOutput struct {
	Token int32
	Text  string
}

TaskOutput is one produced token: its id and the text it renders to. The texts of a task's outputs concatenate to the Text of its Result, except that a matched stop string is delivered as it is decoded and only afterwards trimmed from the result.

type TensorInfo added in v0.5.0

type TensorInfo struct {
	Name string `json:"name"`
	// Type is the ggml type name ("q4_K", "q8_0", "f16", ...).
	Type string `json:"type"`
	// Shape is ne[0..]: the row length first.
	Shape []int64 `json:"ne"`
	// Buffer names the backend buffer holding the tensor: "CPU_REPACK" for
	// tensors the CPU backend repacked for its GEMV/GEMM kernels (rows a
	// multiple of the repack width, a repackable type), "CPU" otherwise.
	Buffer string `json:"buffer"`
	// VecDotType is the activation type the tensor's dot product quantizes
	// to ("q8_K" for K-quants, "q8_0" / "q8_1" for the legacy types).
	VecDotType string `json:"vec_dot_type"`
}

TensorInfo describes one weight tensor of a loaded model and the kernel path the CPU backend gave it.

func (TensorInfo) Repacked added in v0.5.0

func (t TensorInfo) Repacked() bool

Repacked reports whether the CPU backend repacked the tensor for its GEMV/GEMM kernels; tensors that are not repacked run the per-row dot.

type Timings

type Timings struct {
	PromptMS float64 `json:"prompt_ms"`
	DecodeMS float64 `json:"decode_ms"`
}

Timings is the engine's wall-clock split of one generation.

Directories

Path Synopsis
cmd/perfgate command
Command perfgate is the CI entry point of internal/perfgate.
Command perfgate is the CI entry point of internal/perfgate.
perfgate
Package perfgate compares go-llama's throughput against a native llama.cpp build measured on the same machine, and gates CI on the ratio: a change that makes go-llama more than a threshold slower RELATIVE TO NATIVE than the recorded baseline fails.
Package perfgate compares go-llama's throughput against a native llama.cpp build measured on the same machine, and gates CI on the ratio: a change that makes go-llama more than a threshold slower RELATIVE TO NATIVE than the recorded baseline fails.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL