llama

package module
v0.4.2 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 4, 2026 License: MIT Imports: 16 Imported by: 0

README

go-llama

Go Reference CI

llama.cpp in pure Go — run GGUF models anywhere Go runs. No cgo, no shared library, one static binary.

The inference engine is llama.cpp compiled to WebAssembly and then translated to Go — no wasm runtime is involved at run time — built on the llamawasm2go module.

package main

import (
	"fmt"

	llama "github.com/goccy/go-llama"
)

func main() {
	inst, err := llama.New()
	if err != nil {
		panic(err)
	}
	defer inst.Close()

	model, err := inst.LoadModel("model.gguf")
	if err != nil {
		panic(err)
	}
	defer model.Close()

	ctx, err := model.NewContext(llama.ContextParams{NCtx: 2048})
	if err != nil {
		panic(err)
	}
	defer ctx.Close()

	res, err := ctx.Generate("Once upon a time", llama.Params{
		NPredict:    128,
		Temperature: 0.8,
		TopP:        0.95,
	})
	if err != nil {
		panic(err)
	}
	fmt.Println(res.Text)
}

Features

  • Pure Go: works anywhere Go compiles; CGO_ENABLED=0 friendly.
  • Independent instances: llama.New builds an engine with its own linear memory; create several and run them concurrently and in isolation.
  • One model, many contexts: contexts share the weights and keep their own KV cache, which is how to serve independent conversations from one model. One instance can also hold several models — what speculative decoding needs.
  • Sampling: temperature, top-k, top-p, min-p, typical-p, repetition / presence / frequency penalties, seeds, and GBNF grammars.
  • Streaming: Context.Stream calls you back with each piece of text as it is decoded.
  • Interruptible: Context.Interrupt stops a running generation from another goroutine.
  • Speculative decoding, LoRA adapters, chat templates, embeddings, scoring, and state save / load.
  • Configurable sandbox: options on New scope the guest to one directory, hand it an in-memory filesystem, cap its memory, and capture its stdio.
inst, err := llama.New(
	llama.WithPreopenDir("/srv/models"), // the only directory the guest can see
	llama.WithMaxMemory(6<<30),          // fail inside the guest, not in the host
)

Instances, models and contexts

llama.New returns a *Llama — one engine instance, with its own linear memory and C heap. It is fully independent of any other instance, so several can run concurrently on separate goroutines.

Within an instance you load one or more models with LoadModel; each model spawns contexts with NewContext that share its weights and keep their own KV cache. Because a whole instance carries the engine's memory, the common shape is one instance with as many models and contexts as you need — reach for a second instance when you want hard isolation between them.

inst, _ := llama.New()
defer inst.Close()

target, _ := inst.LoadModel("qwen2.5-3b.gguf")
draft, _ := inst.LoadModel("qwen2.5-0.5b.gguf") // same instance: speculative decoding

Close contexts, then models, then the instance — using any handle after its owner is closed is a use-after-free in the engine, and closed handles are refused.

Streaming and interruption

The engine is a single translated module with one C stack, so the goroutine running Generate is the only one that can be inside it. Streaming and interruption reach a running generation from opposite directions around that constraint:

res, err := ctx.Stream("Once upon a time", llama.Params{NPredict: 512},
	func(piece string) { fmt.Print(piece) })

Stream calls onPiece once per decoded token, on the generating goroutine itself — so there is no concurrency and nothing to drop, but the callback must be short and must not call back into the engine. It returns the same complete Result as Generate (the pieces concatenate to Result.Text; a Params.Stop string is delivered as decoded and only then trimmed, so the stream can run a few characters past the returned text). A nil onPiece makes Stream exactly Generate.

Interrupt goes the other way: it writes one aligned word straight into linear memory (never calling into the engine), which the generation loop reads once per token. It is safe to call from any goroutine while a generation runs; Generate then returns what it has with Reason == StopInterrupted.

Memory

wasm32 caps linear memory at 4 GiB, and the model weights plus every context's KV cache live inside it. Target quantized models comfortably under that — roughly 3B parameters at Q4 — and size ContextParams.NCtx accordingly. WithMaxMemory caps growth so an oversized model fails in the guest instead of growing the host process, and WithMemoryReserve reserves up front so a large load does not repeatedly grow and copy.

Performance

The generated Go is compiled by the Go compiler, and on amd64/arm64 most of it ships as assembly derived from that compilation. The SIMD kernels ggml relies on are native NEON on arm64 and SSE on amd64 (the latter under GOAMD64=v2 or higher — set it, or the vector helpers fall back to scalar Go). arm64 is the flagship target, where the dot-product kernels lower to SDOT/SMMLA.

Supply-chain verification

internal/llama.go in this repository is a release artifact of llama-wasm, not hand-written code. It is refreshed with:

make llama LLAMA_WASM_VERSION=v0.1.0

which downloads it and verifies its build-provenance attestation against llama-wasm's release workflow. CI re-runs that verification (make verify) on every push.

Testing

make test        # fetches a tiny GGUF into testdata/ and runs the suite

Point the suite at your own model with GO_LLAMA_TEST_MODEL=/path/to.gguf.

License

MIT (see LICENSE). llama.cpp is MIT; the embedded engine is a derivative work of it.

Documentation

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

This section is empty.

Types

type Build

type Build struct {
	SIMD       bool `json:"simd"`
	Threads    bool `json:"threads"`
	Exceptions bool `json:"exceptions"`
	MaxDevices int  `json:"max_devices"`
}

Build reports how the embedded engine was compiled. Diagnostics only; see BuildInfo.

type Context

type Context struct {
	// contains filtered or unexported fields
}

Context is one inference context over a Model: its own KV cache and its own sampling state. Independent conversations get independent contexts over the same weights.

func (*Context) Close

func (c *Context) Close() error

Close frees the context. The Model outlives it.

func (*Context) Embed

func (c *Context) Embed(text string, normalize bool) ([]float32, error)

Embed returns the embedding of text. The context must have been created with ContextParams.Embeddings set. normalize applies L2 normalisation.

func (*Context) EmbedTokens

func (c *Context) EmbedTokens(tokens []int32, normalize bool) ([]float32, error)

EmbedTokens is Embed from token ids instead of text.

func (*Context) Eval

func (c *Context) Eval(text string, addSpecial, parseSpecial bool) (EvalResult, error)

Eval decodes text into the context's KV cache without sampling — prompt prefill for a later Generate on the same context, or plain evaluation. Positions continue from the cache's current end; Reset starts over. addSpecial and parseSpecial mirror Tokenize.

func (*Context) Generate

func (c *Context) Generate(prompt string, p Params) (Result, error)

Generate runs generation from prompt and returns the whole result.

func (*Context) GenerateWithDraft

func (c *Context) GenerateWithDraft(draft *Context, prompt string, p Params, nDraft int) (Result, error)

GenerateWithDraft is Generate accelerated by speculative decoding: a draft context over a smaller model with the SAME vocabulary proposes up to nDraft tokens per round (0 picks a default) and this context verifies them in one batch. Every emitted token is sampled by THIS context's sampler chain, so the output distribution is exactly Generate's — the draft only trades its cheap decodes for larger verification batches.

Both contexts' KV caches restart from the prompt. Result.NDrafted and Result.NAccepted report how well the draft anticipated the target.

func (*Context) Interrupt

func (c *Context) Interrupt() error

Interrupt stops a running generation at its next token WITHOUT executing any engine code: it writes the interrupt flag straight into linear memory. Calling in instead would need the lock the generation holds, and would share the C stack with it. Generate then returns what it has, with Reason == StopInterrupted.

Safe to call from any goroutine, including while a generation runs. A call when nothing is running is a no-op: the flag is cleared when generation starts.

func (*Context) LoadState

func (c *Context) LoadState(state []byte) error

LoadState restores a state produced by SaveState. The context must be over the same model with compatible parameters, or the engine rejects the payload.

func (*Context) Reset

func (c *Context) Reset() error

Reset drops the KV cache so the next generation starts from nothing.

func (*Context) SaveState

func (c *Context) SaveState() ([]byte, error)

SaveState serializes the context's state — KV cache, sampling RNG and the prompt-prefix history — into a byte slice that LoadState can restore, also in a later process. Save a long system prompt's state once and every future context skips re-decoding it: restore the blob and either continue positionally (Eval-prefill style) or generate with Params.CachePrompt, which picks up the saved prefix immediately.

func (*Context) Score

func (c *Context) Score(text string) (ScoreResult, error)

Score computes the teacher-forced negative log-likelihood of text under the model, the quantity llama.cpp's perplexity tool averages. It decodes text into the context's KV cache; call Reset before generating afterwards.

func (*Context) ScoreChoices added in v0.4.0

func (c *Context) ScoreChoices(choices []string) ([]ScoreResult, error)

ScoreChoices scores candidate continuations of the context's CURRENT cache state: for each choice it returns the negative log-likelihood of the choice's tokens as the continuation of what the context has already decoded. Call it right after decoding the shared stem (Eval, or a Generate whose prompt just ran) — the first token of every choice is scored from the live next-position logits, the rest teacher-forced. Each choice is rolled back out of the KV cache before the next, so the choices never see each other and the context ends exactly where it started. Choices must be non-empty and must not contain newlines (the wire format separates them with '\n').

func (*Context) SetLoRA

func (c *Context) SetLoRA(adapters []LoRAWeight) error

SetLoRA replaces the context's ENTIRE adapter configuration — the set semantics of llama.cpp's own API. An empty (or nil) slice removes every adapter, so SetLoRA(nil) is the spelling of "clear".

func (*Context) Stream

func (c *Context) Stream(prompt string, p Params, onPiece func(string)) (Result, error)

Stream generates from prompt like Generate and calls onPiece with each piece of text as it is decoded, then returns the same complete Result.

onPiece runs inside the generation, on the goroutine that called Stream, so it must not call back into this context or into any other part of the engine. Keep it short: the generation is stopped while it runs.

The pieces concatenate to Result.Text, with one exception: a stop string is delivered as it is decoded and only then trimmed from Result.Text, so with Params.Stop set the stream can run a few characters past the returned text.

A nil onPiece makes this exactly Generate.

type ContextParams

type ContextParams struct {
	// NCtx is the context window in tokens. Zero means the model's training
	// context length. The KV cache scales with it.
	NCtx uint32
	// NBatch and NUBatch are the logical and physical batch sizes. Zero means
	// llama.cpp's defaults.
	NBatch  uint32
	NUBatch uint32
	// NThreads is how many threads ggml may use. It has an effect only in a
	// threads-enabled build (BuildInfo.Threads); the single-threaded wasm
	// clamps to 1.
	NThreads uint32
	// NSeqMax is the number of sequences the context can hold at once. A
	// value > 1 turns on the unified KV cache: NCtx stays the TOTAL cell
	// budget shared by every sequence, and ScoreChoices batches its
	// teacher-forced candidates into one decode (one sequence per candidate,
	// all sharing the stem). Zero or 1 keeps the single-sequence default,
	// where ScoreChoices decodes candidates one at a time.
	NSeqMax uint32
	// Embeddings puts the context in embedding mode, which Context.Embed
	// requires and which disables generation.
	Embeddings bool
	// RopeFreqBase and RopeFreqScale override the model's RoPE
	// configuration, the knobs behind context-length extension schemes.
	// Zero keeps the model's own values.
	RopeFreqBase  float32
	RopeFreqScale float32
}

ContextParams configures a Context. The zero value asks for the model's own training context length and llama.cpp's batch defaults.

type EvalResult

type EvalResult struct {
	NTokens int32 `json:"n_tokens"`
	NPast   int32 `json:"n_past"`
}

EvalResult reports what Eval put into the KV cache: NTokens from this call, NPast the total sequence length now cached.

type FS

type FS = base.FS

FS is the filesystem the guest sees when Config.FS is set. base.NewMemFS returns a ready-made in-memory implementation, so a model can be served out of memory rather than off disk.

type Instance added in v0.3.0

type Instance struct {
	// contains filtered or unexported fields
}

Instance is one engine instance forked from a Snapshot.

func (*Instance) Close added in v0.3.0

func (f *Instance) Close() error

Close releases the fork: its copy-on-write mapping is unmapped, returning every private page. The kept contexts need no individual Close — their backing memory is the mapping itself.

func (*Instance) Context added in v0.3.0

func (f *Instance) Context(name string) *Context

Context returns the kept context by name, or nil.

type Llama

type Llama struct {
	// contains filtered or unexported fields
}

Llama is one engine instance: llama.cpp compiled to wasm, with its own view of linear memory and its own C heap. Load models into it and open contexts over them. Instances are independent — several run concurrently — while within one instance calls are serialised.

Memory an instance never writes is physically shared with the other instances in the process: engines are built over copy-on-write maps of a process-wide image, and instances that load a model an earlier instance already loaded (same options, same path) share the weights the same way instead of loading them again. Set GO_LLAMA_NO_SHARED_IMAGE to force every instance onto a private allocation.

func New

func New(opts ...Option) (*Llama, error)

New brings up an engine instance. The zero-option instance gives the guest the host filesystem, an empty environment and discarded stdio — the filesystem access a Go program reading a model file already has.

func (*Llama) BuildInfo

func (l *Llama) BuildInfo() (Build, error)

BuildInfo reports how the embedded engine was compiled.

func (*Llama) Close

func (l *Llama) Close() error

Close releases the instance, including the copy-on-write memory mapping when the engine is backed by one. Models and contexts opened from it must be closed first; using any of them afterward fails — the engine's memory is gone.

func (*Llama) LoadModel

func (l *Llama) LoadModel(path string) (*Model, error)

LoadModel loads a GGUF model from path into this instance.

path is a guest path. Without WithPreopenDir / WithFS the guest sees the host filesystem, so a relative path is resolved against the working directory first; with a scoped filesystem it is used as given, relative to that root.

type LoRA

type LoRA struct {
	// contains filtered or unexported fields
}

LoRA is a loaded LoRA adapter. It adapts contexts over the model that loaded it (Context.SetLoRA) and must not outlive that model.

func (*LoRA) Close

func (l *LoRA) Close() error

Close frees the adapter. Contexts still using it must clear it first.

type LoRAWeight

type LoRAWeight struct {
	Adapter *LoRA
	Scale   float32
}

LoRAWeight pairs an adapter with the scale to apply it at.

type Message

type Message struct {
	Role    string `json:"role"`
	Content string `json:"content"`
}

Message is one turn of a chat, as ApplyChatTemplate consumes it.

type Model

type Model struct {
	// contains filtered or unexported fields
}

Model is a loaded GGUF model. Several contexts can share one, and several models can be loaded into one Llama instance at once.

func (*Model) ApplyChatTemplate

func (m *Model) ApplyChatTemplate(messages []Message, templateOverride string, addAssistant bool) (string, error)

ApplyChatTemplate renders messages into a prompt with the model's chat template. templateOverride replaces it, and is required when the GGUF carries none; addAssistant appends the generation prefix.

func (*Model) Close

func (m *Model) Close() error

Close frees the model. Contexts created from it must be closed first.

func (*Model) Detokenize

func (m *Model) Detokenize(tokens []int32, renderSpecial bool) (string, error)

Detokenize renders tokens back to text.

func (*Model) Info

func (m *Model) Info() (ModelInfo, error)

Info returns the model's metadata.

func (*Model) LoadLoRA

func (m *Model) LoadLoRA(path string) (*LoRA, error)

LoadLoRA loads a LoRA adapter GGUF for this model. path resolves like LoadModel's.

func (*Model) NewContext

func (m *Model) NewContext(p ContextParams) (*Context, error)

NewContext creates an inference context over the model.

func (*Model) TokenToPiece

func (m *Model) TokenToPiece(token int32, renderSpecial bool) (string, error)

TokenToPiece renders one token. A byte-level token can render to invalid UTF-8 on its own; accumulate pieces before treating the result as text.

func (*Model) Tokenize

func (m *Model) Tokenize(text string, addSpecial, parseSpecial bool) ([]int32, error)

Tokenize splits text into tokens. addSpecial adds the model's BOS/EOS convention; parseSpecial lets special-token text ("<|im_start|>") tokenize as that token rather than as its characters.

type ModelInfo

type ModelInfo struct {
	// Desc is llama.cpp's own one-line description, e.g. "llama 7B Q4_K - Medium".
	Desc string `json:"desc"`
	// NParams is the parameter count; SizeBytes the on-disk tensor size.
	NParams   uint64 `json:"n_params"`
	SizeBytes uint64 `json:"size_bytes"`
	// NCtxTrain is the context length the model was trained with — the
	// largest context worth asking for.
	NCtxTrain int `json:"n_ctx_train"`
	NEmbd     int `json:"n_embd"`
	NLayer    int `json:"n_layer"`
	NHead     int `json:"n_head"`
	NHeadKV   int `json:"n_head_kv"`
	NVocab    int `json:"n_vocab"`
	// HasEncoder / HasDecoder describe the architecture; a plain causal LM
	// has only a decoder.
	HasEncoder bool `json:"has_encoder"`
	HasDecoder bool `json:"has_decoder"`
	// BOSToken / EOSToken are the vocabulary's sentence delimiters, -1
	// when the model defines none. AddBOS is whether the tokenizer
	// prepends BOSToken by convention (what Tokenize's addSpecial obeys).
	BOSToken int32 `json:"bos_token"`
	EOSToken int32 `json:"eos_token"`
	AddBOS   bool  `json:"add_bos"`
	// ChatTemplate is the Jinja-ish template string the GGUF carries, or ""
	// when it has none. Chat needs a template from somewhere: either this or
	// SamplingParams-independent override passed to ApplyChatTemplate.
	ChatTemplate string `json:"chat_template"`
}

ModelInfo describes a loaded model. Everything here comes from the GGUF metadata, so a field is zero when the file does not carry it.

type Option

type Option func(*config)

An Option configures a Llama at New time.

func WithEnv

func WithEnv(env []string) Option

WithEnv sets the environment the guest sees. Without it the guest gets an empty environment: the host process environment is never leaked.

func WithFS

func WithFS(fs FS) Option

WithFS replaces the filesystem backend entirely — feed a model from memory or a virtual tree. It takes precedence over WithPreopenDir. base.NewMemFS returns a ready-made in-memory implementation.

func WithMaxMemory

func WithMaxMemory(bytes uint64) Option

WithMaxMemory caps linear-memory growth, so a model bigger than expected fails inside the guest instead of growing the host process. Zero means the engine's default ceiling: for an instance backed by a copy-on-write mapping that is the 64 GiB of address space the mapping reserves (untouched pages cost nothing), and for a private allocation there is no cap.

func WithMemoryReserve

func WithMemoryReserve(bytes int) Option

WithMemoryReserve pre-reserves linear memory. Loading a multi-gigabyte model otherwise grows memory in steps, copying it forward each time; reserving up front avoids that. Zero uses the engine's default.

func WithPreopenDir

func WithPreopenDir(dir string) Option

WithPreopenDir scopes the guest filesystem to a host directory. Guest paths then resolve inside it, so a model is named relative to it. Without it the guest sees the whole host filesystem, where absolute paths work as written.

func WithStderr

func WithStderr(w io.Writer) Option

func WithStdout

func WithStdout(w io.Writer) Option

WithStdout and WithStderr capture what the guest writes to fd 1 and fd 2. Unset discards, the default, because llama.cpp's own logging is silenced inside the bridge anyway.

type Params

type Params struct {
	// NPredict caps the number of tokens generated. Zero or negative means
	// "as many as the context allows".
	NPredict int
	// Temperature <= 0 selects greedy decoding, which makes generation
	// deterministic and skips the truncation samplers below.
	Temperature float32
	// TopK, TopP, MinP and TypicalP truncate the candidate set before
	// sampling. Zero disables each of them, as does 1.0 for the two
	// probability-mass ones.
	TopK     int
	TopP     float32
	MinP     float32
	TypicalP float32
	// RepeatPenalty, PresencePenalty and FrequencyPenalty discourage
	// repetition over the last RepeatLastN tokens. Zero disables all three;
	// so does RepeatPenalty 1.0, which is llama.cpp's own spelling of "off".
	RepeatPenalty    float32
	RepeatLastN      int
	PresencePenalty  float32
	FrequencyPenalty float32
	// Seed makes sampling reproducible. Zero is a fixed seed, so a
	// temperature above zero still replays identically; set a varying seed to
	// get varying output.
	Seed uint32
	// Mirostat selects the mirostat sampling algorithm: 0 off, 1 v1, 2 v2.
	// When on it replaces the TopK/TopP/MinP/TypicalP truncation samplers,
	// as in llama.cpp. MirostatTau is the target entropy (0 means llama.cpp's
	// 5.0) and MirostatEta the learning rate (0 means 0.1).
	Mirostat    int
	MirostatTau float32
	MirostatEta float32
	// IgnoreEOS keeps generating past end-of-generation tokens by excluding
	// them from sampling, like llama.cpp's --ignore-eos. Generation then runs
	// to NPredict, a stop string, or the context limit.
	IgnoreEOS bool
	// LogitBias adds a bias to specific tokens' logits before sampling.
	// math.Inf(-1) (or any very negative value) forbids a token outright.
	LogitBias map[int32]float32
	// Grammar is a GBNF grammar constraining the output. Empty means none.
	Grammar string
	// Stop ends generation when the output first contains one of these
	// strings — even mid-token — and the text is cut at the match.
	Stop []string
	// CachePrompt treats the prompt as the WHOLE intended context: the
	// longest prefix already in the context's KV cache is kept, whatever the
	// cache holds beyond it is dropped, and only the rest is decoded —
	// llama.cpp server's cache_prompt. Requests that share a long constant
	// preamble (a system prompt, a routing configuration) then pay only for
	// the part that changed; Result.NCached reports the reuse. Off, the
	// prompt appends at the cache's current end, which is what an
	// Eval-prefill continuation expects.
	CachePrompt bool
}

Params configures one generation.

The zero value is greedy decoding with no truncation samplers, no penalties and no token limit: the most predictable thing the model can do, and reproducible run to run. Every field is honoured exactly as written — a zero is a decision ("no top-k", "temperature 0 means greedy"), not "use some default" — so what a caller leaves out cannot be overridden by a default buried in the engine.

type Result

type Result struct {
	// Text is the generated text, with a matched stop string removed.
	Text string `json:"text"`
	// Tokens are the tokens behind Text.
	Tokens []int32 `json:"tokens"`
	// NPrompt is how many tokens the prompt occupied; NDecoded how many were
	// generated.
	NPrompt  int `json:"n_prompt"`
	NDecoded int `json:"n_decoded"`
	// NCached is how many leading prompt tokens a Params.CachePrompt
	// generation reused from the KV cache instead of re-decoding. Zero
	// without CachePrompt.
	NCached int `json:"n_cached"`
	// NDrafted / NAccepted are GenerateWithDraft's speculation counters —
	// how many tokens the draft proposed and how many the target accepted.
	// Zero on a plain Generate.
	NDrafted  int `json:"n_drafted"`
	NAccepted int `json:"n_accepted"`
	// Reason says why generation stopped; Interrupted is a shorthand for
	// Reason == StopInterrupted.
	Reason      StopReason `json:"stop_reason"`
	Interrupted bool       `json:"interrupted"`
	// Timings splits the generation's wall time between the prompt pass
	// and the decode loop.
	Timings Timings `json:"timings"`
}

Result is what Generate returns.

type ScoreResult

type ScoreResult struct {
	NTokens int32   `json:"n_tokens"`
	NScored int32   `json:"n_scored"`
	NLL     float64 `json:"nll"`
}

ScoreResult reports the teacher-forced negative log-likelihood of a text: the model decodes the tokenized text once and NLL sums -log softmax(logits_i)[token_{i+1}] over the NScored predicting positions. Perplexity is math.Exp(NLL / NScored).

type Snapshot added in v0.3.0

type Snapshot struct {
	// contains filtered or unexported fields
}

Snapshot is the copy-on-write image of a prepared instance. It stays valid for the life of the process; forks reference it and share its pages.

func NewSnapshot added in v0.3.0

func NewSnapshot(build func(*SnapshotBuilder) error, opts ...Option) (*Snapshot, error)

NewSnapshot boots an engine on a snapshot image and hands it to build to prepare: load the model, create contexts, decode their prompt prefixes (Params.CachePrompt), and Register the ones forks will use. When build returns, every kept context's threadpool is detached (threads are host constructs and cannot be captured) and the image is sealed.

The builder instance itself is consumed by the build; do not retain the *Llama, models or contexts it produced beyond the callback.

func (*Snapshot) Fork added in v0.3.0

func (s *Snapshot) Fork(nThreads uint32) (*Instance, error)

Fork brings up an instance from the snapshot: the prepared contexts are live immediately, with fresh threadpools of nThreads workers each (nThreads <= 1 keeps them single-threaded). Close the fork when done with it — its private pages are returned then.

type SnapshotBuilder added in v0.3.0

type SnapshotBuilder struct {
	*Llama
	// contains filtered or unexported fields
}

SnapshotBuilder is the instance handed to a NewSnapshot callback: a normal *Llama plus Register, which records the contexts forks will use.

func (*SnapshotBuilder) Register added in v0.3.0

func (b *SnapshotBuilder) Register(name string, c *Context) error

Register records a prepared context under name, making it available from every fork of the snapshot via Instance.Context(name).

type StopReason

type StopReason string

StopReason says why generation ended.

const (
	// StopEOS: the model emitted an end-of-generation token.
	StopEOS StopReason = "eos"
	// StopLength: the token budget ran out (Params.NPredict, or the context
	// filled up).
	StopLength StopReason = "length"
	// StopString: one of Params.Stop matched the tail of the output, which
	// was trimmed off.
	StopString StopReason = "stop"
	// StopInterrupted: Context.Interrupt was called from another goroutine.
	StopInterrupted StopReason = "interrupted"
)

type Timings

type Timings struct {
	PromptMS float64 `json:"prompt_ms"`
	DecodeMS float64 `json:"decode_ms"`
}

Timings is the engine's wall-clock split of one generation.

Directories

Path Synopsis
cmd/perfgate command
Command perfgate is the CI entry point of internal/perfgate.
Command perfgate is the CI entry point of internal/perfgate.
perfgate
Package perfgate compares go-llama's throughput against a native llama.cpp build measured on the same machine, and gates CI on the ratio: a change that makes go-llama more than a threshold slower RELATIVE TO NATIVE than the recorded baseline fails.
Package perfgate compares go-llama's throughput against a native llama.cpp build measured on the same machine, and gates CI on the ratio: a change that makes go-llama more than a threshold slower RELATIVE TO NATIVE than the recorded baseline fails.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL