weft

package module
v0.13.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 10, 2026 License: MIT Imports: 7 Imported by: 0

README

weft

CI Go Reference Version

A modular framework for building AI agents in Go — designed the way the standard library is: small interfaces, context everywhere, functional options, wrapped errors, and zero required configuration. One go get is the whole framework; one import path is the loop alone:

go get github.com/weftgo/weft@v0.13.0       # the framework: every package below, one version
go get github.com/weftgo/weft/core@v0.13.0  # the loop alone: its only dependency is the OTel API
Package What it gives you
weft The agent loop: tools from plain Go functions, parallel tool calls with defined failure semantics, typed streaming events, structured output, approvals, steering, subagents; OpenAI, Anthropic, Google adapters (weft/openai, …), MCP both ways (weft/mcp), reference middleware (weft/mw) and the offline test double (weft/wefttest)
weft/core The same loop as a module of its own, for a service that wants nothing else: core.New, core.Tool, core/wefttest, core/mw. weft re-exports it name for name, so weft.Agent is core.Agent
weft/thread Durable sessions: an append-only conversation tree with branching, compaction, approvals that survive restarts, steering and a bounded pool of child agents (jsonl, SQLite or memory storage)
weft/otel Recording in one line (defer otel.Install()()): every event, transcript and span exported over OpenTelemetry — to a local database, Studio, or any OTLP backend — with per-destination content policy and redaction
weft/obsdb The queryable store those records land in: SQLite locally, ClickHouse hosted (weft/obsdb/clickhouse)
weft/studio The Inspector: runs, sessions, traces and live streams in a web UI, an in-app devtools panel for your own pages, and a playground that re-runs a turn with an edited prompt, model or tools; the weft binary (weft/cmd/weft: weft studio, weft dev) serves it for apps in any language
weft/scope The devtools' scope header: scope.Header(handler, …) sets Weft-Scope on your own chat endpoint so the in-page panel follows the conversation and run a response belongs to
weft/runtime The playground's in-app side: your app executes experiment commands safely — side-effect tools are substituted or parked unless you opt them in

Concurrency is the point, not a feature: a step's tools fan out over goroutines, parallelism is a one-line dial, and tool failures never cancel their siblings.

Status: v0.13.0 — experimental, pre-1.0, released as one module (plus core; see Releases, CHANGELOG.md and, coming from 0.8, MIGRATION-0.9.md). The core's three load-bearing contracts — message model, error model, tool contract — are implemented and tested, and every layer has been through a production-readiness review. Serving, eval and the weft CLI come next — see the roadmap below.

Quick start

A tool is a plain function — the JSON Schema is reflected from the input struct, so the struct is both the contract and the documentation:

type EchoInput struct {
    Msg string `json:"msg" jsonschema:"the message to echo"`
}

echo := weft.Tool("echo", "Echo a message back, uppercased.",
    func(ctx context.Context, in EchoInput) (string, error) {
        return strings.ToUpper(in.Msg), nil
    })

An agent is a value: build it once, run it many times, concurrently. Offline, the model comes from wefttest: a scripted, deterministic stand-in for a real provider (an adapter slots into the same Model seam):

model := wefttest.Script(
    wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"msg":"hello"}`}),
    wefttest.Say("HELLO"),
)
agt := weft.New(
    model,                                   // or anthropic.Model("claude-sonnet-5") — see Providers
    weft.Instructions("You are a support agent."),
    echo,
)

res, err := agt.Generate(ctx, weft.Prompt("Echo hello."))
fmt.Println(res.Text(), res.Usage.Total())

For a one-off tool the input struct can be written inline in the handler signature; a named type reads better and is reusable across tools and tests.

Streaming is Go iteration — typed events, cancel via context, exactly one terminal error:

for ev, err := range agt.Stream(ctx, weft.Prompt("Echo hello.")).Events() {
    if err != nil {
        return err
    }
    switch ev := ev.(type) {
    case weft.TextDelta:
        io.WriteString(w, ev.Text)
    case weft.ToolArgsDelta: // progress: the model is still writing the call
    case weft.ToolStart:
        slog.Info("tool", "name", ev.Name, "seq", ev.Seq)
    case weft.RunFinish:
        slog.Info("done", "steps", ev.Steps, "tokens", ev.Usage.Total())
    }
}

Ending a run is a predicate, and the step budget is a separate safety net:

agt := weft.New(model,
    weft.StopWhen(weft.HasToolCall("submit_answer")), // intended end
    weft.MaxSteps(20),                                // runaway guard → ErrMaxSteps
    submit, search,
)
Structured output

Output[T] constrains the final answer to a struct: a submit_output tool with T's reflected schema is advertised, and the run ends when the model calls it with arguments that decode. An invalid submission is an ordinary tool error the model repairs. GenerateAs returns the value; OutputOf reads it from a streamed run's result.

type Verdict struct {
    Approved bool   `json:"approved"`
    Reason   string `json:"reason" jsonschema:"one sentence"`
}

agt := weft.New(model, weft.Output[Verdict](), lookup)
v, res, err := weft.GenerateAs[Verdict](ctx, agt, weft.Prompt("Review order 42."))

It is tool mode, so it works on every provider; a run that ends in text returns ErrNoOutput with the transcript attached.

Per-tool policy

Trailing options on Tool set policy for that tool alone. The same names on New set the agent-wide default:

run := weft.Tool("run_command", "Run a shell command.", runCommand,
    weft.Timeout(30*time.Second),   // deadline on ctx; expiry is an error result
    weft.MaxResultBytes(0),         // this tool's output arrives whole
    weft.StrictInput(),             // undeclared argument fields are rejected
)
agt := weft.New(model, weft.Timeout(10*time.Second), run, grep)

Bad arguments come back to the model in the schema's own words — field "days": expected integer, got string — so it can map the error to the schema it was shown. Undeclared fields are ignored by default; StrictInput rejects them by name.

The two seams

Behaviour attaches at two chi-style seams; observation at weft.Tap. WrapModel wraps the model, WrapTools wraps every tool call (on the agent, or on one tool). First listed is outermost. Package mw holds the reference set:

agt := weft.New(model,
    weft.WrapModel(
        mw.Log(logger),               // request summary + finish, Debug level
        mw.Fallback(backupModel),     // fail over once Retry has given up on the primary
        mw.Retry(mw.MaxRetries(3)),   // 429/5xx/net errors; retry-after honoured; backoff with jitter
        mw.RepairJSON(),              // close a truncated tool-call argument object once
    ),
    weft.WrapTools(
        mw.Audit(logger),             // every call: run, step, tool, duration, error + cause
        mw.Allow(policy.Permits),     // DENIED: tool "rm" is not allowed — the model sees it
        mw.MapErrors(nil),            // plain errors → INTERNAL: tool "x" failed (cause kept)
    ),
    tools...,
)

Middleware that verifies something puts it on ctx before next, and the handler reads it back through a typed accessor — the same shape as weft.CallFromContext (see ExampleWrapTools_context). A middleware panic is a tool error result, never a run error. Handlers give the model a code to branch on with *weft.ToolError (ORDER_NOT_FOUND: order 42 does not exist; the cause stays in logs); the loop's own failures are coded INVALID_INPUT and NO_SUCH_TOOL. docs/life-of-a-call.md shows where each thing sits; ADR 0006 is the decision.

Seams are the product

The governance features other frameworks ship as processors — PII scrubbing, prompt-injection heuristics, moderation, token limits, response caching — need no processor layer here; each is a closure at one of the two seams, and each carries its own dependencies and policy stances rather than importing yours. mw stays a reference set; the patterns live as tested examples:

  • PII scrub — a tool middleware that masks what results carry (Example_piiScrubMiddleware).
  • Allowlist — shipped as mw.Allow(permits).
  • Token limiter — a model middleware that refuses the call before the provider bills it (Example_tokenLimitMiddleware); weft.UsageLimit covers the measured side.
  • Response cache — a model middleware keyed on the ModelRequest (Example_responseCacheMiddleware); invalidation policy is the caller's.

A named mw package ships only when a pattern needs a dependency or a policy stance weft should own — the mw.RateLimit precedent (rate limiting stays an example because golang.org/x/time would be the module's first dependency). Revisit when a consumer asks for one by name.

Delegating to another agent

A subagent is a tool whose handler runs another agent — the orchestrator-worker pattern with zero new machinery. The child sees only the prompt; its events arrive in the parent's stream wrapped in weft.Nested (Seq from the parent's counter, so the stream stays replayable); its usage rolls into res.Usage and is recorded per call on StepRecord.SubagentUsage:

researcher := weft.New(model, weft.Tool("deep_search", "…", DeepSearch))
orchestrator := weft.New(model,
    weft.Subagent("research", "Research a question in depth.", researcher,
        weft.Timeout(2*time.Minute)),
)

Everything the tool contract offers applies: Timeout bounds the child run, Parallelism bounds concurrent delegations (four research calls in one step run four children), RequireApproval gates the delegation itself. A failed child is data the parent model sees (SUBAGENT_FAILED: …, the child's *RunError on ToolError.Err), a child that ends awaiting approval is SUBAGENT_PENDING — approval-gated tools belong in the orchestrator, not in a child — and a delegation to an agent already running in the call chain is refused (SUBAGENT_CYCLE). ADR 0014 records the mechanics, the lineage ids, and pi's AgentLanes counterpoint.

Composing agents

A plugin is func(deps) weft.Option — a family of tools and its policy closed over its dependencies as one value, composed with weft.Options (and weft.ToolOptions for the per-tool counterpart). Nothing registers itself, so there is no registry, no scopes, no dedup; dependencies are parameters, never globals, and the "registered twice" mistake panics at New as always:

func Orders(svc *OrderService) weft.Option {
    return weft.Options(
        weft.Instructions("You handle orders."),
        svc.Lookup(), svc.Refund(),
        weft.WrapTools(mw.Allow(svc.Permitted)),
    )
}
agt := weft.New(model, base, Orders(orders))
Approval

weft.RequireApproval() on a tool parks its calls: the run ends successfully with them on RunResult.Pending, the step's other tools having run. Resume with the transcript and a decision — over any transport, with no persistence required:

res, _ := agt.Generate(ctx, weft.Prompt("Refund order 42"))
for _, call := range res.Pending { /* ask someone */ }
res, _ = agt.Generate(ctx, weft.Messages(res.Messages...),
    weft.Approve(call.ID), weft.Deny(other.ID, "over the limit"))

Approved calls run (handlers see Call.Approved); denied ones become DENIED: <reason> results the model sees; undecided ones DENIED: no decision. Middleware can park any call by returning an error wrapping ErrApprovalRequired. It is a policy seam, not a security boundary (ADR 0007; examples/approval).

Observability

Set up an OpenTelemetry SDK and every run emits the full span tree — one invoke_agent span per run, one chat span per model call, one execute_tool span per executed tool call, a subagent's run nested under its delegating tool span — with the GenAI semantic attributes (provider, model, tokens, finish reasons). No weft option is needed: the core instruments through the OTel API, which is a no-op until your SDK registers. Exporters and backends are not weft's business.

tp := sdktrace.NewTracerProvider( /* your exporter */ )
defer tp.Shutdown(ctx)
// no weft option: the global provider is picked up, spans appear
agt := weft.New(model, tools...)
// or explicitly, without touching the global:
agt = weft.New(model, append(tools, weft.TracerProvider(tp))...)

For logs, weft.Logger(l) writes one Debug line per phase — run start, run finish, model call, tool call — with ids, the model, durations, usage, and outcomes; never message text or tool arguments. The default (slog.Default, resolved at log time) is silent until your handler enables Debug; slog.New(slog.DiscardHandler) turns the lines off. A handler that bridges slog to OTel correlates the lines with the spans for free: they are logged on the span-carrying context. examples/otel runs the whole thing against the real SDK offline. No prompt, message, tool argument or tool result reaches a span: ids, names, counts, durations, reasons, and error types only. Error text does travel — a failed run or model call records its error as the span's exception, and the log lines carry the error text a tool or model returned, because a log is the caller's (ADR 0016).

Recording runs

The pipeline is the recorder (ADR 0024): defer otel.Install()() writes the local sink (./.weft/weft.db, content on, no network) and every event, delta, transcript record and span leaves as it happens — a crash loses nothing emitted. Caller pairs ride weft.Metadata; a quiet run heartbeats, so a live run reads running and a crashed one interrupted. The recorder's database is package weft/obsdb: the OTLP-shaped model with the derived weft identity, read back through one obsdb.DB interface — runs, sessions, positioned event pages, transcripts, spans by run or trace, public ids — with obsdb/sqlite as the default backend and obsdb/clickhouse as the hosted one (column-compatible with the OTel Collector's ClickHouse exporter, so a stock collector can feed the same database).

defer otel.Install()()                          // the recorder: local sink, content on
db := otel.LocalDB()                            // the same handle Studio reads [D4]
page, _ := db.Runs(ctx, obsdb.RunQuery{})       // identity chain, status derived
rec, _ := db.Run(ctx, page.Runs[0].ID)          // the row + its subagent children
evs, _ := db.Events(ctx, rec.ID, -1, 50)        // positioned events after the cursor; -1 = from the start
Inspecting runs

Module weft/studio is the Inspector, rewritten on obsdb (ADR 0024 S4): the UI, a JSON API (runs with their filters, transcripts, spans, traces, sessions, public ids), the OTLP/HTTP ingest receiver, and the live SSE stream — served as one http.Handler with the UI embedded, no build step, nothing leaves the process. Setup A embeds it beside the app and passes the pipeline's handle for the live lane; a token (studio.Token) walls the API when it leaves loopback. Setup B is the weft binary's weft studio — package weft/cmd/weft, the one place that imports the ClickHouse driver — for any language's OTel app: UI + ingest + the playground + a dev token on 127.0.0.1:7331, --db sqlite://path or clickhouse://user:pass@host:9000/db (studio/README.md). The in-page devtools panel is also on npm as @weftgo/devtools, for bundled apps with no <script> tag: npm install @weftgo/devtools, then import { mount } from "@weftgo/devtools"; mount({ enabled: import.meta.env.DEV }).

mux.Handle("/studio/", http.StripPrefix("/studio",
    studio.Handler(studio.DB(otel.LocalDB()))))   // history + live [D4]
go install github.com/weftgo/weft/cmd/weft@latest
weft studio --db sqlite://.weft/dev.db            # OTel app: point
OTEL_EXPORTER_OTLP_ENDPOINT=http://127.0.0.1:7331 # its exporter here
The weft command

Every subcommand is a thin client of the Studio API or of studio.New — a convenience over options the app can set itself, so a team that skips the CLI loses nothing. Every connection flag mirrors an environment variable one to one (the table below).

Command What it does
weft studio [--addr] [--db] [--token] [--rotate-token] [--manifest] [--open] [--no-playground] Setup B: studio.New with the UI, OTLP ingest, the playground (inert until an app's weft/runtime connects; --no-playground turns it off) and a dev token stable per database. Reuses a Studio already serving the same database on 7331, else takes the next free port in 7331–7340; --addr pins. Writes the discovery file apps find it through (below). The manifest is --manifest, else the nearest weft.json upward (one line says which). --open (default on a terminal) opens the UI with the token in the URL fragment.
weft dev [studio's flags] [--no-watch] [--watch dir] [-- go run ./cmd/app] Studio (as weft studio, in-process, same port policy) plus your app run beside it and restarted on a .go save. Prints one line per start: studio http://127.0.0.1:7331/ · app pid 4242 · runtime rt_… registered. See below.
weft runs [--agent] [--since 2h|RFC3339] [--failed] [--limit] [--json] One row per run (id, agent, status, started, steps) from GET /api/runs; --limit defaults to 50 (0 lists all) and says so on stderr when it hid runs; --json for scripts.
weft open <run id> [--open] [--with-token] Prints the run's page, <url>/runs/<id>, bare — a fixed token may be the panel tokens' signing key, so it stays out of logs; --with-token prints the #token= fragment too; --open hands the browser the link with the token.
weft export <run id> [--format json|jsonl|otlp] GET /api/runs/<id>/export to stdout.
weft export <run id> --wefttest ./testdata [--test TestName] [--force] The run's wefttest replay fixtures unzipped into ./testdata/<TestName>/ (default: the run id), where wefttest.Replay(t, "testdata") reads them; a non-empty target needs --force, which replaces its *.json fixtures (never merges).
weft doctor One line per check of a running Studio, each read from GET /api/meta.
weft version The weft version this binary was built from.
Flag Environment Default
--addr WEFT_STUDIO_ADDR 127.0.0.1:7331 (unpinned)
--db WEFT_DB ./.weft/weft.db
--token WEFT_STUDIO_TOKEN weft studio, weft dev: the database's stable token, <db>.token (never printed); a generated one, printed, for a database with no file
--manifest WEFT_MANIFEST the nearest weft.json upward
--url (runs, open, export, doctor) WEFT_STUDIO_URL http://127.0.0.1:7331

weft dev's --watch and --no-watch shape the dev loop only and have no environment mirror; like --open and --no-playground they are not connection settings. --rotate-token has none either: it is an action, not a setting.

The dev token and the discovery file

The dev token is stable per database: the first start on a SQLite file writes 32 random bytes (base64url) beside it — <db>.token, so ./.weft/weft.db.token by default, mode 0600 — and every later start on that file serves the same token, so a restart keeps the browser tab, the app's WEFT_STUDIO_TOKEN and a second start's reuse probe valid. --rotate-token writes a new one (panel tokens signed with the old one stop verifying); --token / WEFT_STUDIO_TOKEN override it and leave the file alone. The stable token is the panel tokens' signing key, like a fixed one, so it is never printed: the banner names its file, weft open --with-token prints the fragment, --open hands it to the browser. A database with no file (:memory:, ClickHouse) gets a token generated for that process, printed as before.

Once its listener is bound, weft studio (and weft dev) writes studio.json — {"url","token","db","pid","started","version"}, mode 0600, the real port in it — and removes it on a clean exit. It goes to the first of ./.weft/ (when that directory exists: the project's own), $XDG_RUNTIME_DIR/weft/ (Linux), os.UserCacheDir()/weft/ (macOS, Windows, bare containers). otel.Install() and runtime.Install() read the same order when WEFT_STUDIO_URL is unset, so an app with defer otel.Install()() exports to the running Studio, and runtime.Install dials it, with no configuration — one INFO line names the Studio joined. After a default start ./.weft always exists, so the file lands there: run the app from the directory you ran weft studio in, or set WEFT_STUDIO_URL. The file is trusted only when its url is a loopback address (127.0.0.0/8, ::1, localhost), it is fresh (its pid alive, under 24 hours old) and, on unix, it is mode 0600 and owned by you — a studio.json that arrived through a git checkout (0644) or names another host is never used. Anything untrusted is skipped without a word above Debug and the next writer removes it. The writer also drops a .gitignore (*) into ./.weft when it has none, so the database, its token and the file stay out of commits. Two Studios never erase each other's file: the second one's exit puts the first's back. A reusing start writes nothing (the running Studio owns the file) and says so when that file is missing; its port probe sends the stable token only to an address your discovery file names (anything else on the port is probed bare). weft studio stops gracefully on SIGHUP too, so a closed terminal removes the file.

The file is a convenience, never a requirement: WEFT_STUDIO_URL always wins (the file is not read), WEFT_DISCOVERY=off turns the read off, otel.NoEnv() ignores it with the rest of the environment, and a Studio named in code — otel.Studio(url, tok) (which switches the read off: explicit > environment > file) or runtime.Studio(url, tok) — needs no file.

Setup A's handler serves GET <base>/panel-config.json — {"endpoint", "version", "capabilities"} — to a loopback (or AllowOrigins) Host and a same-origin, AllowOrigins or loopback Origin only (loopback by the Host's rule — localhost, *.localhost, 127.0.0.0/8, [::1] — even with AllowOrigins set), a 404 to anyone else, so the devtools panel reads its endpoint instead of inferring it from its own src; api/meta lists the panel-config capability.

weft dev
weft dev -- go run ./cmd/app      # default command: go run .

weft dev starts Studio exactly as weft studio does (same flags, same port policy, the playground on) and runs the command after -- with four environment variables added to yours:

Variable Value What reads it
WEFT_ENV dev (kept when you already set it non-empty) runtime.Install opens its link
WEFT_STUDIO_URL the Studio's URL otel's Studio destination; runtime.Install's default endpoint
WEFT_STUDIO_TOKEN the Studio's token (the real one: stable, fixed or generated) the same two
WEFT_DB the Studio's SQLite file, absolute (not set for ClickHouse) otel.Local("")'s path: the app's local sink and Studio share one file

The last three override whatever your shell had. Everything here is an env var or option the app can set itself — weft studio in one terminal and WEFT_ENV=dev WEFT_STUDIO_URL=… WEFT_STUDIO_TOKEN=… go run ./cmd/app in another is the same thing without the command.

The app runs in its own process group; a .go change under the working directory (--watch dir, repeatable, replaces it; .git, node_modules, vendor, testdata, .weft and dist are skipped) restarts it after 300 ms of quiet: SIGTERM to the group, five seconds, then SIGKILL, then the command again. A build failure is printed and the next save retries; an app that exits on its own is reported with its exit code and the next save restarts it. --no-watch turns the watching off, and weft dev then exits with the app's exit code (127 when the command cannot start at all). A directory that arrives with .go files in it (a checkout, a mv) restarts too; editor lock files (.#name.go) do not; a directory that cannot be watched (inotify's max_user_watches spent) is said in one line. Ctrl-C, SIGTERM and SIGHUP (the terminal closing) stop the app first — the same signal, SIGHUP sent as SIGTERM — then Studio. A SIGKILL of weft dev cannot be caught: on Linux go run still gets SIGTERM (Pdeathsig), but the binary it started — the grandchild — can be orphaned; elsewhere the whole group can.

Each start prints one line: the UI link (bare by default — the stable token and a fixed --token / WEFT_STUDIO_TOKEN stay out of the log; #token= only for a token generated for a database with no file), the app's pid, and the first runtime that registered with this Studio within five seconds — else no runtime registered yet (the app needs runtime.Install; WEFT_ENV=dev is set) and a later line when one does. A Studio bound to every interface (--addr 0.0.0.0:7331) is handed to the app, and printed, as 127.0.0.1.

Reuse follows the port policy, whose probe carries the fixed token (--token / WEFT_STUDIO_TOKEN), else the database's stable token — sent only to an address a trusted discovery file names (from another directory with --db on the same file the probe goes out bare and a second Studio starts on the next port): a Studio already serving the same database is reused (the app gets that URL and token; weft dev stops only the app), so two bare starts in a row reuse. A running Studio walled by another token answers the probe 401 and is skipped — weft dev takes the next port with its own Studio.

go run ./studio/examples/basic records demo runs (a tool call, a subagent, a failure) into an obsdb database and serves Studio on 127.0.0.1:7331; examples/studio-local is setup A's five lines with a live thread session, and its agent's PrepareStep trims a first-step paragraph from the system prompt from step 1 on — the Request pane's "changed by PrepareStep" diff (TestPrepareStepTrimsThePrompt pins the record).

Sessions — weft/thread

The core is stateless on purpose; weft/thread is the layer above it: a conversation as an append-only tree of entries, durable through a Storage backend, with branching, compaction and approvals built on the same tree. A session is a file you can read with jq and back up with cp — one header line, then one line per entry; nothing is ever rewritten or deleted in place.

st, _ := jsonl.Open(dir)                       // or thread.Memory() in tests
s, _ := thread.Create(ctx, st, agent)          // the agent is the session's own

turn, _ := s.Send(ctx, weft.User("Where is order 1234?"))
res, _ := turn.Wait()                          // prompt durable before the run;
                                               // the reply and the turn's ledger after
for ev, err := range turn.Events() { ... }     // forwarded run events, replayable

s.Branch(ctx, entryID)                         // navigate the tree; nothing lost
fork, _ := s.Fork(ctx, entryID)                // a new session, self-contained
s.Close(ctx)                                   // drain, seal, give up the writer's lease
again, _ := thread.Open(ctx, st, s.ID(), agent) // reopen from disk, same context

One Session writes a session, and Close is how it stops. A Session takes its session's writer lease with its first write (Create is one) and holds it until s.Close(ctx), which stops new work, drains the running turn and the queue, and seals the value (thread.ErrClosed). Until then a second Session on the same session opens and reads, and its writes fail with thread.ErrLocked; one that fell behind another writer fails with thread.ErrStale — open the session again. Open itself only reads: it writes nothing, locks nothing and starts no run. A Turn can be waited on three ways — turn.Wait(), turn.WaitContext(ctx), <-turn.Done() — and turn.Outcome() names how it ended; s.WaitIdle(ctx) also waits for the compaction a turn may trigger after it is decided.

Durability is per step (ADR 0011 §7): the turn's messages are appended as they join the run — through weft.OnMessages, the core's transcript observer — so a crash mid-turn loses nothing emitted; the prompt was already durable before the run started. Two backends carry it: thread/jsonl (one file per session) and thread/sqlite (one SQLite file, WAL; thread itself never imports the driver). A crash mid-append leaves at most a torn final line, which the next writer removes before it appends. Across processes and Storage values a second writer gets thread.ErrLocked, while readers never lock, including the live tail:

w := st.(thread.Watcher)
tail, _ := w.Watch(ctx, s.ID(), lastEntryID)       // entries after lastEntryID
for e, err := range tail { … }                     // another process's tail

p, _ := st.List(ctx, thread.Query{
    Meta:        map[string]string{"env": "prod"}, // every pair present, exactly
    TitleSearch: "checkout",                       // the session's current title
})                                                 // newest first; page with Before + BeforeID

docs/thread-operations.md is the operator's page: file layout and backups, locks and leases, shutdown, torn-tail repair, costs and limits, and which errors to retry.

Delegation is bounded and receipted (ADR 0022): a thread/pool admits at most max child runs at work at once — the bound is the Pool value's, FIFO, and a run that is only waiting on its own child holds no slot. p.MustWrap turns any agent into a delegation tool (p.Wrap returns the error instead of panicking) — sync by default, the call waiting for the child session's answer; with pool.Async() the tool result is the receipt and the child runs on — and p.Submit hands background work to a child session directly:

pl := pool.New(4)                                       // at most 4 child runs at work
research := pl.MustWrap("research", "Research a topic.", researcher)
rc, _ := pl.Submit(ctx, s, researcher, "survey the options") // a receipt, at once
rc2, _ := pl.Wait(ctx, s, rc.ID)                        // settled, or parked at an approval
err := pl.Decide(ctx, s, thread.Approve(childCallID))   // records and arms; does not wait
err = pl.Recover(ctx, s)                                // after a restart, once per parent
err = pl.Close(ctx)                                     // cancel what runs, drain

Every child is a session of its own, linked to the parent by lineage, its cost in the parent's Usage.Delegated bucket (not in the parent run's usage). A child that parks at an approval surfaces on the parent's Pending(); pl.Decide records the decisions in the parent and queues the child's resume — follow it with pl.Wait — and the parent's parked call completes with the child's answer. pl.Cancel ends one delegation, pool.Receipts(s) reads the ledger, and pl.Recover reattaches, settles or fails what a dead process left unsettled; it never re-runs a child.

Approvals (ADR 0021) make the core's run boundary durable: a gated call parks as a request entry written with its turn, Pending() survives restarts, and Decide records the decision and resumes on its own — the decision chain (grants, then a bounded Approver, then the park) runs before anything parks:

turn, _ = s.Send(ctx, weft.User("Deploy to prod."))
res, _ = turn.Wait()                       // res.Pending: the gated calls
for _, r := range s.Pending() { notify(r) } // durable, restart-safe

rt, _ := s.Decide(ctx, thread.Approve(id)) // resumes when the boundary completes
follow := turn.Next()                      // the same resume, from the parked turn

s.Grant(ctx, thread.Grant{                 // "always allow go test"
    Tool: "run",
    Args: []thread.Arg{thread.ArgGlob("/command", "go test*")},
})

// Decisions that cross a process boundary are signed.
ring, _ := thread.NewKeyring(thread.Key{ID: "k1", Secret: secret, Active: true})
s, _ = thread.Create(ctx, st, agent,
    thread.WithKeyring(ring),
    thread.RequireSigned(),                       // stored in the header: every Open enforces it
    thread.WithApprover(ask, 30*time.Second))     // the live step and the time it is given
r, _ := s.Request(callID)                         // a challenge bound to this request entry
sd, _ := ring.Sign(r, thread.Approve(callID))     // or key.Sign: one key is one approver
rt, _ = s.DecideSigned(ctx, sd)                   // verified fail-closed, single-use

Grants match tool plus argument predicates (ArgEquals, ArgPrefix, ArgGlob), expire, count uses, revoke, and can deny outright; Quorum(n) needs n distinct approvers (a signed approval counts as its key); RequestExpiry(d) lapses a request nobody decided; RequireSigned() closes the unsigned door for the session's whole life. A signed decision is bound to the request entry it was minted for, so it can never approve a later call that reuses the id. s.Audit() returns the approval trail from the file — an index of what the session recorded, not tamper-evidence. A Send while approvals pend queues behind them.

go run ./thread/examples/approvals parks a call, restarts, decides signed, resumes, and replays a rejected signature — offline, pinned.

Steering (ADR 0019) is what a Send does when the session is busy — the busy policy, per session or per Send:

s, _ := thread.Create(ctx, st, agent, thread.BusyPolicy(thread.Steer))

steer, _ := s.Send(ctx, weft.User("wait — metric units"))  // mid-run
steer.Wait()                                  // nil result: a steer has no run of its own
steer.Outcome()                               // delivered, deferred or dropped
s.Queue()                                     // steers and sends accepted, not yet run
n, _ := s.ClearQueue(ctx)                     // drop them: receipts, Turns end ErrDropped

turn, _ := s.Send(ctx, weft.User("stop, do this instead"),
    thread.As(thread.Interrupt))              // cancel the run, run this next
turn, _ = s.Send(ctx, weft.User("no — this road"),
    thread.As(thread.Rollback))               // …and branch back before it

Steer delivers into the running turn at the loop's drain points — after the tool batch, or at what would have been the final step — through a receipt that is durable from acceptance: queued → delivered | deferred | dropped, entries in the file. A steer that meets an intended end (StopWhen) or an open approval boundary never drains: it defers to a follow-up turn (steer.Next()). A send that waits for a turn of its own — the default Queue policy on a busy session — is durable the same way (an accepted receipt), so a crash loses no accepted message: the next Open restores it to s.Queue(), and it runs ahead of the next Send or at once with s.Continue(ctx). Interrupt cancels the in-flight run — its dangling calls record the interruption text — and denies a parked boundary it supersedes; Rollback also branches the leaf back, so the follow-up answers as though the interrupted turn never happened (its entries stay on their own line: nothing lost). A turn that overflows the window (weft.ErrContextOverflow, mapped by every adapter) compacts — reason overflow — and re-runs once (thread.ReRunOnOverflow(false) to turn it off); a second overflow fails the turn with both errors joined.

Compaction (ADR 0020) keeps long sessions inside the window without losing anything: the older part is summarized behind a fixed marker, the recent part stays raw, and the summarized entries stay in the file. thread.ContextWindow(n) arms the automatic trigger — the provider-reported input of the last step plus an estimated delta against window − Reserve, never a chars-per-token guess — and every layer is replaceable:

s, _ = thread.Create(ctx, st, agent,
    thread.ContextWindow(200_000),       // arms the trigger; ModelWindows per model
    thread.SummaryModel(cheap),          // falls back to the session model
    thread.SummaryFocus("keep file paths"),
    thread.ClearOldToolResults(4),       // stub old tool results before summarizing
    thread.BeforeCompact(hook),          // thread.Proceed() / Cancel() / Replace(c)
)
plan, _ := s.PreviewCompaction(ctx)      // the cut and the summary, no write
s.ApplyCompaction(ctx, plan)             // or s.Compact(ctx) for both
s.Compact(ctx, thread.SummaryInstructions("focus on the API design")) // per-call guidance
s.Uncompact(ctx)                         // branch back — undo is a navigation

thread.NoAutoCompact() turns the automatic trigger off and leaves the manual calls. Compaction is a between-turns operation: Compact, ApplyCompaction and Uncompact fail with thread.ErrBusy while a turn runs. Hooks run without the session's lock and may call the session. Each compaction also emits one informational marker record (kind compaction, scope session: counts and the entry's hash, no messages) through the agent's LoggerProvider when it lands, under the last run of this session that produced the compacted context — held for the next run when none did — which obsdb.DB.Compactions and Studio's run page read back.

go run ./thread/examples/session walks a session through turns, a label, a branch, a fork, a previewed compaction and a reopen from disk, offline through a scripted model.

weft/thread is pre-1.0 and not frozen: its API and its stored format may still change between minor versions, each change recorded in the CHANGELOG with what to write instead. What holds today: a reader fails loudly (ErrNewerFormat) on an entry kind or version it does not know, never skipping it, and golden files pin every format version the current build reads (ADR 0011, format reference).

The manifest — weft.json

One generated, committed, diffable description of every agent and tool (the code stays the only source of truth; the file is output, never input). Gate it with a golden test so it cannot go stale:

func TestManifest(t *testing.T) {
    b, err := weft.Manifest(newAgent())
    if err != nil {
        t.Fatal(err)
    }
    wefttest.Golden(t, "weft.json", b) // regenerate: go test ./... -update
}

A tool or policy change without regenerating fails go test; the diff is the review artifact. (ADR 0012)

Without a weft.json, Studio's Agents page shows the manifests your app's runtime.Install registers (weft studio has the playground on by default), each hash labelled live or remembered; Studio keeps them in memory, so a restart forgets them until the runtime registers again. With --no-playground the Agents page, the playground and the debugger say why they are off (/api/meta's capabilities_off).

A replay from step N may edit what it keeps before it runs: a kept step's user message (step 0's is the turn's prompt), a call's arguments (checked against the tool's schema, refused in the loop's INVALID_INPUT words), a tool result, a call-free reply, or a user message inserted at a step boundary — one transcript_edits list on the command. POST /api/playground/preview takes the same body and answers the exact first request that replay will send, diffed against the one step N recorded, without calling a model or needing a runtime; the replayed run carries weft.edits naming every edit (ADR 0029 decision 8).

MCP: both ways

weft/mcp (a package over the official Go MCP SDK, aliased sdk) is the bridge in both directions, with no adapter layer — the tool contract is the same shape (ADR 0015):

import (
    sdk "github.com/modelcontextprotocol/go-sdk/mcp"
    "github.com/weftgo/weft/mcp"
)

// Consume: a server's tools as ordinary weft tools. The schema bytes
// cross whole (an enum or oneOf reaches the model as sent), the
// calls forward the model's arguments verbatim, and every remote
// failure is a tool result the model sees — data, never a run error.
tools, err := mcp.Tools(ctx, sess, mcp.Prefix("gh_"), mcp.Policy(weft.Timeout(10*time.Second)))
// A tool the bridge cannot import fails that tool, not the listing:
// the good ones are in tools, the skipped ones are named. Warning or
// stop is your call; a listing failure (transport, ctx) is a plain err.
var skipped *mcp.ImportError
if errors.As(err, &skipped) {
    slog.Warn("mcp: tools skipped", "err", skipped)
    err = nil
}
if err != nil {
    return err
}

// Expose: weft tools — or a whole agent, as one named tool with your
// description — on any MCP server.
srv := sdk.NewServer(&sdk.Implementation{Name: "weft", Version: "0"}, nil)
mcp.AddTools(srv, lookup)
mcp.Serve(srv, agent, "Support agent.")   // agent + its tools, under its chain

Two warnings the godoc repeats. A server's tool descriptions are untrusted content — they land in your model's tool list, a surface you did not author; filter with mw.Allow or read Tools()' output before registering it. A foreign tool runs sequentially unless its server marks it readOnlyHint — the conservative reading of an untrusted hint for a tool whose handler you cannot read.

The loop, the seams, the manifest and wefttest treat an imported tool like any other (it is a RawTool); examples/ for both directions live in mcp/examples/ and run offline over in-memory transports. TestRoundTripIsLossless pins the property: export → import keeps the schema and the answers identical.

Providers

First-party adapters wrap the vendors' official Go SDKs — weft never owns an HTTP client — and ship in the framework module. One line per vendor, one adapter for the whole OpenAI-compatible long tail:

import (
    "github.com/weftgo/weft/anthropic"
    "github.com/weftgo/weft/google"
    "github.com/weftgo/weft/openai"
)

openai.Model("gpt-4o-mini")                          // or any compatible server via openai.BaseURL
anthropic.Model("claude-sonnet-5", anthropic.Thinking(true))
google.Model("gemini-2.5-flash")

Reasoning depth is per run: weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkHigh}) — an agent option sets every run's default, a run option overrides it for one — maps to whatever the provider expresses (reasoning_effort, budget_tokens, thinkingBudget); adapters document what they drop. The openai adapter picks the thinking wire form from the base URL; openai.Dialect pins it when detection can't.

Sampling is per run or per step the same way — weft.Params(weft. RequestParams{…}) (Temperature, TopP, MaxTokens, Stop, Seed) folds over the adapter's construction options — and every adapter carries ExtraBody/ExtraHeaders, the caller-wins escape hatch for vendor knobs weft has no option for. Forcing a step's tool calls is weft.ToolChoice (any, a named tool, or none with the catalogue still advertised — the router shape).

Prompt caching (Anthropic): anthropic.PromptCache() marks the request's stable prefix edges — the system block, the final tool definition, the trailing conversation edge — with Anthropic's ephemeral cache_control. Cache writes bill 1.25× and reads 0.1× the base input price, so a long transcript whose prefix repeats across steps saves from the second step on; Usage.CachedInputTokens and Usage.CacheWriteTokens show it measured. The prefix is the caller's to keep stable: a PrepareStep that trims messages invalidates the trailing breakpoint on purpose, and weft.ToolChoiceNone is how you forbid calls on a final step without dropping the tool definitions — and the cache prefix they anchor — from the request.

Every adapter passes the same executable contract (wefttest/conformance): streaming tool-call fragments are assembled into whole calls and — where the provider streams fragments at all (Caps.ToolArgDeltas; Google's calls arrive whole) — also surface live as ToolArgsDelta progress, provider errors pass through unchanged for errors.As, cancellation surfaces as ctx.Err(), a stalled stream fails with ErrStreamIdle while a slow-but-streaming one never does (both are pinned cases), and WEFT_MODEL_REQUESTS=deny refuses every self-built client's call before any network I/O — test suites that must stay offline get loud failures, not surprise bills. (ADR 0013)

The rules that matter

  • Tool error = data; run error = Go error. A failing (or panicking) tool becomes a result the model sees; siblings keep running. Only model failures, cancellation, and step exhaustion reach the caller, as *RunError with the partial transcript attached. (ADR 0002)
  • Messages are role + typed parts, with a versioned JSON contract: every part carries a type discriminator and transcripts round-trip through encoding/json. (ADR 0001)
  • Tools are generic functions whose schema derives from struct tags, shape-compatible with the official Go MCP SDK. (ADR 0003)
  • Concurrent tool events carry a total order (Seq), assigned and emitted atomically, so streams replay exactly. (ADR 0004)
  • Parallel by default, bounded (4); tools always start in call order; weft.Sequential() runs them one at a time for shared state; weft.Parallelism(n) for anything else.
  • String tool outputs are sent verbatim, everything else as JSON.
  • Every run has an id (RunStart, Run.ID(), RunResult.ID); every tool call can learn its own via weft.CallFromContext(ctx).
  • Truncation is visible, never silent: a max_tokens finish is recorded on RunResult.StopReason (the run still succeeds — callers decide what truncated text means); a max_tokens step with tool calls executes none of them — each gets a visible failure and the model retries with a full budget; and oversized tool results are capped (64 KiB by default, weft.MaxResultBytes(n) to change, 0 to disable, per tool or per agent) with a marker the model sees.
  • Two behavioural seams, one observation tap. WrapModel and WrapTools change; Tap sees. A third seam needs an ADR. (ADR 0006)
  • A hung tool never hangs the run: weft.Timeout(d) on a tool or agent turns an overdue call into an error result and moves on.
  • The Model stream contract is enforced: a stream that ends without ModelFinish, continues after it, carries a tool call with an empty ID or name, or panics fails the run wrapping ErrModelContract — a broken adapter cannot corrupt a transcript.
  • One dependency in the core module (weft/core): the OTel API (trace and logs — ADR 0016, ADR 0024), a no-op until an SDK registers — the zero-config instrumentation THE-END-GOAL sanctions as the core's single exception. The vendor SDKs, the OTel SDK and its exporters, the SQLite and ClickHouse drivers are dependencies of the framework module, never of core (ADR 0027).

Layout

facade.go             the framework's root package: generated aliases and wrappers over core
core/                 the loop, a module of its own (go.mod; its one dependency is the OTel API):
  message.go            message model (roles, parts, versioned JSON)
  errors.go, env.go     error model (sentinels, RunError, kill switch)
  tool.go, schema.go    tool contract, per-tool policy, schema reflection
  output.go             structured output (Output, GenerateAs, OutputOf)
  model.go              provider seam (streaming-first Model interface)
  events.go             sealed run-event set
  agent.go, run.go      agent construction options, run/stream/result
  loop.go               the loop: model call → tool chain fan-out → repeat; approval resume
  mw/                   reference middleware: Retry, Fallback, Log, RepairJSON, Allow, Audit, MapErrors
  wefttest/             scripted mock model + the conformance suite
mw/, wefttest/        the framework's facades over core/mw and core/wefttest (generated)
internal/facadegen/   the facade generator (internal/cmd/genfacade runs it under go generate)
openai/               OpenAI Chat Completions (+ compatible servers)
anthropic/            Anthropic Messages (thinking, signatures)
google/               Gemini via genai
thread/ (+ sqlite/)   sessions: the append-only conversation tree, durable Storage backends
otel/                 the observability pipeline: destinations, content policies, heartbeats
obsdb/                the observability database: model, DB interface, sqlite backend, obsdbtest
obsdb/clickhouse/     the hosted backend (collector-compatible schema, materialized views)
studio/               the Inspector on obsdb: UI + JSON API + OTLP ingest + live (web/ is its
                      Bun source)
cmd/weft/             the weft binary: `weft studio` (setup B), dev, runs, open, export, doctor
runtime/              the playground's in-app side: the Studio link, the experiment executor
examples/             runnable examples (getting-started, approval, otel, studio-local;
                      per-adapter: <adapter>/example)
docs/adr/             decision records for the contracts

Two Go modules (ADR 0027): the root is the framework — every directory above except core/ is a package of it — and core/ is the loop alone. A release is two tags, core/vX.Y.Z then vX.Y.Z, and the root requires core at that exact version with no replace; go.work joins them for development. The root package, mw, wefttest and wefttest/conformance are generated facades over core (make generate after changing core's exported API). The whole recorder-and-inspector story is two lines: defer otel.Install()() and studio.Handler(studio.DB(otel.LocalDB())) (Recording runs and Inspecting runs above); weft/runtime adds the playground with one deferred runtime.Install(...) call.

Development

make test   # go test -race ./... in every workspace module
make vet
make lint   # golangci-lint (CI uses .golangci.yml)
make live   # adapter conformance against real keys (-tags live)
make generate  # regenerate the facades after changing core's API
make apidiff   # public API of the framework module vs the last tag (reported)
make apidiff-core  # core vs its last core/v* tag (enforced)
make apidiff-all   # every workspace module vs its own last tag (CI runs it)
make fuzz    # 10 s per fuzz target; FUZZTIME=1m make fuzz for longer
make fmt

Requires Go 1.26 or newer; the current and previous Go releases are supported and both are tested in CI. A fuzz crasher fails CI, its input is uploaded, and the fix PR commits it under core/testdata/fuzz/ as a regression seed — the existing FuzzRepair seed got there that way.

Testing

In reach order: script the dialogue with wefttest.Script(wefttest.ToolCalls(...), wefttest.Say(...)) — offline, deterministic, no key; compare bytes with wefttest.Golden; and when the question is "what does my agent do with what the model actually said", record once and replay forever:

func model(t *testing.T) weft.Model {
    if os.Getenv("WEFT_RECORD") != "" { // the suite's own switch; wefttest never reads it
        return wefttest.Record(t, "testdata/replay", openai.Model("gpt-5", openai.APIKey(key)))
    }
    return wefttest.Replay(t, "testdata/replay")
}

Fixtures are pretty JSON a reviewer reads in a diff — re-recording is the review (ADR 0017). The adapters' own wire-format parsing is proven by wefttest/conformance against recorded .sse fixtures, not by replay (ADR 0013).

API stability is enforced, not aspired to. CI runs scripts/apidiff.sh over both modules, each compared against its own last tag. For core (the last core/v* tag) any incompatible change fails the build. Pre-1.0, a deliberate source-compatible evolution (widening a return type to a superset interface, adding a trailing variadic) can be acknowledged by adding apidiff's exact line to core/.apidiff-allow with a justification; the file is emptied at each tag. Renaming or removing an exported symbol always fails. The framework module holds the pre-freeze layers — thread, obsdb, otel, runtime, studio — so the gate reports its incompatible changes instead of failing, and such a change makes the next tag a minor bump with a breaking CHANGELOG entry.

Roadmap

  1. First provider adapters — done (OpenAI + compatible servers, Anthropic, Google; ADR 0013).
  2. The two middleware seams — done (WrapModel/WrapTools, package mw, the approval boundary; ADR 0006, ADR 0007).
  3. Loop refinements — done (subagents as tools, ModelRetry, usage limits, loop detection, PrepareStep; ADR 0014).
  4. MCP interop — done (consume and expose; ADR 0015). Core observability — done (OTel spans + slog lines; ADR 0016).
  5. The satellites: store — removed (step 5 of ADR 0024; v0.1.3 remains on the module proxy). studio — done (the Inspector, v0.1.0, ADR 0018; rewritten on obsdb, ADR 0024 S4). thread — shipped (sessions, branching, compaction, approvals, steering, pool; ADR 0011 and ADRs 0019–0022; the sandbox, ADR 0023, was abandoned). Pre-1.0 and not frozen: the 2026-10-01 review's fix train shipped as thread v0.9.0 (thread/sqlite v0.3.0), and a field trial in real use comes before any API or format freeze. obsdb, otel, runtime — done (ADR 0024: the recorder, its database, the playground). Next: serve and the eval/prompt/mem/trace modules.

License

MIT — see LICENSE.

Documentation

Overview

Package weft is a modular framework for building AI agents in Go, designed the way the standard library is: small interfaces, context everywhere, functional options, wrapped errors, and zero required configuration.

This package is the framework's front door: it re-exports the agent loop that github.com/weftgo/weft/core implements — tools from plain Go functions, parallel tool calls with defined failure semantics, typed streaming events, structured output, approvals, steering, subagents — so one import gives the loop and one module gives the whole framework:

github.com/weftgo/weft            the loop (this package)
github.com/weftgo/weft/openai     OpenAI and OpenAI-compatible servers
github.com/weftgo/weft/anthropic  Anthropic
github.com/weftgo/weft/google     Google Gemini
github.com/weftgo/weft/mcp        MCP both ways
github.com/weftgo/weft/mw         reference middleware
github.com/weftgo/weft/wefttest   the scripted model for offline tests
github.com/weftgo/weft/thread     durable sessions
github.com/weftgo/weft/otel       recording over OpenTelemetry
github.com/weftgo/weft/obsdb      the store the records land in
github.com/weftgo/weft/studio     the Inspector, devtools panel, playground
github.com/weftgo/weft/runtime    the playground's in-app side

A tool is a plain function; its JSON Schema is derived from the input struct. An agent is a value built once with options and run many times:

echo := weft.Tool("echo", "Echo a message",
	func(ctx context.Context, in struct {
		Msg string `json:"msg"`
	}) (string, error) {
		return "echo: " + in.Msg, nil
	})

agt := weft.New(model, weft.Instructions("You are helpful."), echo)
res, err := agt.Generate(ctx, weft.Prompt("Say hi."))

Every name here is an alias of, or a one-line wrapper around, the same name in core, so values flow between the two without conversion; the godoc of each is the authority. A service that wants the loop alone, with the OpenTelemetry API as its only dependency, imports github.com/weftgo/weft/core instead and writes core.New.

Index

Examples

Constants

View Source
const (
	ContentText      = core.ContentText      // assistant text, text deltas
	ContentReasoning = core.ContentReasoning // reasoning text and deltas
	ContentArgs      = core.ContentArgs      // tool call arguments and arg deltas
	ContentResult    = core.ContentResult    // tool results
	ContentMessages  = core.ContentMessages  // whole transcript batches (weft/otel redacts them per part with the four kinds above)
	ContentPrompt    = core.ContentPrompt    // the composed system text of a prompt record (ADR 0028)
	ContentStop      = core.ContentStop      // one stop sequence of a request record's params (ADR 0028)
)

The seven kinds of content the core ever puts in a record.

View Source
const (
	RoleUser      = core.RoleUser
	RoleAssistant = core.RoleAssistant

	// RoleTool carries the results of one step's tool calls back to the
	// model, one ToolResultPart per call.
	RoleTool = core.RoleTool
)
View Source
const (
	StopEndTurn   = core.StopEndTurn
	StopToolCalls = core.StopToolCalls
	StopMaxTokens = core.StopMaxTokens
)
View Source
const (
	ThinkUnset  = core.ThinkUnset // provider default; nothing is sent
	ThinkOff    = core.ThinkOff   // suppress reasoning where allowed
	ThinkLow    = core.ThinkLow
	ThinkMedium = core.ThinkMedium
	ThinkHigh   = core.ThinkHigh
)
View Source
const (
	ToolChoiceAuto  = core.ToolChoiceAuto  // provider default; nothing is sent
	ToolChoiceAny   = core.ToolChoiceAny   // some tool must be called
	ToolChoiceNamed = core.ToolChoiceNamed // the tool named by Name must be called

	// ToolChoiceNone forbids tool calls while keeping the catalogue
	// advertised. It exists for the prompt-cache interplay: removing
	// tools from the request to stop the model calling them invalidates
	// the cached prefix (ADR 0013's 2026-09-22 amendment), while none
	// keeps the bytes and forbids the calls.
	ToolChoiceNone = core.ToolChoiceNone
)
View Source
const (
	// CodeSubagentFailed marks a child run that failed. The message the
	// model sees carries the child's step and cause; the *RunError
	// itself stays on ToolError.Err, reachable through errors.As, so
	// middleware can branch on it without parsing that text.
	CodeSubagentFailed = core.CodeSubagentFailed

	// CodeSubagentPending marks a child run that ended awaiting an
	// approval decision. Approvals belong in the orchestrator, not in a
	// child: the parent's transcript has nowhere to carry the child's
	// pending call, so the delegation is refused loudly instead of
	// silently (ADR 0014).
	CodeSubagentPending = core.CodeSubagentPending

	// CodeSubagentCycle marks a delegation to an agent already running
	// in this call chain — refused before any model call.
	CodeSubagentCycle = core.CodeSubagentCycle
)

Codes the Subagent tool renders its failures with — model-visible contract, pinned by tests (ADR 0002's table; ADR 0014). Exported like the loop's own codes so policy middleware can branch on them.

View Source
const (
	// CodeInvalidInput marks tool arguments that do not decode into
	// the tool's input struct; the message names the field in the
	// schema's own vocabulary.
	CodeInvalidInput = core.CodeInvalidInput

	// CodeNoSuchTool marks a call naming a tool the agent does not
	// have.
	CodeNoSuchTool = core.CodeNoSuchTool

	// CodeDenied marks a call the approval boundary (Deny, or no
	// decision) or a policy middleware refused — mw.Allow renders it
	// too: to the model, a refused call is a refused call, whoever
	// refused it (ADR 0007).
	CodeDenied = core.CodeDenied
)

Codes the loop renders its own tool failures with — model-visible contract (ADR 0002, "tool errors have codes"), exported so external policy middleware can produce the same wire strings without duplicating them, the same class of constant as SchemaVersion.

View Source
const (
	// ReplayNever is the default: the call has side effects, or nobody
	// has said otherwise. A re-run answers it from the record or parks.
	ReplayNever = core.ReplayNever

	// ReplaySafe vouches that the call is idempotent and side-effect
	// free (a read): a re-run may execute it for real.
	ReplaySafe = core.ReplaySafe
)
View Source
const CodeRetry = core.CodeRetry

CodeRetry is the code ModelRetry renders with — the one code whose result is an instruction rather than a failure report: try the call again with the hint applied.

View Source
const SchemaVersion = core.SchemaVersion

SchemaVersion is the version of the message wire format. The JSON encoding of Message and its parts is a compatibility contract: within a version, field names and shapes change only additively. Persistence and serving layers envelope messages with this number; the core itself never needs it.

Variables

View Source
var (
	// ErrMaxSteps is returned when the model still requests tools after the
	// last allowed step. The partial transcript rides along on RunError.
	ErrMaxSteps = core.ErrMaxSteps

	// ErrUsageLimit is returned when a run's usage exceeds its
	// UsageLimit and the loop would otherwise call the model again. A
	// step that ends the run may overshoot and still succeed: a budget
	// stops further spend, it does not discard finished work.
	ErrUsageLimit = core.ErrUsageLimit

	// ErrModelRetriesExceeded is returned when one tool has produced
	// more consecutive RETRY results than MaxModelRetries allows — a
	// model that cannot self-correct is a run failure, not an infinite
	// loop.
	ErrModelRetriesExceeded = core.ErrModelRetriesExceeded

	// ErrLoopDetected is returned when repeats consecutive steps
	// requested the same set of tool calls (DetectLoops).
	ErrLoopDetected = core.ErrLoopDetected

	// ErrNoSuchTool is returned by Agent.CallTool when the call names a
	// tool the agent does not have. Inside the loop the same condition is
	// folded into an error tool result — data for the model to
	// self-correct, not a run failure.
	ErrNoSuchTool = core.ErrNoSuchTool

	// ErrInvalidToolInput is returned by ToolDef.Invoke and Agent.CallTool
	// when the arguments do not decode into the tool's input type. Like
	// ErrNoSuchTool, the loop turns it into model-visible data.
	ErrInvalidToolInput = core.ErrInvalidToolInput

	// ErrRunConsumed is returned by Run.Events when the event stream has
	// already been consumed; each Run yields exactly one sequence.
	ErrRunConsumed = core.ErrRunConsumed

	// ErrModelContract is wrapped around failures of a Model
	// implementation to honor the stream contract documented on Model:
	// events after ModelFinish, a stream ending without one, a tool call
	// with an empty ID or name, or a panicking stream. It signals an
	// adapter bug, not a model outage — providers' own errors surface
	// unwrapped.
	ErrModelContract = core.ErrModelContract

	// ErrUnsupported is wrapped by a Model that cannot honour part of a
	// request — a FilePart whose media type the provider does not accept,
	// a feature the vendor lacks. It is a run error (the model call
	// fails), so callers can errors.Is on it and fall back to another
	// model.
	ErrUnsupported = core.ErrUnsupported

	// ErrStreamIdle is the stream error an adapter yields when the gap
	// between two chunks exceeds its IdleTimeout. The ctx deadline is the
	// hard limit on a whole call; the idle timeout only catches a stalled
	// stream, so a slow but actively streaming response is never killed.
	// It is a run error like any provider error; callers errors.Is on it
	// provider-agnostically, without knowing which adapter timed out.
	ErrStreamIdle = core.ErrStreamIdle

	// ErrModelRequestsDenied is the stream error every first-party
	// adapter yields from Stream when ModelRequestsAllowed is false — the
	// kill switch for test suites that must never reach the network.
	// wefttest models ignore the switch, so ordinary offline tests are
	// unaffected.
	ErrModelRequestsDenied = core.ErrModelRequestsDenied

	// ErrApprovalRequired marks a tool call that must not run until a
	// human (or an outer system) decides. The loop raises it for tools
	// built with RequireApproval; tool middleware may return an error
	// wrapping it to defer any call. The run then ends successfully with
	// the call on RunResult.Pending; resume with Approve or Deny.
	ErrApprovalRequired = core.ErrApprovalRequired

	// ErrApprovalDenied is the cause on the error result the model sees
	// for a pending call that was denied (Deny, or no decision on
	// resume). It is a tool error — data — never a run error.
	ErrApprovalDenied = core.ErrApprovalDenied

	// ErrDuplicateTool is a run error raised when a step's tool
	// snapshot contains a name twice — the runtime analogue of New's
	// duplicate-name panic. Fix the tool source; the run fails rather
	// than silently dropping the second tool. Agent.CallTool reports
	// the same condition as an error.
	ErrDuplicateTool = core.ErrDuplicateTool

	// ErrNilTool is a run error raised when a tool-source snapshot
	// contains a nil entry — a malformed snapshot, not a tool. The run
	// fails rather than advertising a dereference every adapter would
	// panic on; Agent.CallTool reports the same condition as an error.
	ErrNilTool = core.ErrNilTool

	// ErrInvalidSteer is returned when a steering source (Steering)
	// delivers a message whose role is not RoleUser: the model's own
	// turns come from the model. The run fails at the drain point with
	// nothing from that drain appended; earlier steers stay on the
	// partial transcript riding on RunError (ADR 0019 §3).
	ErrInvalidSteer = core.ErrInvalidSteer

	// ErrContextOverflow is wrapped by every first-party adapter around
	// its provider's error for a request that exceeds the model's
	// context window (ADR 0020 §5). Overflow is a request-shape
	// problem: the same bytes cannot succeed, so mw.Retry never retries
	// it — the session layer (weft/thread, v0.3) routes it to
	// compaction and one re-run instead. The provider's own error stays
	// reachable underneath for errors.As.
	ErrContextOverflow = core.ErrContextOverflow
)

Sentinel errors for named run failures. Branch on them with errors.Is; never match on error strings.

View Source
var ErrInvalidRunOption = core.ErrInvalidRunOption

ErrInvalidRunOption marks a per-run configuration the agent refuses: an unknown tool name in OnlyTools, or a MaxSteps/Parallelism raise (per run they may only lower). Returned wrapped in *RunError before any model call — a run that would misconfigure itself does not start (WEFT-PLAYGROUND §10.1 [D5]).

View Source
var ErrNoOutput = core.ErrNoOutput

ErrNoOutput is returned by GenerateAs and OutputOf when the run ended without a valid submit_output call — the model answered in text, hit the step budget, or never produced arguments that decode into Out.

Functions

func Manifest

func Manifest(agents ...*Agent) ([]byte, error)

Manifest renders the agents as their `weft.json` document: one generated, committed, diffable description of every agent and tool. Studio, docs, review, and compatibility checks read a file instead of a live process; the code stays the only source of truth. The file is output, never input — nothing is configured from it. Generate it in a golden test (wefttest.Golden) so a stale file fails `go test`; run that test with `-update` to regenerate.

Tools are listed in registration order, agents in argument order, and encoding/json sorts map keys, so the bytes are deterministic for the same agents. Unnamed and duplicate agent names are errors: the file is a review artifact, and `agent_1` in a diff is noise.

Example

The manifest is generated output: one committed, diffable description of every agent and tool. Gate it with a golden test so it cannot go stale.

package main

import (
	"context"
	"fmt"
	"log"
	"strings"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	agt := weft.New(wefttest.Script(wefttest.Say("ok")),
		weft.Name("support-bot"),
		weft.Instructions("You are a support agent."),
		weft.Tool("refund_order", "Refund a customer's order.",
			func(_ context.Context, _ struct {
				OrderID string `json:"order_id"`
			}) (string, error) {
				return "refunded", nil
			}),
	)
	b, err := weft.Manifest(agt)
	if err != nil {
		log.Fatal(err)
	}
	lines := strings.Split(string(b), "\n")
	fmt.Println(lines[0])
	fmt.Println(lines[1])
}
Output:
{
  "weft": 1,

func MetadataFromContext added in v0.6.0

func MetadataFromContext(ctx context.Context) map[string]string

MetadataFromContext returns a copy of the metadata in force on ctx — the run's own merged over its ancestors'. nil when there is none.

func ModelRequestsAllowed

func ModelRequestsAllowed() bool

ModelRequestsAllowed reports whether adapters may call a provider. It is false when WEFT_MODEL_REQUESTS=deny — the guard for test suites that must never reach the network (Pydantic AI's ALLOW_MODEL_REQUESTS). First-party adapters check it at the top of Stream and yield ErrModelRequestsDenied when it is false, before any network I/O — but only for a client they built themselves from credentials; a client the caller injected through the adapter's Client(c) option is a test double by construction and stays reachable (ADR 0013's kill-switch clause). wefttest models ignore the switch entirely.

The environment is read on every call, not cached: a model call is network-bound so one getenv is noise, test suites can toggle the switch per test with t.Setenv, and the package keeps no state at all.

func ModelRetry added in v0.2.0

func ModelRetry(hint string) error

ModelRetry is a tool error asking the model to try the call again with the hint applied: it renders as "RETRY: <hint>" and the loop continues. Return it when the arguments are well-formed but wrong in a way the model can fix — an ambiguous date, an id that needs a prefix. The loop counts RETRY results per tool name per run and fails the run with ErrModelRetriesExceeded when a tool exceeds MaxModelRetries (default 3); a successful result for that tool resets its count. Middleware-produced retries count too — the loop sees the code through the chain, not who returned it.

Example

A retry hint: the model sees "RETRY: <hint>", fixes the arguments, and the loop continues; a tool that cannot be satisfied fails the run after MaxModelRetries consecutive asks.

package main

import (
	"context"
	"fmt"
	"log"
	"sync/atomic"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	var calls atomic.Int32
	parse := weft.Tool("parse_date", "Parse a date.",
		func(_ context.Context, in struct {
			D string `json:"d" jsonschema:"the date, ISO-8601"`
		}) (string, error) {
			if in.D != "2026-09-19" {
				return "", weft.ModelRetry("date must be ISO-8601, e.g. 2026-09-19")
			}
			_ = calls.Add(1)
			return "2026-09-19", nil
		})
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "parse_date", Args: `{"d":"tomorrow"}`}),
		wefttest.ToolCalls(wefttest.Call{Name: "parse_date", Args: `{"d":"2026-09-19"}`}),
		wefttest.Say("Parsed."),
	), parse)
	res, err := agt.Generate(context.Background(), weft.Prompt("When is it?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println("first attempt:", res.Steps[0].Results[0].Content)
	fmt.Println("second attempt:", res.Steps[1].Results[0].Content)
}
Output:
first attempt: RETRY: date must be ISO-8601, e.g. 2026-09-19
second attempt: 2026-09-19

func OutputOf added in v0.2.0

func OutputOf[Out any](res *RunResult) (Out, error)

OutputOf decodes the structured output recorded in a finished run — the last submit_output call with a non-error result — for callers that streamed the run and hold its RunResult. It returns ErrNoOutput when no such call exists.

Types

type Agent

type Agent = core.Agent

Agent is an immutable, reusable value: a model, a system instruction, a tool set, and an execution policy. Build it once with New; run it many times, concurrently if you like — runs share no state, and the registered tool set is frozen at construction (see ToolDef).

Example (Conversation)

Continuing a conversation: feed the transcript back with the next question. The agent value is unchanged — the history lives in the messages you pass, never in the agent.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	agt := weft.New(wefttest.Script(
		wefttest.Say("Order 1234? It shipped yesterday."),
		wefttest.Say("Order 5678? Still pending."),
	))
	res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 1234?"))
	if err != nil {
		log.Fatal(err)
	}
	res2, err := agt.Generate(context.Background(),
		weft.Messages(res.Messages...), weft.Prompt("And order 5678?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
	fmt.Println(res2.Text())
}
Output:
Order 1234? It shipped yesterday.
Order 5678? Still pending.
Example (Generate)
package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	echo := weft.Tool("echo", "Echo a message.",
		func(_ context.Context, in struct {
			Msg string `json:"msg"`
		}) (string, error) {
			return "echo: " + in.Msg, nil
		})
	model := wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"msg":"hello"}`}),
		wefttest.Say("I echoed your message."),
	)
	agt := weft.New(model, weft.Instructions("You echo things."), echo)

	res, err := agt.Generate(context.Background(), weft.Prompt("Echo hello."))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
	fmt.Println("steps:", res.NumSteps(), "tokens:", res.Usage.Total())
}
Output:
I echoed your message.
steps: 2 tokens: 30
Example (Stream)
package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	roll := weft.Tool("roll_dice", "Roll a six-sided die.",
		func(_ context.Context, _ struct{}) (int, error) {
			return 4, nil // deterministic for the example
		})
	model := wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "roll_dice"}),
		wefttest.Say("You rolled a 4!"),
	)
	agt := weft.New(model, roll)

	for ev, err := range agt.Stream(context.Background(), weft.Prompt("Roll a die.")).Events() {
		if err != nil {
			log.Fatal(err)
		}
		switch ev := ev.(type) {
		case weft.ToolStart:
			fmt.Println("tool:", ev.Name)
		case weft.TextDelta:
			fmt.Println("text:", ev.Text)
		case weft.RunFinish:
			fmt.Println("done in", ev.Steps, "steps")
		}
	}
}
Output:
tool: roll_dice
text: You rolled a 4!
done in 2 steps

func AgentFromContext added in v0.3.6

func AgentFromContext(ctx context.Context) *Agent

AgentFromContext returns the agent running on ctx — the head of the ancestry chain: the agent whose run is producing the events a Tap or OnRunEnd observer receives, and that a tool handler is running for. For a subagent's child run it is the child (the chain grows by one per nesting level, ADR 0014). Nil outside a run — a manually dispatched Agent.CallTool, a handler invoked directly — so observers must not depend on it. Only the head is exposed, read-only: an Agent is immutable after New, so nothing can be changed through it. The consumer it ships for is the store's Record (ADR 0010), which describes the run by the agent that ran it (manifest hash, logger).

func New

func New(m Model, opts ...Option) *Agent

New builds an Agent. Nil models panic — including typed nils such as var m *someModel; New(m), which would otherwise crash much later inside a run goroutine. Everything else has a working default.

type AttemptInfo added in v0.10.0

type AttemptInfo = core.AttemptInfo

AttemptInfo is one provider request inside a model call, as the code that made it saw it: a retry middleware's try, a fallback's model, an adapter's request. Fields the reporter does not know stay zero. The attempt's number is not the caller's to give: the reporter numbers the model call's attempts 1, 2, … in the order they are reported, so numbers stay unique however many layers report.

type Call

type Call = core.Call

Call identifies the tool invocation a handler is serving. Retrieve it with CallFromContext — for audit logs, per-call idempotency keys, or progress reporting that must name its call.

func CallFromContext

func CallFromContext(ctx context.Context) (Call, bool)

CallFromContext returns the Call a tool handler is serving. ok is false when ctx did not come from the agent loop (a direct Invoke, for example).

type ContentKind added in v0.6.0

type ContentKind = core.ContentKind

ContentKind names what a content field holds. weft/otel's Redact receives it — for event fields, deltas, and each part of a transcript batch; the core only uses it in StripContent's table, to say which field of which event is content (ADR 0024 S1.1, [D2]).

type Event

type Event = core.Event

Event is the sealed set of run progress events, yielded by Run.Events in emission order. New event types may be added additively; external types cannot join, so switches over events stay exhaustively lintable.

On the wire every event carries a "type" discriminator (run_start, step_start, text_delta, reasoning_delta, tool_args_delta, tool_start, tool_finish, step_finish, steered, run_finish, nested) and UnmarshalEvent restores it — the same rule and the same compatibility contract as the message parts (ADR 0004).

Every event except RunStart carries RunID: concurrent runs on one agent emit interleaved streams, and a per-run Seq counter is unique only within its run, so RunID is what attributes an event to its run.

Events are snapshots. Their fields — including the Args byte slices — do not alias the run's transcript; a consumer may retain or write into them freely.

func StripContent added in v0.6.0

func StripContent(ev Event) Event

StripContent returns ev with every content field emptied (the table below) — what a content-off destination receives. Nested recurses. The shape is kept: ids, names, counts, positions and usage survive, so a stripped record still attributes and orders; it is no longer replay-grade, by design.

The request, prompt and tools records (ADR 0028) are not events and have their own rule, which weft/otel applies: a content-off destination keeps the request record with its params.stop emptied (hashes, names and numbers survive) and drops the prompt and tools records, as it drops transcript batches.

Event             Emptied                       Kept
text_delta        text                          run_id
reasoning_delta   text                          run_id
tool_args_delta   args                          run_id, name
tool_start        args (becomes null)           seq, call_id, name
tool_finish       content                       seq, call_id, name, is_error
steered           messages (becomes [])         seq, step
run_finish        each pending call's args      usage, steps, pending ids and names
run_start         nothing (no content)          all, instructions_hash included
step_start        nothing (no content)          all
step_finish       nothing (no content)          all
nested            recurses into event           the envelope
Example

StripContent empties every content field of an event — what a content-off destination receives. Identity survives; content does not.

package main

import (
	"encoding/json"
	"fmt"

	"github.com/weftgo/weft"
)

func main() {
	ev := weft.ToolFinish{RunID: "r", Seq: 3, CallID: "c1", Name: "lookup", Content: `{"status":"shipped"}`}
	b, _ := json.Marshal(weft.StripContent(ev))
	fmt.Println(string(b))
}
Output:
{"type":"tool_finish","run_id":"r","seq":3,"call_id":"c1","name":"lookup","content":"","is_error":false}

func UnmarshalEvent

func UnmarshalEvent(b []byte) (Event, error)

UnmarshalEvent decodes one wire event, dispatching on its "type" discriminator. An unknown or missing type is an error, never a silent drop: a recorded stream must replay exactly what was emitted.

Example

Recorded event streams decode back into typed events: store writes them, the Inspector replays them.

package main

import (
	"fmt"
	"log"

	"github.com/weftgo/weft"
)

func main() {
	ev, err := weft.UnmarshalEvent([]byte(`{"type":"tool_start","seq":5,"call_id":"c1","name":"echo","args":{"m":"x"}}`))
	if err != nil {
		log.Fatal(err)
	}
	start := ev.(weft.ToolStart)
	fmt.Println(start.Name, start.CallID, start.Seq, string(start.Args))
}
Output:
echo c1 5 {"m":"x"}

type FilePart

type FilePart = core.FilePart

FilePart is a file the user supplies to the model: an image, a PDF, audio. Exactly one of Data (inline; base64 on the wire via []byte's default encoding) or URL is set. Adapters map it to the vendor's image/document block; an adapter that cannot carry this MediaType fails the model call with an error wrapping ErrUnsupported. The core never reads the bytes. The exactly-one rule is documented, not enforced here — the adapter is the layer that knows what it can send, and it returns ErrUnsupported for a part with both or neither set.

type InstructionsOption added in v0.6.0

type InstructionsOption = core.InstructionsOption

InstructionsOption is accepted by both New and Stream/Generate (the ThinkingOption shape): the system prompt is as per-question as reasoning depth. On an agent it is every run's default; on a run it replaces that default for the run alone.

func Instructions

func Instructions(text string) InstructionsOption

Instructions sets the agent's system prompt. As an Option it is every run's default; as a RunOption it replaces the prompt for one run — the playground's prompt experiment, one run wide, no second agent built (WEFT-PLAYGROUND §10.1 [D5]). Manifest keeps reporting the agent's construction-time prompt.

type MaxStepsOption added in v0.6.0

type MaxStepsOption = core.MaxStepsOption

MaxStepsOption is accepted by both New and Stream/Generate (the ThinkingOption shape), with one rule on the run side: a run may only lower the agent's budget.

func MaxSteps

func MaxSteps(n int) MaxStepsOption

MaxSteps is the safety budget: the most model calls a run may make (default 10). Exceeding it fails the run with ErrMaxSteps — a runaway loop is a failure to surface, never a quiet success. Use StopWhen for the intended end of a run. As an Option it bounds every run; as a RunOption it may only LOWER the agent's value for the run — a raise fails the run with ErrInvalidRunOption before any model call, because a per-run raise is not a budget, it is the budget escaping (WEFT-PLAYGROUND §10.1 [D5]). Values below 1 are ignored.

type Message

type Message = core.Message

Message is one turn in a conversation: a role plus an ordered list of content parts. A step's tool results are collected on a single RoleTool message; provider adapters fan out or merge as their wire format requires.

On the wire every part carries a "type" discriminator ("text", "tool_call", "tool_result", "reasoning"), so a Message round-trips through encoding/json losslessly.

func Assistant

func Assistant(text string) Message

Assistant returns an assistant message with a single text part.

func Repair

func Repair(msgs []Message) []Message

Repair makes a transcript valid model input: every tool call has a result (missing ones become visible error results), results with no call are dropped, everything else is untouched. The loop applies it to the input of every run; store and runtime call it before persisting. It is pure (the input is never mutated) and idempotent: Repair(Repair(m)) equals Repair(m). nil input yields nil; an empty non-nil input yields an empty non-nil transcript.

Example

A partial transcript (the run was interrupted mid-step) becomes valid model input: the missing result is synthesised, visibly.

package main

import (
	"encoding/json"
	"fmt"

	"github.com/weftgo/weft"
)

func main() {
	msgs := []weft.Message{
		weft.User("Where is order 1234?"),
		{Role: weft.RoleAssistant, Content: []weft.Part{
			weft.ToolCallPart{ID: "c1", Name: "lookup_order", Args: json.RawMessage(`{"order_id":"1234"}`)},
		}},
	}
	for _, m := range weft.Repair(msgs) {
		fmt.Println(m.Role)
	}
}
Output:
user
assistant
tool

func User

func User(text string) Message

User returns a user message with a single text part.

func UserParts

func UserParts(parts ...Part) Message

UserParts returns a user message with the given parts, for prompts that mix text and files: UserParts(TextPart{"What is this?"}, FilePart{MediaType: "image/png", URL: u}). The parts are copied; the caller's slice is not retained.

Example

Provider reasoning round-trips: it is preserved in the transcript, placed before the text of the same assistant turn. A prompt can mix text and files with UserParts; the core carries the bytes and the adapter maps them to the provider's image block.

package main

import (
	"encoding/json"
	"fmt"
	"log"

	"github.com/weftgo/weft"
)

func main() {
	msg := weft.UserParts(
		weft.TextPart{Text: "What is this?"},
		weft.FilePart{MediaType: "image/png", URL: "https://example.com/cat.png"},
	)
	b, err := json.Marshal(msg)
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(string(b))
}
Output:
{"role":"user","content":[{"type":"text","text":"What is this?"},{"type":"file","media_type":"image/png","url":"https://example.com/cat.png"}]}

type Model

type Model = core.Model

Model is the provider seam. Implementations stream one step's output as events; first-party adapters wrap the vendors' official Go SDKs rather than re-implementing HTTP.

The stream contract:

  • Events are yielded in order: any number of ModelTextDelta, ModelReasoningDelta, and ModelToolCall values — optionally interleaved with ModelToolCallDelta progress as argument fragments stream — then exactly one ModelFinish.
  • Failure is reported as a single terminal yield of (nil, err); no events follow it.
  • The sequence honors ctx: when ctx is done, the model yields (nil, ctx.Err()) if it has not finished already.

The loop enforces this contract: a stream that ends without ModelFinish, continues after it, yields a tool call with an empty ID or name, two tool calls sharing an ID in one step, or panics fails the run with an error wrapping ErrModelContract. A contract-violating adapter cannot corrupt a transcript silently.

func Unwrap added in v0.3.7

func Unwrap(m Model) Model

Unwrap reports the Model one level inside m — the model a middleware wrapper wraps — or nil when m does not implement the optional

interface{ Unwrap() Model }

convention (the same shape as InfoOf's). Model middleware declares its inner model with it — func (w *wrapper) Unwrap() core.Model { return w.next } — so a caller walks a chain one middleware at a time, without knowing the wrapper types:

for m := agt.Model(); m != nil; m = core.Unwrap(m) { … }

The walk ends at the first non-wrapper, which is the model New was given unless user middleware wrapped something of its own.

type ModelEvent

type ModelEvent = core.ModelEvent

ModelEvent is the sealed set of events a model yields during one step. Tool calls arrive whole — assembling providers' streamed argument fragments is the adapter's job, which is what makes the core's streaming uniform across providers.

type ModelFinish

type ModelFinish = core.ModelFinish

ModelFinish closes a step with its stop reason and token usage.

type ModelInfo

type ModelInfo = core.ModelInfo

ModelInfo identifies a model for telemetry (§8.1) and the manifest. A Model that can report it implements the optional

interface{ Info() ModelInfo }

which the loop detects and surfaces on RunStart.Model; Model itself stays one method. Middleware that wraps a Model should forward Info (TODO §4.1).

func InfoOf added in v0.2.0

func InfoOf(m Model) ModelInfo

InfoOf reports m's ModelInfo when it implements the optional

interface{ Info() ModelInfo }

and the zero ModelInfo otherwise. Model middleware forwards identity with it: func (w *wrapper) Info() core.ModelInfo { return core.InfoOf(w.next) }.

type ModelMiddleware added in v0.2.0

type ModelMiddleware = core.ModelMiddleware

ModelMiddleware wraps a Model, the chi shape: it sees every ModelRequest the loop builds and every event the inner model yields, and may retry, substitute, log, or rewrite. Implementations should forward Info (see InfoOf) so RunStart.Model and the manifest still name the underlying model, and declare the model they wrap with Unwrap (see the Unwrap helper) so a caller can walk the chain one middleware at a time.

type ModelReasoningDelta

type ModelReasoningDelta = core.ModelReasoningDelta

ModelReasoningDelta is an increment of provider reasoning (Anthropic thinking, Gemini thought summaries). Signature is the provider's opaque token for the block, if any; adapters set it on the delta that completes a block. The core stores and forwards reasoning and never reads it.

Block boundaries: a delta carrying a non-empty Signature closes the current reasoning block; the next reasoning delta opens a new one. Providers send a block's signature last (Anthropic's signature_delta ends a thinking block; Gemini's per-part signature is emitted after the part's text), so one ReasoningPart per provider block survives the round trip. Reasoning without any signature accumulates into a single block — nothing downstream can send unsigned blocks back anyway, so their boundaries are not load-bearing.

type ModelRequest

type ModelRequest = core.ModelRequest

ModelRequest is everything a model needs for one step: the system instruction, the transcript so far, and the callable tools.

Read-only: the loop builds each request with fresh copies of the Messages and Tools slices, so appending to them or reassigning their elements cannot reach the run or the agent. The values inside remain shared — message Content parts, and ToolDef fields frozen at construction — so implementations must still not modify them, and must clone anything they retain beyond the call.

type ModelTextDelta

type ModelTextDelta = core.ModelTextDelta

ModelTextDelta is an increment of assistant text.

type ModelToolCall

type ModelToolCall = core.ModelToolCall

ModelToolCall is one complete tool invocation request. ID is the provider's call identifier, echoed back on the matching ToolResultPart; it must be unique among one step's calls (results and approval decisions key on it). Signature is the provider's opaque token attached to the call itself (Gemini attaches thought signatures to functionCall parts and requires them returned on the same part); adapters that do not have one leave it empty.

type ModelToolCallDelta added in v0.2.0

type ModelToolCallDelta = core.ModelToolCallDelta

ModelToolCallDelta is an increment of a streamed tool call's arguments — progress only: the assembled call still arrives whole as a ModelToolCall before ModelFinish. Adapters whose providers stream argument fragments (OpenAI-compatible function.arguments pieces, Anthropic input_json_delta) yield these so consumers can show the model "writing" a call instead of dead air; adapters whose calls arrive whole (Google) simply yield none. Index is the provider's fragment key where one exists (OpenAI's delta index); Name is the best-known name so far — for many providers only the first fragment of a call carries it.

type Nested added in v0.2.0

type Nested = core.Nested

Nested wraps one event of a child run started by a Subagent tool. CallID is the parent's tool call that owns the child run; Seq is from the parent's counter, so the parent stream stays totally ordered with the child's events in place. Event is any child event — including a Nested from a grandchild — and its own RunID and Seq are the child's. A child's RunStart..RunFinish all arrive inside the parent's ToolStart..ToolFinish for the delegating call, and none arrive after its ToolFinish (the late-event rule, ADR 0004). On the wire the type is "nested" and UnmarshalEvent restores the inner event recursively.

type Option

type Option = core.Option

Option configures an Agent at construction. Options are small values returned by Instructions, MaxSteps, Parallelism, Sequential, StopWhen, Output, the PolicyOptions (MaxResultBytes, Timeout, StrictInput), and the Tool constructor.

func Content added in v0.6.0

func Content(capture bool) Option

Content sets whether this agent's runs put content into their records, overriding whatever the logger in force asks for. The zero Option (no Content call) means "as the logger says": capture is resolved at each emission, not at New, because agents are usually built before the observability pipeline installs and the global provider delegates.

Content(false): never, even if a destination wants it.
Content(true):  always, even with no destination asking (tests).

Either way, spans carry no content (ADR 0016 O7); this option governs record bodies only. The core reads no environment variable — turning capture on per destination is weft/otel's job, through the same standard Enabled channel the core consults.

func DetectLoops added in v0.2.0

func DetectLoops(repeats int) Option

DetectLoops fails a run with ErrLoopDetected when repeats consecutive steps request the same set of tool calls — the same names with the same arguments, in any order. Varying arguments are not a loop: a corrected retry has a different signature. Results are not part of the signature: a tool whose output carries a timestamp must not hide a loop, and a result-inclusive hash would miss it. Off by default (values below 2 are ignored); the CLI scaffold writes DetectLoops(5).

Example

A stuck model repeating one request: DetectLoops fails the run loudly instead of burning the step budget.

package main

import (
	"context"
	"errors"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	echo := weft.Tool("echo", "", func(_ context.Context, _ struct{}) (string, error) {
		return "ok", nil
	})
	turns := []wefttest.Turn{}
	for range 3 {
		turns = append(turns, wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"i":1}`}))
	}
	agt := weft.New(wefttest.Script(turns...), echo, weft.DetectLoops(3))
	_, err := agt.Generate(context.Background(), weft.Prompt("q"))
	fmt.Println(errors.Is(err, weft.ErrLoopDetected))
}
Output:
true

func Logger added in v0.2.0

func Logger(l *slog.Logger) Option

Logger sets the logger the agent's runs report to, at Debug level: one line when a run starts and ends, one per model call, one per tool call — run and call ids, the model, durations, usage, stop reasons and outcomes; never message text, tool arguments or tool results. Error text is the one exception: a failed run, model call or tool call logs the error it returned, because a log is the caller's. nil (the default) means slog.Default, which is silent until its handler enables Debug, so weft logs nothing in a program that did not ask; slog.New(slog.DiscardHandler) turns the lines off outright. The lines carry the span-carrying context, so a handler that bridges to OTel correlates them with the spans for free. mw.Log and mw.Audit are separate: middleware the caller places, at the level the caller chooses.

func LoggerProvider added in v0.6.0

func LoggerProvider(lp log.LoggerProvider) Option

LoggerProvider sets the OpenTelemetry logger provider the agent's runs emit records to. Without it, runs use the global provider (go.opentelemetry.io/otel/log/global), a no-op until an SDK registers one — so a program that sets up an SDK gets weft's records with no weft option at all, and capture stays off until a destination asks for it. Tests and dependency-injected programs pass their own provider here instead of touching the global. Records are the run's events, deltas and transcript batches as OTel log records (ADR 0024); spans still go to the tracer provider (TracerProvider).

func MaxModelRetries added in v0.2.0

func MaxModelRetries(n int) Option

MaxModelRetries sets how many consecutive RETRY results (see ModelRetry) one tool may produce in a run before the run fails with ErrModelRetriesExceeded (default 3). The count is per tool name — two parallel calls to the same tool both retrying count as two — and a successful result for the tool resets it. Calls resumed under Approve do not feed the counter: they run before step 0 and belong to no step (ADR 0007). Values below 1 are ignored.

func Name

func Name(name string) Option

Name names the agent: it appears on RunStart.Agent and in the manifest, which requires it (core.Manifest errors on unnamed agents). Empty values are ignored.

func OnRunEnd added in v0.3.5

func OnRunEnd(fn func(ctx context.Context, res *RunResult, err error)) Option

OnRunEnd registers an observer the loop calls exactly once per run, after RunFinish is delivered or the RunError is built and before Run returns — the outcome signal a tap cannot carry: a failed run emits no event after its last delivered one (ADR 0004), so failure is invisible to Tap, and neither middleware seam wraps the run. res is the run's result and err its error: on success err is nil and res is complete; on failure err is the *RunError and res its partial transcript (ADR 0002) — the same value RunError.Result holds; on cancellation err is the RunError wrapping the ctx error. For a Subagent's child run it fires inside the parent's tool call, like the child's events, and res.ID names the child. A run that panics (a PrepareStep function is arbitrary user code) never fires it: the run crashed, it did not end — the panic still reaches the caller. Like Tap it observes and cannot change anything (ADR 0006's note: not a third seam — a seam wraps a call, this observes an outcome); like Tap a panic in it is contained and counted (TapPanics). Several OnRunEnd options run in registration order. ctx is the run's span-carrying context. The one consumer this ships for is the store (TODO §11): store.Record pairs a Tap for the event stream with OnRunEnd for the result and the failure.

Example

OnRunEnd is the outcome observer: unlike Tap it sees the run's end even when the run fails, because a failed run emits no event after its last delivered one. The store's Record option pairs the two.

package main

import (
	"context"
	"errors"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	agt := weft.New(
		wefttest.Script(wefttest.Fail(errors.New("provider down"))),
		weft.OnRunEnd(func(_ context.Context, res *weft.RunResult, err error) {
			fmt.Printf("run ended: has id=%v failed=%v steps=%d\n", res.ID != "", err != nil, res.NumSteps())
		}),
	)
	_, _ = agt.Generate(context.Background(), weft.Prompt("hi"))
}
Output:
run ended: has id=true failed=true steps=0

func Options added in v0.2.0

func Options(opts ...Option) Option

Options composes several options into one, applied in order — a family of tools and policy closed over its dependencies as a single value. A plugin is `func(deps) core.Option`, and dependencies are parameters, never globals; nothing registers itself, so there is no registry, no scopes, no dedup. Nil entries are ignored; duplicate tool names still panic at New. Nesting composes by construction.

Example

Composing agents: a plugin is func(deps) weft.Option — a family of tools and its policy closed over its dependencies as one value. Dependencies are parameters, never globals.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	orders := func(deps *string) weft.Option {
		return weft.Options(
			weft.Instructions("You handle orders."),
			weft.Tool("lookup_order", "Look up an order by id.",
				func(_ context.Context, in struct {
					ID string `json:"id" jsonschema:"the order id"`
				}) (string, error) {
					return "order " + in.ID + ": " + *deps, nil
				}),
			weft.MaxResultBytes(1024),
		)
	}
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":"42"}`}),
		wefttest.Say("Done."),
	), orders(new(string)))
	res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 42?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
}
Output:
order 42:

func Output added in v0.2.0

func Output[Out any]() Option

Output constrains the run's final answer to Out. It registers a tool named submit_output whose input schema is reflected from Out, exactly as Tool reflects a handler's input, and stops the run once the model has called it with arguments that decode. Invalid arguments come back to the model as an ErrInvalidToolInput result naming the field, so repair is the ordinary tool-error loop — nothing special happens. Read the value with GenerateAs, or OutputOf after Stream:

type Verdict struct {
    Approved bool   `json:"approved"`
    Reason   string `json:"reason" jsonschema:"one sentence"`
}

agt := core.New(model, core.Output[Verdict](), lookup)
v, res, err := core.GenerateAs[Verdict](ctx, agt, core.Prompt("Review order 42."))

Tool mode works on every provider; adapters with a native JSON-schema mode may use it later behind the same API. Output panics if Out is not a struct (or a pointer to one), for the reason Tool does.

func PrepareStep added in v0.2.0

func PrepareStep(fn func(ctx context.Context, step int, req ModelRequest) (ModelRequest, error)) Option

PrepareStep installs a function the loop calls before every model call, with the request it built for step: the agent's instructions, the transcript so far, the step's tool snapshot, the run's thinking level. The function returns the request the step uses — trimmed messages, a subset of tools, a rewritten system — or an error that fails the run. What it returns is what the step advertises and dispatches against: a tool it removes cannot be called that step, and a ToolDef it adds can. PromptSnippets are composed after it, from the tools it returns, so removing a tool removes its snippet without the function having to know snippets exist. The request is a copy the function may mutate freely: message parts, argument bytes, and tool definitions are cloned before the chain runs, so in-place writes reach neither the transcript nor the agent's frozen registry. The transcript in RunResult is never affected; only the request is. It runs before the model seam, so WrapModel middleware sees the prepared request. Several PrepareStep options run in order, each receiving the previous one's result. A nil function is ignored. Like ToolSource, it is one of the two knobs that can break a prompt-cache prefix — trim deliberately.

Example

Phased tool exposure: the first step plans with read-only tools; the second acts. PrepareStep is the one loop knob — what it returns is what the step both advertises and dispatches against.

package main

import (
	"context"
	"fmt"
	"log"
	"slices"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	lookup := weft.Tool("lookup", "Look up an order.", func(_ context.Context, _ struct{}) (string, error) {
		return "order 1234: broken item", nil
	})
	refund := weft.Tool("refund", "Refund an order.", func(_ context.Context, _ struct{}) (string, error) {
		return "refunded", nil
	})
	phase := func(_ context.Context, step int, req weft.ModelRequest) (weft.ModelRequest, error) {
		if step == 0 { // investigate before acting
			req.Tools = slices.DeleteFunc(req.Tools, func(t *weft.ToolDef) bool { return t.Name == "refund" })
		}
		return req, nil
	}
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "refund"}), // not advertised in step 0
		wefttest.ToolCalls(wefttest.Call{Name: "refund"}), // now it is
		wefttest.Say("Done."),
	), lookup, refund, weft.PrepareStep(phase))
	res, err := agt.Generate(context.Background(), weft.Prompt("Refund order 1234."))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
	fmt.Println(res.Steps[1].Results[0].Content)
}
Output:
NO_SUCH_TOOL: no tool named "refund"
refunded

func StopWhen

func StopWhen(conds ...StopCondition) Option

StopWhen adds stop conditions; the run ends when any one is met. Without StopWhen a run ends when the model replies without requesting tools. Compose the built-ins — HasToolCall, StepCountIs — or write your own:

core.StopWhen(core.HasToolCall("submit_answer"))
core.StopWhen(core.StopFunc(func(steps []core.StepRecord) bool { ... }))

Stop conditions are the intended end of a run; MaxSteps is the safety budget behind them.

func Tap

func Tap(fn func(ctx context.Context, ev Event)) Option

Tap registers an observer that sees every event of every run, including runs made with Generate, synchronously and in emission order on the emitting goroutine. It must be fast and must not block: it runs under the event-ordering lock, so a slow tap delays every tool event of its step and blocks the emitting tool goroutines — the same consumer-speed coupling Run.Events documents. Slow observation (a database write, a network sink) must not happen inside the tap: hand each event to a queue and drain it on your own goroutine; ExampleTap_async is the tested pattern. Taps run in registration order; a panic in one is recovered, counted (see TapPanics), and dropped, so a broken observer cannot break a run. Taps observe and cannot change anything — behaviour attaches at the two middleware seams. ctx is the run's context, which carries the run's span, so a tap that starts its own spans parents them under it for free.

Example

Tap observes every event of every run — including Generate, which has no stream to range over. Taps see; the middleware seams change.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	var calls int
	agt := weft.New(
		wefttest.Script(
			wefttest.ToolCalls(wefttest.Call{Name: "roll_dice"}),
			wefttest.Say("rolled"),
		),
		weft.Tap(func(_ context.Context, ev weft.Event) {
			if _, ok := ev.(weft.ToolStart); ok {
				calls++
			}
		}),
		weft.Tool("roll_dice", "Roll a die.",
			func(_ context.Context, _ struct{}) (int, error) { return 4, nil }),
	)
	if _, err := agt.Generate(context.Background(), weft.Prompt("Roll.")); err != nil {
		log.Fatal(err)
	}
	fmt.Println("tool calls:", calls)
}
Output:
tool calls: 1
Example (Async)

Slow observers (a database write, a network sink) must not run inside a Tap: taps are synchronous and run under the step's event-ordering lock, so a slow tap delays every tool event of its step. The pattern: hand each event to a queue inside the tap — never block — and drain it on your own goroutine, which may be as slow as it likes. A bounded queue with a visible drop counter keeps a stuck drain from wedging the run.

package main

import (
	"context"
	"fmt"
	"log"
	"sync/atomic"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	events := make(chan weft.Event, 1024)
	var dropped atomic.Int64
	done := make(chan int)
	go func() {
		starts := 0
		for ev := range events {
			if _, ok := ev.(weft.ToolStart); ok {
				starts++
			}
		}
		done <- starts
	}()

	agt := weft.New(
		wefttest.Script(
			wefttest.ToolCalls(wefttest.Call{Name: "ping"}),
			wefttest.Say("done"),
		),
		weft.Tap(func(_ context.Context, ev weft.Event) {
			select {
			case events <- ev: // fast: hand to the belt
			default: // full belt: drop, but visibly
				dropped.Add(1)
			}
		}),
		weft.Tool("ping", "", func(_ context.Context, _ struct{}) (string, error) {
			return "pong", nil
		}),
	)
	if _, err := agt.Generate(context.Background(), weft.Prompt("hi")); err != nil {
		log.Fatal(err)
	}
	close(events)
	fmt.Println("tool starts:", <-done, "dropped:", dropped.Load())
}
Output:
tool starts: 1 dropped: 0

func ToolSource

func ToolSource(fn func() []*ToolDef) Option

ToolSource replaces the tool set the loop advertises and dispatches against with the given function's return value, fetched fresh exactly once per step (and once per manual Agent.CallTool): the step's advertisement and its dispatch both resolve against that one snapshot, so what the model was shown is exactly what runs. The seam is for registries that change while the agent runs (plugins installed mid-run, MCP servers polled per step); a tool registered mid-step becomes callable on the next step's fetch. The Agent stays immutable: the source is a value; synchronization and freshness of the list belong to the source's owner. A snapshot with a duplicate name fails the run with ErrDuplicateTool and one with a nil entry with ErrNilTool (the runtime analogues of New's duplicate-name panic) rather than silently dropping the second tool or advertising a dereference. A nil function (the default) keeps the static construction-time list — Tool/option registration is then the only source of tools, byte-identical to an agent without a source. Manifest and Agent.Tools still report the static construction-time set: a manifest describes the code, not the registry behind a source.

func TracerProvider added in v0.2.0

func TracerProvider(tp trace.TracerProvider) Option

TracerProvider sets the OpenTelemetry tracer provider the agent's runs report spans to. Without it, runs use the global provider (otel.GetTracerProvider), which is a no-op until an SDK registers one — so a program that sets up an SDK gets weft's spans with no option at all. Tests and dependency-injected programs pass their own provider here instead of touching the global. Every run reports one invoke_agent span, one chat span per model call, and one execute_tool span per executed tool call, with the GenAI semantic attributes (ADR 0016); no message or tool-argument content is ever put on a span.

func UsageLimit added in v0.2.0

func UsageLimit(max Usage) Option

UsageLimit bounds a run's total token usage — the run's own model calls plus every subagent's (RunResult.Usage). A field left zero is unlimited. The limit is checked after each step, before the loop makes another model call: a step that ends the run — final answer, StopWhen, pending approvals — may overshoot and still succeed, because a budget's job is to stop further spend, not to discard finished work. Exceeding it fails the run with ErrUsageLimit and the partial transcript on RunError.Result. There is no default limit: the right value is workload-specific, and MaxSteps is the default budget. Usage.Total() is not a separate limit — a caller who wants a total sets both fields.

Example

A token budget: exceeded, the run fails with ErrUsageLimit before the next model call; the partial transcript rides on RunError.Result.

package main

import (
	"context"
	"errors"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	echo := weft.Tool("echo", "", func(_ context.Context, _ struct{}) (string, error) {
		return "ok", nil
	})
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "echo"}),
		wefttest.Say("never reached"),
	), echo, weft.UsageLimit(weft.Usage{OutputTokens: 4}))
	_, err := agt.Generate(context.Background(), weft.Prompt("again"))
	var re *weft.RunError
	if errors.As(err, &re) {
		fmt.Println(errors.Is(err, weft.ErrUsageLimit), "steps kept:", len(re.Result.Steps))
	}
}
Output:
true steps kept: 1

func WrapModel added in v0.2.0

func WrapModel(mw ...ModelMiddleware) Option

WrapModel installs model middleware around the agent's model. The first middleware listed is the outermost — WrapModel(a, b) calls a(b(model)) — and successive WrapModel options append inward. The chain is built once, at New, and sees each step's ModelRequest as the loop built it. The reference set is in package mw: Retry, Fallback, Log, RepairJSON. Nil entries are ignored.

Example

Model middleware wraps the agent's model, chi-style: the first listed is the outermost. The reference set lives in package mw.

package main

import (
	"context"
	"fmt"
	"iter"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	logged := func(next weft.Model) weft.Model {
		return loggingModel{next: next}
	}
	agt := weft.New(wefttest.Script(wefttest.Say("hello")), weft.WrapModel(logged))
	res, err := agt.Generate(context.Background(), weft.Prompt("hi"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
}

type loggingModel struct{ next weft.Model }

func (m loggingModel) Info() weft.ModelInfo { return weft.InfoOf(m.next) }

func (m loggingModel) Stream(ctx context.Context, req weft.ModelRequest) iter.Seq2[weft.ModelEvent, error] {
	fmt.Printf("model call: %d messages\n", len(req.Messages))
	return m.next.Stream(ctx, req)
}
Output:
model call: 1 messages
hello

type OutputDecoder added in v0.3.0

type OutputDecoder[Out any] = core.OutputDecoder[Out]

OutputDecoder turns a streaming run's submit_output argument deltas into a filling-in Out — the UI half of structured output: a consumer rendering a form as the model writes it. It is a decoder value, not an event-stream wrapper: feed it the events you already consume and render what comes back.

dec := core.NewOutputDecoder[Form]()
for ev, err := range run.Events() {
    if err != nil { return err }
    if p, ok := dec.Feed(ev); ok { render(p) }
}
form, err := dec.Result()

The decoder keys on the submit_output stream identity ToolArgsDelta carries — the tool name and step boundaries — resets its buffer on StepStart and on a fresh submit_output ToolStart, and closes it on the matching ToolFinish. Nested events are ignored: a subagent's structured output is its own decoder's job. Decode is lenient and prefix-shaped: after each delta it attempts the longest closed prefix of the arguments so far (see partial_json.go) and reports it when it changed and decoded — fields not yet present stay zero, a garbage mid-stream prefix keeps the last good partial, and errors surface only at Result, which follows the OutputOf rule (the last submit_output call with a non-error result) and returns ErrNoOutput when none finished. Never model-visible: the decoder reads the stream and writes nothing back.

Cost: every delta rescans the buffered arguments (the close is linear in what has arrived), so a submission's decode cost grows with the square of its size — the price of a fresh partial on every delta. At tool-argument scale (a few KiB) it is noise; a UI feeding very large submissions can trade freshness for linearity by calling Feed less often — the decoder keeps the last good partial across the deltas it skips.

Example

An OutputDecoder renders structured output while it streams: feed it the events you already consume and draw the filling-in form.

package main

import (
	"context"
	"encoding/json"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	type Form struct {
		Name  string `json:"name"`
		Count int    `json:"count"`
	}
	turn := wefttest.Raw(
		weft.ModelToolCallDelta{Index: 0, Name: "submit_output", Args: `{"name":"Ada",`},
		weft.ModelToolCallDelta{Index: 0, Name: "submit_output", Args: `"count":3}`},
		weft.ModelToolCall{ID: "c1", Name: "submit_output", Args: json.RawMessage(`{"name":"Ada","count":3}`)},
		weft.ModelFinish{Reason: weft.StopToolCalls, Usage: weft.Usage{InputTokens: 3, OutputTokens: 2}},
	)
	run := weft.New(wefttest.Script(turn), weft.Output[Form]()).Stream(context.Background(), weft.Prompt("fill the form"))
	dec := weft.NewOutputDecoder[Form]()
	for ev, err := range run.Events() {
		if err != nil {
			log.Fatal(err)
		}
		if p, ok := dec.Feed(ev); ok {
			fmt.Printf("render: name=%q count=%d\n", p.Name, p.Count)
		}
	}
	form, err := dec.Result()
	if err != nil {
		log.Fatal(err)
	}
	fmt.Printf("final:  name=%q count=%d\n", form.Name, form.Count)
}
Output:
render: name="Ada" count=0
render: name="Ada" count=3
final:  name="Ada" count=3

func NewOutputDecoder added in v0.3.0

func NewOutputDecoder[Out any]() *OutputDecoder[Out]

NewOutputDecoder returns a decoder ready to feed a run's events.

type ParallelismOption added in v0.6.0

type ParallelismOption = core.ParallelismOption

ParallelismOption is accepted by both New and Stream/Generate (the ThinkingOption shape), with the MaxStepsOption rule: a run may only lower the agent's width.

func Parallelism

func Parallelism(n int) ParallelismOption

Parallelism sets the maximum number of a step's tool calls executing at once (default 4). As an Option it bounds every run; as a RunOption it may only LOWER the agent's value for the run — a raise fails with ErrInvalidRunOption before any model call (the MaxSteps rule; the per-run knob is for safety, not for escape). Values below 1 are ignored.

type ParamsOption added in v0.3.0

type ParamsOption = core.ParamsOption

ParamsOption is accepted by both New and Stream/Generate: sampling knobs are as per-question as a forced choice (the ThinkingOption shape).

func Params added in v0.3.0

func Params(p RequestParams) ParamsOption

Params sets the agent's default sampling knobs — Temperature, TopP, MaxTokens, Stop, Seed — for every run's model calls; as a RunOption it overrides that default for one run:

agt := core.New(m, core.Params(core.RequestParams{Temperature: ptr(0.2)}))
agt.Generate(ctx, core.Params(core.RequestParams{Temperature: ptr(0.9)}), core.Prompt(q))

A nil or empty field keeps the adapter's construction-time default for that knob (its Temperature option, and so on); a run-level Params replaces the agent's struct whole, never merging field by field — set every knob the override should carry. A PrepareStep function can edit the request's Params per step ("cold for classification steps, creative for drafting" is a two-line function). Adapters drop knobs their provider lacks, and the provider's own limits apply (google narrows Seed to int32).

Example

Params sets per-step sampling: a PrepareStep function turns the temperature down for the classifying step and back for drafting — one struct, edited per step, no second Model construction.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	p := func(f float64) *float64 { return &f }
	classify := weft.Tool("classify", "Classify the request.",
		func(_ context.Context, _ struct{}) (string, error) { return "billing", nil })
	agt := weft.New(
		wefttest.Script(
			wefttest.ToolCalls(wefttest.Call{Name: "classify"}),
			wefttest.Say("A billing question, answered at temperature 0.9."),
		),
		weft.Params(weft.RequestParams{Temperature: p(0.9)}),
		weft.PrepareStep(func(_ context.Context, step int, req weft.ModelRequest) (weft.ModelRequest, error) {
			if step == 0 {
				req.Params.Temperature = p(0) // cold for classification
			}
			return req, nil
		}),
		classify,
	)
	res, err := agt.Generate(context.Background(), weft.Prompt("Why did my invoice double?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
}
Output:
A billing question, answered at temperature 0.9.

type Part

type Part = core.Part

Part is one content part of a Message. The set of part types is closed: text, tool calls, tool results, reasoning, and files today; approval parts are planned additions that will join this interface.

type PolicyOption added in v0.2.0

type PolicyOption = core.PolicyOption

PolicyOption is accepted by both New and Tool. On an agent it is the default policy for every tool call; on a tool it overrides the agent's default for that tool. The manifest records both levels.

func MaxResultBytes

func MaxResultBytes(n int) PolicyOption

MaxResultBytes sets the maximum size of one tool result's text, in bytes (default 64 KiB). The loop caps longer results — successes, failures, and panics alike — cutting on a rune boundary and appending a marker the model can see, so it knows the output is partial. MaxResultBytes(0) removes the cap; negative values are ignored. On a tool it overrides the agent's cap for that tool alone — a per-tool MaxResultBytes(0) lifts the cap for a tool whose output must arrive whole. Agent.CallTool returns uncapped output: the cap is a run policy, applied by the loop.

func Sequential

func Sequential() PolicyOption

Sequential restricts a step's tool calls to run one at a time, in call order: each tool finishes before the next starts — the safe setting for tools with shared state. Under any parallelism, tools *start* in call order; Sequential additionally serializes their execution. It also sets ModelRequest.SequentialTools, so adapters ask the provider not to emit parallel batches in the first place.

On a tool, Sequential is a barrier: the dispatcher lets every in-flight call of the step finish, runs this call alone, then resumes the step's parallelism for the calls after it. Other tools keep running in parallel; only this one is serialized.

Example

A Sequential tool is a barrier: it runs alone. The step's calls in flight finish first, and the calls after it wait — under any parallelism. Results stay in call order either way.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	write := weft.Tool("write", "Append to the ledger.", func(_ context.Context, _ struct{}) (string, error) {
		return "written", nil
	}, weft.Sequential())
	check := weft.Tool("check", "Verify the ledger.", func(_ context.Context, _ struct{}) (string, error) {
		return "verified", nil
	})
	model := wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "check"}, wefttest.Call{Name: "write"}, wefttest.Call{Name: "check"}),
		wefttest.Say("done"),
	)
	res, err := weft.New(model, write, check, weft.Parallelism(4)).Generate(context.Background(), weft.Prompt("x"))
	if err != nil {
		log.Fatal(err)
	}
	for _, r := range res.Steps[0].Results {
		fmt.Println(r.Name, "->", r.Content)
	}
}
Output:
check -> verified
write -> written
check -> verified

func StrictInput added in v0.2.0

func StrictInput() PolicyOption

StrictInput rejects tool arguments that carry fields the input struct does not declare, as an ErrInvalidToolInput result naming the field. The default is lenient — undeclared fields are ignored, the encoding/json default — because models routinely add stray keys and a rejection costs a round trip, not accuracy. Use it where an ignored field would be a silent misinterpretation of the call. On the agent it applies to every tool; on a tool to that tool alone. RawTool handlers are unaffected: they receive the raw arguments.

func Timeout added in v0.2.0

func Timeout(d time.Duration) PolicyOption

Timeout bounds one tool call. On a tool it is that tool's deadline — Timeout(0) on a tool removes the agent's default for that tool alone, the timeout analogue of MaxResultBytes(0); on the agent it is the default for every tool without its own, and non-positive values are ignored there. The handler's ctx carries the deadline; when it expires the loop records an error result ("tool X timed out after 10s") the model sees and moves on, abandoning the handler's goroutine — handlers must honour ctx to release their resources. Without any Timeout only the run's ctx bounds a call.

Example

Per-tool policy: trailing options on Tool override the agent's defaults for that tool alone. Here a slow tool times out into an error result the model sees, and the run carries on.

package main

import (
	"context"
	"fmt"
	"log"
	"time"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	slow := weft.Tool("slow", "Takes a while.",
		func(ctx context.Context, _ struct{}) (string, error) {
			<-ctx.Done() // a well-behaved handler honours the deadline
			return "", ctx.Err()
		},
		weft.Timeout(10*time.Millisecond))
	model := wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "slow"}),
		wefttest.Say("It did not answer in time."),
	)
	agt := weft.New(model, slow)

	res, err := agt.Generate(context.Background(), weft.Prompt("Try the slow tool."))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
	fmt.Println(res.Text())
}
Output:
tool "slow" timed out after 10ms
It did not answer in time.

func WrapTools added in v0.2.0

func WrapTools(mw ...ToolMiddleware) PolicyOption

WrapTools installs tool middleware. On the agent it wraps every tool call the loop (and Agent.CallTool) dispatches; on a tool it wraps that tool alone, inside the agent's chain: agent middleware → tool middleware → decode → handler. The first middleware listed is the outermost, so WrapTools(a, b) runs a(b(call)), and successive WrapTools options append inward. Middleware sees CallFromContext and may decorate ctx for the handler (see ExampleWrapTools_context); it returns the result text or an error — a *ToolError to give the model a code, an error wrapping ErrApprovalRequired to park the call on RunResult.Pending. The reference set is in package mw: Allow, Audit, MapErrors. Nil entries are ignored.

Example

Tool middleware wraps every call the loop dispatches. A middleware that returns an error produces an error result the model sees; a *weft.ToolError gives it a code.

package main

import (
	"context"
	"fmt"
	"log"
	"strings"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	readOnly := func(next weft.ToolCaller) weft.ToolCaller {
		return func(ctx context.Context, call weft.ToolCallPart) (string, error) {
			if strings.HasPrefix(call.Name, "delete_") {
				return "", &weft.ToolError{Code: "DENIED", Message: "this agent is read-only"}
			}
			return next(ctx, call)
		}
	}
	del := weft.Tool("delete_order", "", func(_ context.Context, _ struct{}) (string, error) { return "deleted", nil })
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "delete_order"}),
		wefttest.Say("I cannot do that."),
	), del, weft.WrapTools(readOnly))
	res, err := agt.Generate(context.Background(), weft.Prompt("delete order 1"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
	fmt.Println(res.Text())
}
Output:
DENIED: this agent is read-only
I cannot do that.
Example (Context)

The context-decoration convention: middleware that has verified something (here, who is calling) adds it to ctx before next, and the handler reads it back through a typed accessor — the same shape as weft.CallFromContext. Decorate only with data the middleware has verified; derive business values in the handler after decode.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	authUser := func(next weft.ToolCaller) weft.ToolCaller {
		return func(ctx context.Context, call weft.ToolCallPart) (string, error) {
			user, err := verifyUser(ctx) // a session lookup, a token check, ...
			if err != nil {
				return "", &weft.ToolError{Code: "UNAUTHENTICATED", Message: "sign in first", Err: err}
			}
			return next(withUser(ctx, user), call)
		}
	}
	myOrders := weft.Tool("my_orders", "List the caller's orders.",
		func(ctx context.Context, _ struct{}) (string, error) {
			u, ok := userFromContext(ctx)
			if !ok {
				return "", weft.Errorf("UNAUTHENTICATED", "no user on the call")
			}
			return "orders for " + u.Name, nil
		})
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "my_orders"}),
		wefttest.Say("Here they are."),
	), myOrders, weft.WrapTools(authUser))
	res, err := agt.Generate(context.Background(), weft.Prompt("show my orders"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
}

type exampleUser struct{ Name string }

type userKey struct{}

func withUser(ctx context.Context, u exampleUser) context.Context {
	return context.WithValue(ctx, userKey{}, u)
}

func userFromContext(ctx context.Context) (exampleUser, bool) {
	u, ok := ctx.Value(userKey{}).(exampleUser)
	return u, ok
}

func verifyUser(context.Context) (exampleUser, error) { return exampleUser{Name: "ada"}, nil }
Output:
orders for ada

type RawPair added in v0.10.0

type RawPair = core.RawPair

RawPair is one attempt's wire bodies: the request as sent and the response as received, as the reporter holds them. The bytes are content (the request carries the prompt), governed by the content policy wherever they are recorded.

type ReasoningDelta

type ReasoningDelta = core.ReasoningDelta

ReasoningDelta is an increment of provider reasoning, in the order the model produced it relative to TextDelta. Signatures are not streamed; they are on the ReasoningPart of the transcript.

type ReasoningPart

type ReasoningPart = core.ReasoningPart

ReasoningPart is one provider reasoning block surfaced by providers that expose it (Anthropic thinking blocks, Gemini thought parts). Weft preserves it in the transcript but does not act on it. Signature is the provider's opaque token for the block (Anthropic rejects thinking sent back without its signature); adapters echo it unchanged. A step yields one ReasoningPart per provider block — a delta carrying a signature closes the block (see ModelReasoningDelta) — placed before the TextPart of the same assistant message, in the order the model produced them.

Example
package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	model := wefttest.Script(
		wefttest.Think("The user greets; reply in kind.", wefttest.Say("Hello!")),
	)
	res, err := weft.New(model).Generate(context.Background(), weft.Prompt("Hi."))
	if err != nil {
		log.Fatal(err)
	}
	for _, p := range res.Messages[1].Content {
		fmt.Printf("%T\n", p)
	}
}
Output:
core.ReasoningPart
core.TextPart

type ReplayPolicy added in v0.7.0

type ReplayPolicy = core.ReplayPolicy

ReplayPolicy is a tool's side-effect class: what a re-run of a recorded conversation (a playground experiment, a replay fixture) may do with a call to this tool. The zero value, ReplayNever, is what an unannotated tool counts as — nobody has vouched for it, so its calls are substituted with the recorded result or parked for a human decision, never silently re-fired (WEFT-PLAYGROUND.md §6 rule 3; the one class a refund belongs to).

type Reporter added in v0.10.0

type Reporter = core.Reporter

Reporter is the reporting path from the model chain into the loop's own record (ADR 0016): middleware and adapters inside the chain tell the run's observer about attempts and wire bodies, which the loop cannot see from outside the chain. It is reporting, not a seam (ADR 0006): no report alters a step, a retry, a tool call or a model choice. A report never returns an error, never blocks and never panics into the caller — a panicking tracer or log handler is contained and counted in Agent.TapPanics, the report dropped. A report made after its model call ended is dropped, best-effort: a goroutine the chain left behind that reports while the call is ending may still land one attempt under the ended chat span. The zero Reporter, and the one ReportFromContext returns outside a run's model call, discards every report; a Reporter is safe for concurrent use.

Layers that report must not double-report one provider request. A model or middleware that reports its own attempts says so with an optional method, ReportsAttempts() bool, returning true; a reporting layer above it (mw.Retry, mw.Fallback) walks the Unwrap chain, finds the marker and stays silent. A layer that unwraps to a self-reporting model but may not stream through it (a router) returns false, which ends the walk. The marker is a convention, not a core type (ADR 0013).

func ReportFromContext added in v0.10.0

func ReportFromContext(ctx context.Context) Reporter

ReportFromContext returns the reporter of the model call whose context ctx is (or derives from): the loop puts one on the context it hands to the model chain, so a ModelMiddleware or a Model adapter reaches it from the ctx of its Stream. Outside a run's model call — a tool handler, a bare Model.Stream, a run started on the chain's context (it is masked at run start), any context the loop did not hand the chain — it is a no-op Reporter, never nil, so callers do not check. Using it is optional for adapters (ADR 0013).

type RequestParams added in v0.3.0

type RequestParams = core.RequestParams

RequestParams is per-step sampling: the knobs a caller turns between "cold for classification, creative for drafting". Every field is a pointer or slice so the three states stay distinguishable — nil or empty keeps the adapter's construction default (its Temperature, TopP, MaxTokens, Stop, or Seed option), set overrides it for this request alone, and the adapter never replaces a construction value with a zero. A set pointer to 0 is a value (Temperature of exactly 0 is sent), with one provider exception: an anthropic MaxTokens of 0 falls to the adapter's default, because the API requires a positive value. A negative MaxTokens fails the run at the step that carries it — no provider accepts one, and the loop names the bug rather than letting each adapter improvise. A run-level Params option replaces the agent's struct whole, it does not merge field by field; PrepareStep can edit it per step. Adapters drop knobs their provider lacks (Seed on anthropic) under the "adapters document what they drop" rule.

type Role

type Role = core.Role

Role is the author of a Message.

There is no system role: the system instruction is agent-level (Instructions) and travels on ModelRequest.System, so a transcript never carries it and adapters never have to merge it.

type Run

type Run = core.Run

Run is a handle to one streaming execution. Create it with Agent.Stream, then either range over Events (exactly once) or call Wait, which runs the agent to completion and reports the final result.

type RunError

type RunError = core.RunError

RunError reports a step-scoped failure: the model stream failed, the context was canceled, or the step budget ran out. Err is the cause — use errors.Is/As on it. Result carries the transcript up to the failure, so partial work is never lost.

type RunFinish

type RunFinish = core.RunFinish

RunFinish is always the final event of a successful run and carries the run's total usage and step count.

type RunOption

type RunOption = core.RunOption

RunOption configures a single run.

func Approve added in v0.2.0

func Approve(callID string) RunOption

Approve resumes a call left on RunResult.Pending by an earlier run: pass the earlier transcript with Messages and the decision, and the loop executes the call — through the ordinary tool chain, with Call.Approved set — before its next model call. Ids that are not pending are ignored.

func Deny added in v0.2.0

func Deny(callID, reason string) RunOption

Deny resolves a pending call without running it: the model sees an error result reading "DENIED: <reason>" and the loop continues. Pending calls given neither Approve nor Deny are denied with the reason "no decision" ("DENIED: no decision").

func Messages

func Messages(msgs ...Message) RunOption

Messages adds existing messages (a session transcript, few-shot examples) to the run's input.

func Metadata added in v0.6.0

func Metadata(kv map[string]string) RunOption

Metadata returns the RunOption attaching caller key/value pairs to the run: every span and every record of the run carries them, and so do the runs of its subagents (the pairs ride the context). Several Metadata options merge in order; a later key wins. Keys under "weft." are the weft modules' namespace by convention (thread writes weft.session.id); the core does not police callers.

Limits, because metadata rides every span and record: at most 64 keys, a key at most 128 bytes, a value at most 1024 bytes. An entry over a limit — or with an empty key — is dropped, never truncated, and counted on the run's invoke_agent span (weft.metadata.dropped); under the key cap the "weft." keys are kept first, then the rest in sorted order, so a large tag set never costs a run its session identity. Read the merged, limited view back with MetadataFromContext, e.g. inside a Tap, a tool handler, or a Subagent's child run.

Example

Metadata attaches caller key/value pairs to one run: every span and record of the run carries them, and a Subagent's child run inherits them through the context. Keys under "weft." are the weft modules' namespace — thread stamps weft.session.id this way.

package main

import (
	"context"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	agt := weft.New(wefttest.Script(wefttest.Say("ok")),
		weft.Tap(func(ctx context.Context, ev weft.Event) {
			if _, ok := ev.(weft.RunStart); !ok {
				return
			}
			md := weft.MetadataFromContext(ctx)
			fmt.Println("tenant =", md["tenant"], "session =", md["weft.session.id"])
		}))
	_, _ = agt.Generate(context.Background(),
		weft.Metadata(map[string]string{
			"tenant":          "acme",
			"weft.session.id": "s_01",
		}),
		weft.Prompt("hello"))
}
Output:
tenant = acme session = s_01

func OnMessages added in v0.5.0

func OnMessages(fn func(ctx context.Context, step int, msgs []Message)) RunOption

OnMessages returns the RunOption registering an observer the loop calls whenever messages join the run's transcript: the assistant message a step produced (reasoning, text and calls in their final shape, signatures included), the tool message that follows its calls, and the messages a steering drain delivered. msgs is exactly what joined, in transcript order, as a deep copy — retaining or mutating it changes nothing the run sees — and step is the step the messages belong to (a steer's messages name the step whose drain delivered them). The calls are synchronous on the run's goroutine, in transcript order, so a consumer that appends each batch to durable storage persists a mid-run crash's worth of exact transcript; like Tap an observer must be fast and must not block (the run waits for it), and like Tap a panic in one is contained and counted (TapPanics), never breaking the run. Observers cannot change anything — the transcript is the run's; behaviour attaches at the two seams. Several OnMessages options run in registration order. A Subagent's child run does not inherit them: a child's transcript belongs to whoever runs the child (the same rule as Steering). This is TODO §5.12's shape (b), the answer to "reconstruct messages from events (lossy: signatures, block boundaries) or wait for the run to end": the exact bytes, as they join.

func OnlyTools added in v0.6.0

func OnlyTools(names ...string) RunOption

OnlyTools narrows this run to the named tools among the agent's registered ones — the ToolSource snapshot when one exists, fetched fresh per step as ever. Narrowing only: the playground cannot add a tool, because a new tool is code. A name the agent does not have fails the run with ErrInvalidRunOption before any model call. The step's advertisement and its dispatch resolve against the same narrowed snapshot, so what the model was shown is exactly what runs; calls to a tool a PrepareStep function dropped fail as unknown, as today. Manifest and Agent.Tools keep reporting the static set: they describe the code, not one run's experiment (WEFT-PLAYGROUND §10.1 [D5]). With no names, the run keeps the agent's full set. Several OnlyTools options add up: the run keeps every tool any of them names.

Example

The playground's per-run configuration (WEFT-PLAYGROUND §10.1): one run of an immutable agent, changed without rebuilding it. OnlyTools narrows to registered tools; UseModel swaps in an allowed alternate.

package main

import (
	"context"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	lookup := weft.Tool("lookup", "Look up an order.", func(ctx context.Context, in struct {
		ID string `json:"id"`
	}) (string, error) {
		return `{"status":"shipped"}`, nil
	})
	agt := weft.New(wefttest.Script(wefttest.Say("order shipped")),
		weft.Name("support"), lookup)
	_, _ = agt.Generate(context.Background(),
		weft.Prompt("where is order 4411?"),
		weft.OnlyTools("lookup"),                       // narrowing; unknown name → ErrInvalidRunOption
		weft.Instructions("Answer in one short line."), // this run's prompt
	)
}

func ParkAllExcept added in v0.8.0

func ParkAllExcept(names ...string) RunOption

ParkAllExcept parks, at the approval boundary (ADR 0007), every tool call of the run whose tool is not named — ParkOn turned around: the caller lists what may run, and everything else waits for a decision. It is the rule for a run that must not fire a side effect nobody vouched for (a playground re-run, WEFT-PLAYGROUND §6 rule 3; ADR 0024 D7), where a list of tools to park cannot be complete:

  • The rule is applied by name to the tool each call resolves to in its step's dispatch snapshot, so a tool only a ToolSource supplies — absent from Agent.Tools and the manifest — parks like any other.
  • It reaches the runs started inside this run: a Subagent's child run (any run on a tool call's context) applies the same rule to its own tools, where ParkOn stops at the run it was given to. A child that parks ends as a child approval boundary always has — the delegating call's result is SUBAGENT_PENDING (ADR 0014); the parent does not park. Names are matched in parent and child alike, so name a tool only if every tool of that name down the delegation may run. A child run's own ParkAllExcept can narrow the inherited list, never widen it.

A parked call is ParkOn's parked call: its ToolStart and no ToolFinish, the run ending successfully with it on RunResult.Pending, Approve/Deny/Resolve on the resuming run deciding it — an approved call runs once even when the resume carries the rule again, and the next call to the tool parks again. The rules only add up toward parking: a tool ParkOn names or built with RequireApproval parks whether or not it is named here, and several ParkAllExcept options let through only the names all of them list. OnlyTools is independent — it decides what is offered, this decides what of it runs unasked. A name no tool carries is not an error (a ToolSource's names are not known up front) and with no names every call parks — except an agent's own Output submission (submit_output on an agent built with Output, in this run or a child's): it is the run's answer, not a side effect, so it never needs naming; ParkOn can still park it.

Example

ParkAllExcept is the default-deny park rule: the caller names what may run, and every other tool call parks — including a tool only a ToolSource supplies, which no list built from Agent.Tools could name.

package main

import (
	"context"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	lookup := weft.Tool("lookup", "Look up an order.", func(context.Context, struct{}) (string, error) {
		return "shipped", nil
	})
	wire := weft.Tool("wire_money", "Send a payment.", func(context.Context, struct{}) (string, error) {
		return "sent", nil
	})
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(
			wefttest.Call{Name: "lookup", ID: "c1"},
			wefttest.Call{Name: "wire_money", ID: "c2"},
		),
		wefttest.Say("paid"),
	), weft.ToolSource(func() []*weft.ToolDef { return []*weft.ToolDef{lookup, wire} }))
	res, err := agt.Generate(context.Background(),
		weft.Prompt("pay invoice 4411"), weft.ParkAllExcept("lookup"))
	if err != nil {
		return
	}
	fmt.Println("ran:", res.Steps[0].Results[0].Name)
	fmt.Println("pending:", res.Pending[0].Name)
	// A human decides; the next run resumes under the same rule:
	_, _ = agt.Generate(context.Background(),
		weft.Messages(res.Messages...), weft.Approve("c2"), weft.ParkAllExcept("lookup"))
}
Output:
ran: lookup
pending: wire_money

func ParkOn added in v0.6.0

func ParkOn(tools ...string) RunOption

ParkOn parks a call to any of the named tools at the approval boundary (ADR 0007), exactly as if the tool had been built with RequireApproval: the call gets its ToolStart and no ToolFinish, the run ends successfully with the call on RunResult.Pending, and Approve/Deny/Resolve on a resuming run decide it. This is how a breakpoint or side-effect parking reaches a runtime-started run without touching the agent, which is immutable after New (ADR 0024 D7, WEFT-PLAYGROUND §10.1). With no names, nothing parks. Names are not validated — a name no tool carries parks nothing — and the set covers this run only: a Subagent's child run does not inherit it. ParkAllExcept is the default-deny form, for when the tools to park cannot all be named.

Example

ParkOn parks the named tool's calls at the approval boundary, exactly as RequireApproval would — the breakpoint that reaches a run without touching the immutable agent. Approve resumes it.

package main

import (
	"context"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	refund := weft.Tool("refund", "Refund an order.", func(ctx context.Context, in struct{}) (string, error) {
		return "refunded", nil
	})
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "refund", ID: "c1"}),
		wefttest.Say("refunded"),
	), refund)
	res, err := agt.Generate(context.Background(),
		weft.Prompt("refund order 4411"), weft.ParkOn("refund"))
	if err != nil {
		return
	}
	fmt.Println("pending:", len(res.Pending))
	// A human decides; the next run resumes:
	_, _ = agt.Generate(context.Background(),
		weft.Messages(res.Messages...), weft.Approve("c1"))
}
Output:
pending: 1

func Prompt

func Prompt(text string) RunOption

Prompt adds a user message to the run's input.

func Resolve added in v0.3.0

func Resolve(callID, content string) RunOption

Resolve resumes a pending call with a result computed outside the process — the human-as-tool-executor shape: run the query in prod, paste what happened, and the next model call sees it. The content becomes the call's ToolResultPart verbatim; the handler never runs, no execute_tool span or ToolStart/ToolFinish is emitted (nothing executed), and Call.Approved is never set. MaxResultBytes applies as to any result. Composes with Approve and Deny in one resuming call; the last option for an id wins.

Resolve on a call that is not pending in the resumed transcript is a loud run error at step 0 — a deliberate asymmetry with Approve and Deny, which ignore unknown ids: those are yes/no marks over an id set, while Resolve carries a payload the caller expects the model to see, and dropping it silently is the one thing the error model forbids (ADR 0007's 2026-09-22 amendment).

func ResolveError added in v0.3.0

func ResolveError(callID, content string) RunOption

ResolveError is Resolve with the result marked as an error: the model sees the content on an error result, the shape a failed execution would have produced.

func RunID

func RunID(id string) RunOption

RunID sets the run's identifier instead of generating one — for replays, idempotent retries, and correlating with an outer system's own ids. Empty values are ignored.

func Steering added in v0.4.0

func Steering(fn SteerFunc) RunOption

Steering installs a steering source for this run: a pull hook the loop drains at two fixed points — after a step's tool batch, once every call of the batch has its result, and at a final step, where a delivered message redirects the run into one more step. Delivered messages are appended to the transcript before the next step's PrepareStep chain runs, so request rewrites see them, and reported as a Steered event between that step's StepFinish and the next StepStart.

The hook is never drained when the run ends at the approval boundary or through a StopWhen condition: those ends stay ends, and the source keeps its messages for a follow-up. A redirect at a final point consumes a step and goes through the same continuation checks as any continuation (MaxSteps, UsageLimit, DetectLoops); if they fail, the steer is in RunError.Result.Messages, delivered but unanswered.

It is a run option on purpose (ADR 0019): the run it steers is the one its queue belongs to, and a child run started by a Subagent tool does not inherit it — forwarding a steer to a child is a session decision made explicitly. A delivered message with any role other than RoleUser fails the run with ErrInvalidSteer: the model's own turns come from the model.

Example

Steering delivers a user's message to a running turn at a safe point: after the tool batch (every call paired with its result), or at what would have been the final step, which the steer redirects into one more step. The delivered message is ordinary transcript — the model, the record, and the next turn all see it — reported as a Steered event between StepFinish and the next StepStart (ADR 0019).

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	lookup := weft.Tool("lookup", "Look up an order.",
		func(_ context.Context, _ struct{}) (string, error) { return "shipped yesterday", nil })
	src := wefttest.NewSteers().At(0, weft.User("That is order 1234 — I meant 5678."))
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "lookup"}),
		wefttest.Say("Order 5678 is still pending."),
	), lookup)
	res, err := agt.Generate(context.Background(), weft.Prompt("Where is my order?"), src.Option())
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.NumSteps(), "steps")
	fmt.Println(res.Text())
	last := res.Messages[len(res.Messages)-2] // the steer, an ordinary user message
	fmt.Println(last.Role, last.Text())
}
Output:
2 steps
Order 5678 is still pending.
user That is order 1234 — I meant 5678.

func UseModel added in v0.6.0

func UseModel(m Model) RunOption

UseModel replaces the agent's model for this run, rebuilding the WrapModel chain over it (first registered = outermost, the New rule): the run's model calls go through the same middleware over m. Pass a model the runtime registered as an allowed alternate; a nil model is ignored. RunStart.Model and the chat spans report the run's model (middleware forwards Info); Agent.Model and the manifest keep naming the agent's own (WEFT-PLAYGROUND §10.1 [D5]).

type RunResult

type RunResult = core.RunResult

RunResult is the outcome of a completed run: its id, the full transcript (including the input messages), one record per step, and summed usage.

func GenerateAs added in v0.2.0

func GenerateAs[Out any](ctx context.Context, a *Agent, opts ...RunOption) (Out, *RunResult, error)

GenerateAs runs the agent and returns its structured output, decoded from the last valid submit_output call. The agent must have been built with Output[Out]; a run that ends without a valid submission returns ErrNoOutput alongside the result, so the transcript is still inspectable. Run errors are *RunError as for Generate.

Example

Structured output: Output constrains the final answer to a struct, and GenerateAs returns it decoded. An invalid submission is an ordinary tool error the model repairs; a valid one ends the run.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	type Verdict struct {
		Approved bool   `json:"approved"`
		Reason   string `json:"reason" jsonschema:"one sentence"`
	}
	model := wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "submit_output", Args: `{"approved":true,"reason":"within policy"}`}),
	)
	agt := weft.New(model, weft.Instructions("Review refund requests."), weft.Output[Verdict]())

	v, res, err := weft.GenerateAs[Verdict](context.Background(), agt, weft.Prompt("Refund order 42?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(v.Approved, v.Reason)
	fmt.Println("steps:", res.NumSteps())
}
Output:
true within policy
steps: 1

type RunStart

type RunStart = core.RunStart

RunStart is always the first event of a run and carries its id and, when reported, the model's identity and the agent's name.

InstructionsHash is the lowercase hex sha256 of the run's raw configured instructions — the agent's Instructions, or the run's override, before PrepareStep and before PromptSnippets are composed in (ADR 0028 §4). The loop always sets it: a run with no instructions carries the hash of the empty string. It is a hash, not content, so it survives StripContent; an event built elsewhere may leave it empty, and it is then absent on the wire.

type Schema

type Schema = core.Schema

Schema is the subset of JSON Schema (draft 2020-12) that weft derives from tool input structs. Its shape tracks what the official Go MCP SDK derives via google/jsonschema-go, so tools cross over to MCP without conversion.

func ParseSchema added in v0.2.0

func ParseSchema(b json.RawMessage) (*Schema, error)

ParseSchema reads a JSON Schema document from outside Go into a *Schema: the structured fields it knows are populated (Type, Properties, Required, … — the manifest and any reader that walks the tree see them), and the document's own bytes are kept and re-emitted by Schema.MarshalJSON, so an enum or a oneOf the Schema type cannot express still reaches the model exactly as written. The top-level type must be an object — providers and MCP both require it, and failing here, at import, beats failing at the first model call.

The structured view is lenient: a keyword whose shape the Schema type cannot hold — a boolean additionalProperties, a type array such as ["string","null"], tuple or boolean items, a non-string description — leaves that field zero (an unconstrained node) and is not an error, because the bytes carry it whole and the view is for readers, not the model. Only the document itself is checked: invalid JSON, trailing data, or a non-object top level return an error naming the problem.

type SteerFunc added in v0.4.0

type SteerFunc = core.SteerFunc

SteerFunc returns the messages to deliver at a safe point, or nil. It must not block: drain a queue, do not wait on one — the loop calls it between steps, so a source that waits stalls the run. A source with nothing to deliver returns nil. A panic in a SteerFunc fails the run like a panic in a PrepareStep function (the deferred recover ends the run span and re-panics): it is arbitrary user code the caller owns.

The returned messages become ordinary transcript messages — the transcript, the record, and the next turn all see them — and are reported as a Steered event. Ownership passes to the run, like the messages given to Messages: the source must not reuse or mutate them.

type SteerPoint added in v0.4.0

type SteerPoint = core.SteerPoint

SteerPoint tells a steering source where the run is. RunID is the run being steered; Step is the step that just finished; Final is true when the run would otherwise end here, so a non-empty return redirects the run into one more step instead of ending it.

type Steered added in v0.4.0

type Steered = core.Steered

Steered reports messages the run's steering source delivered at the step's drain point — after the tool batch, or at a final step whose steer redirected the run into one more step. It sits between that step's StepFinish and the next StepStart, numbered from the run's Seq counter like every Seq-carrying event. Messages are the delivered values (RoleUser) as a snapshot: they do not alias the run's transcript (ADR 0019).

type StepFinish

type StepFinish = core.StepFinish

StepFinish reports that step Index is complete: the model call finished and, when the step requested tools, they have run — it follows the step's ToolFinish events and precedes the stop-condition check (docs/life-of-a-call.md). Raw is the provider's own stop reason when Reason was approximated (see ModelFinish.Raw); empty when the mapping was exact.

LatencyMS and TTFTMS are the step's model call timed by the loop as it consumes the stream (ADR 0016's 2026-10-07 A4 note), in whole milliseconds rounded up, so a measured interval is never 0 and 0 (absent on the wire) means not measured — an event recorded before the fields existed, or one built by hand. LatencyMS runs from the call's start (the chain's Stream) to the stream's end (the ModelFinish); TTFTMS from the same start to the first TextDelta or ToolArgsDelta — reasoning does not count — and is 0 when the call yielded neither (a non-streaming adapter, a script's bare tool call). The timing spans the whole model chain: a retry's backoff and a fallback's failed tries are inside it. Both are observations, not behaviour: nothing in the loop reads them.

type StepRecord

type StepRecord = core.StepRecord

StepRecord captures everything one model step produced: its text, the tool calls it requested, and the results of executing them in call order.

type StepStart

type StepStart = core.StepStart

StepStart reports that the model is being called for step Index.

type StopCondition

type StopCondition = core.StopCondition

StopCondition decides, after a step's tool calls have run, whether the run is complete. It sees every step so far; the last element is the step just finished. Returning true ends the run successfully without another model call. The built-ins — HasToolCall, StepCountIs — also implement fmt.Stringer so the manifest (TODO §2.9) can name them; adapt an ordinary function with StopFunc.

func HasToolCall

func HasToolCall(names ...string) StopCondition

HasToolCall stops the run once the step just finished called any of the named tools — the "final answer tool" pattern.

func StepCountIs

func StepCountIs(n int) StopCondition

StepCountIs stops the run after exactly n steps, successfully — unlike MaxSteps, which treats reaching the budget as a failure.

type StopFunc

type StopFunc = core.StopFunc

StopFunc adapts an ordinary function to a StopCondition, the http.HandlerFunc shape.

type StopReason

type StopReason = core.StopReason

StopReason is why a model step ended.

type TextDelta

type TextDelta = core.TextDelta

TextDelta is an increment of assistant text.

type TextPart

type TextPart = core.TextPart

TextPart is a span of user or assistant text.

type ThinkingConfig added in v0.2.0

type ThinkingConfig = core.ThinkingConfig

ThinkingConfig is the per-run reasoning request: a level on the neutral scale, plus a token budget for providers whose depth control is a cap (Anthropic budget_tokens, Gemini thinkingBudget). A level without a Budget leaves the depth to the provider (Anthropic adaptive thinking, Gemini's own level mapping); a Budget without a level pins it. Off wins over a Budget when both are set.

type ThinkingLevel added in v0.2.0

type ThinkingLevel = core.ThinkingLevel

ThinkingLevel is a provider-neutral reasoning-effort scale. The zero value, ThinkUnset, sends nothing and keeps the provider default; every other value asks the adapter to express that depth on the wire in whatever form the provider has — reasoning_effort, a thinking object, a token budget. Adapters map what the provider can express and document what they drop (TODO §5.14).

type ThinkingOption added in v0.2.0

type ThinkingOption = core.ThinkingOption

ThinkingOption is accepted by both New and Stream/Generate: reasoning depth is a per-question concern, not a per-agent one. On an agent it is the default for every run; on a run it overrides that default.

func Thinking added in v0.2.0

func Thinking(cfg ThinkingConfig) ThinkingOption

Thinking sets the reasoning level for the agent's model calls. As an Option it is every run's default; as a RunOption it overrides that default for one run — the quick-ask shape: fast by default, think on demand, without rebuilding the agent.

agt := core.New(m, core.Thinking(core.ThinkingConfig{Level: core.ThinkOff}))
agt.Generate(ctx, core.Thinking(core.ThinkingConfig{Level: core.ThinkHigh}), core.Prompt(q))

The zero Level keeps the provider default; adapters map what the provider can express and document what they drop.

Example

Fast by default, think on demand: the agent option sets every run's default, a run option overrides it for that run alone. The scripted model records what each run asked for.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	model := wefttest.Script(wefttest.Say("ok"), wefttest.Say("ok"))
	agt := weft.New(model, weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkOff}))
	deep := weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkHigh, Budget: 2048})
	if _, err := agt.Generate(context.Background(), deep, weft.Prompt("hard")); err != nil {
		log.Fatal(err)
	}
	if _, err := agt.Generate(context.Background(), weft.Prompt("quick")); err != nil {
		log.Fatal(err)
	}
	for _, req := range model.Requests() {
		fmt.Println(req.Thinking == weft.ThinkingConfig{Level: weft.ThinkHigh, Budget: 2048},
			req.Thinking.Level == weft.ThinkOff)
	}
}
Output:
true false
false true

type ToolArgsDelta added in v0.2.0

type ToolArgsDelta = core.ToolArgsDelta

ToolArgsDelta reports an increment of a tool call's arguments as the model streams them — the model is "writing" the call, which can take a while for large arguments (generated code, long documents). It is progress only: the call has not been made, and ToolStart still arrives when it executes. Name is the best-known name so far; a provider that streams fragments of several calls interleaves their deltas, distinguished by name where the provider supplies one.

type ToolCallPart

type ToolCallPart = core.ToolCallPart

ToolCallPart is a tool invocation requested by the model. Args is the raw JSON the model produced; the loop unmarshals it into the tool's input type before invoking the handler. Signature is the provider's opaque token attached to the call itself (Gemini's thought signatures ride functionCall parts and must return on the same part); empty for providers without one.

type ToolCaller added in v0.2.0

type ToolCaller = core.ToolCaller

ToolCaller is one link of the tool-call chain: it takes a call and returns the result text the model will see, or an error. The innermost caller decodes the arguments and runs the handler; every ToolMiddleware wraps one.

type ToolChoiceConfig added in v0.3.0

type ToolChoiceConfig = core.ToolChoiceConfig

ToolChoiceConfig constrains what a step's model call may emit. The zero value is the provider default. Name is required when Mode is ToolChoiceNamed and must be empty under every other mode; the loop fails the run on a mismatch rather than sending a malformed choice (a programming error, not a sentinel condition).

type ToolChoiceMode added in v0.3.0

type ToolChoiceMode = core.ToolChoiceMode

ToolChoiceMode selects how the provider must shape a step's tool calls. The zero value, ToolChoiceAuto, keeps the provider default and sends nothing; every other value asks the adapter to express the constraint in the provider's own tool_choice form.

type ToolChoiceOption added in v0.3.0

type ToolChoiceOption = core.ToolChoiceOption

ToolChoiceOption is accepted by both New and Stream/Generate: which tool calls a step must make is a per-question concern as much as a per-agent one (the ThinkingOption shape).

func ToolChoice added in v0.3.0

func ToolChoice(cfg ToolChoiceConfig) ToolChoiceOption

ToolChoice forces the agent's model calls to include (or forbear from) tool calls. As an Option it is every run's default; as a RunOption it overrides that default for one run:

// a router: the first step must call classify
agt := core.New(m, classify, core.ToolChoice(core.ToolChoiceConfig{Mode: core.ToolChoiceNamed, Name: "classify"}))
// a final step that must answer in text, cache prefix intact
agt.Generate(ctx, core.ToolChoice(core.ToolChoiceConfig{Mode: core.ToolChoiceNone}), core.Prompt(q))

Modes: Any (some tool must be called), Named (Name must be — the router and eval-harness shape), None (no call may be made; the catalogue stays advertised, so a prompt-cache prefix on the tool definitions survives), Auto (the zero value: provider default, nothing sent). A PrepareStep function can rewrite the request's ToolChoice per step — force classify on step 0, then auto — the same way it rewrites tools. A non-auto choice the step's tool snapshot cannot satisfy (no tools, a name not advertised, a name under another mode) fails the run with a descriptive error.

Example

ToolChoice forces a step's tool calls — the router shape: classify must be the first call, the rest of the run is unconstrained. One PrepareStep function rewrites the request's ToolChoice per step; the agent-level option would force every step instead.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	classify := weft.Tool("classify", "Classify the request.",
		func(_ context.Context, _ struct{}) (string, error) { return "billing", nil })
	agt := weft.New(
		wefttest.Script(
			wefttest.ToolCalls(wefttest.Call{Name: "classify"}),
			wefttest.Say("This is a billing question."),
		),
		weft.PrepareStep(func(_ context.Context, step int, req weft.ModelRequest) (weft.ModelRequest, error) {
			if step == 0 {
				req.ToolChoice = weft.ToolChoiceConfig{Mode: weft.ToolChoiceNamed, Name: "classify"}
			}
			return req, nil
		}),
		classify,
	)
	res, err := agt.Generate(context.Background(), weft.Prompt("Why did my invoice double?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
}
Output:
This is a billing question.

type ToolDef

type ToolDef = core.ToolDef

ToolDef is a named, schema-described tool the model can call. Build one with the generic Tool constructor; the agent loop (and any manual dispatcher) executes it through Invoke.

A ToolDef is mutable until it is registered with New and frozen thereafter: New keeps a deep copy, so mutating the value that was passed in — or a copy returned by Agent.Tools — never reaches the agent or a running run.

func RawTool

func RawTool(name, description string, schema *Schema, fn func(ctx context.Context, args json.RawMessage) (string, error), opts ...ToolOption) *ToolDef

RawTool defines a tool from an explicit schema instead of reflection — for tools defined outside Go source: plugin manifests, MCP remotes, gateways. fn receives the model's raw JSON arguments verbatim: no unmarshalling, no ErrInvalidToolInput — validation belongs to fn, and StrictInput has no effect. A nil schema becomes the empty object schema, so the tool accepts any object input. Panics match Tool: empty name, nil fn. The manifest records no source line (the definition is not in Go source). Trailing options set per-tool policy as for Tool.

func Subagent added in v0.2.0

func Subagent(name, description string, child *Agent, opts ...ToolOption) *ToolDef

Subagent defines a tool that delegates to another agent. The model calls it with one argument, prompt; the child runs on a fresh transcript holding only that prompt, on the parent call's context, and the tool's result is the child's final text (or, for a child built with Output, the submitted JSON). The child's events arrive in the parent's stream wrapped in Nested; its usage is added to the parent's RunResult.Usage and recorded on StepRecord.SubagentUsage.

A subagent is an ordinary tool: Timeout bounds the child run, MaxResultBytes caps its answer, RequireApproval gates the delegation, Sequential makes it a barrier, WrapTools wraps the delegation once (the child's own seams govern inside it), and the manifest lists it.

A child run that fails is a tool error the parent model sees (SUBAGENT_FAILED), never a parent run error; a child that ends awaiting approval is SUBAGENT_PENDING; a child that is already running above this call is refused with SUBAGENT_CYCLE. Subagent panics if child is nil.

Example

Delegating to another agent: a subagent is a tool whose handler runs another agent on the prompt alone. The child's events arrive wrapped in Nested; its usage rolls into the parent's total.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	researcher := weft.New(wefttest.Script(
		wefttest.Say("order 1234 shipped yesterday"),
	))
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "research", Args: `{"prompt":"where is order 1234?"}`}),
		wefttest.Say("Researched."),
	), weft.Subagent("research", "Research a question in depth.", researcher))
	res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 1234?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
	fmt.Println("total tokens:", res.Usage.Total())
}
Output:
order 1234 shipped yesterday
total tokens: 45

func Tool

func Tool[In, Out any](name, description string, fn func(ctx context.Context, in In) (Out, error), opts ...ToolOption) *ToolDef

Tool defines a tool from a plain function. In and Out are inferred from the handler and the input schema is reflected from In's struct tags, so the compiler checks the handler's shape and nothing is written twice:

type RefundInput struct {
    OrderID string `json:"order_id" jsonschema:"the order to refund"`
    Reason  string `json:"reason,omitempty"`
}

core.Tool("refund_order", "Refund a customer's order",
    func(ctx context.Context, in RefundInput) (Receipt, error) {
        return billing.Refund(ctx, in.OrderID, in.Reason)
    },
    core.Timeout(10*time.Second))

What the model sees as the result is the text of a string Out, and the JSON encoding of any other Out. Inside the handler, CallFromContext reports which call is running. Trailing options set per-tool policy: Timeout, MaxResultBytes, StrictInput.

The handler shape mirrors the official Go MCP SDK's AddTool[In, Out], so a weft tool can be exposed over MCP without an adapter layer.

Tool panics if name is empty, fn is nil, or In is not a struct (or a pointer to one): providers and MCP require an object at the top level of a tool schema, and a scalar there would fail every real call.

Example

A tool is a plain function; the input schema is derived from the struct.

package main

import (
	"context"
	"encoding/json"
	"fmt"
	"log"

	"github.com/weftgo/weft"
)

func main() {
	type WeatherInput struct {
		City string `json:"city" jsonschema:"the city to look up"`
		Days *int   `json:"days,omitempty"`
	}
	getWeather := weft.Tool("get_weather", "Get a forecast.",
		func(_ context.Context, in WeatherInput) (string, error) {
			return "sunny in " + in.City, nil
		})

	b, err := json.MarshalIndent(getWeather.InputSchema, "", "  ")
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(getWeather.Name)
	fmt.Println(string(b))
}
Output:
get_weather
{
  "type": "object",
  "properties": {
    "city": {
      "type": "string",
      "description": "the city to look up"
    },
    "days": {
      "type": "integer"
    }
  },
  "required": [
    "city"
  ]
}

type ToolError added in v0.2.0

type ToolError = core.ToolError

ToolError is a tool failure with a stable code the model can branch on. Code is SCREAMING_SNAKE by convention ("ORDER_NOT_FOUND"); Message is what the model reads; Err is the internal cause — available to tool middleware and audit logs through errors.As/Unwrap, and never shown to the model. A handler returning *ToolError produces the result "<CODE>: <Message>"; the loop renders its own failures with codes too: INVALID_INPUT (arguments that do not decode) and NO_SUCH_TOOL. Codes are not validated. Plain errors keep rendering as err.Error(); install mw.MapErrors to code them centrally.

Example

A *ToolError carries a code the model can branch on and a cause it never sees.

package main

import (
	"context"
	"errors"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	lookup := weft.Tool("lookup_order", "", func(_ context.Context, in struct {
		ID string `json:"id"`
	}) (string, error) {
		return "", &weft.ToolError{
			Code:    "ORDER_NOT_FOUND",
			Message: "order " + in.ID + " does not exist",
			Err:     errors.New("pg: no rows in result set"), // for logs and middleware only
		}
	})
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":"42"}`}),
		wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":42}`}),
		wefttest.Say("No such order."),
	), lookup)
	res, err := agt.Generate(context.Background(), weft.Prompt("order 42?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
	fmt.Println(res.Steps[1].Results[0].Content)
}
Output:
ORDER_NOT_FOUND: order 42 does not exist
INVALID_INPUT: tool "lookup_order": field "id": expected string, got number

func Errorf added in v0.2.0

func Errorf(code, format string, args ...any) *ToolError

Errorf builds a *ToolError with a formatted Message. A %w verb sets Err as fmt.Errorf would — with several %w verbs every cause stays reachable through errors.Is — so the cause is available to middleware while the model sees only the formatted text.

type ToolFinish

type ToolFinish = core.ToolFinish

ToolFinish reports that a tool invocation completed, successfully or not. Content is the tool's JSON output, or the failure text when IsError is set — the same value the model sees on the matching ToolResultPart, so a UI can render results as they land.

type ToolMiddleware added in v0.2.0

type ToolMiddleware = core.ToolMiddleware

ToolMiddleware wraps a ToolCaller, the chi shape: it may act before next (deny, decorate ctx, log), after it (map errors, audit), or instead of it. Panic containment stays outside the chain — a panic in middleware is still a tool error result, never a run error.

type ToolOption added in v0.2.0

type ToolOption = core.ToolOption

ToolOption configures one tool at definition time. The options that make sense at both levels — MaxResultBytes, Timeout, StrictInput — are PolicyOptions: on the agent they set the default for every tool, on a tool they override that default for the tool alone.

func Origin added in v0.10.0

func Origin(name string) ToolOption

Origin names where a tool came from, for the observability record only: it is the tools record's source (ADR 0028 §5), and changes nothing the model sees or the loop does. A tool is "local" by default; Subagent sets "subagent" and weft/mcp's Tools sets "mcp". Any other string is recorded verbatim; the last Origin wins.

func PromptSnippet added in v0.2.0

func PromptSnippet(text string) ToolOption

PromptSnippet attaches system-prompt lines to a tool. The loop appends the snippets of the tools it advertises to the agent's Instructions for every model call — one paragraph per tool, in registration order, separated by blank lines — so usage rules live with the tool instead of in a central prompt that drifts as the set grows. Empty snippets add nothing. The manifest records the snippet.

Example

PromptSnippet keeps a tool's usage rules next to the tool; the loop appends them to the instructions of every model call.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	search := weft.Tool("search", "Search the docs.", func(_ context.Context, _ struct{}) (string, error) { return "", nil },
		weft.PromptSnippet("Cite the search result you used."))
	model := wefttest.Script(wefttest.Say("ok"))
	if _, err := weft.New(model, weft.Instructions("You answer questions."), search).Generate(context.Background(), weft.Prompt("hi")); err != nil {
		log.Fatal(err)
	}
	fmt.Println(model.Requests()[0].System)
}
Output:
You answer questions.

Cite the search result you used.

func Replay added in v0.7.0

func Replay(p ReplayPolicy) ToolOption

Replay sets the tool's ReplayPolicy. Only ReplaySafe needs to be said: ReplayNever is the zero value and the default, and any other value counts as never rather than being trusted (an unknown class is never's, the safe default). Like every tool option the last one wins, so Replay(ReplayNever) after an earlier Replay(ReplaySafe) — shared defaults, then this tool's own word — is never's. The manifest records the class.

Example

Replay declares a tool's side-effect class for re-runs: safe vouches the call is idempotent (a re-run may execute it for real); every unannotated tool counts as never — substituted or parked, never silently re-fired (WEFT-PLAYGROUND.md §6 rule 3).

package main

import (
	"context"
	"fmt"

	"github.com/weftgo/weft"
)

func main() {
	lookup := weft.Tool("lookup_order", "Look up an order.", func(_ context.Context, _ struct{}) (string, error) {
		return "shipped", nil
	}, weft.Replay(weft.ReplaySafe))
	refund := weft.Tool("refund", "Refund an order.", func(_ context.Context, _ struct{}) (string, error) {
		return "refunded", nil
	})
	fmt.Println("lookup:", lookup.ReplayPolicy())
	fmt.Println("refund:", refund.ReplayPolicy())
}
Output:
lookup: safe
refund: never

func RequireApproval added in v0.2.0

func RequireApproval() ToolOption

RequireApproval marks a tool whose calls never run without a decision. When the model calls it the loop executes the step's other tools, then ends the run successfully with the call on RunResult.Pending (and RunFinish.Pending) and no result in the transcript. Resume with the transcript plus Approve or Deny:

res, _ := agt.Generate(ctx, core.Prompt("Refund order 42"))
for _, call := range res.Pending { /* ask someone */ }
res, _ = agt.Generate(ctx, core.Messages(res.Messages...), core.Approve(call.ID))

Approved calls run through the ordinary chain with Call.Approved set; denied ones (and pending calls given no decision) become error results the model sees. This is a policy and UX seam, not a security boundary: the boundary is the sandbox a tool runs in.

Example

A RequireApproval tool parks its calls: the run ends successfully with them on Pending, and a later run resumes with a decision.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	refund := weft.Tool("refund", "Refund an order.", func(_ context.Context, in struct {
		Order string `json:"order"`
	}) (string, error) {
		return "refunded " + in.Order, nil
	}, weft.RequireApproval())
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{ID: "c1", Name: "refund", Args: `{"order":"42"}`}),
		wefttest.Say("Done."),
	), refund)

	res, err := agt.Generate(context.Background(), weft.Prompt("refund order 42"))
	if err != nil {
		log.Fatal(err)
	}
	for _, call := range res.Pending {
		fmt.Printf("awaiting approval: %s %s\n", call.Name, call.Args)
	}

	// Someone decided. Resume with the transcript and the decision.
	res, err = agt.Generate(context.Background(), weft.Messages(res.Messages...), weft.Approve("c1"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
}
Output:
awaiting approval: refund {"order":"42"}
Done.

func ToolOptions added in v0.2.0

func ToolOptions(opts ...ToolOption) ToolOption

ToolOptions composes several tool options into one, applied in order — the Tool-defining counterpart of Options, so a package of policy (Timeout, MaxResultBytes, StrictInput, a WrapTools chain) can be named and reused:

productPolicy := core.ToolOptions(core.Timeout(5*time.Second), core.StrictInput())
core.Tool("lookup", "…", fn, productPolicy)

Nil entries are ignored.

Example

ToolOptions composes tool options into one named value, so a package of per-tool policy travels under one name.

package main

import (
	"context"
	"fmt"
	"log"
	"time"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	productPolicy := weft.ToolOptions(
		weft.Timeout(5*time.Second),
		weft.MaxResultBytes(1024),
	)
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":"42"}`}),
		wefttest.Say("Done."),
	), weft.Tool("lookup_order", "Look up an order by id.",
		func(_ context.Context, in struct {
			ID string `json:"id" jsonschema:"the order id"`
		}) (string, error) {
			return "order " + in.ID + " shipped", nil
		}, productPolicy))
	res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 42?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
}
Output:
order 42 shipped

type ToolResultPart

type ToolResultPart = core.ToolResultPart

ToolResultPart is the outcome of one tool call, returned to the model as data. Content is the JSON encoding of the tool's output, or the failure message when IsError is set. Tool failures never abort a run; the model sees them and can recover.

type ToolStart

type ToolStart = core.ToolStart

ToolStart reports that a tool invocation began. Events from tools running in parallel interleave: pair them by CallID and order by Seq, a per-run counter assigned at emission that totally orders the stream. A call parked by the approval boundary has a ToolStart and no ToolFinish; it is listed on RunFinish.Pending instead.

type Usage

type Usage = core.Usage

Usage is token accounting for one step or one whole run.

The totals are inclusive: CachedInputTokens and CacheWriteTokens are subsets of InputTokens (tokens billed as input — read from or written to a provider prompt cache), ReasoningTokens is a subset of OutputTokens (provider-side reasoning the model burned). The splits are reporting, not budget bases: UsageLimit and Total keep reading the two totals, so a cached-heavy run budgets identically to an uncached one with the same totals. All three are omitempty on the wire — an event from before they existed round-trips unchanged (SchemaVersion stays 1, ADR 0001/0004).

Directories

Path Synopsis
Package anthropic is the weft adapter for the Anthropic Messages API, wrapping the official anthropic-sdk-go.
Package anthropic is the weft adapter for the Anthropic Messages API, wrapping the official anthropic-sdk-go.
example command
Command example runs a two-step agent conversation against the real Anthropic API.
Command example runs a two-step agent conversation against the real Anthropic API.
cmd
weft command
Command weft is the framework's one binary (plan B1): setup B's local Studio and a terminal over the Studio API, for scripts and CI.
Command weft is the framework's one binary (plan B1): setup B's local Studio and a terminal over the Studio API, for scripts and CI.
core module
examples
approval command
Command approval shows the approval boundary end to end, offline: a tool marked RequireApproval parks its call, the run ends with the call on Pending, a person decides, and a second run resumes with the transcript plus the decision.
Command approval shows the approval boundary end to end, offline: a tool marked RequireApproval parks its call, the run ends with the call on Pending, a person decides, and a second run resumes with the transcript plus the decision.
getting-started command
Command getting-started runs a complete weft agent offline: a scripted model (from wefttest), one tool, and the full event stream.
Command getting-started runs a complete weft agent offline: a scripted model (from wefttest), one tool, and the full event stream.
Package google is the weft adapter for the Gemini API via the official google.golang.org/genai SDK.
Package google is the weft adapter for the Gemini API via the official google.golang.org/genai SDK.
example command
Command example runs a two-step agent conversation against the real Gemini API.
Command example runs a two-step agent conversation against the real Gemini API.
internal
adapterkit
Package adapterkit holds the helpers every first-party adapter needs but no vendor SDK touches: schema rendering, the terminal-error rule (with the context-overflow mapping and the marker table mw's retry classifier reads), and the FilePart exactly-one guard.
Package adapterkit holds the helpers every first-party adapter needs but no vendor SDK touches: schema rendering, the terminal-error rule (with the context-overflow mapping and the marker table mw's retry classifier reads), and the FilePart exactly-one guard.
cmd/genfacade command
Command genfacade writes a facade package (internal/facadegen): the framework's github.com/weftgo/weft, weft/mw, weft/wefttest and weft/wefttest/conformance re-export the core module's packages of the same name.
Command genfacade writes a facade package (internal/facadegen): the framework's github.com/weftgo/weft, weft/mw, weft/wefttest and weft/wefttest/conformance re-export the core module's packages of the same name.
discovery
Package discovery is how an app finds the running Studio with no configuration (plan B3): `weft studio` and `weft dev` write a small file, studio.json, naming the Studio they serve, and weft/otel and weft/runtime read it when nothing else names one.
Package discovery is how an app finds the running Studio with no configuration (plan B3): `weft studio` and `weft dev` write a small file, studio.json, naming the Studio they serve, and weft/otel and weft/runtime read it when nothing else names one.
doctor
Package doctor is `weft doctor` (plan B5): it asks a running Studio what it is and how it is wired, and prints one line per check.
Package doctor is `weft doctor` (plan B5): it asks a running Studio what it is and how it is wired, and prints one line per check.
facadegen
Package facadegen writes a facade package: one that re-exports another package's whole exported API under a new import path, so github.com/weftgo/weft (the framework) offers the loop that github.com/weftgo/weft/core (the slim module) implements, under the same names.
Package facadegen writes a facade package: one that re-exports another package's whole exported API under a new import path, so github.com/weftgo/weft (the framework) offers the loop that github.com/weftgo/weft/core (the slim module) implements, under the same names.
listen
Package listen is the Studio command's port policy (plan B2): one default port, 7331, stable and reusable, never silently different.
Package listen is the Studio command's port policy (plan B2): one default port, 7331, stable and reusable, never silently different.
mcp
Package mcp is the bridge between weft and the Model Context Protocol, both directions over the official Go SDK (github.com/modelcontextprotocol/go-sdk, aliased `sdk` in examples — this package keeps the name `mcp`):
Package mcp is the bridge between weft and the Model Context Protocol, both directions over the official Go SDK (github.com/modelcontextprotocol/go-sdk, aliased `sdk` in examples — this package keeps the name `mcp`):
examples/client command
Command client consumes MCP servers as weft tools: it connects two in-memory servers concurrently under one deadline, imports their tools with per-server prefixes, registers them on a scripted agent, runs it, and prints the manifest — the imported tools' raw schemas show in it.
Command client consumes MCP servers as weft tools: it connects two in-memory servers concurrently under one deadline, imports their tools with per-server prefixes, registers them on a scripted agent, runs it, and prints the manifest — the imported tools' raw schemas show in it.
examples/server command
Command server exposes a weft agent as an MCP server over stdio — the getting-started agent plus its lookup tool, both directions of §7 in one process:
Command server exposes a weft agent as an MCP server over stdio — the getting-started agent plus its lookup tool, both directions of §7 in one process:
Package mw is the framework's import path for the reference middleware that github.com/weftgo/weft/core/mw implements: model middleware for weft.WrapModel (Retry, Fallback, Log, RepairJSON) and tool middleware for weft.WrapTools (Allow, Audit, MapErrors).
Package mw is the framework's import path for the reference middleware that github.com/weftgo/weft/core/mw implements: model middleware for weft.WrapModel (Retry, Fallback, Log, RepairJSON) and tool middleware for weft.WrapTools (Allow, Audit, MapErrors).
Package obsdb is the observability database: what weft/otel's local sink writes and Weft Studio reads, in one module so neither owns the schema (ADR 0024, S3).
Package obsdb is the observability database: what weft/otel's local sink writes and Weft Studio reads, in one module so neither owns the schema (ADR 0024, S3).
clickhouse
Package clickhouse — see doc.go for the package documentation.
Package clickhouse — see doc.go for the package documentation.
internal/reqread
Package reqread holds what obsdb's two backends share to answer DB.Prompt, DB.Tools and DB.Catalogs over the records they read: the lookups by hash and the one-per-hash catalog list.
Package reqread holds what obsdb's two backends share to answer DB.Prompt, DB.Tools and DB.Catalogs over the records they read: the lookups by hash and the one-per-hash catalog list.
obsdbtest
Package obsdbtest is the shared conformance table for obsdb.DB backends — the executable form of S3.5's promises, the storetest pattern.
Package obsdbtest is the shared conformance table for obsdb.DB backends — the executable form of S3.5's promises, the storetest pattern.
sqlite
Package sqlite is obsdb's default backend: the observability schema in a single SQLite file on the CGO-free modernc.org/sqlite driver (thread/sqlite's choice).
Package sqlite is obsdb's default backend: the observability schema in a single SQLite file on the CGO-free modernc.org/sqlite driver (thread/sqlite's choice).
Package openai is the weft adapter for the OpenAI Chat Completions API and OpenAI-compatible servers (gateways, local models), wrapping the official openai-go SDK.
Package openai is the weft adapter for the OpenAI Chat Completions API and OpenAI-compatible servers (gateways, local models), wrapping the official openai-go SDK.
example command
Command example runs a two-step agent conversation against the real OpenAI API (or any OPENAI_BASE_URL server).
Command example runs a two-step agent conversation against the real OpenAI API (or any OPENAI_BASE_URL server).
Package otel wires weft's observability to one or more OpenTelemetry destinations at once — the local sink, Weft Studio, Datadog, Langfuse, any OTLP endpoint or your own exporters — each with its own signals and content policy (ADR 0024, S2).
Package otel wires weft's observability to one or more OpenTelemetry destinations at once — the local sink, Weft Studio, Datadog, Langfuse, any OTLP endpoint or your own exporters — each with its own signals and content policy (ADR 0024, S2).
Package runtime is the playground's in-app side (WEFT-PLAYGROUND.md §10.2, [D6]): it registers the agents your code built with the Studio the app is already observed by, receives experiment commands over the runtime link, and executes them as real runs of those agents — your tools, your model keys, your process.
Package runtime is the playground's in-app side (WEFT-PLAYGROUND.md §10.2, [D6]): it registers the agents your code built with the Studio the app is already observed by, receives experiment commands over the runtime link, and executes them as real runs of those agents — your tools, your model keys, your process.
examples/local command
Command local is setup A's playground in one process: an embedded Studio on loopback, the app's own pipeline exporting into it, one agent with a scripted model and two tools, and the runtime link — WEFT-PLAYGROUND.md §7's P0 slice, drivable with curl.
Command local is setup A's playground in one process: an embedded Studio on loopback, the app's own pipeline exporting into it, one agent with a scripted model and two tools, and the runtime link — WEFT-PLAYGROUND.md §7's P0 slice, drivable with curl.
Package scope is the Go side of the devtools' Scope (plan §13.3): the unit every discovery rung, deep link and API call carries — one conversation's public id and, optionally, the session, flow and run inside it — and the net/http middleware that hands it to the page.
Package scope is the Go side of the devtools' Scope (plan §13.3): the unit every discovery rung, deep link and API call carries — one conversation's public id and, optionally, the session, flow and run inside it — and the net/http middleware that hands it to the page.
store module
Package studio is the Inspector: the UI, the JSON API, the live stream and the OTLP receiver over one observability database (weft/obsdb), served by Go alone.
Package studio is the Inspector: the UI, the JSON API, the live stream and the OTLP receiver over one observability database (weft/obsdb), served by Go alone.
examples/basic command
Command basic records demo runs — a tool call, a subagent, a failure — into an obsdb sqlite database and serves Studio on 127.0.0.1:7331:
Command basic records demo runs — a tool call, a subagent, a failure — into an obsdb sqlite database and serves Studio on 127.0.0.1:7331:
ingest
Package ingest is Studio's OTLP/HTTP receiver (S4.4): POST /v1/traces and /v1/logs, protobuf and JSON, optional gzip, a 16 MiB limit on the decompressed body, and the publish-then-write pipeline —
Package ingest is Studio's OTLP/HTTP receiver (S4.4): POST /v1/traces and /v1/logs, protobuf and JSON, optional gzip, a 16 MiB limit on the decompressed body, and the publish-then-write pipeline —
runtime
Package runtime is the runtime link's server side (WEFT-PLAYGROUND §10.3, S4.2): the registry of connected runtimes and the three routes a weft/runtime client speaks to — register, the SSE command stream, acks.
Package runtime is the runtime link's server side (WEFT-PLAYGROUND §10.3, S4.2): the registry of connected runtimes and the three routes a weft/runtime client speaks to — register, the SSE command stream, acks.
cmd module
Package thread gives weft sessions: a conversation as an append-only tree of entries, durable through a Storage backend, with turns, branching, compaction, approvals and delegation all built on that one tree.
Package thread gives weft sessions: a conversation as an append-only tree of entries, durable through a Storage backend, with turns, branching, compaction, approvals and delegation all built on that one tree.
backend
Package backend is for authors of thread.Storage backends: it resolves the open options an application passes — thread.Salvage, thread.FsyncOnFlush, thread.NoLock, thread.OpenLogger — into the configuration a backend acts on.
Package backend is for authors of thread.Storage backends: it resolves the open options an application passes — thread.Salvage, thread.FsyncOnFlush, thread.NoLock, thread.OpenLogger — into the configuration a backend acts on.
examples/approvals command
Command approvals walks one weft/thread session through the approval flow (ADR 0021): a gated call parks, the process "restarts" — the session is reopened from the JSONL file — a decision arrives signed over the challenge the session minted, and the conversation resumes under it.
Command approvals walks one weft/thread session through the approval flow (ADR 0021): a gated call parks, the process "restarts" — the session is reopened from the JSONL file — a decision arrives signed over the challenge the session minted, and the conversation resumes under it.
examples/refund-plan command
Command refund-plan demonstrates the orchestration ledger on a weft/thread session: a refund-support chat whose policy state lives in the session file as custom entries, so a process that dies mid-flow is replaced by one that resumes exactly where the file says — the last refund_plan entry.
Command refund-plan demonstrates the orchestration ledger on a weft/thread session: a refund-support chat whose policy state lives in the session file as custom entries, so a process that dies mid-flow is replaced by one that resumes exactly where the file says — the last refund_plan entry.
examples/session command
Command session walks one weft/thread session through its whole life: two turns, a label, a branch off the first answer, a fork of the branch, a manual compaction with preview, a close, and a reopen from disk — everything on a JSONL backend, everything offline through a scripted model, every id deterministic.
Command session walks one weft/thread session through its whole life: two turns, a label, a branch off the first answer, a fork of the branch, a manual compaction with preview, a close, and a reopen from disk — everything on a JSONL backend, everything offline through a scripted model, every id deterministic.
internal/carry
Package carry builds the context a run continues on when it runs on behalf of another one — a resume over a parked boundary, a deferred steer's follow-up, an async pool child: cancellation and deadline from one context, and, for the keys that context does not carry, the values of the context the work began on.
Package carry builds the context a run continues on when it runs on behalf of another one — a resume over a parked boundary, a deferred steer's follow-up, an async pool child: cancellation and deadline from one context, and, for the keys that context does not carry, the values of the context the work began on.
internal/opencfg
Package opencfg holds the resolved form of thread's open options — the one piece of the option vocabulary that thread (which declares the options), its in-tree backends (Memory) and thread/backend (which publishes the resolution to backend authors) all need, kept here so that none of them has to import another to share it.
Package opencfg holds the resolved form of thread's open options — the one piece of the option vocabulary that thread (which declares the options), its in-tree backends (Memory) and thread/backend (which publishes the resolution to backend authors) all need, kept here so that none of them has to import another to share it.
internal/rules
Package rules holds the small rules every thread.Storage backend must answer identically: the List limit, the metadata and title filters, the paging cursor, and the line discipline of a session's bytes.
Package rules holds the small rules every thread.Storage backend must answer identically: the List limit, the metadata and title filters, the paging cursor, and the line discipline of a session's bytes.
jsonl
Package jsonl is the default durable thread.Storage: one directory, one <id>.jsonl file per session — the header line first, then one entry per line in append order (ADR 0011).
Package jsonl is the default durable thread.Storage: one directory, one <id>.jsonl file per session — the header line first, then one entry per line in append order (ADR 0011).
pool
Package pool provides bounded concurrent child runs for thread sessions (ADR 0022): a FIFO semaphore per Pool value, subagent tools whose children run as sessions of their own, linked to the parent session and call, and receipts recording every delegation's journey.
Package pool provides bounded concurrent child runs for thread sessions (ADR 0022): a FIFO semaphore per Pool value, subagent tools whose children run as sessions of their own, linked to the parent session and call, and receipts recording every delegation's journey.
sqlite
Package sqlite is the thread's second durable backend: every session in one SQLite file on the CGO-free modernc.org/sqlite driver — the store's choice (store/sqlite, Crush's before it), reused so one dependency serves both modules — with WAL and embedded migrations.
Package sqlite is the thread's second durable backend: every session in one SQLite file on the CGO-free modernc.org/sqlite driver — the store's choice (store/sqlite, Crush's before it), reused so one dependency serves both modules — with WAL and embedded migrations.
threadtest
Package threadtest is the shared conformance table for thread.Storage backends — the executable form of ADR 0011 §5's promises.
Package threadtest is the shared conformance table for thread.Storage backends — the executable form of ADR 0011 §5's promises.
Package version is weft's one version string: the framework module's release tag, stamped into Studio's api/meta, the devtools panel bundle, the otel resource, the runtime's registration and the binaries' version output.
Package version is weft's one version string: the framework module's release tag, stamped into Studio's api/meta, the devtools panel bundle, the otel resource, the runtime's registration and the binaries' version output.
Package wefttest is the framework's import path for the offline test double that github.com/weftgo/weft/core/wefttest implements: a scripted, deterministic Model (Script, Say, ToolCalls, Fail, …), record-and-replay against a real provider, and helpers for events and golden files.
Package wefttest is the framework's import path for the offline test double that github.com/weftgo/weft/core/wefttest implements: a scripted, deterministic Model (Script, Say, ToolCalls, Fail, …), record-and-replay against a real provider, and helpers for events and golden files.
conformance
Package conformance is the framework's import path for the adapter conformance suite that github.com/weftgo/weft/core/wefttest/conformance implements: Run drives a Model through the behaviours every provider adapter must share.
Package conformance is the framework's import path for the adapter conformance suite that github.com/weftgo/weft/core/wefttest/conformance implements: Run drives a Model through the behaviours every provider adapter must share.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL