weft

package module
v0.3.4 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 26, 2026 License: MIT Imports: 29 Imported by: 0

README

weft

CI Go Reference Version

A thin, opinionated core for building AI agents in Go — designed the way the standard library is: small interfaces, context everywhere, functional options, wrapped errors, and zero required configuration.

Weft runs the agent loop — call a model, execute its tool calls in parallel with defined failure semantics, stream typed events while work is in flight — and nothing else. Concurrency is the point, not a feature: a step's tools fan out over goroutines, parallelism is a one-line dial, and tool failures never cancel their siblings.

Status: v0.3.0 — experimental, pre-1.0. The three load-bearing contracts — message model, error model, tool contract — are implemented and tested; the provider adapters (OpenAI + compatible servers, Anthropic, Google) wrap the vendors' official Go SDKs; the two middleware seams (WrapModel/WrapTools, package mw), the approval boundary (now with externally-computed results: Resolve), subagents as tools, MCP interop both ways (weft/mcp), observability (OTel spans, slog lines), wefttest record/replay, and the parity-round controls — tool-choice forcing, per-step sampling params, richer Usage splits, anthropic prompt caching, and the streaming OutputDecoder — are in; see docs/adr/ and the roadmap below. The surrounding modules (serving, ops, devtools, cli) come next, in that order of demand.

Quick start

A tool is a plain function — the JSON Schema is reflected from the input struct, so the struct is both the contract and the documentation:

type EchoInput struct {
    Msg string `json:"msg" jsonschema:"the message to echo"`
}

echo := weft.Tool("echo", "Echo a message back, uppercased.",
    func(ctx context.Context, in EchoInput) (string, error) {
        return strings.ToUpper(in.Msg), nil
    })

An agent is a value: build it once, run it many times, concurrently. Offline, the model comes from wefttest: a scripted, deterministic stand-in for a real provider (an adapter slots into the same Model seam):

model := wefttest.Script(
    wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"msg":"hello"}`}),
    wefttest.Say("HELLO"),
)
agt := weft.New(
    model,                                   // or anthropic.Model("claude-sonnet-5") — see Providers
    weft.Instructions("You are a support agent."),
    echo,
)

res, err := agt.Generate(ctx, weft.Prompt("Echo hello."))
fmt.Println(res.Text(), res.Usage.Total())

For a one-off tool the input struct can be written inline in the handler signature; a named type reads better and is reusable across tools and tests.

Streaming is Go iteration — typed events, cancel via context, exactly one terminal error:

for ev, err := range agt.Stream(ctx, weft.Prompt("Echo hello.")).Events() {
    if err != nil {
        return err
    }
    switch ev := ev.(type) {
    case weft.TextDelta:
        io.WriteString(w, ev.Text)
    case weft.ToolArgsDelta: // progress: the model is still writing the call
    case weft.ToolStart:
        slog.Info("tool", "name", ev.Name, "seq", ev.Seq)
    case weft.RunFinish:
        slog.Info("done", "steps", ev.Steps, "tokens", ev.Usage.Total())
    }
}

Ending a run is a predicate, and the step budget is a separate safety net:

agt := weft.New(model,
    weft.StopWhen(weft.HasToolCall("submit_answer")), // intended end
    weft.MaxSteps(20),                                // runaway guard → ErrMaxSteps
    submit, search,
)
Structured output

Output[T] constrains the final answer to a struct: a submit_output tool with T's reflected schema is advertised, and the run ends when the model calls it with arguments that decode. An invalid submission is an ordinary tool error the model repairs. GenerateAs returns the value; OutputOf reads it from a streamed run's result.

type Verdict struct {
    Approved bool   `json:"approved"`
    Reason   string `json:"reason" jsonschema:"one sentence"`
}

agt := weft.New(model, weft.Output[Verdict](), lookup)
v, res, err := weft.GenerateAs[Verdict](ctx, agt, weft.Prompt("Review order 42."))

It is tool mode, so it works on every provider; a run that ends in text returns ErrNoOutput with the transcript attached.

Per-tool policy

Trailing options on Tool set policy for that tool alone. The same names on New set the agent-wide default:

run := weft.Tool("run_command", "Run a shell command.", runCommand,
    weft.Timeout(30*time.Second),   // deadline on ctx; expiry is an error result
    weft.MaxResultBytes(0),         // this tool's output arrives whole
    weft.StrictInput(),             // undeclared argument fields are rejected
)
agt := weft.New(model, weft.Timeout(10*time.Second), run, grep)

Bad arguments come back to the model in the schema's own words — field "days": expected integer, got string — so it can map the error to the schema it was shown. Undeclared fields are ignored by default; StrictInput rejects them by name.

The two seams

Behaviour attaches at two chi-style seams; observation at weft.Tap. WrapModel wraps the model, WrapTools wraps every tool call (on the agent, or on one tool). First listed is outermost. Package mw holds the reference set:

agt := weft.New(model,
    weft.WrapModel(
        mw.Log(logger),               // request summary + finish, Debug level
        mw.Fallback(backupModel),     // fail over once Retry has given up on the primary
        mw.Retry(mw.MaxRetries(3)),   // 429/5xx/net errors; retry-after honoured; backoff with jitter
        mw.RepairJSON(),              // close a truncated tool-call argument object once
    ),
    weft.WrapTools(
        mw.Audit(logger),             // every call: run, step, tool, duration, error + cause
        mw.Allow(policy.Permits),     // DENIED: tool "rm" is not allowed — the model sees it
        mw.MapErrors(nil),            // plain errors → INTERNAL: tool "x" failed (cause kept)
    ),
    tools...,
)

Middleware that verifies something puts it on ctx before next, and the handler reads it back through a typed accessor — the same shape as weft.CallFromContext (see ExampleWrapTools_context). A middleware panic is a tool error result, never a run error. Handlers give the model a code to branch on with *weft.ToolError (ORDER_NOT_FOUND: order 42 does not exist; the cause stays in logs); the loop's own failures are coded INVALID_INPUT and NO_SUCH_TOOL. docs/life-of-a-call.md shows where each thing sits; ADR 0006 is the decision.

Seams are the product

The governance features other frameworks ship as processors — PII scrubbing, prompt-injection heuristics, moderation, token limits, response caching — need no processor layer here; each is a closure at one of the two seams, and each carries its own dependencies and policy stances rather than importing yours. mw stays a reference set; the patterns live as tested examples:

  • PII scrub — a tool middleware that masks what results carry (Example_piiScrubMiddleware).
  • Allowlist — shipped as mw.Allow(permits).
  • Token limiter — a model middleware that refuses the call before the provider bills it (Example_tokenLimitMiddleware); weft.UsageLimit covers the measured side.
  • Response cache — a model middleware keyed on the ModelRequest (Example_responseCacheMiddleware); invalidation policy is the caller's.

A named mw package ships only when a pattern needs a dependency or a policy stance weft should own — the mw.RateLimit precedent (rate limiting stays an example because golang.org/x/time would be the module's first dependency). Revisit when a consumer asks for one by name.

Delegating to another agent

A subagent is a tool whose handler runs another agent — the orchestrator-worker pattern with zero new machinery. The child sees only the prompt; its events arrive in the parent's stream wrapped in weft.Nested (Seq from the parent's counter, so the stream stays replayable); its usage rolls into res.Usage and is recorded per call on StepRecord.SubagentUsage:

researcher := weft.New(model, weft.Tool("deep_search", "…", DeepSearch))
orchestrator := weft.New(model,
    weft.Subagent("research", "Research a question in depth.", researcher,
        weft.Timeout(2*time.Minute)),
)

Everything the tool contract offers applies: Timeout bounds the child run, Parallelism bounds concurrent delegations (four research calls in one step run four children), RequireApproval gates the delegation itself. A failed child is data the parent model sees (SUBAGENT_FAILED: …, the child's *RunError on ToolError.Err), a child that ends awaiting approval is SUBAGENT_PENDING — approval-gated tools belong in the orchestrator, not in a child — and a delegation to an agent already running in the call chain is refused (SUBAGENT_CYCLE). ADR 0014 records the mechanics, the lineage ids, and pi's AgentLanes counterpoint.

Composing agents

A plugin is func(deps) weft.Option — a family of tools and its policy closed over its dependencies as one value, composed with weft.Options (and weft.ToolOptions for the per-tool counterpart). Nothing registers itself, so there is no registry, no scopes, no dedup; dependencies are parameters, never globals, and the "registered twice" mistake panics at New as always:

func Orders(svc *OrderService) weft.Option {
    return weft.Options(
        weft.Instructions("You handle orders."),
        svc.Lookup(), svc.Refund(),
        weft.WrapTools(mw.Allow(svc.Permitted)),
    )
}
agt := weft.New(model, base, Orders(orders))
Approval

weft.RequireApproval() on a tool parks its calls: the run ends successfully with them on RunResult.Pending, the step's other tools having run. Resume with the transcript and a decision — over any transport, with no persistence required:

res, _ := agt.Generate(ctx, weft.Prompt("Refund order 42"))
for _, call := range res.Pending { /* ask someone */ }
res, _ = agt.Generate(ctx, weft.Messages(res.Messages...),
    weft.Approve(call.ID), weft.Deny(other.ID, "over the limit"))

Approved calls run (handlers see Call.Approved); denied ones become DENIED: <reason> results the model sees; undecided ones DENIED: no decision. Middleware can park any call by returning an error wrapping ErrApprovalRequired. It is a policy seam, not a security boundary (ADR 0007; examples/approval).

Observability

Set up an OpenTelemetry SDK and every run emits the full span tree — one invoke_agent span per run, one chat span per model call, one execute_tool span per executed tool call, a subagent's run nested under its delegating tool span — with the GenAI semantic attributes (provider, model, tokens, finish reasons). No weft option is needed: the core instruments through the OTel API, which is a no-op until your SDK registers. Exporters and backends are not weft's business.

tp := sdktrace.NewTracerProvider( /* your exporter */ )
defer tp.Shutdown(ctx)
// no weft option: the global provider is picked up, spans appear
agt := weft.New(model, tools...)
// or explicitly, without touching the global:
agt = weft.New(model, append(tools, weft.TracerProvider(tp))...)

For logs, weft.Logger(l) writes one Debug line per phase — run start, run finish, model call, tool call — with ids, the model, durations, usage, and outcomes; never message text or tool arguments. The default (slog.Default, resolved at log time) is silent until your handler enables Debug; slog.New(slog.DiscardHandler) turns the lines off. A handler that bridges slog to OTel correlates the lines with the spans for free: they are logged on the span-carrying context. examples/otel runs the whole thing against the real SDK offline. No prompt, message, tool argument or tool result reaches a span: ids, names, counts, durations, reasons, and error types only. Error text does travel — a failed run or model call records its error as the span's exception, and the log lines carry the error text a tool or model returned, because a log is the caller's (ADR 0016).

The manifest — weft.json

One generated, committed, diffable description of every agent and tool (the code stays the only source of truth; the file is output, never input). Gate it with a golden test so it cannot go stale:

func TestManifest(t *testing.T) {
    b, err := weft.Manifest(newAgent())
    if err != nil {
        t.Fatal(err)
    }
    wefttest.Golden(t, "weft.json", b) // regenerate: go test ./... -update
}

A tool or policy change without regenerating fails go test; the diff is the review artifact. (ADR 0012)

MCP: both ways

weft/mcp (its own module over the official Go MCP SDK, aliased sdk) is the bridge in both directions, with no adapter layer — the tool contract is the same shape (ADR 0015):

import (
    sdk "github.com/modelcontextprotocol/go-sdk/mcp"
    "github.com/weftgo/weft/mcp"
)

// Consume: a server's tools as ordinary weft tools. The schema bytes
// cross whole (an enum or oneOf reaches the model as sent), the
// calls forward the model's arguments verbatim, and every remote
// failure is a tool result the model sees — data, never a run error.
tools, err := mcp.Tools(ctx, sess, mcp.Prefix("gh_"), mcp.Policy(weft.Timeout(10*time.Second)))
// A tool the bridge cannot import fails that tool, not the listing:
// the good ones are in tools, the skipped ones are named. Warning or
// stop is your call; a listing failure (transport, ctx) is a plain err.
var skipped *mcp.ImportError
if errors.As(err, &skipped) {
    slog.Warn("mcp: tools skipped", "err", skipped)
    err = nil
}
if err != nil {
    return err
}

// Expose: weft tools — or a whole agent, as one named tool with your
// description — on any MCP server.
srv := sdk.NewServer(&sdk.Implementation{Name: "weft", Version: "0"}, nil)
mcp.AddTools(srv, lookup)
mcp.Serve(srv, agent, "Support agent.")   // agent + its tools, under its chain

Two warnings the godoc repeats. A server's tool descriptions are untrusted content — they land in your model's tool list, a surface you did not author; filter with mw.Allow or read Tools()' output before registering it. A foreign tool runs sequentially unless its server marks it readOnlyHint — the conservative reading of an untrusted hint for a tool whose handler you cannot read.

The loop, the seams, the manifest and wefttest treat an imported tool like any other (it is a RawTool); examples/ for both directions live in mcp/examples/ and run offline over in-memory transports. TestRoundTripIsLossless pins the property: export → import keeps the schema and the answers identical.

Providers

First-party adapters wrap the vendors' official Go SDKs — weft never owns an HTTP client — and are versioned as their own modules. One line per vendor, one adapter for the whole OpenAI-compatible long tail:

import (
    "github.com/weftgo/weft/anthropic"
    "github.com/weftgo/weft/google"
    "github.com/weftgo/weft/openai"
)

openai.Model("gpt-4o-mini")                          // or any compatible server via openai.BaseURL
anthropic.Model("claude-sonnet-5", anthropic.Thinking(true))
google.Model("gemini-2.5-flash")

Reasoning depth is per run: weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkHigh}) — an agent option sets every run's default, a run option overrides it for one — maps to whatever the provider expresses (reasoning_effort, budget_tokens, thinkingBudget); adapters document what they drop. The openai adapter picks the thinking wire form from the base URL; openai.Dialect pins it when detection can't.

Sampling is per run or per step the same way — weft.Params(weft. RequestParams{…}) (Temperature, TopP, MaxTokens, Stop, Seed) folds over the adapter's construction options — and every adapter carries ExtraBody/ExtraHeaders, the caller-wins escape hatch for vendor knobs weft has no option for. Forcing a step's tool calls is weft.ToolChoice (any, a named tool, or none with the catalogue still advertised — the router shape).

Prompt caching (Anthropic): anthropic.PromptCache() marks the request's stable prefix edges — the system block, the final tool definition, the trailing conversation edge — with Anthropic's ephemeral cache_control. Cache writes bill 1.25× and reads 0.1× the base input price, so a long transcript whose prefix repeats across steps saves from the second step on; Usage.CachedInputTokens and Usage.CacheWriteTokens show it measured. The prefix is the caller's to keep stable: a PrepareStep that trims messages invalidates the trailing breakpoint on purpose, and weft.ToolChoiceNone is how you forbid calls on a final step without dropping the tool definitions — and the cache prefix they anchor — from the request.

Every adapter passes the same executable contract (wefttest/conformance): streaming tool-call fragments are assembled into whole calls and — where the provider streams fragments at all (Caps.ToolArgDeltas; Google's calls arrive whole) — also surface live as ToolArgsDelta progress, provider errors pass through unchanged for errors.As, cancellation surfaces as ctx.Err(), a stalled stream fails with ErrStreamIdle while a slow-but-streaming one never does (both are pinned cases), and WEFT_MODEL_REQUESTS=deny refuses every self-built client's call before any network I/O — test suites that must stay offline get loud failures, not surprise bills. (ADR 0013)

The rules that matter

  • Tool error = data; run error = Go error. A failing (or panicking) tool becomes a result the model sees; siblings keep running. Only model failures, cancellation, and step exhaustion reach the caller, as *RunError with the partial transcript attached. (ADR 0002)
  • Messages are role + typed parts, with a versioned JSON contract: every part carries a type discriminator and transcripts round-trip through encoding/json. (ADR 0001)
  • Tools are generic functions whose schema derives from struct tags, shape-compatible with the official Go MCP SDK. (ADR 0003)
  • Concurrent tool events carry a total order (Seq), assigned and emitted atomically, so streams replay exactly. (ADR 0004)
  • Parallel by default, bounded (4); tools always start in call order; weft.Sequential() runs them one at a time for shared state; weft.Parallelism(n) for anything else.
  • String tool outputs are sent verbatim, everything else as JSON.
  • Every run has an id (RunStart, Run.ID(), RunResult.ID); every tool call can learn its own via weft.CallFromContext(ctx).
  • Truncation is visible, never silent: a max_tokens finish is recorded on RunResult.StopReason (the run still succeeds — callers decide what truncated text means); a max_tokens step with tool calls executes none of them — each gets a visible failure and the model retries with a full budget; and oversized tool results are capped (64 KiB by default, weft.MaxResultBytes(n) to change, 0 to disable, per tool or per agent) with a marker the model sees.
  • Two behavioural seams, one observation tap. WrapModel and WrapTools change; Tap sees. A third seam needs an ADR. (ADR 0006)
  • A hung tool never hangs the run: weft.Timeout(d) on a tool or agent turns an overdue call into an error result and moves on.
  • The Model stream contract is enforced: a stream that ends without ModelFinish, continues after it, carries a tool call with an empty ID or name, or panics fails the run wrapping ErrModelContract — a broken adapter cannot corrupt a transcript.
  • One dependency in the core module: the OTel API package (ADR 0016), a no-op until an SDK registers — the zero-config instrumentation THE-END-GOAL sanctions as the core's single exception. The vendor SDKs live in the adapter modules (their own go.mod); the OTel SDK itself lives only in examples/otel; exporters stay in a satellite.

Layout

doc.go, message.go    message model (roles, parts, versioned JSON)
errors.go, env.go     error model (sentinels, RunError, kill switch)
tool.go, schema.go    tool contract, per-tool policy, schema reflection
output.go             structured output (Output, GenerateAs, OutputOf)
model.go              provider seam (streaming-first Model interface)
events.go             sealed run-event set
agent.go, run.go      agent construction options, run/stream/result
loop.go               the loop: model call → tool chain fan-out → repeat; approval resume
mw/                   reference middleware: Retry, Fallback, Log, RepairJSON, Allow, Audit, MapErrors
wefttest/             scripted mock model + the conformance suite
openai/               OpenAI Chat Completions (+ compatible servers)
anthropic/            Anthropic Messages (thinking, signatures)
google/               Gemini via genai
examples/             runnable core example (per-adapter: <adapter>/example)
docs/adr/             decision records for the contracts

Development

make test   # go test -race ./... in every workspace module
make vet
make lint   # golangci-lint (CI uses .golangci.yml)
make live   # adapter conformance against real keys (-tags live)
make apidiff  # public API of the root module vs the last tag (CI runs it)
make fuzz    # 10 s per fuzz target; FUZZTIME=1m make fuzz for longer
make fmt

Requires Go 1.26 or newer; the current and previous Go releases are supported and both are tested in CI. A fuzz crasher fails CI, its input is uploaded, and the fix PR commits it under testdata/fuzz/ as a regression seed — the existing FuzzRepair seed got there that way.

Testing

In reach order: script the dialogue with wefttest.Script(wefttest.ToolCalls(...), wefttest.Say(...)) — offline, deterministic, no key; compare bytes with wefttest.Golden; and when the question is "what does my agent do with what the model actually said", record once and replay forever:

func model(t *testing.T) weft.Model {
    if os.Getenv("WEFT_RECORD") != "" { // the suite's own switch; wefttest never reads it
        return wefttest.Record(t, "testdata/replay", openai.New(key, "gpt-5"))
    }
    return wefttest.Replay(t, "testdata/replay")
}

Fixtures are pretty JSON a reviewer reads in a diff — re-recording is the review (ADR 0017). The adapters' own wire-format parsing is proven by wefttest/conformance against recorded .sse fixtures, not by replay (ADR 0013).

API stability is enforced, not aspired to. CI runs scripts/apidiff.sh: the root module's exported API is compared against the last v* tag and any incompatible change fails the build. Pre-1.0, a deliberate source-compatible evolution (widening a return type to a superset interface, adding a trailing variadic) can be acknowledged by adding apidiff's exact line to .apidiff-allow with a justification; the file is emptied at each tag. Renaming or removing an exported symbol always fails.

Roadmap

  1. First provider adapters — done (OpenAI + compatible servers, Anthropic, Google; ADR 0013).
  2. The two middleware seams — done (WrapModel/WrapTools, package mw, the approval boundary; ADR 0006, ADR 0007).
  3. Loop refinements — done (subagents as tools, ModelRetry, usage limits, loop detection, PrepareStep; ADR 0014).
  4. MCP interop — done (consume and expose; ADR 0015). Core observability — done (OTel spans + slog lines; ADR 0016).
  5. The satellites: runtime (sessions, approvals), store, serve, studio, and the eval/prompt/mem/trace modules.

License

MIT — see LICENSE.

Documentation

Overview

Package weft is a thin, opinionated core for building agents in Go.

Weft runs the agent loop — call a model, execute its tool calls in parallel with defined failure semantics, stream typed events while work is in flight — and nothing else. It is designed the way the standard library is: small interfaces, context everywhere, functional options, wrapped errors, and no required configuration.

A tool is a plain function; its JSON Schema is derived from the input struct. An agent is a value built once with options and run many times:

echo := weft.Tool("echo", "Echo a message",
	func(ctx context.Context, in struct {
		Msg string `json:"msg"`
	}) (string, error) {
		return "echo: " + in.Msg, nil
	})

agt := weft.New(model, weft.Instructions("You are helpful."), echo)
res, err := agt.Generate(ctx, weft.Prompt("Say hi."))

A run ends when the model replies without tool calls or a StopWhen condition is met; MaxSteps is the safety budget behind both. Streaming is a range loop over typed events (Agent.Stream). Tool failures are data the model sees; only model failures, cancellation, and the step budget reach the caller, as *RunError. Output constrains the final answer to a struct (GenerateAs), and trailing options on Tool — Timeout, MaxResultBytes, StrictInput — set per-tool policy. Thinking sets the reasoning depth (agent default, run override).

The model seam (Model) is streaming-first; provider adapters translate vendor wire formats into weft's events. The wefttest package provides a scriptable Model for offline tests.

Index

Examples

Constants

View Source
const (
	// CodeSubagentFailed marks a child run that failed. The message the
	// model sees carries the child's step and cause; the *RunError
	// itself stays on ToolError.Err, reachable through errors.As, so
	// middleware can branch on it without parsing that text.
	CodeSubagentFailed = "SUBAGENT_FAILED"
	// CodeSubagentPending marks a child run that ended awaiting an
	// approval decision. Approvals belong in the orchestrator, not in a
	// child: the parent's transcript has nowhere to carry the child's
	// pending call, so the delegation is refused loudly instead of
	// silently (ADR 0014).
	CodeSubagentPending = "SUBAGENT_PENDING"
	// CodeSubagentCycle marks a delegation to an agent already running
	// in this call chain — refused before any model call.
	CodeSubagentCycle = "SUBAGENT_CYCLE"
)

Codes the Subagent tool renders its failures with — model-visible contract, pinned by tests (ADR 0002's table; ADR 0014). Exported like the loop's own codes so policy middleware can branch on them.

View Source
const (
	// CodeInvalidInput marks tool arguments that do not decode into
	// the tool's input struct; the message names the field in the
	// schema's own vocabulary.
	CodeInvalidInput = "INVALID_INPUT"
	// CodeNoSuchTool marks a call naming a tool the agent does not
	// have.
	CodeNoSuchTool = "NO_SUCH_TOOL"
	// CodeDenied marks a call the approval boundary (Deny, or no
	// decision) or a policy middleware refused — mw.Allow renders it
	// too: to the model, a refused call is a refused call, whoever
	// refused it (ADR 0007).
	CodeDenied = "DENIED"
)

Codes the loop renders its own tool failures with — model-visible contract (ADR 0002, "tool errors have codes"), exported so external policy middleware can produce the same wire strings without duplicating them, the same class of constant as SchemaVersion.

View Source
const CodeRetry = "RETRY"

CodeRetry is the code ModelRetry renders with — the one code whose result is an instruction rather than a failure report: try the call again with the hint applied.

View Source
const SchemaVersion = 1

SchemaVersion is the version of the message wire format. The JSON encoding of Message and its parts is a compatibility contract: within a version, field names and shapes change only additively. Persistence and serving layers envelope messages with this number; the core itself never needs it.

Variables

View Source
var (
	// ErrMaxSteps is returned when the model still requests tools after the
	// last allowed step. The partial transcript rides along on RunError.
	ErrMaxSteps = errors.New("weft: run exceeded the maximum number of steps")

	// ErrUsageLimit is returned when a run's usage exceeds its
	// UsageLimit and the loop would otherwise call the model again. A
	// step that ends the run may overshoot and still succeed: a budget
	// stops further spend, it does not discard finished work.
	ErrUsageLimit = errors.New("weft: run exceeded its usage limit")

	// ErrModelRetriesExceeded is returned when one tool has produced
	// more consecutive RETRY results than MaxModelRetries allows — a
	// model that cannot self-correct is a run failure, not an infinite
	// loop.
	ErrModelRetriesExceeded = errors.New("weft: a tool asked the model to retry too many times")

	// ErrLoopDetected is returned when repeats consecutive steps
	// requested the same set of tool calls (DetectLoops).
	ErrLoopDetected = errors.New("weft: the model repeated the same tool calls too many times")

	// ErrNoSuchTool is returned by Agent.CallTool when the call names a
	// tool the agent does not have. Inside the loop the same condition is
	// folded into an error tool result — data for the model to
	// self-correct, not a run failure.
	ErrNoSuchTool = errors.New("weft: no tool with that name")

	// ErrInvalidToolInput is returned by ToolDef.Invoke and Agent.CallTool
	// when the arguments do not decode into the tool's input type. Like
	// ErrNoSuchTool, the loop turns it into model-visible data.
	ErrInvalidToolInput = errors.New("weft: tool input is not valid for its schema")

	// ErrRunConsumed is returned by Run.Events when the event stream has
	// already been consumed; each Run yields exactly one sequence.
	ErrRunConsumed = errors.New("weft: run events already consumed")

	// ErrModelContract is wrapped around failures of a Model
	// implementation to honor the stream contract documented on Model:
	// events after ModelFinish, a stream ending without one, a tool call
	// with an empty ID or name, or a panicking stream. It signals an
	// adapter bug, not a model outage — providers' own errors surface
	// unwrapped.
	ErrModelContract = errors.New("weft: model violated the stream contract")

	// ErrUnsupported is wrapped by a Model that cannot honour part of a
	// request — a FilePart whose media type the provider does not accept,
	// a feature the vendor lacks. It is a run error (the model call
	// fails), so callers can errors.Is on it and fall back to another
	// model.
	ErrUnsupported = errors.New("weft: request uses a feature the model does not support")

	// ErrStreamIdle is the stream error an adapter yields when the gap
	// between two chunks exceeds its IdleTimeout. The ctx deadline is the
	// hard limit on a whole call; the idle timeout only catches a stalled
	// stream, so a slow but actively streaming response is never killed.
	// It is a run error like any provider error; callers errors.Is on it
	// provider-agnostically, without knowing which adapter timed out.
	ErrStreamIdle = errors.New("weft: model stream idle timeout")

	// ErrModelRequestsDenied is the stream error every first-party
	// adapter yields from Stream when ModelRequestsAllowed is false — the
	// kill switch for test suites that must never reach the network.
	// wefttest models ignore the switch, so ordinary offline tests are
	// unaffected.
	ErrModelRequestsDenied = errors.New("weft: model requests denied by WEFT_MODEL_REQUESTS")

	// ErrApprovalRequired marks a tool call that must not run until a
	// human (or an outer system) decides. The loop raises it for tools
	// built with RequireApproval; tool middleware may return an error
	// wrapping it to defer any call. The run then ends successfully with
	// the call on RunResult.Pending; resume with Approve or Deny.
	ErrApprovalRequired = errors.New("weft: tool call requires approval")

	// ErrApprovalDenied is the cause on the error result the model sees
	// for a pending call that was denied (Deny, or no decision on
	// resume). It is a tool error — data — never a run error.
	ErrApprovalDenied = errors.New("weft: tool call denied")

	// ErrDuplicateTool is a run error raised when a step's tool
	// snapshot contains a name twice — the runtime analogue of New's
	// duplicate-name panic. Fix the tool source; the run fails rather
	// than silently dropping the second tool. Agent.CallTool reports
	// the same condition as an error.
	ErrDuplicateTool = errors.New("weft: duplicate tool name from tool source")

	// ErrNilTool is a run error raised when a tool-source snapshot
	// contains a nil entry — a malformed snapshot, not a tool. The run
	// fails rather than advertising a dereference every adapter would
	// panic on; Agent.CallTool reports the same condition as an error.
	ErrNilTool = errors.New("weft: nil tool in tool source snapshot")
)

Sentinel errors for named run failures. Branch on them with errors.Is; never match on error strings.

View Source
var ErrNoOutput = errors.New("weft: run ended without a structured output")

ErrNoOutput is returned by GenerateAs and OutputOf when the run ended without a valid submit_output call — the model answered in text, hit the step budget, or never produced arguments that decode into Out.

Functions

func Manifest

func Manifest(agents ...*Agent) ([]byte, error)

Manifest renders the agents as their `weft.json` document: one generated, committed, diffable description of every agent and tool. Studio, docs, review, and compatibility checks read a file instead of a live process; the code stays the only source of truth. The file is output, never input — nothing is configured from it. Generate it in a golden test (wefttest.Golden) so a stale file fails `go test`; run that test with `-update` to regenerate.

Tools are listed in registration order, agents in argument order, and encoding/json sorts map keys, so the bytes are deterministic for the same agents. Unnamed and duplicate agent names are errors: the file is a review artifact, and `agent_1` in a diff is noise.

Example

The manifest is generated output: one committed, diffable description of every agent and tool. Gate it with a golden test so it cannot go stale.

package main

import (
	"context"
	"fmt"
	"log"
	"strings"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	agt := weft.New(wefttest.Script(wefttest.Say("ok")),
		weft.Name("support-bot"),
		weft.Instructions("You are a support agent."),
		weft.Tool("refund_order", "Refund a customer's order.",
			func(_ context.Context, _ struct {
				OrderID string `json:"order_id"`
			}) (string, error) {
				return "refunded", nil
			}),
	)
	b, err := weft.Manifest(agt)
	if err != nil {
		log.Fatal(err)
	}
	lines := strings.Split(string(b), "\n")
	fmt.Println(lines[0])
	fmt.Println(lines[1])
}
Output:
{
  "weft": 1,

func ModelRequestsAllowed

func ModelRequestsAllowed() bool

ModelRequestsAllowed reports whether adapters may call a provider. It is false when WEFT_MODEL_REQUESTS=deny — the guard for test suites that must never reach the network (Pydantic AI's ALLOW_MODEL_REQUESTS). First-party adapters check it at the top of Stream and yield ErrModelRequestsDenied when it is false, before any network I/O — but only for a client they built themselves from credentials; a client the caller injected through the adapter's Client(c) option is a test double by construction and stays reachable (ADR 0013's kill-switch clause). wefttest models ignore the switch entirely.

The environment is read on every call, not cached: a model call is network-bound so one getenv is noise, test suites can toggle the switch per test with t.Setenv, and the package keeps no state at all.

func ModelRetry added in v0.2.0

func ModelRetry(hint string) error

ModelRetry is a tool error asking the model to try the call again with the hint applied: it renders as "RETRY: <hint>" and the loop continues. Return it when the arguments are well-formed but wrong in a way the model can fix — an ambiguous date, an id that needs a prefix. The loop counts RETRY results per tool name per run and fails the run with ErrModelRetriesExceeded when a tool exceeds MaxModelRetries (default 3); a successful result for that tool resets its count. Middleware-produced retries count too — the loop sees the code through the chain, not who returned it.

Example

A retry hint: the model sees "RETRY: <hint>", fixes the arguments, and the loop continues; a tool that cannot be satisfied fails the run after MaxModelRetries consecutive asks.

package main

import (
	"context"
	"fmt"
	"log"
	"sync/atomic"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	var calls atomic.Int32
	parse := weft.Tool("parse_date", "Parse a date.",
		func(_ context.Context, in struct {
			D string `json:"d" jsonschema:"the date, ISO-8601"`
		}) (string, error) {
			if in.D != "2026-09-19" {
				return "", weft.ModelRetry("date must be ISO-8601, e.g. 2026-09-19")
			}
			_ = calls.Add(1)
			return "2026-09-19", nil
		})
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "parse_date", Args: `{"d":"tomorrow"}`}),
		wefttest.ToolCalls(wefttest.Call{Name: "parse_date", Args: `{"d":"2026-09-19"}`}),
		wefttest.Say("Parsed."),
	), parse)
	res, err := agt.Generate(context.Background(), weft.Prompt("When is it?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println("first attempt:", res.Steps[0].Results[0].Content)
	fmt.Println("second attempt:", res.Steps[1].Results[0].Content)
}
Output:
first attempt: RETRY: date must be ISO-8601, e.g. 2026-09-19
second attempt: 2026-09-19

func OutputOf added in v0.2.0

func OutputOf[Out any](res *RunResult) (Out, error)

OutputOf decodes the structured output recorded in a finished run — the last submit_output call with a non-error result — for callers that streamed the run and hold its RunResult. It returns ErrNoOutput when no such call exists.

Types

type Agent

type Agent struct {
	// contains filtered or unexported fields
}

Agent is an immutable, reusable value: a model, a system instruction, a tool set, and an execution policy. Build it once with New; run it many times, concurrently if you like — runs share no state, and the registered tool set is frozen at construction (see ToolDef).

Example (Conversation)

Continuing a conversation: feed the transcript back with the next question. The agent value is unchanged — the history lives in the messages you pass, never in the agent.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	agt := weft.New(wefttest.Script(
		wefttest.Say("Order 1234? It shipped yesterday."),
		wefttest.Say("Order 5678? Still pending."),
	))
	res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 1234?"))
	if err != nil {
		log.Fatal(err)
	}
	res2, err := agt.Generate(context.Background(),
		weft.Messages(res.Messages...), weft.Prompt("And order 5678?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
	fmt.Println(res2.Text())
}
Output:
Order 1234? It shipped yesterday.
Order 5678? Still pending.
Example (Generate)
package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	echo := weft.Tool("echo", "Echo a message.",
		func(_ context.Context, in struct {
			Msg string `json:"msg"`
		}) (string, error) {
			return "echo: " + in.Msg, nil
		})
	model := wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"msg":"hello"}`}),
		wefttest.Say("I echoed your message."),
	)
	agt := weft.New(model, weft.Instructions("You echo things."), echo)

	res, err := agt.Generate(context.Background(), weft.Prompt("Echo hello."))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
	fmt.Println("steps:", res.NumSteps(), "tokens:", res.Usage.Total())
}
Output:
I echoed your message.
steps: 2 tokens: 30
Example (Stream)
package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	roll := weft.Tool("roll_dice", "Roll a six-sided die.",
		func(_ context.Context, _ struct{}) (int, error) {
			return 4, nil // deterministic for the example
		})
	model := wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "roll_dice"}),
		wefttest.Say("You rolled a 4!"),
	)
	agt := weft.New(model, roll)

	for ev, err := range agt.Stream(context.Background(), weft.Prompt("Roll a die.")).Events() {
		if err != nil {
			log.Fatal(err)
		}
		switch ev := ev.(type) {
		case weft.ToolStart:
			fmt.Println("tool:", ev.Name)
		case weft.TextDelta:
			fmt.Println("text:", ev.Text)
		case weft.RunFinish:
			fmt.Println("done in", ev.Steps, "steps")
		}
	}
}
Output:
tool: roll_dice
text: You rolled a 4!
done in 2 steps

func New

func New(m Model, opts ...Option) *Agent

New builds an Agent. Nil models panic — including typed nils such as var m *someModel; New(m), which would otherwise crash much later inside a run goroutine. Everything else has a working default.

func (*Agent) CallTool

func (a *Agent) CallTool(ctx context.Context, call ToolCallPart) (string, error)

CallTool dispatches one tool call by name the way the loop does — through the agent's WrapTools chain and the tool's own — and returns the result text the model would see. It is the seam for manual dispatchers. Unlike the loop it does not contain failures or apply run policy: an unknown name returns an error wrapping ErrNoSuchTool, undecodable arguments one wrapping ErrInvalidToolInput, a RequireApproval tool one wrapping ErrApprovalRequired, and handler errors, middleware errors, and panics propagate. A tool source whose snapshot carries a duplicate name returns an error wrapping ErrDuplicateTool. Timeouts and result caps are not applied; StrictInput is, at both levels, as in the loop. No span or log line is produced: manual dispatchers own their context and their own reporting (ADR 0016).

func (*Agent) Generate

func (a *Agent) Generate(ctx context.Context, opts ...RunOption) (*RunResult, error)

Generate runs the agent to completion and returns the final result. Internally it is Stream with the events folded away; errors are returned as *RunError, with the partial transcript attached.

func (*Agent) Name added in v0.2.0

func (a *Agent) Name() string

Name returns the agent's name as set by the Name option; empty when the agent is unnamed. Manifest requires a name; Serve (weft/mcp) names the tool it exposes after the agent through this accessor.

func (*Agent) Stream

func (a *Agent) Stream(ctx context.Context, opts ...RunOption) *Run

Stream starts a run and returns its handle. The run is lazy: nothing executes until Events is consumed. Canceling ctx aborts the model call and any in-flight tools.

func (*Agent) TapPanics added in v0.2.0

func (a *Agent) TapPanics() int64

TapPanics reports how many tap invocations have panicked and been contained since construction. A rising counter means an observer is broken; runs are unaffected by design.

func (*Agent) Tools

func (a *Agent) Tools() []*ToolDef

Tools returns copies of the registered tool definitions, in registration order. Mutating them — fields or schema trees — does not affect the agent; the registered set is frozen at New.

type Call

type Call struct {
	RunID  string // the run this call belongs to
	Step   int    // zero-based index of the step that requested it
	CallID string // the provider's call identifier
	Name   string // the tool name
	// Approved is set when the call was parked by the approval boundary
	// and is now running under an Approve decision — middleware that
	// defers calls with ErrApprovalRequired reads it to let the
	// approved call through.
	Approved bool
}

Call identifies the tool invocation a handler is serving. Retrieve it with CallFromContext — for audit logs, per-call idempotency keys, or progress reporting that must name its call.

func CallFromContext

func CallFromContext(ctx context.Context) (c Call, ok bool)

CallFromContext returns the Call a tool handler is serving. ok is false when ctx did not come from the agent loop (a direct Invoke, for example).

type Event

type Event interface {
	// contains filtered or unexported methods
}

Event is the sealed set of run progress events, yielded by Run.Events in emission order. New event types may be added additively; external types cannot join, so switches over events stay exhaustively lintable.

On the wire every event carries a "type" discriminator (run_start, step_start, text_delta, reasoning_delta, tool_args_delta, tool_start, tool_finish, step_finish, run_finish, nested) and UnmarshalEvent restores it — the same rule and the same compatibility contract as the message parts (ADR 0004).

Every event except RunStart carries RunID: concurrent runs on one agent emit interleaved streams, and a per-run Seq counter is unique only within its run, so RunID is what attributes an event to its run.

Events are snapshots. Their fields — including the Args byte slices — do not alias the run's transcript; a consumer may retain or write into them freely.

func UnmarshalEvent

func UnmarshalEvent(b []byte) (Event, error)

UnmarshalEvent decodes one wire event, dispatching on its "type" discriminator. An unknown or missing type is an error, never a silent drop: a recorded stream must replay exactly what was emitted.

Example

Recorded event streams decode back into typed events: store writes them, the Inspector replays them.

package main

import (
	"fmt"
	"log"

	"github.com/weftgo/weft"
)

func main() {
	ev, err := weft.UnmarshalEvent([]byte(`{"type":"tool_start","seq":5,"call_id":"c1","name":"echo","args":{"m":"x"}}`))
	if err != nil {
		log.Fatal(err)
	}
	start := ev.(weft.ToolStart)
	fmt.Println(start.Name, start.CallID, start.Seq, string(start.Args))
}
Output:
echo c1 5 {"m":"x"}

type FilePart

type FilePart struct {
	MediaType string `json:"media_type"`
	Data      []byte `json:"data,omitempty"`
	URL       string `json:"url,omitempty"`
}

FilePart is a file the user supplies to the model: an image, a PDF, audio. Exactly one of Data (inline; base64 on the wire via []byte's default encoding) or URL is set. Adapters map it to the vendor's image/document block; an adapter that cannot carry this MediaType fails the model call with an error wrapping ErrUnsupported. The core never reads the bytes. The exactly-one rule is documented, not enforced here — the adapter is the layer that knows what it can send, and it returns ErrUnsupported for a part with both or neither set.

func (FilePart) MarshalJSON

func (p FilePart) MarshalJSON() ([]byte, error)

MarshalJSON encodes the part with its "type" discriminator.

type Message

type Message struct {
	Role    Role   `json:"role"`
	Content []Part `json:"content"`
}

Message is one turn in a conversation: a role plus an ordered list of content parts. A step's tool results are collected on a single RoleTool message; provider adapters fan out or merge as their wire format requires.

On the wire every part carries a "type" discriminator ("text", "tool_call", "tool_result", "reasoning"), so a Message round-trips through encoding/json losslessly.

func Assistant

func Assistant(text string) Message

Assistant returns an assistant message with a single text part.

func Repair

func Repair(msgs []Message) []Message

Repair makes a transcript valid model input: every tool call has a result (missing ones become visible error results), results with no call are dropped, everything else is untouched. The loop applies it to the input of every run; store and runtime call it before persisting. It is pure (the input is never mutated) and idempotent: Repair(Repair(m)) equals Repair(m). nil input yields nil; an empty non-nil input yields an empty non-nil transcript.

Example

A partial transcript (the run was interrupted mid-step) becomes valid model input: the missing result is synthesised, visibly.

package main

import (
	"encoding/json"
	"fmt"

	"github.com/weftgo/weft"
)

func main() {
	msgs := []weft.Message{
		weft.User("Where is order 1234?"),
		{Role: weft.RoleAssistant, Content: []weft.Part{
			weft.ToolCallPart{ID: "c1", Name: "lookup_order", Args: json.RawMessage(`{"order_id":"1234"}`)},
		}},
	}
	for _, m := range weft.Repair(msgs) {
		fmt.Println(m.Role)
	}
}
Output:
user
assistant
tool

func User

func User(text string) Message

User returns a user message with a single text part.

func UserParts

func UserParts(parts ...Part) Message

UserParts returns a user message with the given parts, for prompts that mix text and files: UserParts(TextPart{"What is this?"}, FilePart{MediaType: "image/png", URL: u}). The parts are copied; the caller's slice is not retained.

Example

Provider reasoning round-trips: it is preserved in the transcript, placed before the text of the same assistant turn. A prompt can mix text and files with UserParts; the core carries the bytes and the adapter maps them to the provider's image block.

package main

import (
	"encoding/json"
	"fmt"
	"log"

	"github.com/weftgo/weft"
)

func main() {
	msg := weft.UserParts(
		weft.TextPart{Text: "What is this?"},
		weft.FilePart{MediaType: "image/png", URL: "https://example.com/cat.png"},
	)
	b, err := json.Marshal(msg)
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(string(b))
}
Output:
{"role":"user","content":[{"type":"text","text":"What is this?"},{"type":"file","media_type":"image/png","url":"https://example.com/cat.png"}]}

func (Message) Text

func (m Message) Text() string

Text returns the concatenation of the message's text parts.

func (*Message) UnmarshalJSON

func (m *Message) UnmarshalJSON(b []byte) error

UnmarshalJSON decodes a message, dispatching each content part on its "type" discriminator. An unknown part type is an error: the wire format is versioned, and silently dropping content would corrupt transcripts.

type Model

type Model interface {
	Stream(ctx context.Context, req ModelRequest) iter.Seq2[ModelEvent, error]
}

Model is the provider seam. Implementations stream one step's output as events; first-party adapters wrap the vendors' official Go SDKs rather than re-implementing HTTP.

The stream contract:

  • Events are yielded in order: any number of ModelTextDelta, ModelReasoningDelta, and ModelToolCall values — optionally interleaved with ModelToolCallDelta progress as argument fragments stream — then exactly one ModelFinish.
  • Failure is reported as a single terminal yield of (nil, err); no events follow it.
  • The sequence honors ctx: when ctx is done, the model yields (nil, ctx.Err()) if it has not finished already.

The loop enforces this contract: a stream that ends without ModelFinish, continues after it, yields a tool call with an empty ID or name, two tool calls sharing an ID in one step, or panics fails the run with an error wrapping ErrModelContract. A contract-violating adapter cannot corrupt a transcript silently.

type ModelEvent

type ModelEvent interface {
	// contains filtered or unexported methods
}

ModelEvent is the sealed set of events a model yields during one step. Tool calls arrive whole — assembling providers' streamed argument fragments is the adapter's job, which is what makes the core's streaming uniform across providers.

type ModelFinish

type ModelFinish struct {
	Reason StopReason
	Usage  Usage
	// Raw is the provider's own stop reason when Reason had to be
	// approximated ("refusal", "pause_turn", "content_filter", ...).
	// Empty when the mapping was exact. The core never interprets it; it
	// is recorded on the step (StepRecord.RawStopReason) and the
	// StepFinish event so callers can see a refusal without a tap.
	Raw string
}

ModelFinish closes a step with its stop reason and token usage.

type ModelInfo

type ModelInfo struct {
	Provider string `json:"provider"` // "openai", "anthropic", "wefttest"
	Name     string `json:"name"`     // the vendor's model id
}

ModelInfo identifies a model for telemetry (§8.1) and the manifest. A Model that can report it implements the optional

interface{ Info() ModelInfo }

which the loop detects and surfaces on RunStart.Model; Model itself stays one method. Middleware that wraps a Model should forward Info (TODO §4.1).

func InfoOf added in v0.2.0

func InfoOf(m Model) ModelInfo

InfoOf reports m's ModelInfo when it implements the optional

interface{ Info() ModelInfo }

and the zero ModelInfo otherwise. Model middleware forwards identity with it: func (w *wrapper) Info() weft.ModelInfo { return weft.InfoOf(w.next) }.

type ModelMiddleware added in v0.2.0

type ModelMiddleware func(next Model) Model

ModelMiddleware wraps a Model, the chi shape: it sees every ModelRequest the loop builds and every event the inner model yields, and may retry, substitute, log, or rewrite. Implementations should forward Info (see InfoOf) so RunStart.Model and the manifest still name the underlying model.

type ModelReasoningDelta

type ModelReasoningDelta struct {
	Text      string
	Signature string
}

ModelReasoningDelta is an increment of provider reasoning (Anthropic thinking, Gemini thought summaries). Signature is the provider's opaque token for the block, if any; adapters set it on the delta that completes a block. The core stores and forwards reasoning and never reads it.

Block boundaries: a delta carrying a non-empty Signature closes the current reasoning block; the next reasoning delta opens a new one. Providers send a block's signature last (Anthropic's signature_delta ends a thinking block; Gemini's per-part signature is emitted after the part's text), so one ReasoningPart per provider block survives the round trip. Reasoning without any signature accumulates into a single block — nothing downstream can send unsigned blocks back anyway, so their boundaries are not load-bearing.

type ModelRequest

type ModelRequest struct {
	System   string
	Messages []Message
	Tools    []*ToolDef
	// Thinking asks the model to reason at the given level for this
	// step. The zero value keeps the provider default (and any
	// construction-time adapter option, such as anthropic.Thinking);
	// the loop fills it from the agent's Thinking option, which a
	// run-level Thinking overrides.
	Thinking ThinkingConfig
	// SequentialTools asks the provider not to emit parallel tool-call
	// batches. The zero value keeps the provider default; the loop sets
	// it to true exactly under Sequential() (and Parallelism(1)), so the
	// model does not emit batches the execution policy would serialize
	// anyway. Adapters mirror it in the provider's parallel-tool-calls
	// setting.
	SequentialTools bool
	// ToolChoice constrains what the model may emit this step: some
	// tool, a named tool, or none — the router-agent and hardened-Output
	// knob. The zero value keeps the provider default; the loop fills it
	// from the agent's ToolChoice option, which a run-level ToolChoice
	// overrides, and a PrepareStep function can rewrite it per step
	// (force classify on step 0, then auto). It constrains what the
	// provider is asked to emit, never execution: a call the provider
	// emits anyway runs, because the model was shown the tool.
	ToolChoice ToolChoiceConfig
	// Params overrides the adapter's construction-time sampling knobs
	// (Temperature, TopP, MaxTokens, Stop, Seed) for this request alone;
	// see RequestParams for the nil-keeps-default fold. The zero value
	// keeps every construction default, so a request without it
	// serialises exactly as v0.2.0's did.
	Params RequestParams
}

ModelRequest is everything a model needs for one step: the system instruction, the transcript so far, and the callable tools.

Read-only: the loop builds each request with fresh copies of the Messages and Tools slices, so appending to them or reassigning their elements cannot reach the run or the agent. The values inside remain shared — message Content parts, and ToolDef fields frozen at construction — so implementations must still not modify them, and must clone anything they retain beyond the call.

type ModelTextDelta

type ModelTextDelta struct {
	Text string
}

ModelTextDelta is an increment of assistant text.

type ModelToolCall

type ModelToolCall struct {
	ID        string
	Name      string
	Args      json.RawMessage
	Signature string
}

ModelToolCall is one complete tool invocation request. ID is the provider's call identifier, echoed back on the matching ToolResultPart; it must be unique among one step's calls (results and approval decisions key on it). Signature is the provider's opaque token attached to the call itself (Gemini attaches thought signatures to functionCall parts and requires them returned on the same part); adapters that do not have one leave it empty.

type ModelToolCallDelta added in v0.2.0

type ModelToolCallDelta struct {
	Index int
	Name  string
	Args  string
}

ModelToolCallDelta is an increment of a streamed tool call's arguments — progress only: the assembled call still arrives whole as a ModelToolCall before ModelFinish. Adapters whose providers stream argument fragments (OpenAI-compatible function.arguments pieces, Anthropic input_json_delta) yield these so consumers can show the model "writing" a call instead of dead air; adapters whose calls arrive whole (Google) simply yield none. Index is the provider's fragment key where one exists (OpenAI's delta index); Name is the best-known name so far — for many providers only the first fragment of a call carries it.

type Nested added in v0.2.0

type Nested struct {
	RunID  string `json:"run_id"`
	Seq    int64  `json:"seq"`
	CallID string `json:"call_id"`
	Event  Event  `json:"event"`
}

Nested wraps one event of a child run started by a Subagent tool. CallID is the parent's tool call that owns the child run; Seq is from the parent's counter, so the parent stream stays totally ordered with the child's events in place. Event is any child event — including a Nested from a grandchild — and its own RunID and Seq are the child's. A child's RunStart..RunFinish all arrive inside the parent's ToolStart..ToolFinish for the delegating call, and none arrive after its ToolFinish (the late-event rule, ADR 0004). On the wire the type is "nested" and UnmarshalEvent restores the inner event recursively.

func (Nested) MarshalJSON added in v0.2.0

func (e Nested) MarshalJSON() ([]byte, error)

MarshalJSON encodes the event with its "type" discriminator. The inner event marshals through its own MarshalJSON, so its discriminator is present and decoding recurses (UnmarshalEvent inside Nested.UnmarshalJSON).

func (*Nested) UnmarshalJSON added in v0.2.0

func (e *Nested) UnmarshalJSON(b []byte) error

UnmarshalJSON decodes the envelope and the inner event recursively — an unknown inner type is an error exactly as at the top level.

type Option

type Option interface {
	// contains filtered or unexported methods
}

Option configures an Agent at construction. Options are small values returned by Instructions, MaxSteps, Parallelism, Sequential, StopWhen, Output, the PolicyOptions (MaxResultBytes, Timeout, StrictInput), and the Tool constructor.

func DetectLoops added in v0.2.0

func DetectLoops(repeats int) Option

DetectLoops fails a run with ErrLoopDetected when repeats consecutive steps request the same set of tool calls — the same names with the same arguments, in any order. Varying arguments are not a loop: a corrected retry has a different signature. Results are not part of the signature: a tool whose output carries a timestamp must not hide a loop, and a result-inclusive hash would miss it. Off by default (values below 2 are ignored); the CLI scaffold writes DetectLoops(5).

Example

A stuck model repeating one request: DetectLoops fails the run loudly instead of burning the step budget.

package main

import (
	"context"
	"errors"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	echo := weft.Tool("echo", "", func(_ context.Context, _ struct{}) (string, error) {
		return "ok", nil
	})
	turns := []wefttest.Turn{}
	for range 3 {
		turns = append(turns, wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"i":1}`}))
	}
	agt := weft.New(wefttest.Script(turns...), echo, weft.DetectLoops(3))
	_, err := agt.Generate(context.Background(), weft.Prompt("q"))
	fmt.Println(errors.Is(err, weft.ErrLoopDetected))
}
Output:
true

func Instructions

func Instructions(text string) Option

Instructions sets the agent's system prompt.

func Logger added in v0.2.0

func Logger(l *slog.Logger) Option

Logger sets the logger the agent's runs report to, at Debug level: one line when a run starts and ends, one per model call, one per tool call — run and call ids, the model, durations, usage, stop reasons and outcomes; never message text, tool arguments or tool results. Error text is the one exception: a failed run, model call or tool call logs the error it returned, because a log is the caller's. nil (the default) means slog.Default, which is silent until its handler enables Debug, so weft logs nothing in a program that did not ask; slog.New(slog.DiscardHandler) turns the lines off outright. The lines carry the span-carrying context, so a handler that bridges to OTel correlates them with the spans for free. mw.Log and mw.Audit are separate: middleware the caller places, at the level the caller chooses.

Example

ExampleLogger shows the whole log surface: one line per phase, at Debug, ids and counts only.

package main

import (
	"context"
	"log/slog"
	"os"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	scrub := func(_ []string, a slog.Attr) slog.Attr {
		switch a.Key {
		case slog.TimeKey, "dur": // timestamps and durations are not reproducible
			return slog.Attr{}
		}
		return a
	}
	logger := slog.New(slog.NewTextHandler(os.Stdout, &slog.HandlerOptions{
		Level:       slog.LevelDebug,
		ReplaceAttr: scrub,
	}))
	echo := weft.Tool("echo", "Echo the message.", func(_ context.Context, in struct {
		Msg string `json:"msg"`
	}) (string, error) {
		return "echo: " + in.Msg, nil
	})
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"msg":"hi"}`}),
		wefttest.Say("done"),
	), weft.Logger(logger), weft.Name("demo"), echo)
	if _, err := agt.Generate(context.Background(), weft.RunID("demo"), weft.Prompt("hi")); err != nil {
		return
	}
}
Output:
level=DEBUG msg="run start" run=demo agent=demo provider=wefttest model=script
level=DEBUG msg="model call" run=demo step=0 provider=wefttest model=script reason=tool_calls input_tokens=10 output_tokens=5 tool_calls=1
level=DEBUG msg="tool call" run=demo step=0 call=call_1 tool=echo result_bytes=8
level=DEBUG msg="model call" run=demo step=1 provider=wefttest model=script reason=stop input_tokens=10 output_tokens=5 tool_calls=0
level=DEBUG msg="run finish" run=demo steps=2 input_tokens=20 output_tokens=10 stop=stop

func MaxModelRetries added in v0.2.0

func MaxModelRetries(n int) Option

MaxModelRetries sets how many consecutive RETRY results (see ModelRetry) one tool may produce in a run before the run fails with ErrModelRetriesExceeded (default 3). The count is per tool name — two parallel calls to the same tool both retrying count as two — and a successful result for the tool resets it. Calls resumed under Approve do not feed the counter: they run before step 0 and belong to no step (ADR 0007). Values below 1 are ignored.

func MaxSteps

func MaxSteps(n int) Option

MaxSteps is the safety budget: the most model calls a run may make (default 10). Exceeding it fails the run with ErrMaxSteps — a runaway loop is a failure to surface, never a quiet success. Use StopWhen for the intended end of a run. Values below 1 are ignored.

func Name

func Name(name string) Option

Name names the agent: it appears on RunStart.Agent and in the manifest, which requires it (weft.Manifest errors on unnamed agents). Empty values are ignored.

func Options added in v0.2.0

func Options(opts ...Option) Option

Options composes several options into one, applied in order — a family of tools and policy closed over its dependencies as a single value. A plugin is `func(deps) weft.Option`, and dependencies are parameters, never globals; nothing registers itself, so there is no registry, no scopes, no dedup. Nil entries are ignored; duplicate tool names still panic at New. Nesting composes by construction.

Example

Composing agents: a plugin is func(deps) weft.Option — a family of tools and its policy closed over its dependencies as one value. Dependencies are parameters, never globals.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	orders := func(deps *string) weft.Option {
		return weft.Options(
			weft.Instructions("You handle orders."),
			weft.Tool("lookup_order", "Look up an order by id.",
				func(_ context.Context, in struct {
					ID string `json:"id" jsonschema:"the order id"`
				}) (string, error) {
					return "order " + in.ID + ": " + *deps, nil
				}),
			weft.MaxResultBytes(1024),
		)
	}
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":"42"}`}),
		wefttest.Say("Done."),
	), orders(new(string)))
	res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 42?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
}
Output:
order 42:

func Output added in v0.2.0

func Output[Out any]() Option

Output constrains the run's final answer to Out. It registers a tool named submit_output whose input schema is reflected from Out, exactly as Tool reflects a handler's input, and stops the run once the model has called it with arguments that decode. Invalid arguments come back to the model as an ErrInvalidToolInput result naming the field, so repair is the ordinary tool-error loop — nothing special happens. Read the value with GenerateAs, or OutputOf after Stream:

type Verdict struct {
    Approved bool   `json:"approved"`
    Reason   string `json:"reason" jsonschema:"one sentence"`
}

agt := weft.New(model, weft.Output[Verdict](), lookup)
v, res, err := weft.GenerateAs[Verdict](ctx, agt, weft.Prompt("Review order 42."))

Tool mode works on every provider; adapters with a native JSON-schema mode may use it later behind the same API. Output panics if Out is not a struct (or a pointer to one), for the reason Tool does.

func Parallelism

func Parallelism(n int) Option

Parallelism sets the maximum number of a step's tool calls executing at once (default 4). Values below 1 are ignored.

func PrepareStep added in v0.2.0

func PrepareStep(fn func(ctx context.Context, step int, req ModelRequest) (ModelRequest, error)) Option

PrepareStep installs a function the loop calls before every model call, with the request it built for step: the agent's instructions, the transcript so far, the step's tool snapshot, the run's thinking level. The function returns the request the step uses — trimmed messages, a subset of tools, a rewritten system — or an error that fails the run. What it returns is what the step advertises and dispatches against: a tool it removes cannot be called that step, and a ToolDef it adds can. PromptSnippets are composed after it, from the tools it returns, so removing a tool removes its snippet without the function having to know snippets exist. The request is a copy the function may mutate freely: message parts, argument bytes, and tool definitions are cloned before the chain runs, so in-place writes reach neither the transcript nor the agent's frozen registry. The transcript in RunResult is never affected; only the request is. It runs before the model seam, so WrapModel middleware sees the prepared request. Several PrepareStep options run in order, each receiving the previous one's result. A nil function is ignored. Like ToolSource, it is one of the two knobs that can break a prompt-cache prefix — trim deliberately.

Example

Phased tool exposure: the first step plans with read-only tools; the second acts. PrepareStep is the one loop knob — what it returns is what the step both advertises and dispatches against.

package main

import (
	"context"
	"fmt"
	"log"
	"slices"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	lookup := weft.Tool("lookup", "Look up an order.", func(_ context.Context, _ struct{}) (string, error) {
		return "order 1234: broken item", nil
	})
	refund := weft.Tool("refund", "Refund an order.", func(_ context.Context, _ struct{}) (string, error) {
		return "refunded", nil
	})
	phase := func(_ context.Context, step int, req weft.ModelRequest) (weft.ModelRequest, error) {
		if step == 0 { // investigate before acting
			req.Tools = slices.DeleteFunc(req.Tools, func(t *weft.ToolDef) bool { return t.Name == "refund" })
		}
		return req, nil
	}
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "refund"}), // not advertised in step 0
		wefttest.ToolCalls(wefttest.Call{Name: "refund"}), // now it is
		wefttest.Say("Done."),
	), lookup, refund, weft.PrepareStep(phase))
	res, err := agt.Generate(context.Background(), weft.Prompt("Refund order 1234."))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
	fmt.Println(res.Steps[1].Results[0].Content)
}
Output:
NO_SUCH_TOOL: no tool named "refund"
refunded

func StopWhen

func StopWhen(conds ...StopCondition) Option

StopWhen adds stop conditions; the run ends when any one is met. Without StopWhen a run ends when the model replies without requesting tools. Compose the built-ins — HasToolCall, StepCountIs — or write your own:

weft.StopWhen(weft.HasToolCall("submit_answer"))
weft.StopWhen(weft.StopFunc(func(steps []weft.StepRecord) bool { ... }))

Stop conditions are the intended end of a run; MaxSteps is the safety budget behind them.

func Tap

func Tap(fn func(ctx context.Context, ev Event)) Option

Tap registers an observer that sees every event of every run, including runs made with Generate, synchronously and in emission order on the emitting goroutine. It must be fast and must not block: it runs under the event-ordering lock, so a slow tap delays every tool event of its step and blocks the emitting tool goroutines — the same consumer-speed coupling Run.Events documents. Slow observation (a database write, a network sink) must not happen inside the tap: hand each event to a queue and drain it on your own goroutine; ExampleTap_async is the tested pattern. Taps run in registration order; a panic in one is recovered, counted (see TapPanics), and dropped, so a broken observer cannot break a run. Taps observe and cannot change anything — behaviour attaches at the two middleware seams. ctx is the run's context, which carries the run's span, so a tap that starts its own spans parents them under it for free.

Example

Tap observes every event of every run — including Generate, which has no stream to range over. Taps see; the middleware seams change.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	var calls int
	agt := weft.New(
		wefttest.Script(
			wefttest.ToolCalls(wefttest.Call{Name: "roll_dice"}),
			wefttest.Say("rolled"),
		),
		weft.Tap(func(_ context.Context, ev weft.Event) {
			if _, ok := ev.(weft.ToolStart); ok {
				calls++
			}
		}),
		weft.Tool("roll_dice", "Roll a die.",
			func(_ context.Context, _ struct{}) (int, error) { return 4, nil }),
	)
	if _, err := agt.Generate(context.Background(), weft.Prompt("Roll.")); err != nil {
		log.Fatal(err)
	}
	fmt.Println("tool calls:", calls)
}
Output:
tool calls: 1
Example (Async)

Slow observers (a database write, a network sink) must not run inside a Tap: taps are synchronous and run under the step's event-ordering lock, so a slow tap delays every tool event of its step. The pattern: hand each event to a queue inside the tap — never block — and drain it on your own goroutine, which may be as slow as it likes. A bounded queue with a visible drop counter keeps a stuck drain from wedging the run.

package main

import (
	"context"
	"fmt"
	"log"
	"sync/atomic"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	events := make(chan weft.Event, 1024)
	var dropped atomic.Int64
	done := make(chan int)
	go func() {
		starts := 0
		for ev := range events {
			if _, ok := ev.(weft.ToolStart); ok {
				starts++
			}
		}
		done <- starts
	}()

	agt := weft.New(
		wefttest.Script(
			wefttest.ToolCalls(wefttest.Call{Name: "ping"}),
			wefttest.Say("done"),
		),
		weft.Tap(func(_ context.Context, ev weft.Event) {
			select {
			case events <- ev: // fast: hand to the belt
			default: // full belt: drop, but visibly
				dropped.Add(1)
			}
		}),
		weft.Tool("ping", "", func(_ context.Context, _ struct{}) (string, error) {
			return "pong", nil
		}),
	)
	if _, err := agt.Generate(context.Background(), weft.Prompt("hi")); err != nil {
		log.Fatal(err)
	}
	close(events)
	fmt.Println("tool starts:", <-done, "dropped:", dropped.Load())
}
Output:
tool starts: 1 dropped: 0

func ToolSource

func ToolSource(fn func() []*ToolDef) Option

ToolSource replaces the tool set the loop advertises and dispatches against with the given function's return value, fetched fresh exactly once per step (and once per manual Agent.CallTool): the step's advertisement and its dispatch both resolve against that one snapshot, so what the model was shown is exactly what runs. The seam is for registries that change while the agent runs (plugins installed mid-run, MCP servers polled per step); a tool registered mid-step becomes callable on the next step's fetch. The Agent stays immutable: the source is a value; synchronization and freshness of the list belong to the source's owner. A snapshot with a duplicate name fails the run with ErrDuplicateTool and one with a nil entry with ErrNilTool (the runtime analogues of New's duplicate-name panic) rather than silently dropping the second tool or advertising a dereference. A nil function (the default) keeps the static construction-time list — Tool/option registration is then the only source of tools, byte-identical to an agent without a source. Manifest and Agent.Tools still report the static construction-time set: a manifest describes the code, not the registry behind a source.

func TracerProvider added in v0.2.0

func TracerProvider(tp trace.TracerProvider) Option

TracerProvider sets the OpenTelemetry tracer provider the agent's runs report spans to. Without it, runs use the global provider (otel.GetTracerProvider), which is a no-op until an SDK registers one — so a program that sets up an SDK gets weft's spans with no option at all. Tests and dependency-injected programs pass their own provider here instead of touching the global. Every run reports one invoke_agent span, one chat span per model call, and one execute_tool span per executed tool call, with the GenAI semantic attributes (ADR 0016); no message or tool-argument content is ever put on a span.

Example

ExampleTracerProvider prints the span tree of a two-step run, as an exporter would render it.

package main

import (
	"context"
	"fmt"
	"strings"
	"sync"
	"testing"

	"go.opentelemetry.io/otel/attribute"
	"go.opentelemetry.io/otel/codes"
	"go.opentelemetry.io/otel/trace"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

// recProvider records every span started through it: name, kind,
// parent, attributes, status, recorded exceptions, and end state.
type recProvider struct {
	trace.TracerProvider

	mu    sync.Mutex
	spans []*recSpan
	next  uint64
}

func newRecProvider() *recProvider { return &recProvider{} }

func (p *recProvider) Tracer(string, ...trace.TracerOption) trace.Tracer {
	return &recTracer{p: p}
}

// find returns the one span with the exact name, failing the test when
// it is missing or ambiguous — every assertion here is against a tree
// whose span names are unique per test, so start order never matters.
func (p *recProvider) find(t *testing.T, name string) *recSpan {
	t.Helper()
	p.mu.Lock()
	defer p.mu.Unlock()
	var found *recSpan
	for _, s := range p.spans {
		if s.name == name {
			if found != nil {
				t.Fatalf("span %q recorded more than once", name)
			}
			found = s
		}
	}
	if found == nil {
		var names []string
		for _, s := range p.spans {
			names = append(names, s.name)
		}
		t.Fatalf("span %q not recorded; have %v", name, names)
	}
	return found
}

type recTracer struct {
	trace.Tracer
	p *recProvider
}

func (t *recTracer) Start(ctx context.Context, name string, opts ...trace.SpanStartOption) (context.Context, trace.Span) {
	cfg := trace.NewSpanStartConfig(opts...)
	p := t.p
	p.mu.Lock()
	p.next++
	n := p.next
	parent := trace.SpanContextFromContext(ctx)
	traceID := parent.TraceID()
	if !parent.IsValid() {
		var b [16]byte
		b[14], b[15] = byte(n>>8), byte(n)
		traceID = trace.TraceID(b)
	}
	var sid [8]byte
	sid[7] = byte(n)
	sc := trace.NewSpanContext(trace.SpanContextConfig{
		TraceID:    traceID,
		SpanID:     trace.SpanID(sid),
		TraceFlags: trace.FlagsSampled,
	})
	s := &recSpan{name: name, kind: cfg.SpanKind(), sc: sc, parent: parent, recording: true}
	p.spans = append(p.spans, s)
	p.mu.Unlock()
	return trace.ContextWithSpan(ctx, s), s
}

type recSpan struct {
	trace.Span

	mu        sync.Mutex
	name      string
	kind      trace.SpanKind
	sc        trace.SpanContext
	parent    trace.SpanContext
	attrs     []attribute.KeyValue
	status    codes.Code
	statusMsg string
	events    []string
	ended     bool
	recording bool
}

func (s *recSpan) End(...trace.SpanEndOption) {
	s.mu.Lock()
	defer s.mu.Unlock()
	s.ended, s.recording = true, false
}

func (s *recSpan) SpanContext() trace.SpanContext { return s.sc }

func (s *recSpan) IsRecording() bool {
	s.mu.Lock()
	defer s.mu.Unlock()
	return s.recording
}

func (s *recSpan) SetAttributes(attrs ...attribute.KeyValue) {
	s.mu.Lock()
	defer s.mu.Unlock()
	s.attrs = append(s.attrs, attrs...)
}

func (s *recSpan) SetStatus(c codes.Code, msg string) {
	s.mu.Lock()
	defer s.mu.Unlock()
	s.status, s.statusMsg = c, msg
}

func (s *recSpan) RecordError(err error, _ ...trace.EventOption) {
	s.mu.Lock()
	defer s.mu.Unlock()
	s.events = append(s.events, "exception: "+err.Error())
}

func (s *recSpan) Name() string {
	s.mu.Lock()
	defer s.mu.Unlock()
	return s.name
}

// attrsMap flattens the span's attributes for exact-set assertions.
func (s *recSpan) attrsMap() map[string]string {
	s.mu.Lock()
	defer s.mu.Unlock()
	out := map[string]string{}
	for _, kv := range s.attrs {
		out[string(kv.Key)] = kv.Value.String()
	}
	return out
}

func (s *recSpan) state() (ended bool, status codes.Code, statusMsg string, events []string) {
	s.mu.Lock()
	defer s.mu.Unlock()
	return s.ended, s.status, s.statusMsg, append([]string(nil), s.events...)
}

// spanEcho is the standard tool of the span tests: one string in, one
// string out, nothing async.
var spanEcho = weft.Tool("echo", "Echo the message.", func(_ context.Context, in struct {
	Msg string `json:"msg"`
}) (string, error) {
	return "echo: " + in.Msg, nil
})

// spanScript is the standard two-step script: one echo call, then a
// final reply. wefttest fixes usage at 10 input / 5 output tokens per
// step, so the token attributes are assertable.
func spanScript() *wefttest.Model {
	return wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"msg":"hi"}`}),
		wefttest.Say("done"),
	)
}

func main() {
	tp := newRecProvider()
	agt := weft.New(spanScript(), weft.TracerProvider(tp), weft.Name("demo"), spanEcho)
	if _, err := agt.Generate(context.Background(), weft.Prompt("hi")); err != nil {
		return
	}
	tp.mu.Lock()
	bySpan := map[[8]byte]*recSpan{}
	for _, s := range tp.spans {
		bySpan[s.sc.SpanID()] = s
	}
	var print func(s *recSpan, depth int)
	print = func(s *recSpan, depth int) {
		_, status, _, _ := s.state()
		fmt.Printf("%s%s [%v]\n", strings.Repeat("  ", depth), s.name, status)
		for _, child := range tp.spans {
			if child.parent.SpanID() == s.sc.SpanID() {
				print(child, depth+1)
			}
		}
	}
	for _, s := range tp.spans {
		if !s.parent.IsValid() {
			print(s, 0)
		}
	}
	tp.mu.Unlock()
}
Output:
invoke_agent demo [Ok]
  chat script [Ok]
  execute_tool echo [Ok]
  chat script [Ok]

func UsageLimit added in v0.2.0

func UsageLimit(max Usage) Option

UsageLimit bounds a run's total token usage — the run's own model calls plus every subagent's (RunResult.Usage). A field left zero is unlimited. The limit is checked after each step, before the loop makes another model call: a step that ends the run — final answer, StopWhen, pending approvals — may overshoot and still succeed, because a budget's job is to stop further spend, not to discard finished work. Exceeding it fails the run with ErrUsageLimit and the partial transcript on RunError.Result. There is no default limit: the right value is workload-specific, and MaxSteps is the default budget. Usage.Total() is not a separate limit — a caller who wants a total sets both fields.

Example

A token budget: exceeded, the run fails with ErrUsageLimit before the next model call; the partial transcript rides on RunError.Result.

package main

import (
	"context"
	"errors"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	echo := weft.Tool("echo", "", func(_ context.Context, _ struct{}) (string, error) {
		return "ok", nil
	})
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "echo"}),
		wefttest.Say("never reached"),
	), echo, weft.UsageLimit(weft.Usage{OutputTokens: 4}))
	_, err := agt.Generate(context.Background(), weft.Prompt("again"))
	var re *weft.RunError
	if errors.As(err, &re) {
		fmt.Println(errors.Is(err, weft.ErrUsageLimit), "steps kept:", len(re.Result.Steps))
	}
}
Output:
true steps kept: 1

func WrapModel added in v0.2.0

func WrapModel(mw ...ModelMiddleware) Option

WrapModel installs model middleware around the agent's model. The first middleware listed is the outermost — WrapModel(a, b) calls a(b(model)) — and successive WrapModel options append inward. The chain is built once, at New, and sees each step's ModelRequest as the loop built it. The reference set is in package mw: Retry, Fallback, Log, RepairJSON. Nil entries are ignored.

Example

Model middleware wraps the agent's model, chi-style: the first listed is the outermost. The reference set lives in package mw.

package main

import (
	"context"
	"fmt"
	"iter"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	logged := func(next weft.Model) weft.Model {
		return loggingModel{next: next}
	}
	agt := weft.New(wefttest.Script(wefttest.Say("hello")), weft.WrapModel(logged))
	res, err := agt.Generate(context.Background(), weft.Prompt("hi"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
}

type loggingModel struct{ next weft.Model }

func (m loggingModel) Info() weft.ModelInfo { return weft.InfoOf(m.next) }

func (m loggingModel) Stream(ctx context.Context, req weft.ModelRequest) iter.Seq2[weft.ModelEvent, error] {
	fmt.Printf("model call: %d messages\n", len(req.Messages))
	return m.next.Stream(ctx, req)
}
Output:
model call: 1 messages
hello

type OutputDecoder added in v0.3.0

type OutputDecoder[Out any] struct {
	// contains filtered or unexported fields
}

OutputDecoder turns a streaming run's submit_output argument deltas into a filling-in Out — the UI half of structured output: a consumer rendering a form as the model writes it. It is a decoder value, not an event-stream wrapper: feed it the events you already consume and render what comes back.

dec := weft.NewOutputDecoder[Form]()
for ev, err := range run.Events() {
    if err != nil { return err }
    if p, ok := dec.Feed(ev); ok { render(p) }
}
form, err := dec.Result()

The decoder keys on the submit_output stream identity ToolArgsDelta carries — the tool name and step boundaries — resets its buffer on StepStart and on a fresh submit_output ToolStart, and closes it on the matching ToolFinish. Nested events are ignored: a subagent's structured output is its own decoder's job. Decode is lenient and prefix-shaped: after each delta it attempts the longest closed prefix of the arguments so far (see partial_json.go) and reports it when it changed and decoded — fields not yet present stay zero, a garbage mid-stream prefix keeps the last good partial, and errors surface only at Result, which follows the OutputOf rule (the last submit_output call with a non-error result) and returns ErrNoOutput when none finished. Never model-visible: the decoder reads the stream and writes nothing back.

Cost: every delta rescans the buffered arguments (the close is linear in what has arrived), so a submission's decode cost grows with the square of its size — the price of a fresh partial on every delta. At tool-argument scale (a few KiB) it is noise; a UI feeding very large submissions can trade freshness for linearity by calling Feed less often — the decoder keeps the last good partial across the deltas it skips.

Example

An OutputDecoder renders structured output while it streams: feed it the events you already consume and draw the filling-in form.

package main

import (
	"context"
	"encoding/json"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	type Form struct {
		Name  string `json:"name"`
		Count int    `json:"count"`
	}
	turn := wefttest.Raw(
		weft.ModelToolCallDelta{Index: 0, Name: "submit_output", Args: `{"name":"Ada",`},
		weft.ModelToolCallDelta{Index: 0, Name: "submit_output", Args: `"count":3}`},
		weft.ModelToolCall{ID: "c1", Name: "submit_output", Args: json.RawMessage(`{"name":"Ada","count":3}`)},
		weft.ModelFinish{Reason: weft.StopToolCalls, Usage: weft.Usage{InputTokens: 3, OutputTokens: 2}},
	)
	run := weft.New(wefttest.Script(turn), weft.Output[Form]()).Stream(context.Background(), weft.Prompt("fill the form"))
	dec := weft.NewOutputDecoder[Form]()
	for ev, err := range run.Events() {
		if err != nil {
			log.Fatal(err)
		}
		if p, ok := dec.Feed(ev); ok {
			fmt.Printf("render: name=%q count=%d\n", p.Name, p.Count)
		}
	}
	form, err := dec.Result()
	if err != nil {
		log.Fatal(err)
	}
	fmt.Printf("final:  name=%q count=%d\n", form.Name, form.Count)
}
Output:
render: name="Ada" count=0
render: name="Ada" count=3
final:  name="Ada" count=3

func NewOutputDecoder added in v0.3.0

func NewOutputDecoder[Out any]() *OutputDecoder[Out]

NewOutputDecoder returns a decoder ready to feed a run's events.

func (*OutputDecoder[Out]) Feed added in v0.3.0

func (d *OutputDecoder[Out]) Feed(ev Event) (partial Out, ok bool)

Feed consumes one run event. It ignores everything except this run's StepStart, the submit_output ToolStart (a whole-call arrival, and the marker of a fresh call), the submit_output ToolArgsDelta stream, and the submit_output ToolFinish; after an args delta it returns the best-effort partial Out and ok == true when the partial changed. A second submit_output call within the step starts the buffer over — Result follows the last call, the OutputOf rule; concatenating two distinct calls' arguments could only ever produce bytes that do not decode.

func (*OutputDecoder[Out]) Result added in v0.3.0

func (d *OutputDecoder[Out]) Result() (Out, error)

Result decodes the completed submission — the last submit_output call with a non-error result, the OutputOf rule — and returns ErrNoOutput when no valid call finished.

type ParamsOption added in v0.3.0

type ParamsOption interface {
	Option
	RunOption
}

ParamsOption is accepted by both New and Stream/Generate: sampling knobs are as per-question as a forced choice (the ThinkingOption shape).

func Params added in v0.3.0

func Params(p RequestParams) ParamsOption

Params sets the agent's default sampling knobs — Temperature, TopP, MaxTokens, Stop, Seed — for every run's model calls; as a RunOption it overrides that default for one run:

agt := weft.New(m, weft.Params(weft.RequestParams{Temperature: ptr(0.2)}))
agt.Generate(ctx, weft.Params(weft.RequestParams{Temperature: ptr(0.9)}), weft.Prompt(q))

A nil or empty field keeps the adapter's construction-time default for that knob (its Temperature option, and so on); a run-level Params replaces the agent's struct whole, never merging field by field — set every knob the override should carry. A PrepareStep function can edit the request's Params per step ("cold for classification steps, creative for drafting" is a two-line function). Adapters drop knobs their provider lacks, and the provider's own limits apply (google narrows Seed to int32).

Example

Params sets per-step sampling: a PrepareStep function turns the temperature down for the classifying step and back for drafting — one struct, edited per step, no second Model construction.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	p := func(f float64) *float64 { return &f }
	classify := weft.Tool("classify", "Classify the request.",
		func(_ context.Context, _ struct{}) (string, error) { return "billing", nil })
	agt := weft.New(
		wefttest.Script(
			wefttest.ToolCalls(wefttest.Call{Name: "classify"}),
			wefttest.Say("A billing question, answered at temperature 0.9."),
		),
		weft.Params(weft.RequestParams{Temperature: p(0.9)}),
		weft.PrepareStep(func(_ context.Context, step int, req weft.ModelRequest) (weft.ModelRequest, error) {
			if step == 0 {
				req.Params.Temperature = p(0) // cold for classification
			}
			return req, nil
		}),
		classify,
	)
	res, err := agt.Generate(context.Background(), weft.Prompt("Why did my invoice double?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
}
Output:
A billing question, answered at temperature 0.9.

type Part

type Part interface {
	// contains filtered or unexported methods
}

Part is one content part of a Message. The set of part types is closed: text, tool calls, tool results, reasoning, and files today; approval parts are planned additions that will join this interface.

type PolicyOption added in v0.2.0

type PolicyOption interface {
	Option
	ToolOption
}

PolicyOption is accepted by both New and Tool. On an agent it is the default policy for every tool call; on a tool it overrides the agent's default for that tool. The manifest records both levels.

func MaxResultBytes

func MaxResultBytes(n int) PolicyOption

MaxResultBytes sets the maximum size of one tool result's text, in bytes (default 64 KiB). The loop caps longer results — successes, failures, and panics alike — cutting on a rune boundary and appending a marker the model can see, so it knows the output is partial. MaxResultBytes(0) removes the cap; negative values are ignored. On a tool it overrides the agent's cap for that tool alone — a per-tool MaxResultBytes(0) lifts the cap for a tool whose output must arrive whole. Agent.CallTool returns uncapped output: the cap is a run policy, applied by the loop.

func Sequential

func Sequential() PolicyOption

Sequential restricts a step's tool calls to run one at a time, in call order: each tool finishes before the next starts — the safe setting for tools with shared state. Under any parallelism, tools *start* in call order; Sequential additionally serializes their execution. It also sets ModelRequest.SequentialTools, so adapters ask the provider not to emit parallel batches in the first place.

On a tool, Sequential is a barrier: the dispatcher lets every in-flight call of the step finish, runs this call alone, then resumes the step's parallelism for the calls after it. Other tools keep running in parallel; only this one is serialized.

Example

A Sequential tool is a barrier: it runs alone. The step's calls in flight finish first, and the calls after it wait — under any parallelism. Results stay in call order either way.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	write := weft.Tool("write", "Append to the ledger.", func(_ context.Context, _ struct{}) (string, error) {
		return "written", nil
	}, weft.Sequential())
	check := weft.Tool("check", "Verify the ledger.", func(_ context.Context, _ struct{}) (string, error) {
		return "verified", nil
	})
	model := wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "check"}, wefttest.Call{Name: "write"}, wefttest.Call{Name: "check"}),
		wefttest.Say("done"),
	)
	res, err := weft.New(model, write, check, weft.Parallelism(4)).Generate(context.Background(), weft.Prompt("x"))
	if err != nil {
		log.Fatal(err)
	}
	for _, r := range res.Steps[0].Results {
		fmt.Println(r.Name, "->", r.Content)
	}
}
Output:
check -> verified
write -> written
check -> verified

func StrictInput added in v0.2.0

func StrictInput() PolicyOption

StrictInput rejects tool arguments that carry fields the input struct does not declare, as an ErrInvalidToolInput result naming the field. The default is lenient — undeclared fields are ignored, the encoding/json default — because models routinely add stray keys and a rejection costs a round trip, not accuracy. Use it where an ignored field would be a silent misinterpretation of the call. On the agent it applies to every tool; on a tool to that tool alone. RawTool handlers are unaffected: they receive the raw arguments.

func Timeout added in v0.2.0

func Timeout(d time.Duration) PolicyOption

Timeout bounds one tool call. On a tool it is that tool's deadline — Timeout(0) on a tool removes the agent's default for that tool alone, the timeout analogue of MaxResultBytes(0); on the agent it is the default for every tool without its own, and non-positive values are ignored there. The handler's ctx carries the deadline; when it expires the loop records an error result ("tool X timed out after 10s") the model sees and moves on, abandoning the handler's goroutine — handlers must honour ctx to release their resources. Without any Timeout only the run's ctx bounds a call.

Example

Per-tool policy: trailing options on Tool override the agent's defaults for that tool alone. Here a slow tool times out into an error result the model sees, and the run carries on.

package main

import (
	"context"
	"fmt"
	"log"
	"time"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	slow := weft.Tool("slow", "Takes a while.",
		func(ctx context.Context, _ struct{}) (string, error) {
			<-ctx.Done() // a well-behaved handler honours the deadline
			return "", ctx.Err()
		},
		weft.Timeout(10*time.Millisecond))
	model := wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "slow"}),
		wefttest.Say("It did not answer in time."),
	)
	agt := weft.New(model, slow)

	res, err := agt.Generate(context.Background(), weft.Prompt("Try the slow tool."))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
	fmt.Println(res.Text())
}
Output:
tool "slow" timed out after 10ms
It did not answer in time.

func WrapTools added in v0.2.0

func WrapTools(mw ...ToolMiddleware) PolicyOption

WrapTools installs tool middleware. On the agent it wraps every tool call the loop (and Agent.CallTool) dispatches; on a tool it wraps that tool alone, inside the agent's chain: agent middleware → tool middleware → decode → handler. The first middleware listed is the outermost, so WrapTools(a, b) runs a(b(call)), and successive WrapTools options append inward. Middleware sees CallFromContext and may decorate ctx for the handler (see ExampleWrapTools_context); it returns the result text or an error — a *ToolError to give the model a code, an error wrapping ErrApprovalRequired to park the call on RunResult.Pending. The reference set is in package mw: Allow, Audit, MapErrors. Nil entries are ignored.

Example

Tool middleware wraps every call the loop dispatches. A middleware that returns an error produces an error result the model sees; a *weft.ToolError gives it a code.

package main

import (
	"context"
	"fmt"
	"log"
	"strings"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	readOnly := func(next weft.ToolCaller) weft.ToolCaller {
		return func(ctx context.Context, call weft.ToolCallPart) (string, error) {
			if strings.HasPrefix(call.Name, "delete_") {
				return "", &weft.ToolError{Code: "DENIED", Message: "this agent is read-only"}
			}
			return next(ctx, call)
		}
	}
	del := weft.Tool("delete_order", "", func(_ context.Context, _ struct{}) (string, error) { return "deleted", nil })
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "delete_order"}),
		wefttest.Say("I cannot do that."),
	), del, weft.WrapTools(readOnly))
	res, err := agt.Generate(context.Background(), weft.Prompt("delete order 1"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
	fmt.Println(res.Text())
}
Output:
DENIED: this agent is read-only
I cannot do that.
Example (Context)

The context-decoration convention: middleware that has verified something (here, who is calling) adds it to ctx before next, and the handler reads it back through a typed accessor — the same shape as weft.CallFromContext. Decorate only with data the middleware has verified; derive business values in the handler after decode.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	authUser := func(next weft.ToolCaller) weft.ToolCaller {
		return func(ctx context.Context, call weft.ToolCallPart) (string, error) {
			user, err := verifyUser(ctx) // a session lookup, a token check, ...
			if err != nil {
				return "", &weft.ToolError{Code: "UNAUTHENTICATED", Message: "sign in first", Err: err}
			}
			return next(withUser(ctx, user), call)
		}
	}
	myOrders := weft.Tool("my_orders", "List the caller's orders.",
		func(ctx context.Context, _ struct{}) (string, error) {
			u, ok := userFromContext(ctx)
			if !ok {
				return "", weft.Errorf("UNAUTHENTICATED", "no user on the call")
			}
			return "orders for " + u.Name, nil
		})
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "my_orders"}),
		wefttest.Say("Here they are."),
	), myOrders, weft.WrapTools(authUser))
	res, err := agt.Generate(context.Background(), weft.Prompt("show my orders"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
}

type exampleUser struct{ Name string }

type userKey struct{}

func withUser(ctx context.Context, u exampleUser) context.Context {
	return context.WithValue(ctx, userKey{}, u)
}

func userFromContext(ctx context.Context) (exampleUser, bool) {
	u, ok := ctx.Value(userKey{}).(exampleUser)
	return u, ok
}

func verifyUser(context.Context) (exampleUser, error) { return exampleUser{Name: "ada"}, nil }
Output:
orders for ada

type ReasoningDelta

type ReasoningDelta struct {
	RunID string `json:"run_id"`
	Text  string `json:"text"`
}

ReasoningDelta is an increment of provider reasoning, in the order the model produced it relative to TextDelta. Signatures are not streamed; they are on the ReasoningPart of the transcript.

func (ReasoningDelta) MarshalJSON

func (e ReasoningDelta) MarshalJSON() ([]byte, error)

MarshalJSON encodes the event with its "type" discriminator.

type ReasoningPart

type ReasoningPart struct {
	Text      string `json:"text"`
	Signature string `json:"signature,omitempty"`
}

ReasoningPart is one provider reasoning block surfaced by providers that expose it (Anthropic thinking blocks, Gemini thought parts). Weft preserves it in the transcript but does not act on it. Signature is the provider's opaque token for the block (Anthropic rejects thinking sent back without its signature); adapters echo it unchanged. A step yields one ReasoningPart per provider block — a delta carrying a signature closes the block (see ModelReasoningDelta) — placed before the TextPart of the same assistant message, in the order the model produced them.

Example
package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	model := wefttest.Script(
		wefttest.Think("The user greets; reply in kind.", wefttest.Say("Hello!")),
	)
	res, err := weft.New(model).Generate(context.Background(), weft.Prompt("Hi."))
	if err != nil {
		log.Fatal(err)
	}
	for _, p := range res.Messages[1].Content {
		fmt.Printf("%T\n", p)
	}
}
Output:
weft.ReasoningPart
weft.TextPart

func (ReasoningPart) MarshalJSON

func (p ReasoningPart) MarshalJSON() ([]byte, error)

MarshalJSON encodes the part with its "type" discriminator.

type RequestParams added in v0.3.0

type RequestParams struct {
	Temperature *float64
	TopP        *float64
	MaxTokens   *int
	Stop        []string
	Seed        *int64
}

RequestParams is per-step sampling: the knobs a caller turns between "cold for classification, creative for drafting". Every field is a pointer or slice so the three states stay distinguishable — nil or empty keeps the adapter's construction default (its Temperature, TopP, MaxTokens, Stop, or Seed option), set overrides it for this request alone, and the adapter never replaces a construction value with a zero. A set pointer to 0 is a value (Temperature of exactly 0 is sent), with one provider exception: an anthropic MaxTokens of 0 falls to the adapter's default, because the API requires a positive value. A negative MaxTokens fails the run at the step that carries it — no provider accepts one, and the loop names the bug rather than letting each adapter improvise. A run-level Params option replaces the agent's struct whole, it does not merge field by field; PrepareStep can edit it per step. Adapters drop knobs their provider lacks (Seed on anthropic) under the "adapters document what they drop" rule.

type Role

type Role string

Role is the author of a Message.

There is no system role: the system instruction is agent-level (Instructions) and travels on ModelRequest.System, so a transcript never carries it and adapters never have to merge it.

const (
	RoleUser      Role = "user"
	RoleAssistant Role = "assistant"
	// RoleTool carries the results of one step's tool calls back to the
	// model, one ToolResultPart per call.
	RoleTool Role = "tool"
)

type Run

type Run struct {
	// contains filtered or unexported fields
}

Run is a handle to one streaming execution. Create it with Agent.Stream, then either range over Events (exactly once) or call Wait, which runs the agent to completion and reports the final result.

func (*Run) Close added in v0.2.0

func (r *Run) Close()

Close releases the run's resources, canceling it if still running. It is for abandoned runs: a run you will consume needs no Close — Events and Wait release everything themselves. Safe to call any number of times, before or after consumption.

func (*Run) Events

func (r *Run) Events() iter.Seq2[Event, error]

Events returns the run's event stream. It is single-use; a second call yields only ErrRunConsumed.

Events arrive in emission order (see ToolStart for the concurrent-tool ordering rule). Delivery is a direct hand-off over an unbuffered channel, and tool events are emitted while holding the step's event-ordering lock — so a slow consumer does not merely receive late: it delays event emission and gates the start of the step's subsequent tools. For latency-sensitive parallel tools, consume promptly (range over Events in a dedicated goroutine that buffers) or use Generate, which needs no consumer.

A failed run delivers its error exactly once as the final element; a successful run ends with RunFinish. Breaking out of the range cancels the run.

func (*Run) ID

func (r *Run) ID() string

ID returns the run's identifier. It is fixed at Stream time, so it can be logged or handed to a client before the first event is consumed.

func (*Run) Wait

func (r *Run) Wait() (*RunResult, error)

Wait blocks until the run finishes and returns its result. If Events has not been consumed, Wait runs the agent itself, discarding events; it is safe to call after ranging over Events, or from another goroutine while ranging over them.

type RunError

type RunError struct {
	Step   int
	Err    error
	Result *RunResult
}

RunError reports a step-scoped failure: the model stream failed, the context was canceled, or the step budget ran out. Err is the cause — use errors.Is/As on it. Result carries the transcript up to the failure, so partial work is never lost.

func (*RunError) Error

func (e *RunError) Error() string

func (*RunError) Unwrap

func (e *RunError) Unwrap() error

type RunFinish

type RunFinish struct {
	RunID string `json:"run_id"`
	Usage Usage  `json:"usage"`
	Steps int    `json:"steps"`
	// Pending mirrors RunResult.Pending: calls the run ended on without
	// executing, awaiting Approve/Deny. Such a call had its ToolStart
	// and no ToolFinish.
	Pending []ToolCallPart `json:"pending,omitempty"`
}

RunFinish is always the final event of a successful run and carries the run's total usage and step count.

func (RunFinish) MarshalJSON

func (e RunFinish) MarshalJSON() ([]byte, error)

MarshalJSON encodes the event with its "type" discriminator.

type RunOption

type RunOption interface {
	// contains filtered or unexported methods
}

RunOption configures a single run.

func Approve added in v0.2.0

func Approve(callID string) RunOption

Approve resumes a call left on RunResult.Pending by an earlier run: pass the earlier transcript with Messages and the decision, and the loop executes the call — through the ordinary tool chain, with Call.Approved set — before its next model call. Ids that are not pending are ignored.

func Deny added in v0.2.0

func Deny(callID, reason string) RunOption

Deny resolves a pending call without running it: the model sees an error result reading "DENIED: <reason>" and the loop continues. Pending calls given neither Approve nor Deny are denied with the reason "no decision" ("DENIED: no decision").

func Messages

func Messages(msgs ...Message) RunOption

Messages adds existing messages (a session transcript, few-shot examples) to the run's input.

func Prompt

func Prompt(text string) RunOption

Prompt adds a user message to the run's input.

func Resolve added in v0.3.0

func Resolve(callID, content string) RunOption

Resolve resumes a pending call with a result computed outside the process — the human-as-tool-executor shape: run the query in prod, paste what happened, and the next model call sees it. The content becomes the call's ToolResultPart verbatim; the handler never runs, no execute_tool span or ToolStart/ToolFinish is emitted (nothing executed), and Call.Approved is never set. MaxResultBytes applies as to any result. Composes with Approve and Deny in one resuming call; the last option for an id wins.

Resolve on a call that is not pending in the resumed transcript is a loud run error at step 0 — a deliberate asymmetry with Approve and Deny, which ignore unknown ids: those are yes/no marks over an id set, while Resolve carries a payload the caller expects the model to see, and dropping it silently is the one thing the error model forbids (ADR 0007's 2026-09-22 amendment).

func ResolveError added in v0.3.0

func ResolveError(callID, content string) RunOption

ResolveError is Resolve with the result marked as an error: the model sees the content on an error result, the shape a failed execution would have produced.

func RunID

func RunID(id string) RunOption

RunID sets the run's identifier instead of generating one — for replays, idempotent retries, and correlating with an outer system's own ids. Empty values are ignored.

type RunResult

type RunResult struct {
	ID string
	// StopReason is the last step's finish reason. StopMaxTokens here
	// means the final reply was cut off by the output-token limit — the
	// run still succeeds, and the caller decides what truncated text
	// means.
	StopReason StopReason
	Messages   []Message
	Steps      []StepRecord
	Usage      Usage
	// Pending lists the tool calls of the last step that await an
	// approval decision (RequireApproval, or middleware returning
	// ErrApprovalRequired). The run ended successfully without running
	// them and the transcript carries no result for them; resume with
	// Messages(res.Messages...) plus Approve/Deny per call.
	Pending []ToolCallPart
}

RunResult is the outcome of a completed run: its id, the full transcript (including the input messages), one record per step, and summed usage.

func GenerateAs added in v0.2.0

func GenerateAs[Out any](ctx context.Context, a *Agent, opts ...RunOption) (Out, *RunResult, error)

GenerateAs runs the agent and returns its structured output, decoded from the last valid submit_output call. The agent must have been built with Output[Out]; a run that ends without a valid submission returns ErrNoOutput alongside the result, so the transcript is still inspectable. Run errors are *RunError as for Generate.

Example

Structured output: Output constrains the final answer to a struct, and GenerateAs returns it decoded. An invalid submission is an ordinary tool error the model repairs; a valid one ends the run.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	type Verdict struct {
		Approved bool   `json:"approved"`
		Reason   string `json:"reason" jsonschema:"one sentence"`
	}
	model := wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "submit_output", Args: `{"approved":true,"reason":"within policy"}`}),
	)
	agt := weft.New(model, weft.Instructions("Review refund requests."), weft.Output[Verdict]())

	v, res, err := weft.GenerateAs[Verdict](context.Background(), agt, weft.Prompt("Refund order 42?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(v.Approved, v.Reason)
	fmt.Println("steps:", res.NumSteps())
}
Output:
true within policy
steps: 1

func (*RunResult) NumSteps

func (r *RunResult) NumSteps() int

NumSteps returns how many model calls the run made.

func (*RunResult) Text

func (r *RunResult) Text() string

Text returns the final assistant text — the text parts of the last assistant message.

type RunStart

type RunStart struct {
	ID    string    `json:"id"`
	Model ModelInfo `json:"model"`
	Agent string    `json:"agent,omitempty"`
}

RunStart is always the first event of a run and carries its id and, when reported, the model's identity and the agent's name.

func (RunStart) MarshalJSON

func (e RunStart) MarshalJSON() ([]byte, error)

MarshalJSON encodes the event with its "type" discriminator.

type Schema

type Schema struct {
	Type        string             `json:"type,omitempty"`
	Format      string             `json:"format,omitempty"`
	Description string             `json:"description,omitempty"`
	Properties  map[string]*Schema `json:"properties,omitempty"`
	// AdditionalProperties types a map's values (JSON Schema draft
	// 2020-12), so map[string]int stops being "some object". The
	// boolean false form is not expressible: reflection always has a
	// value type, and hand-written RawTool schemas that need to lock
	// properties down must say so in Description.
	AdditionalProperties *Schema  `json:"additionalProperties,omitempty"`
	Required             []string `json:"required,omitempty"`
	Items                *Schema  `json:"items,omitempty"`
	// contains filtered or unexported fields
}

Schema is the subset of JSON Schema (draft 2020-12) that weft derives from tool input structs. Its shape tracks what the official Go MCP SDK derives via google/jsonschema-go, so tools cross over to MCP without conversion.

func ParseSchema added in v0.2.0

func ParseSchema(b json.RawMessage) (*Schema, error)

ParseSchema reads a JSON Schema document from outside Go into a *Schema: the structured fields it knows are populated (Type, Properties, Required, … — the manifest and any reader that walks the tree see them), and the document's own bytes are kept and re-emitted by Schema.MarshalJSON, so an enum or a oneOf the Schema type cannot express still reaches the model exactly as written. The top-level type must be an object — providers and MCP both require it, and failing here, at import, beats failing at the first model call.

The structured view is lenient: a keyword whose shape the Schema type cannot hold — a boolean additionalProperties, a type array such as ["string","null"], tuple or boolean items, a non-string description — leaves that field zero (an unconstrained node) and is not an error, because the bytes carry it whole and the view is for readers, not the model. Only the document itself is checked: invalid JSON, trailing data, or a non-object top level return an error naming the problem.

func (*Schema) MarshalJSON added in v0.2.0

func (s *Schema) MarshalJSON() ([]byte, error)

MarshalJSON emits the schema's parsed bytes verbatim when it came from ParseSchema — the manifest's input_schema, the adapters' tool parameters, and any other reader see exactly what the foreign server sent — and the plain struct encoding otherwise (a reflected schema has no bytes to honour, so existing goldens are unchanged).

type StepFinish

type StepFinish struct {
	RunID  string     `json:"run_id"`
	Index  int        `json:"index"`
	Reason StopReason `json:"reason"`
	Usage  Usage      `json:"usage"`
	Raw    string     `json:"raw,omitempty"`
}

StepFinish reports that step Index is complete: the model call finished and, when the step requested tools, they have run — it follows the step's ToolFinish events and precedes the stop-condition check (docs/life-of-a-call.md). Raw is the provider's own stop reason when Reason was approximated (see ModelFinish.Raw); empty when the mapping was exact.

func (StepFinish) MarshalJSON

func (e StepFinish) MarshalJSON() ([]byte, error)

MarshalJSON encodes the event with its "type" discriminator.

type StepRecord

type StepRecord struct {
	Index int
	// StopReason is the mapped reason (stop, tool_calls, max_tokens);
	// RawStopReason is the provider's own value when the mapping was
	// approximated ("refusal", "content_filter", ...) — see
	// ModelFinish.Raw.
	StopReason    StopReason
	RawStopReason string
	Usage         Usage
	Text          string
	ToolCalls     []ToolCallPart
	Results       []ToolResultPart
	// SubagentUsage is the usage of each child run this step started,
	// keyed by the parent's call id — including a failed child's partial
	// usage. It is already included in RunResult.Usage; Usage above is
	// the step's own model call only. Nil when the step ran no subagent.
	SubagentUsage map[string]Usage
}

StepRecord captures everything one model step produced: its text, the tool calls it requested, and the results of executing them in call order.

type StepStart

type StepStart struct {
	RunID string `json:"run_id"`
	Index int    `json:"index"`
}

StepStart reports that the model is being called for step Index.

func (StepStart) MarshalJSON

func (e StepStart) MarshalJSON() ([]byte, error)

MarshalJSON encodes the event with its "type" discriminator.

type StopCondition

type StopCondition interface {
	Stop(steps []StepRecord) bool
}

StopCondition decides, after a step's tool calls have run, whether the run is complete. It sees every step so far; the last element is the step just finished. Returning true ends the run successfully without another model call. The built-ins — HasToolCall, StepCountIs — also implement fmt.Stringer so the manifest (TODO §2.9) can name them; adapt an ordinary function with StopFunc.

func HasToolCall

func HasToolCall(names ...string) StopCondition

HasToolCall stops the run once the step just finished called any of the named tools — the "final answer tool" pattern.

func StepCountIs

func StepCountIs(n int) StopCondition

StepCountIs stops the run after exactly n steps, successfully — unlike MaxSteps, which treats reaching the budget as a failure.

type StopFunc

type StopFunc func(steps []StepRecord) bool

StopFunc adapts an ordinary function to a StopCondition, the http.HandlerFunc shape.

func (StopFunc) Stop

func (f StopFunc) Stop(steps []StepRecord) bool

Stop ends the run when f says so.

type StopReason

type StopReason string

StopReason is why a model step ended.

const (
	StopEndTurn   StopReason = "stop"
	StopToolCalls StopReason = "tool_calls"
	StopMaxTokens StopReason = "max_tokens"
)

type TextDelta

type TextDelta struct {
	RunID string `json:"run_id"`
	Text  string `json:"text"`
}

TextDelta is an increment of assistant text.

func (TextDelta) MarshalJSON

func (e TextDelta) MarshalJSON() ([]byte, error)

MarshalJSON encodes the event with its "type" discriminator.

type TextPart

type TextPart struct {
	Text string `json:"text"`
}

TextPart is a span of user or assistant text.

func (TextPart) MarshalJSON

func (p TextPart) MarshalJSON() ([]byte, error)

MarshalJSON encodes the part with its "type" discriminator.

type ThinkingConfig added in v0.2.0

type ThinkingConfig struct {
	Level  ThinkingLevel
	Budget int64
}

ThinkingConfig is the per-run reasoning request: a level on the neutral scale, plus a token budget for providers whose depth control is a cap (Anthropic budget_tokens, Gemini thinkingBudget). A level without a Budget leaves the depth to the provider (Anthropic adaptive thinking, Gemini's own level mapping); a Budget without a level pins it. Off wins over a Budget when both are set.

type ThinkingLevel added in v0.2.0

type ThinkingLevel int

ThinkingLevel is a provider-neutral reasoning-effort scale. The zero value, ThinkUnset, sends nothing and keeps the provider default; every other value asks the adapter to express that depth on the wire in whatever form the provider has — reasoning_effort, a thinking object, a token budget. Adapters map what the provider can express and document what they drop (TODO §5.14).

const (
	ThinkUnset ThinkingLevel = iota // provider default; nothing is sent
	ThinkOff                        // suppress reasoning where allowed
	ThinkLow
	ThinkMedium
	ThinkHigh
)

type ThinkingOption added in v0.2.0

type ThinkingOption interface {
	Option
	RunOption
}

ThinkingOption is accepted by both New and Stream/Generate: reasoning depth is a per-question concern, not a per-agent one. On an agent it is the default for every run; on a run it overrides that default.

func Thinking added in v0.2.0

func Thinking(cfg ThinkingConfig) ThinkingOption

Thinking sets the reasoning level for the agent's model calls. As an Option it is every run's default; as a RunOption it overrides that default for one run — the quick-ask shape: fast by default, think on demand, without rebuilding the agent.

agt := weft.New(m, weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkOff}))
agt.Generate(ctx, weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkHigh}), weft.Prompt(q))

The zero Level keeps the provider default; adapters map what the provider can express and document what they drop.

Example

Fast by default, think on demand: the agent option sets every run's default, a run option overrides it for that run alone. The scripted model records what each run asked for.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	model := wefttest.Script(wefttest.Say("ok"), wefttest.Say("ok"))
	agt := weft.New(model, weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkOff}))
	deep := weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkHigh, Budget: 2048})
	if _, err := agt.Generate(context.Background(), deep, weft.Prompt("hard")); err != nil {
		log.Fatal(err)
	}
	if _, err := agt.Generate(context.Background(), weft.Prompt("quick")); err != nil {
		log.Fatal(err)
	}
	for _, req := range model.Requests() {
		fmt.Println(req.Thinking == weft.ThinkingConfig{Level: weft.ThinkHigh, Budget: 2048},
			req.Thinking.Level == weft.ThinkOff)
	}
}
Output:
true false
false true

type ToolArgsDelta added in v0.2.0

type ToolArgsDelta struct {
	RunID string `json:"run_id"`
	Name  string `json:"name"`
	Args  string `json:"args"`
}

ToolArgsDelta reports an increment of a tool call's arguments as the model streams them — the model is "writing" the call, which can take a while for large arguments (generated code, long documents). It is progress only: the call has not been made, and ToolStart still arrives when it executes. Name is the best-known name so far; a provider that streams fragments of several calls interleaves their deltas, distinguished by name where the provider supplies one.

func (ToolArgsDelta) MarshalJSON added in v0.2.0

func (e ToolArgsDelta) MarshalJSON() ([]byte, error)

MarshalJSON encodes the event with its "type" discriminator.

type ToolCallPart

type ToolCallPart struct {
	ID        string          `json:"id"`
	Name      string          `json:"name"`
	Args      json.RawMessage `json:"args"`
	Signature string          `json:"signature,omitempty"`
}

ToolCallPart is a tool invocation requested by the model. Args is the raw JSON the model produced; the loop unmarshals it into the tool's input type before invoking the handler. Signature is the provider's opaque token attached to the call itself (Gemini's thought signatures ride functionCall parts and must return on the same part); empty for providers without one.

func (ToolCallPart) MarshalJSON

func (p ToolCallPart) MarshalJSON() ([]byte, error)

MarshalJSON encodes the part with its "type" discriminator.

type ToolCaller added in v0.2.0

type ToolCaller func(ctx context.Context, call ToolCallPart) (string, error)

ToolCaller is one link of the tool-call chain: it takes a call and returns the result text the model will see, or an error. The innermost caller decodes the arguments and runs the handler; every ToolMiddleware wraps one.

type ToolChoiceConfig added in v0.3.0

type ToolChoiceConfig struct {
	Mode ToolChoiceMode
	Name string
}

ToolChoiceConfig constrains what a step's model call may emit. The zero value is the provider default. Name is required when Mode is ToolChoiceNamed and must be empty under every other mode; the loop fails the run on a mismatch rather than sending a malformed choice (a programming error, not a sentinel condition).

type ToolChoiceMode added in v0.3.0

type ToolChoiceMode string

ToolChoiceMode selects how the provider must shape a step's tool calls. The zero value, ToolChoiceAuto, keeps the provider default and sends nothing; every other value asks the adapter to express the constraint in the provider's own tool_choice form.

const (
	ToolChoiceAuto  ToolChoiceMode = ""     // provider default; nothing is sent
	ToolChoiceAny   ToolChoiceMode = "any"  // some tool must be called
	ToolChoiceNamed ToolChoiceMode = "tool" // the tool named by Name must be called
	// ToolChoiceNone forbids tool calls while keeping the catalogue
	// advertised. It exists for the prompt-cache interplay: removing
	// tools from the request to stop the model calling them invalidates
	// the cached prefix (ADR 0013's 2026-09-22 amendment), while none
	// keeps the bytes and forbids the calls.
	ToolChoiceNone ToolChoiceMode = "none"
)

type ToolChoiceOption added in v0.3.0

type ToolChoiceOption interface {
	Option
	RunOption
}

ToolChoiceOption is accepted by both New and Stream/Generate: which tool calls a step must make is a per-question concern as much as a per-agent one (the ThinkingOption shape).

func ToolChoice added in v0.3.0

func ToolChoice(cfg ToolChoiceConfig) ToolChoiceOption

ToolChoice forces the agent's model calls to include (or forbear from) tool calls. As an Option it is every run's default; as a RunOption it overrides that default for one run:

// a router: the first step must call classify
agt := weft.New(m, classify, weft.ToolChoice(weft.ToolChoiceConfig{Mode: weft.ToolChoiceNamed, Name: "classify"}))
// a final step that must answer in text, cache prefix intact
agt.Generate(ctx, weft.ToolChoice(weft.ToolChoiceConfig{Mode: weft.ToolChoiceNone}), weft.Prompt(q))

Modes: Any (some tool must be called), Named (Name must be — the router and eval-harness shape), None (no call may be made; the catalogue stays advertised, so a prompt-cache prefix on the tool definitions survives), Auto (the zero value: provider default, nothing sent). A PrepareStep function can rewrite the request's ToolChoice per step — force classify on step 0, then auto — the same way it rewrites tools. A non-auto choice the step's tool snapshot cannot satisfy (no tools, a name not advertised, a name under another mode) fails the run with a descriptive error.

Example

ToolChoice forces a step's tool calls — the router shape: classify must be the first call, the rest of the run is unconstrained. One PrepareStep function rewrites the request's ToolChoice per step; the agent-level option would force every step instead.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	classify := weft.Tool("classify", "Classify the request.",
		func(_ context.Context, _ struct{}) (string, error) { return "billing", nil })
	agt := weft.New(
		wefttest.Script(
			wefttest.ToolCalls(wefttest.Call{Name: "classify"}),
			wefttest.Say("This is a billing question."),
		),
		weft.PrepareStep(func(_ context.Context, step int, req weft.ModelRequest) (weft.ModelRequest, error) {
			if step == 0 {
				req.ToolChoice = weft.ToolChoiceConfig{Mode: weft.ToolChoiceNamed, Name: "classify"}
			}
			return req, nil
		}),
		classify,
	)
	res, err := agt.Generate(context.Background(), weft.Prompt("Why did my invoice double?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
}
Output:
This is a billing question.

type ToolDef

type ToolDef struct {
	Name         string  `json:"name"`
	Description  string  `json:"description,omitempty"`
	InputSchema  *Schema `json:"input_schema,omitempty"`
	OutputSchema *Schema `json:"output_schema,omitempty"`
	// contains filtered or unexported fields
}

ToolDef is a named, schema-described tool the model can call. Build one with the generic Tool constructor; the agent loop (and any manual dispatcher) executes it through Invoke.

A ToolDef is mutable until it is registered with New and frozen thereafter: New keeps a deep copy, so mutating the value that was passed in — or a copy returned by Agent.Tools — never reaches the agent or a running run.

func RawTool

func RawTool(name, description string, schema *Schema, fn func(ctx context.Context, args json.RawMessage) (string, error), opts ...ToolOption) *ToolDef

RawTool defines a tool from an explicit schema instead of reflection — for tools defined outside Go source: plugin manifests, MCP remotes, gateways. fn receives the model's raw JSON arguments verbatim: no unmarshalling, no ErrInvalidToolInput — validation belongs to fn, and StrictInput has no effect. A nil schema becomes the empty object schema, so the tool accepts any object input. Panics match Tool: empty name, nil fn. The manifest records no source line (the definition is not in Go source). Trailing options set per-tool policy as for Tool.

func Subagent added in v0.2.0

func Subagent(name, description string, child *Agent, opts ...ToolOption) *ToolDef

Subagent defines a tool that delegates to another agent. The model calls it with one argument, prompt; the child runs on a fresh transcript holding only that prompt, on the parent call's context, and the tool's result is the child's final text (or, for a child built with Output, the submitted JSON). The child's events arrive in the parent's stream wrapped in Nested; its usage is added to the parent's RunResult.Usage and recorded on StepRecord.SubagentUsage.

A subagent is an ordinary tool: Timeout bounds the child run, MaxResultBytes caps its answer, RequireApproval gates the delegation, Sequential makes it a barrier, WrapTools wraps the delegation once (the child's own seams govern inside it), and the manifest lists it.

A child run that fails is a tool error the parent model sees (SUBAGENT_FAILED), never a parent run error; a child that ends awaiting approval is SUBAGENT_PENDING; a child that is already running above this call is refused with SUBAGENT_CYCLE. Subagent panics if child is nil.

Example

Delegating to another agent: a subagent is a tool whose handler runs another agent on the prompt alone. The child's events arrive wrapped in Nested; its usage rolls into the parent's total.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	researcher := weft.New(wefttest.Script(
		wefttest.Say("order 1234 shipped yesterday"),
	))
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "research", Args: `{"prompt":"where is order 1234?"}`}),
		wefttest.Say("Researched."),
	), weft.Subagent("research", "Research a question in depth.", researcher))
	res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 1234?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
	fmt.Println("total tokens:", res.Usage.Total())
}
Output:
order 1234 shipped yesterday
total tokens: 45

func Tool

func Tool[In, Out any](name, description string, fn func(ctx context.Context, in In) (Out, error), opts ...ToolOption) *ToolDef

Tool defines a tool from a plain function. In and Out are inferred from the handler and the input schema is reflected from In's struct tags, so the compiler checks the handler's shape and nothing is written twice:

type RefundInput struct {
    OrderID string `json:"order_id" jsonschema:"the order to refund"`
    Reason  string `json:"reason,omitempty"`
}

weft.Tool("refund_order", "Refund a customer's order",
    func(ctx context.Context, in RefundInput) (Receipt, error) {
        return billing.Refund(ctx, in.OrderID, in.Reason)
    },
    weft.Timeout(10*time.Second))

What the model sees as the result is the text of a string Out, and the JSON encoding of any other Out. Inside the handler, CallFromContext reports which call is running. Trailing options set per-tool policy: Timeout, MaxResultBytes, StrictInput.

The handler shape mirrors the official Go MCP SDK's AddTool[In, Out], so a weft tool can be exposed over MCP without an adapter layer.

Tool panics if name is empty, fn is nil, or In is not a struct (or a pointer to one): providers and MCP require an object at the top level of a tool schema, and a scalar there would fail every real call.

Example

A tool is a plain function; the input schema is derived from the struct.

package main

import (
	"context"
	"encoding/json"
	"fmt"
	"log"

	"github.com/weftgo/weft"
)

func main() {
	type WeatherInput struct {
		City string `json:"city" jsonschema:"the city to look up"`
		Days *int   `json:"days,omitempty"`
	}
	getWeather := weft.Tool("get_weather", "Get a forecast.",
		func(_ context.Context, in WeatherInput) (string, error) {
			return "sunny in " + in.City, nil
		})

	b, err := json.MarshalIndent(getWeather.InputSchema, "", "  ")
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(getWeather.Name)
	fmt.Println(string(b))
}
Output:
get_weather
{
  "type": "object",
  "properties": {
    "city": {
      "type": "string",
      "description": "the city to look up"
    },
    "days": {
      "type": "integer"
    }
  },
  "required": [
    "city"
  ]
}

func (*ToolDef) Invoke

func (t *ToolDef) Invoke(ctx context.Context, args json.RawMessage) (string, error)

Invoke runs the tool with raw JSON arguments: it unmarshals into the handler's input type, calls the handler, and returns the result text the model will see (a string output verbatim, anything else JSON-encoded). A decode failure returns an error wrapping ErrInvalidToolInput. Invoke applies the tool's own StrictInput setting; timeouts and result caps are run policy, applied by the loop.

func (*ToolDef) PromptSnippet added in v0.2.0

func (t *ToolDef) PromptSnippet() string

PromptSnippet reports the tool's PromptSnippet text; empty when unset.

func (*ToolDef) RequiresApproval added in v0.2.0

func (t *ToolDef) RequiresApproval() bool

RequiresApproval reports whether the tool was built with RequireApproval. The loop reads it to park the call; AddTools (weft/mcp) refuses such a tool at registration, because MCP has no approval channel and running it unapproved would bypass the gate.

type ToolError added in v0.2.0

type ToolError struct {
	Code    string
	Message string
	Err     error
}

ToolError is a tool failure with a stable code the model can branch on. Code is SCREAMING_SNAKE by convention ("ORDER_NOT_FOUND"); Message is what the model reads; Err is the internal cause — available to tool middleware and audit logs through errors.As/Unwrap, and never shown to the model. A handler returning *ToolError produces the result "<CODE>: <Message>"; the loop renders its own failures with codes too: INVALID_INPUT (arguments that do not decode) and NO_SUCH_TOOL. Codes are not validated. Plain errors keep rendering as err.Error(); install mw.MapErrors to code them centrally.

Example

A *ToolError carries a code the model can branch on and a cause it never sees.

package main

import (
	"context"
	"errors"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	lookup := weft.Tool("lookup_order", "", func(_ context.Context, in struct {
		ID string `json:"id"`
	}) (string, error) {
		return "", &weft.ToolError{
			Code:    "ORDER_NOT_FOUND",
			Message: "order " + in.ID + " does not exist",
			Err:     errors.New("pg: no rows in result set"), // for logs and middleware only
		}
	})
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":"42"}`}),
		wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":42}`}),
		wefttest.Say("No such order."),
	), lookup)
	res, err := agt.Generate(context.Background(), weft.Prompt("order 42?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
	fmt.Println(res.Steps[1].Results[0].Content)
}
Output:
ORDER_NOT_FOUND: order 42 does not exist
INVALID_INPUT: tool "lookup_order": field "id": expected string, got number

func Errorf added in v0.2.0

func Errorf(code, format string, args ...any) *ToolError

Errorf builds a *ToolError with a formatted Message. A %w verb sets Err as fmt.Errorf would — with several %w verbs every cause stays reachable through errors.Is — so the cause is available to middleware while the model sees only the formatted text.

func (*ToolError) Error added in v0.2.0

func (e *ToolError) Error() string

Error renders the model-visible form: "CODE: Message". Err is not included.

func (*ToolError) Unwrap added in v0.2.0

func (e *ToolError) Unwrap() error

Unwrap returns the internal cause, for errors.Is/As in middleware and logs.

type ToolFinish

type ToolFinish struct {
	RunID   string `json:"run_id"`
	Seq     int64  `json:"seq"`
	CallID  string `json:"call_id"`
	Name    string `json:"name"`
	Content string `json:"content"`
	IsError bool   `json:"is_error"`
}

ToolFinish reports that a tool invocation completed, successfully or not. Content is the tool's JSON output, or the failure text when IsError is set — the same value the model sees on the matching ToolResultPart, so a UI can render results as they land.

func (ToolFinish) MarshalJSON

func (e ToolFinish) MarshalJSON() ([]byte, error)

MarshalJSON encodes the event with its "type" discriminator.

type ToolMiddleware added in v0.2.0

type ToolMiddleware func(next ToolCaller) ToolCaller

ToolMiddleware wraps a ToolCaller, the chi shape: it may act before next (deny, decorate ctx, log), after it (map errors, audit), or instead of it. Panic containment stays outside the chain — a panic in middleware is still a tool error result, never a run error.

type ToolOption added in v0.2.0

type ToolOption interface {
	// contains filtered or unexported methods
}

ToolOption configures one tool at definition time. The options that make sense at both levels — MaxResultBytes, Timeout, StrictInput — are PolicyOptions: on the agent they set the default for every tool, on a tool they override that default for the tool alone.

func PromptSnippet added in v0.2.0

func PromptSnippet(text string) ToolOption

PromptSnippet attaches system-prompt lines to a tool. The loop appends the snippets of the tools it advertises to the agent's Instructions for every model call — one paragraph per tool, in registration order, separated by blank lines — so usage rules live with the tool instead of in a central prompt that drifts as the set grows. Empty snippets add nothing. The manifest records the snippet.

Example

PromptSnippet keeps a tool's usage rules next to the tool; the loop appends them to the instructions of every model call.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	search := weft.Tool("search", "Search the docs.", func(_ context.Context, _ struct{}) (string, error) { return "", nil },
		weft.PromptSnippet("Cite the search result you used."))
	model := wefttest.Script(wefttest.Say("ok"))
	if _, err := weft.New(model, weft.Instructions("You answer questions."), search).Generate(context.Background(), weft.Prompt("hi")); err != nil {
		log.Fatal(err)
	}
	fmt.Println(model.Requests()[0].System)
}
Output:
You answer questions.

Cite the search result you used.

func RequireApproval added in v0.2.0

func RequireApproval() ToolOption

RequireApproval marks a tool whose calls never run without a decision. When the model calls it the loop executes the step's other tools, then ends the run successfully with the call on RunResult.Pending (and RunFinish.Pending) and no result in the transcript. Resume with the transcript plus Approve or Deny:

res, _ := agt.Generate(ctx, weft.Prompt("Refund order 42"))
for _, call := range res.Pending { /* ask someone */ }
res, _ = agt.Generate(ctx, weft.Messages(res.Messages...), weft.Approve(call.ID))

Approved calls run through the ordinary chain with Call.Approved set; denied ones (and pending calls given no decision) become error results the model sees. This is a policy and UX seam, not a security boundary: the boundary is the sandbox a tool runs in.

Example

A RequireApproval tool parks its calls: the run ends successfully with them on Pending, and a later run resumes with a decision.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	refund := weft.Tool("refund", "Refund an order.", func(_ context.Context, in struct {
		Order string `json:"order"`
	}) (string, error) {
		return "refunded " + in.Order, nil
	}, weft.RequireApproval())
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{ID: "c1", Name: "refund", Args: `{"order":"42"}`}),
		wefttest.Say("Done."),
	), refund)

	res, err := agt.Generate(context.Background(), weft.Prompt("refund order 42"))
	if err != nil {
		log.Fatal(err)
	}
	for _, call := range res.Pending {
		fmt.Printf("awaiting approval: %s %s\n", call.Name, call.Args)
	}

	// Someone decided. Resume with the transcript and the decision.
	res, err = agt.Generate(context.Background(), weft.Messages(res.Messages...), weft.Approve("c1"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Text())
}
Output:
awaiting approval: refund {"order":"42"}
Done.

func ToolOptions added in v0.2.0

func ToolOptions(opts ...ToolOption) ToolOption

ToolOptions composes several tool options into one, applied in order — the Tool-defining counterpart of Options, so a package of policy (Timeout, MaxResultBytes, StrictInput, a WrapTools chain) can be named and reused:

productPolicy := weft.ToolOptions(weft.Timeout(5*time.Second), weft.StrictInput())
weft.Tool("lookup", "…", fn, productPolicy)

Nil entries are ignored.

Example

ToolOptions composes tool options into one named value, so a package of per-tool policy travels under one name.

package main

import (
	"context"
	"fmt"
	"log"
	"time"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	productPolicy := weft.ToolOptions(
		weft.Timeout(5*time.Second),
		weft.MaxResultBytes(1024),
	)
	agt := weft.New(wefttest.Script(
		wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":"42"}`}),
		wefttest.Say("Done."),
	), weft.Tool("lookup_order", "Look up an order by id.",
		func(_ context.Context, in struct {
			ID string `json:"id" jsonschema:"the order id"`
		}) (string, error) {
			return "order " + in.ID + " shipped", nil
		}, productPolicy))
	res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 42?"))
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(res.Steps[0].Results[0].Content)
}
Output:
order 42 shipped

type ToolResultPart

type ToolResultPart struct {
	CallID  string `json:"call_id"`
	Name    string `json:"name"`
	Content string `json:"content"`
	IsError bool   `json:"is_error"`
}

ToolResultPart is the outcome of one tool call, returned to the model as data. Content is the JSON encoding of the tool's output, or the failure message when IsError is set. Tool failures never abort a run; the model sees them and can recover.

func (ToolResultPart) MarshalJSON

func (p ToolResultPart) MarshalJSON() ([]byte, error)

MarshalJSON encodes the part with its "type" discriminator.

type ToolStart

type ToolStart struct {
	RunID  string          `json:"run_id"`
	Seq    int64           `json:"seq"`
	CallID string          `json:"call_id"`
	Name   string          `json:"name"`
	Args   json.RawMessage `json:"args"`
}

ToolStart reports that a tool invocation began. Events from tools running in parallel interleave: pair them by CallID and order by Seq, a per-run counter assigned at emission that totally orders the stream. A call parked by the approval boundary has a ToolStart and no ToolFinish; it is listed on RunFinish.Pending instead.

func (ToolStart) MarshalJSON

func (e ToolStart) MarshalJSON() ([]byte, error)

MarshalJSON encodes the event with its "type" discriminator.

type Usage

type Usage struct {
	InputTokens       int64 `json:"input_tokens"`
	OutputTokens      int64 `json:"output_tokens"`
	CachedInputTokens int64 `json:"cached_input_tokens,omitempty"` // ⊆ InputTokens: read from a prompt cache
	CacheWriteTokens  int64 `json:"cache_write_tokens,omitempty"`  // ⊆ InputTokens: written to a prompt cache (anthropic)
	ReasoningTokens   int64 `json:"reasoning_tokens,omitempty"`    // ⊆ OutputTokens
}

Usage is token accounting for one step or one whole run.

The totals are inclusive: CachedInputTokens and CacheWriteTokens are subsets of InputTokens (tokens billed as input — read from or written to a provider prompt cache), ReasoningTokens is a subset of OutputTokens (provider-side reasoning the model burned). The splits are reporting, not budget bases: UsageLimit and Total keep reading the two totals, so a cached-heavy run budgets identically to an uncached one with the same totals. All three are omitempty on the wire — an event from before they existed round-trips unchanged (SchemaVersion stays 1, ADR 0001/0004).

func (Usage) Add

func (u Usage) Add(o Usage) Usage

Add returns the element-wise sum of u and o.

func (Usage) Total

func (u Usage) Total() int64

Total returns the sum of input and output tokens.

Directories

Path Synopsis
anthropic module
core module
examples
approval command
Command approval shows the approval boundary end to end, offline: a tool marked RequireApproval parks its call, the run ends with the call on Pending, a person decides, and a second run resumes with the transcript plus the decision.
Command approval shows the approval boundary end to end, offline: a tool marked RequireApproval parks its call, the run ends with the call on Pending, a person decides, and a second run resumes with the transcript plus the decision.
getting-started command
Command getting-started runs a complete weft agent offline: a scripted model (from wefttest), one tool, and the full event stream.
Command getting-started runs a complete weft agent offline: a scripted model (from wefttest), one tool, and the full event stream.
google module
internal
adapterkit
Package adapterkit holds the helpers every first-party adapter needs but no vendor SDK touches: schema rendering, the terminal-error rule, and the FilePart exactly-one guard.
Package adapterkit holds the helpers every first-party adapter needs but no vendor SDK touches: schema rendering, the terminal-error rule, and the FilePart exactly-one guard.
jsonclose
Package jsonclose closes truncated JSON documents: the state machine both the OutputDecoder's prefix salvage (partial_json.go) and mw.RepairJSON need.
Package jsonclose closes truncated JSON documents: the state machine both the OutputDecoder's prefix salvage (partial_json.go) and mw.RepairJSON need.
jsonconflict
Package jsonconflict holds a fixture type for schema tests that need an embedded field to collide with a struct's own JSON name.
Package jsonconflict holds a fixture type for schema tests that need an embedded field to collide with a struct's own JSON name.
mcp module
Package mw holds the reference middleware for weft's two seams: model middleware (weft.WrapModel) and tool middleware (weft.WrapTools).
Package mw holds the reference middleware for weft's two seams: model middleware (weft.WrapModel) and tool middleware (weft.WrapTools).
obsdb module
clickhouse module
openai module
otel module
runtime module
store module
studio module
cmd module
thread module
sqlite module
Package wefttest provides a scriptable weft.Model, in the spirit of httptest: write the agent's dialogue as a sequence of scripted model steps and test agents offline, deterministically, with no network.
Package wefttest provides a scriptable weft.Model, in the spirit of httptest: write the agent's dialogue as a sequence of scripted model steps and test agents offline, deterministically, with no network.
conformance
Package conformance is the provider adapter contract as an executable table.
Package conformance is the provider adapter contract as an executable table.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL