Documentation
¶
Overview ¶
Package weft is a thin, opinionated core for building agents in Go.
Weft runs the agent loop — call a model, execute its tool calls in parallel with defined failure semantics, stream typed events while work is in flight — and nothing else. It is designed the way the standard library is: small interfaces, context everywhere, functional options, wrapped errors, and no required configuration.
A tool is a plain function; its JSON Schema is derived from the input struct. An agent is a value built once with options and run many times:
echo := weft.Tool("echo", "Echo a message",
func(ctx context.Context, in struct {
Msg string `json:"msg"`
}) (string, error) {
return "echo: " + in.Msg, nil
})
agt := weft.New(model, weft.Instructions("You are helpful."), echo)
res, err := agt.Generate(ctx, weft.Prompt("Say hi."))
A run ends when the model replies without tool calls or a StopWhen condition is met; MaxSteps is the safety budget behind both. Streaming is a range loop over typed events (Agent.Stream). Tool failures are data the model sees; only model failures, cancellation, and the step budget reach the caller, as *RunError. Output constrains the final answer to a struct (GenerateAs), and trailing options on Tool — Timeout, MaxResultBytes, StrictInput — set per-tool policy. Thinking sets the reasoning depth (agent default, run override).
The model seam (Model) is streaming-first; provider adapters translate vendor wire formats into weft's events. The wefttest package provides a scriptable Model for offline tests.
Index ¶
- Constants
- Variables
- func Manifest(agents ...*Agent) ([]byte, error)
- func ModelRequestsAllowed() bool
- func ModelRetry(hint string) error
- func OutputOf[Out any](res *RunResult) (Out, error)
- type Agent
- func (a *Agent) CallTool(ctx context.Context, call ToolCallPart) (string, error)
- func (a *Agent) Generate(ctx context.Context, opts ...RunOption) (*RunResult, error)
- func (a *Agent) Name() string
- func (a *Agent) Stream(ctx context.Context, opts ...RunOption) *Run
- func (a *Agent) TapPanics() int64
- func (a *Agent) Tools() []*ToolDef
- type Call
- type Event
- type FilePart
- type Message
- type Model
- type ModelEvent
- type ModelFinish
- type ModelInfo
- type ModelMiddleware
- type ModelReasoningDelta
- type ModelRequest
- type ModelTextDelta
- type ModelToolCall
- type ModelToolCallDelta
- type Nested
- type Option
- func DetectLoops(repeats int) Option
- func Instructions(text string) Option
- func Logger(l *slog.Logger) Option
- func MaxModelRetries(n int) Option
- func MaxSteps(n int) Option
- func Name(name string) Option
- func OnRunEnd(fn func(ctx context.Context, res *RunResult, err error)) Option
- func Options(opts ...Option) Option
- func Output[Out any]() Option
- func Parallelism(n int) Option
- func PrepareStep(fn func(ctx context.Context, step int, req ModelRequest) (ModelRequest, error)) Option
- func StopWhen(conds ...StopCondition) Option
- func Tap(fn func(ctx context.Context, ev Event)) Option
- func ToolSource(fn func() []*ToolDef) Option
- func TracerProvider(tp trace.TracerProvider) Option
- func UsageLimit(max Usage) Option
- func WrapModel(mw ...ModelMiddleware) Option
- type OutputDecoder
- type ParamsOption
- type Part
- type PolicyOption
- type ReasoningDelta
- type ReasoningPart
- type RequestParams
- type Role
- type Run
- type RunError
- type RunFinish
- type RunOption
- type RunResult
- type RunStart
- type Schema
- type StepFinish
- type StepRecord
- type StepStart
- type StopCondition
- type StopFunc
- type StopReason
- type TextDelta
- type TextPart
- type ThinkingConfig
- type ThinkingLevel
- type ThinkingOption
- type ToolArgsDelta
- type ToolCallPart
- type ToolCaller
- type ToolChoiceConfig
- type ToolChoiceMode
- type ToolChoiceOption
- type ToolDef
- type ToolError
- type ToolFinish
- type ToolMiddleware
- type ToolOption
- type ToolResultPart
- type ToolStart
- type Usage
Examples ¶
- Agent (Conversation)
- Agent (Generate)
- Agent (Stream)
- DetectLoops
- GenerateAs
- Logger
- Manifest
- ModelRetry
- OnRunEnd
- Options
- OutputDecoder
- Params
- PrepareStep
- PromptSnippet
- ReasoningPart
- Repair
- RequireApproval
- Sequential
- Subagent
- Tap
- Tap (Async)
- Thinking
- Timeout
- Tool
- ToolChoice
- ToolError
- ToolOptions
- TracerProvider
- UnmarshalEvent
- UsageLimit
- UserParts
- WrapModel
- WrapTools
- WrapTools (Context)
Constants ¶
const ( // CodeSubagentFailed marks a child run that failed. The message the // model sees carries the child's step and cause; the *RunError // itself stays on ToolError.Err, reachable through errors.As, so // middleware can branch on it without parsing that text. CodeSubagentFailed = "SUBAGENT_FAILED" // CodeSubagentPending marks a child run that ended awaiting an // approval decision. Approvals belong in the orchestrator, not in a // child: the parent's transcript has nowhere to carry the child's // pending call, so the delegation is refused loudly instead of // silently (ADR 0014). CodeSubagentPending = "SUBAGENT_PENDING" // CodeSubagentCycle marks a delegation to an agent already running // in this call chain — refused before any model call. CodeSubagentCycle = "SUBAGENT_CYCLE" )
Codes the Subagent tool renders its failures with — model-visible contract, pinned by tests (ADR 0002's table; ADR 0014). Exported like the loop's own codes so policy middleware can branch on them.
const ( // CodeInvalidInput marks tool arguments that do not decode into // the tool's input struct; the message names the field in the // schema's own vocabulary. CodeInvalidInput = "INVALID_INPUT" // CodeNoSuchTool marks a call naming a tool the agent does not // have. CodeNoSuchTool = "NO_SUCH_TOOL" // CodeDenied marks a call the approval boundary (Deny, or no // decision) or a policy middleware refused — mw.Allow renders it // too: to the model, a refused call is a refused call, whoever // refused it (ADR 0007). CodeDenied = "DENIED" )
Codes the loop renders its own tool failures with — model-visible contract (ADR 0002, "tool errors have codes"), exported so external policy middleware can produce the same wire strings without duplicating them, the same class of constant as SchemaVersion.
const CodeRetry = "RETRY"
CodeRetry is the code ModelRetry renders with — the one code whose result is an instruction rather than a failure report: try the call again with the hint applied.
const SchemaVersion = 1
SchemaVersion is the version of the message wire format. The JSON encoding of Message and its parts is a compatibility contract: within a version, field names and shapes change only additively. Persistence and serving layers envelope messages with this number; the core itself never needs it.
Variables ¶
var ( // ErrMaxSteps is returned when the model still requests tools after the // last allowed step. The partial transcript rides along on RunError. ErrMaxSteps = errors.New("weft: run exceeded the maximum number of steps") // ErrUsageLimit is returned when a run's usage exceeds its // UsageLimit and the loop would otherwise call the model again. A // step that ends the run may overshoot and still succeed: a budget // stops further spend, it does not discard finished work. ErrUsageLimit = errors.New("weft: run exceeded its usage limit") // ErrModelRetriesExceeded is returned when one tool has produced // more consecutive RETRY results than MaxModelRetries allows — a // model that cannot self-correct is a run failure, not an infinite // loop. ErrModelRetriesExceeded = errors.New("weft: a tool asked the model to retry too many times") // ErrLoopDetected is returned when repeats consecutive steps // requested the same set of tool calls (DetectLoops). ErrLoopDetected = errors.New("weft: the model repeated the same tool calls too many times") // ErrNoSuchTool is returned by Agent.CallTool when the call names a // tool the agent does not have. Inside the loop the same condition is // folded into an error tool result — data for the model to // self-correct, not a run failure. ErrNoSuchTool = errors.New("weft: no tool with that name") // ErrInvalidToolInput is returned by ToolDef.Invoke and Agent.CallTool // when the arguments do not decode into the tool's input type. Like // ErrNoSuchTool, the loop turns it into model-visible data. ErrInvalidToolInput = errors.New("weft: tool input is not valid for its schema") // ErrRunConsumed is returned by Run.Events when the event stream has // already been consumed; each Run yields exactly one sequence. ErrRunConsumed = errors.New("weft: run events already consumed") // ErrModelContract is wrapped around failures of a Model // implementation to honor the stream contract documented on Model: // events after ModelFinish, a stream ending without one, a tool call // with an empty ID or name, or a panicking stream. It signals an // adapter bug, not a model outage — providers' own errors surface // unwrapped. ErrModelContract = errors.New("weft: model violated the stream contract") // ErrUnsupported is wrapped by a Model that cannot honour part of a // request — a FilePart whose media type the provider does not accept, // a feature the vendor lacks. It is a run error (the model call // fails), so callers can errors.Is on it and fall back to another // model. ErrUnsupported = errors.New("weft: request uses a feature the model does not support") // ErrStreamIdle is the stream error an adapter yields when the gap // between two chunks exceeds its IdleTimeout. The ctx deadline is the // hard limit on a whole call; the idle timeout only catches a stalled // stream, so a slow but actively streaming response is never killed. // It is a run error like any provider error; callers errors.Is on it // provider-agnostically, without knowing which adapter timed out. ErrStreamIdle = errors.New("weft: model stream idle timeout") // ErrModelRequestsDenied is the stream error every first-party // adapter yields from Stream when ModelRequestsAllowed is false — the // kill switch for test suites that must never reach the network. // wefttest models ignore the switch, so ordinary offline tests are // unaffected. ErrModelRequestsDenied = errors.New("weft: model requests denied by WEFT_MODEL_REQUESTS") // ErrApprovalRequired marks a tool call that must not run until a // human (or an outer system) decides. The loop raises it for tools // built with RequireApproval; tool middleware may return an error // wrapping it to defer any call. The run then ends successfully with // the call on RunResult.Pending; resume with Approve or Deny. ErrApprovalRequired = errors.New("weft: tool call requires approval") // ErrApprovalDenied is the cause on the error result the model sees // for a pending call that was denied (Deny, or no decision on // resume). It is a tool error — data — never a run error. ErrApprovalDenied = errors.New("weft: tool call denied") // ErrDuplicateTool is a run error raised when a step's tool // snapshot contains a name twice — the runtime analogue of New's // duplicate-name panic. Fix the tool source; the run fails rather // than silently dropping the second tool. Agent.CallTool reports // the same condition as an error. ErrDuplicateTool = errors.New("weft: duplicate tool name from tool source") // ErrNilTool is a run error raised when a tool-source snapshot // contains a nil entry — a malformed snapshot, not a tool. The run // fails rather than advertising a dereference every adapter would // panic on; Agent.CallTool reports the same condition as an error. ErrNilTool = errors.New("weft: nil tool in tool source snapshot") )
Sentinel errors for named run failures. Branch on them with errors.Is; never match on error strings.
var ErrNoOutput = errors.New("weft: run ended without a structured output")
ErrNoOutput is returned by GenerateAs and OutputOf when the run ended without a valid submit_output call — the model answered in text, hit the step budget, or never produced arguments that decode into Out.
Functions ¶
func Manifest ¶
Manifest renders the agents as their `weft.json` document: one generated, committed, diffable description of every agent and tool. Studio, docs, review, and compatibility checks read a file instead of a live process; the code stays the only source of truth. The file is output, never input — nothing is configured from it. Generate it in a golden test (wefttest.Golden) so a stale file fails `go test`; run that test with `-update` to regenerate.
Tools are listed in registration order, agents in argument order, and encoding/json sorts map keys, so the bytes are deterministic for the same agents. Unnamed and duplicate agent names are errors: the file is a review artifact, and `agent_1` in a diff is noise.
Example ¶
The manifest is generated output: one committed, diffable description of every agent and tool. Gate it with a golden test so it cannot go stale.
package main
import (
"context"
"fmt"
"log"
"strings"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
agt := weft.New(wefttest.Script(wefttest.Say("ok")),
weft.Name("support-bot"),
weft.Instructions("You are a support agent."),
weft.Tool("refund_order", "Refund a customer's order.",
func(_ context.Context, _ struct {
OrderID string `json:"order_id"`
}) (string, error) {
return "refunded", nil
}),
)
b, err := weft.Manifest(agt)
if err != nil {
log.Fatal(err)
}
lines := strings.Split(string(b), "\n")
fmt.Println(lines[0])
fmt.Println(lines[1])
}
Output: { "weft": 1,
func ModelRequestsAllowed ¶
func ModelRequestsAllowed() bool
ModelRequestsAllowed reports whether adapters may call a provider. It is false when WEFT_MODEL_REQUESTS=deny — the guard for test suites that must never reach the network (Pydantic AI's ALLOW_MODEL_REQUESTS). First-party adapters check it at the top of Stream and yield ErrModelRequestsDenied when it is false, before any network I/O — but only for a client they built themselves from credentials; a client the caller injected through the adapter's Client(c) option is a test double by construction and stays reachable (ADR 0013's kill-switch clause). wefttest models ignore the switch entirely.
The environment is read on every call, not cached: a model call is network-bound so one getenv is noise, test suites can toggle the switch per test with t.Setenv, and the package keeps no state at all.
func ModelRetry ¶ added in v0.2.0
ModelRetry is a tool error asking the model to try the call again with the hint applied: it renders as "RETRY: <hint>" and the loop continues. Return it when the arguments are well-formed but wrong in a way the model can fix — an ambiguous date, an id that needs a prefix. The loop counts RETRY results per tool name per run and fails the run with ErrModelRetriesExceeded when a tool exceeds MaxModelRetries (default 3); a successful result for that tool resets its count. Middleware-produced retries count too — the loop sees the code through the chain, not who returned it.
Example ¶
A retry hint: the model sees "RETRY: <hint>", fixes the arguments, and the loop continues; a tool that cannot be satisfied fails the run after MaxModelRetries consecutive asks.
package main
import (
"context"
"fmt"
"log"
"sync/atomic"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
var calls atomic.Int32
parse := weft.Tool("parse_date", "Parse a date.",
func(_ context.Context, in struct {
D string `json:"d" jsonschema:"the date, ISO-8601"`
}) (string, error) {
if in.D != "2026-09-19" {
return "", weft.ModelRetry("date must be ISO-8601, e.g. 2026-09-19")
}
_ = calls.Add(1)
return "2026-09-19", nil
})
agt := weft.New(wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "parse_date", Args: `{"d":"tomorrow"}`}),
wefttest.ToolCalls(wefttest.Call{Name: "parse_date", Args: `{"d":"2026-09-19"}`}),
wefttest.Say("Parsed."),
), parse)
res, err := agt.Generate(context.Background(), weft.Prompt("When is it?"))
if err != nil {
log.Fatal(err)
}
fmt.Println("first attempt:", res.Steps[0].Results[0].Content)
fmt.Println("second attempt:", res.Steps[1].Results[0].Content)
}
Output: first attempt: RETRY: date must be ISO-8601, e.g. 2026-09-19 second attempt: 2026-09-19
Types ¶
type Agent ¶
type Agent struct {
// contains filtered or unexported fields
}
Agent is an immutable, reusable value: a model, a system instruction, a tool set, and an execution policy. Build it once with New; run it many times, concurrently if you like — runs share no state, and the registered tool set is frozen at construction (see ToolDef).
Example (Conversation) ¶
Continuing a conversation: feed the transcript back with the next question. The agent value is unchanged — the history lives in the messages you pass, never in the agent.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
agt := weft.New(wefttest.Script(
wefttest.Say("Order 1234? It shipped yesterday."),
wefttest.Say("Order 5678? Still pending."),
))
res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 1234?"))
if err != nil {
log.Fatal(err)
}
res2, err := agt.Generate(context.Background(),
weft.Messages(res.Messages...), weft.Prompt("And order 5678?"))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Text())
fmt.Println(res2.Text())
}
Output: Order 1234? It shipped yesterday. Order 5678? Still pending.
Example (Generate) ¶
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
echo := weft.Tool("echo", "Echo a message.",
func(_ context.Context, in struct {
Msg string `json:"msg"`
}) (string, error) {
return "echo: " + in.Msg, nil
})
model := wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"msg":"hello"}`}),
wefttest.Say("I echoed your message."),
)
agt := weft.New(model, weft.Instructions("You echo things."), echo)
res, err := agt.Generate(context.Background(), weft.Prompt("Echo hello."))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Text())
fmt.Println("steps:", res.NumSteps(), "tokens:", res.Usage.Total())
}
Output: I echoed your message. steps: 2 tokens: 30
Example (Stream) ¶
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
roll := weft.Tool("roll_dice", "Roll a six-sided die.",
func(_ context.Context, _ struct{}) (int, error) {
return 4, nil // deterministic for the example
})
model := wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "roll_dice"}),
wefttest.Say("You rolled a 4!"),
)
agt := weft.New(model, roll)
for ev, err := range agt.Stream(context.Background(), weft.Prompt("Roll a die.")).Events() {
if err != nil {
log.Fatal(err)
}
switch ev := ev.(type) {
case weft.ToolStart:
fmt.Println("tool:", ev.Name)
case weft.TextDelta:
fmt.Println("text:", ev.Text)
case weft.RunFinish:
fmt.Println("done in", ev.Steps, "steps")
}
}
}
Output: tool: roll_dice text: You rolled a 4! done in 2 steps
func New ¶
New builds an Agent. Nil models panic — including typed nils such as var m *someModel; New(m), which would otherwise crash much later inside a run goroutine. Everything else has a working default.
func (*Agent) CallTool ¶
CallTool dispatches one tool call by name the way the loop does — through the agent's WrapTools chain and the tool's own — and returns the result text the model would see. It is the seam for manual dispatchers. Unlike the loop it does not contain failures or apply run policy: an unknown name returns an error wrapping ErrNoSuchTool, undecodable arguments one wrapping ErrInvalidToolInput, a RequireApproval tool one wrapping ErrApprovalRequired, and handler errors, middleware errors, and panics propagate. A tool source whose snapshot carries a duplicate name returns an error wrapping ErrDuplicateTool. Timeouts and result caps are not applied; StrictInput is, at both levels, as in the loop. No span or log line is produced: manual dispatchers own their context and their own reporting (ADR 0016).
func (*Agent) Generate ¶
Generate runs the agent to completion and returns the final result. Internally it is Stream with the events folded away; errors are returned as *RunError, with the partial transcript attached.
func (*Agent) Name ¶ added in v0.2.0
Name returns the agent's name as set by the Name option; empty when the agent is unnamed. Manifest requires a name; Serve (weft/mcp) names the tool it exposes after the agent through this accessor.
func (*Agent) Stream ¶
Stream starts a run and returns its handle. The run is lazy: nothing executes until Events is consumed. Canceling ctx aborts the model call and any in-flight tools.
type Call ¶
type Call struct {
RunID string // the run this call belongs to
Step int // zero-based index of the step that requested it
CallID string // the provider's call identifier
Name string // the tool name
// Approved is set when the call was parked by the approval boundary
// and is now running under an Approve decision — middleware that
// defers calls with ErrApprovalRequired reads it to let the
// approved call through.
Approved bool
}
Call identifies the tool invocation a handler is serving. Retrieve it with CallFromContext — for audit logs, per-call idempotency keys, or progress reporting that must name its call.
type Event ¶
type Event interface {
// contains filtered or unexported methods
}
Event is the sealed set of run progress events, yielded by Run.Events in emission order. New event types may be added additively; external types cannot join, so switches over events stay exhaustively lintable.
On the wire every event carries a "type" discriminator (run_start, step_start, text_delta, reasoning_delta, tool_args_delta, tool_start, tool_finish, step_finish, run_finish, nested) and UnmarshalEvent restores it — the same rule and the same compatibility contract as the message parts (ADR 0004).
Every event except RunStart carries RunID: concurrent runs on one agent emit interleaved streams, and a per-run Seq counter is unique only within its run, so RunID is what attributes an event to its run.
Events are snapshots. Their fields — including the Args byte slices — do not alias the run's transcript; a consumer may retain or write into them freely.
func UnmarshalEvent ¶
UnmarshalEvent decodes one wire event, dispatching on its "type" discriminator. An unknown or missing type is an error, never a silent drop: a recorded stream must replay exactly what was emitted.
Example ¶
Recorded event streams decode back into typed events: store writes them, the Inspector replays them.
package main
import (
"fmt"
"log"
"github.com/weftgo/weft"
)
func main() {
ev, err := weft.UnmarshalEvent([]byte(`{"type":"tool_start","seq":5,"call_id":"c1","name":"echo","args":{"m":"x"}}`))
if err != nil {
log.Fatal(err)
}
start := ev.(weft.ToolStart)
fmt.Println(start.Name, start.CallID, start.Seq, string(start.Args))
}
Output: echo c1 5 {"m":"x"}
type FilePart ¶
type FilePart struct {
MediaType string `json:"media_type"`
Data []byte `json:"data,omitempty"`
URL string `json:"url,omitempty"`
}
FilePart is a file the user supplies to the model: an image, a PDF, audio. Exactly one of Data (inline; base64 on the wire via []byte's default encoding) or URL is set. Adapters map it to the vendor's image/document block; an adapter that cannot carry this MediaType fails the model call with an error wrapping ErrUnsupported. The core never reads the bytes. The exactly-one rule is documented, not enforced here — the adapter is the layer that knows what it can send, and it returns ErrUnsupported for a part with both or neither set.
func (FilePart) MarshalJSON ¶
MarshalJSON encodes the part with its "type" discriminator.
type Message ¶
Message is one turn in a conversation: a role plus an ordered list of content parts. A step's tool results are collected on a single RoleTool message; provider adapters fan out or merge as their wire format requires.
On the wire every part carries a "type" discriminator ("text", "tool_call", "tool_result", "reasoning"), so a Message round-trips through encoding/json losslessly.
func Repair ¶
Repair makes a transcript valid model input: every tool call has a result (missing ones become visible error results), results with no call are dropped, everything else is untouched. The loop applies it to the input of every run; store and runtime call it before persisting. It is pure (the input is never mutated) and idempotent: Repair(Repair(m)) equals Repair(m). nil input yields nil; an empty non-nil input yields an empty non-nil transcript.
Example ¶
A partial transcript (the run was interrupted mid-step) becomes valid model input: the missing result is synthesised, visibly.
package main
import (
"encoding/json"
"fmt"
"github.com/weftgo/weft"
)
func main() {
msgs := []weft.Message{
weft.User("Where is order 1234?"),
{Role: weft.RoleAssistant, Content: []weft.Part{
weft.ToolCallPart{ID: "c1", Name: "lookup_order", Args: json.RawMessage(`{"order_id":"1234"}`)},
}},
}
for _, m := range weft.Repair(msgs) {
fmt.Println(m.Role)
}
}
Output: user assistant tool
func UserParts ¶
UserParts returns a user message with the given parts, for prompts that mix text and files: UserParts(TextPart{"What is this?"}, FilePart{MediaType: "image/png", URL: u}). The parts are copied; the caller's slice is not retained.
Example ¶
Provider reasoning round-trips: it is preserved in the transcript, placed before the text of the same assistant turn. A prompt can mix text and files with UserParts; the core carries the bytes and the adapter maps them to the provider's image block.
package main
import (
"encoding/json"
"fmt"
"log"
"github.com/weftgo/weft"
)
func main() {
msg := weft.UserParts(
weft.TextPart{Text: "What is this?"},
weft.FilePart{MediaType: "image/png", URL: "https://example.com/cat.png"},
)
b, err := json.Marshal(msg)
if err != nil {
log.Fatal(err)
}
fmt.Println(string(b))
}
Output: {"role":"user","content":[{"type":"text","text":"What is this?"},{"type":"file","media_type":"image/png","url":"https://example.com/cat.png"}]}
func (*Message) UnmarshalJSON ¶
UnmarshalJSON decodes a message, dispatching each content part on its "type" discriminator. An unknown part type is an error: the wire format is versioned, and silently dropping content would corrupt transcripts.
type Model ¶
type Model interface {
Stream(ctx context.Context, req ModelRequest) iter.Seq2[ModelEvent, error]
}
Model is the provider seam. Implementations stream one step's output as events; first-party adapters wrap the vendors' official Go SDKs rather than re-implementing HTTP.
The stream contract:
- Events are yielded in order: any number of ModelTextDelta, ModelReasoningDelta, and ModelToolCall values — optionally interleaved with ModelToolCallDelta progress as argument fragments stream — then exactly one ModelFinish.
- Failure is reported as a single terminal yield of (nil, err); no events follow it.
- The sequence honors ctx: when ctx is done, the model yields (nil, ctx.Err()) if it has not finished already.
The loop enforces this contract: a stream that ends without ModelFinish, continues after it, yields a tool call with an empty ID or name, two tool calls sharing an ID in one step, or panics fails the run with an error wrapping ErrModelContract. A contract-violating adapter cannot corrupt a transcript silently.
type ModelEvent ¶
type ModelEvent interface {
// contains filtered or unexported methods
}
ModelEvent is the sealed set of events a model yields during one step. Tool calls arrive whole — assembling providers' streamed argument fragments is the adapter's job, which is what makes the core's streaming uniform across providers.
type ModelFinish ¶
type ModelFinish struct {
Reason StopReason
Usage Usage
// Raw is the provider's own stop reason when Reason had to be
// approximated ("refusal", "pause_turn", "content_filter", ...).
// Empty when the mapping was exact. The core never interprets it; it
// is recorded on the step (StepRecord.RawStopReason) and the
// StepFinish event so callers can see a refusal without a tap.
Raw string
}
ModelFinish closes a step with its stop reason and token usage.
type ModelInfo ¶
type ModelInfo struct {
Provider string `json:"provider"` // "openai", "anthropic", "wefttest"
Name string `json:"name"` // the vendor's model id
}
ModelInfo identifies a model for telemetry (§8.1) and the manifest. A Model that can report it implements the optional
interface{ Info() ModelInfo }
which the loop detects and surfaces on RunStart.Model; Model itself stays one method. Middleware that wraps a Model should forward Info (TODO §4.1).
type ModelMiddleware ¶ added in v0.2.0
ModelMiddleware wraps a Model, the chi shape: it sees every ModelRequest the loop builds and every event the inner model yields, and may retry, substitute, log, or rewrite. Implementations should forward Info (see InfoOf) so RunStart.Model and the manifest still name the underlying model.
type ModelReasoningDelta ¶
ModelReasoningDelta is an increment of provider reasoning (Anthropic thinking, Gemini thought summaries). Signature is the provider's opaque token for the block, if any; adapters set it on the delta that completes a block. The core stores and forwards reasoning and never reads it.
Block boundaries: a delta carrying a non-empty Signature closes the current reasoning block; the next reasoning delta opens a new one. Providers send a block's signature last (Anthropic's signature_delta ends a thinking block; Gemini's per-part signature is emitted after the part's text), so one ReasoningPart per provider block survives the round trip. Reasoning without any signature accumulates into a single block — nothing downstream can send unsigned blocks back anyway, so their boundaries are not load-bearing.
type ModelRequest ¶
type ModelRequest struct {
System string
Messages []Message
Tools []*ToolDef
// Thinking asks the model to reason at the given level for this
// step. The zero value keeps the provider default (and any
// construction-time adapter option, such as anthropic.Thinking);
// the loop fills it from the agent's Thinking option, which a
// run-level Thinking overrides.
Thinking ThinkingConfig
// SequentialTools asks the provider not to emit parallel tool-call
// batches. The zero value keeps the provider default; the loop sets
// it to true exactly under Sequential() (and Parallelism(1)), so the
// model does not emit batches the execution policy would serialize
// anyway. Adapters mirror it in the provider's parallel-tool-calls
// setting.
SequentialTools bool
// ToolChoice constrains what the model may emit this step: some
// tool, a named tool, or none — the router-agent and hardened-Output
// knob. The zero value keeps the provider default; the loop fills it
// from the agent's ToolChoice option, which a run-level ToolChoice
// overrides, and a PrepareStep function can rewrite it per step
// (force classify on step 0, then auto). It constrains what the
// provider is asked to emit, never execution: a call the provider
// emits anyway runs, because the model was shown the tool.
ToolChoice ToolChoiceConfig
// Params overrides the adapter's construction-time sampling knobs
// (Temperature, TopP, MaxTokens, Stop, Seed) for this request alone;
// see RequestParams for the nil-keeps-default fold. The zero value
// keeps every construction default, so a request without it
// serialises exactly as v0.2.0's did.
Params RequestParams
}
ModelRequest is everything a model needs for one step: the system instruction, the transcript so far, and the callable tools.
Read-only: the loop builds each request with fresh copies of the Messages and Tools slices, so appending to them or reassigning their elements cannot reach the run or the agent. The values inside remain shared — message Content parts, and ToolDef fields frozen at construction — so implementations must still not modify them, and must clone anything they retain beyond the call.
type ModelTextDelta ¶
type ModelTextDelta struct {
Text string
}
ModelTextDelta is an increment of assistant text.
type ModelToolCall ¶
type ModelToolCall struct {
ID string
Name string
Args json.RawMessage
Signature string
}
ModelToolCall is one complete tool invocation request. ID is the provider's call identifier, echoed back on the matching ToolResultPart; it must be unique among one step's calls (results and approval decisions key on it). Signature is the provider's opaque token attached to the call itself (Gemini attaches thought signatures to functionCall parts and requires them returned on the same part); adapters that do not have one leave it empty.
type ModelToolCallDelta ¶ added in v0.2.0
ModelToolCallDelta is an increment of a streamed tool call's arguments — progress only: the assembled call still arrives whole as a ModelToolCall before ModelFinish. Adapters whose providers stream argument fragments (OpenAI-compatible function.arguments pieces, Anthropic input_json_delta) yield these so consumers can show the model "writing" a call instead of dead air; adapters whose calls arrive whole (Google) simply yield none. Index is the provider's fragment key where one exists (OpenAI's delta index); Name is the best-known name so far — for many providers only the first fragment of a call carries it.
type Nested ¶ added in v0.2.0
type Nested struct {
RunID string `json:"run_id"`
Seq int64 `json:"seq"`
CallID string `json:"call_id"`
Event Event `json:"event"`
}
Nested wraps one event of a child run started by a Subagent tool. CallID is the parent's tool call that owns the child run; Seq is from the parent's counter, so the parent stream stays totally ordered with the child's events in place. Event is any child event — including a Nested from a grandchild — and its own RunID and Seq are the child's. A child's RunStart..RunFinish all arrive inside the parent's ToolStart..ToolFinish for the delegating call, and none arrive after its ToolFinish (the late-event rule, ADR 0004). On the wire the type is "nested" and UnmarshalEvent restores the inner event recursively.
func (Nested) MarshalJSON ¶ added in v0.2.0
MarshalJSON encodes the event with its "type" discriminator. The inner event marshals through its own MarshalJSON, so its discriminator is present and decoding recurses (UnmarshalEvent inside Nested.UnmarshalJSON).
func (*Nested) UnmarshalJSON ¶ added in v0.2.0
UnmarshalJSON decodes the envelope and the inner event recursively — an unknown inner type is an error exactly as at the top level.
type Option ¶
type Option interface {
// contains filtered or unexported methods
}
Option configures an Agent at construction. Options are small values returned by Instructions, MaxSteps, Parallelism, Sequential, StopWhen, Output, the PolicyOptions (MaxResultBytes, Timeout, StrictInput), and the Tool constructor.
func DetectLoops ¶ added in v0.2.0
DetectLoops fails a run with ErrLoopDetected when repeats consecutive steps request the same set of tool calls — the same names with the same arguments, in any order. Varying arguments are not a loop: a corrected retry has a different signature. Results are not part of the signature: a tool whose output carries a timestamp must not hide a loop, and a result-inclusive hash would miss it. Off by default (values below 2 are ignored); the CLI scaffold writes DetectLoops(5).
Example ¶
A stuck model repeating one request: DetectLoops fails the run loudly instead of burning the step budget.
package main
import (
"context"
"errors"
"fmt"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
echo := weft.Tool("echo", "", func(_ context.Context, _ struct{}) (string, error) {
return "ok", nil
})
turns := []wefttest.Turn{}
for range 3 {
turns = append(turns, wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"i":1}`}))
}
agt := weft.New(wefttest.Script(turns...), echo, weft.DetectLoops(3))
_, err := agt.Generate(context.Background(), weft.Prompt("q"))
fmt.Println(errors.Is(err, weft.ErrLoopDetected))
}
Output: true
func Instructions ¶
Instructions sets the agent's system prompt.
func Logger ¶ added in v0.2.0
Logger sets the logger the agent's runs report to, at Debug level: one line when a run starts and ends, one per model call, one per tool call — run and call ids, the model, durations, usage, stop reasons and outcomes; never message text, tool arguments or tool results. Error text is the one exception: a failed run, model call or tool call logs the error it returned, because a log is the caller's. nil (the default) means slog.Default, which is silent until its handler enables Debug, so weft logs nothing in a program that did not ask; slog.New(slog.DiscardHandler) turns the lines off outright. The lines carry the span-carrying context, so a handler that bridges to OTel correlates them with the spans for free. mw.Log and mw.Audit are separate: middleware the caller places, at the level the caller chooses.
Example ¶
ExampleLogger shows the whole log surface: one line per phase, at Debug, ids and counts only.
package main
import (
"context"
"log/slog"
"os"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
scrub := func(_ []string, a slog.Attr) slog.Attr {
switch a.Key {
case slog.TimeKey, "dur": // timestamps and durations are not reproducible
return slog.Attr{}
}
return a
}
logger := slog.New(slog.NewTextHandler(os.Stdout, &slog.HandlerOptions{
Level: slog.LevelDebug,
ReplaceAttr: scrub,
}))
echo := weft.Tool("echo", "Echo the message.", func(_ context.Context, in struct {
Msg string `json:"msg"`
}) (string, error) {
return "echo: " + in.Msg, nil
})
agt := weft.New(wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"msg":"hi"}`}),
wefttest.Say("done"),
), weft.Logger(logger), weft.Name("demo"), echo)
if _, err := agt.Generate(context.Background(), weft.RunID("demo"), weft.Prompt("hi")); err != nil {
return
}
}
Output: level=DEBUG msg="run start" run=demo agent=demo provider=wefttest model=script level=DEBUG msg="model call" run=demo step=0 provider=wefttest model=script reason=tool_calls input_tokens=10 output_tokens=5 tool_calls=1 level=DEBUG msg="tool call" run=demo step=0 call=call_1 tool=echo result_bytes=8 level=DEBUG msg="model call" run=demo step=1 provider=wefttest model=script reason=stop input_tokens=10 output_tokens=5 tool_calls=0 level=DEBUG msg="run finish" run=demo steps=2 input_tokens=20 output_tokens=10 stop=stop
func MaxModelRetries ¶ added in v0.2.0
MaxModelRetries sets how many consecutive RETRY results (see ModelRetry) one tool may produce in a run before the run fails with ErrModelRetriesExceeded (default 3). The count is per tool name — two parallel calls to the same tool both retrying count as two — and a successful result for the tool resets it. Calls resumed under Approve do not feed the counter: they run before step 0 and belong to no step (ADR 0007). Values below 1 are ignored.
func MaxSteps ¶
MaxSteps is the safety budget: the most model calls a run may make (default 10). Exceeding it fails the run with ErrMaxSteps — a runaway loop is a failure to surface, never a quiet success. Use StopWhen for the intended end of a run. Values below 1 are ignored.
func Name ¶
Name names the agent: it appears on RunStart.Agent and in the manifest, which requires it (weft.Manifest errors on unnamed agents). Empty values are ignored.
func OnRunEnd ¶ added in v0.3.5
OnRunEnd registers an observer the loop calls exactly once per run, after RunFinish is delivered or the RunError is built and before Run returns — the outcome signal a tap cannot carry: a failed run emits no event after its last delivered one (ADR 0004), so failure is invisible to Tap, and neither middleware seam wraps the run. res is the run's result and err its error: on success err is nil and res is complete; on failure err is the *RunError and res its partial transcript (ADR 0002) — the same value RunError.Result holds; on cancellation err is the RunError wrapping the ctx error. For a Subagent's child run it fires inside the parent's tool call, like the child's events, and res.ID names the child. A run that panics (a PrepareStep function is arbitrary user code) never fires it: the run crashed, it did not end — the panic still reaches the caller. Like Tap it observes and cannot change anything (ADR 0006's note: not a third seam — a seam wraps a call, this observes an outcome); like Tap a panic in it is contained and counted (TapPanics). Several OnRunEnd options run in registration order. ctx is the run's span-carrying context. The one consumer this ships for is the store (TODO §11): store.Record pairs a Tap for the event stream with OnRunEnd for the result and the failure.
Example ¶
OnRunEnd is the outcome observer: unlike Tap it sees the run's end even when the run fails, because a failed run emits no event after its last delivered one. The store's Record option pairs the two.
package main
import (
"context"
"errors"
"fmt"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
agt := weft.New(
wefttest.Script(wefttest.Fail(errors.New("provider down"))),
weft.OnRunEnd(func(_ context.Context, res *weft.RunResult, err error) {
fmt.Printf("run ended: has id=%v failed=%v steps=%d\n", res.ID != "", err != nil, res.NumSteps())
}),
)
_, _ = agt.Generate(context.Background(), weft.Prompt("hi"))
}
Output: run ended: has id=true failed=true steps=0
func Options ¶ added in v0.2.0
Options composes several options into one, applied in order — a family of tools and policy closed over its dependencies as a single value. A plugin is `func(deps) weft.Option`, and dependencies are parameters, never globals; nothing registers itself, so there is no registry, no scopes, no dedup. Nil entries are ignored; duplicate tool names still panic at New. Nesting composes by construction.
Example ¶
Composing agents: a plugin is func(deps) weft.Option — a family of tools and its policy closed over its dependencies as one value. Dependencies are parameters, never globals.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
orders := func(deps *string) weft.Option {
return weft.Options(
weft.Instructions("You handle orders."),
weft.Tool("lookup_order", "Look up an order by id.",
func(_ context.Context, in struct {
ID string `json:"id" jsonschema:"the order id"`
}) (string, error) {
return "order " + in.ID + ": " + *deps, nil
}),
weft.MaxResultBytes(1024),
)
}
agt := weft.New(wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":"42"}`}),
wefttest.Say("Done."),
), orders(new(string)))
res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 42?"))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Steps[0].Results[0].Content)
}
Output: order 42:
func Output ¶ added in v0.2.0
Output constrains the run's final answer to Out. It registers a tool named submit_output whose input schema is reflected from Out, exactly as Tool reflects a handler's input, and stops the run once the model has called it with arguments that decode. Invalid arguments come back to the model as an ErrInvalidToolInput result naming the field, so repair is the ordinary tool-error loop — nothing special happens. Read the value with GenerateAs, or OutputOf after Stream:
type Verdict struct {
Approved bool `json:"approved"`
Reason string `json:"reason" jsonschema:"one sentence"`
}
agt := weft.New(model, weft.Output[Verdict](), lookup)
v, res, err := weft.GenerateAs[Verdict](ctx, agt, weft.Prompt("Review order 42."))
Tool mode works on every provider; adapters with a native JSON-schema mode may use it later behind the same API. Output panics if Out is not a struct (or a pointer to one), for the reason Tool does.
func Parallelism ¶
Parallelism sets the maximum number of a step's tool calls executing at once (default 4). Values below 1 are ignored.
func PrepareStep ¶ added in v0.2.0
func PrepareStep(fn func(ctx context.Context, step int, req ModelRequest) (ModelRequest, error)) Option
PrepareStep installs a function the loop calls before every model call, with the request it built for step: the agent's instructions, the transcript so far, the step's tool snapshot, the run's thinking level. The function returns the request the step uses — trimmed messages, a subset of tools, a rewritten system — or an error that fails the run. What it returns is what the step advertises and dispatches against: a tool it removes cannot be called that step, and a ToolDef it adds can. PromptSnippets are composed after it, from the tools it returns, so removing a tool removes its snippet without the function having to know snippets exist. The request is a copy the function may mutate freely: message parts, argument bytes, and tool definitions are cloned before the chain runs, so in-place writes reach neither the transcript nor the agent's frozen registry. The transcript in RunResult is never affected; only the request is. It runs before the model seam, so WrapModel middleware sees the prepared request. Several PrepareStep options run in order, each receiving the previous one's result. A nil function is ignored. Like ToolSource, it is one of the two knobs that can break a prompt-cache prefix — trim deliberately.
Example ¶
Phased tool exposure: the first step plans with read-only tools; the second acts. PrepareStep is the one loop knob — what it returns is what the step both advertises and dispatches against.
package main
import (
"context"
"fmt"
"log"
"slices"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
lookup := weft.Tool("lookup", "Look up an order.", func(_ context.Context, _ struct{}) (string, error) {
return "order 1234: broken item", nil
})
refund := weft.Tool("refund", "Refund an order.", func(_ context.Context, _ struct{}) (string, error) {
return "refunded", nil
})
phase := func(_ context.Context, step int, req weft.ModelRequest) (weft.ModelRequest, error) {
if step == 0 { // investigate before acting
req.Tools = slices.DeleteFunc(req.Tools, func(t *weft.ToolDef) bool { return t.Name == "refund" })
}
return req, nil
}
agt := weft.New(wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "refund"}), // not advertised in step 0
wefttest.ToolCalls(wefttest.Call{Name: "refund"}), // now it is
wefttest.Say("Done."),
), lookup, refund, weft.PrepareStep(phase))
res, err := agt.Generate(context.Background(), weft.Prompt("Refund order 1234."))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Steps[0].Results[0].Content)
fmt.Println(res.Steps[1].Results[0].Content)
}
Output: NO_SUCH_TOOL: no tool named "refund" refunded
func StopWhen ¶
func StopWhen(conds ...StopCondition) Option
StopWhen adds stop conditions; the run ends when any one is met. Without StopWhen a run ends when the model replies without requesting tools. Compose the built-ins — HasToolCall, StepCountIs — or write your own:
weft.StopWhen(weft.HasToolCall("submit_answer"))
weft.StopWhen(weft.StopFunc(func(steps []weft.StepRecord) bool { ... }))
Stop conditions are the intended end of a run; MaxSteps is the safety budget behind them.
func Tap ¶
Tap registers an observer that sees every event of every run, including runs made with Generate, synchronously and in emission order on the emitting goroutine. It must be fast and must not block: it runs under the event-ordering lock, so a slow tap delays every tool event of its step and blocks the emitting tool goroutines — the same consumer-speed coupling Run.Events documents. Slow observation (a database write, a network sink) must not happen inside the tap: hand each event to a queue and drain it on your own goroutine; ExampleTap_async is the tested pattern. Taps run in registration order; a panic in one is recovered, counted (see TapPanics), and dropped, so a broken observer cannot break a run. Taps observe and cannot change anything — behaviour attaches at the two middleware seams. ctx is the run's context, which carries the run's span, so a tap that starts its own spans parents them under it for free.
Example ¶
Tap observes every event of every run — including Generate, which has no stream to range over. Taps see; the middleware seams change.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
var calls int
agt := weft.New(
wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "roll_dice"}),
wefttest.Say("rolled"),
),
weft.Tap(func(_ context.Context, ev weft.Event) {
if _, ok := ev.(weft.ToolStart); ok {
calls++
}
}),
weft.Tool("roll_dice", "Roll a die.",
func(_ context.Context, _ struct{}) (int, error) { return 4, nil }),
)
if _, err := agt.Generate(context.Background(), weft.Prompt("Roll.")); err != nil {
log.Fatal(err)
}
fmt.Println("tool calls:", calls)
}
Output: tool calls: 1
Example (Async) ¶
Slow observers (a database write, a network sink) must not run inside a Tap: taps are synchronous and run under the step's event-ordering lock, so a slow tap delays every tool event of its step. The pattern: hand each event to a queue inside the tap — never block — and drain it on your own goroutine, which may be as slow as it likes. A bounded queue with a visible drop counter keeps a stuck drain from wedging the run.
package main
import (
"context"
"fmt"
"log"
"sync/atomic"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
events := make(chan weft.Event, 1024)
var dropped atomic.Int64
done := make(chan int)
go func() {
starts := 0
for ev := range events {
if _, ok := ev.(weft.ToolStart); ok {
starts++
}
}
done <- starts
}()
agt := weft.New(
wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "ping"}),
wefttest.Say("done"),
),
weft.Tap(func(_ context.Context, ev weft.Event) {
select {
case events <- ev: // fast: hand to the belt
default: // full belt: drop, but visibly
dropped.Add(1)
}
}),
weft.Tool("ping", "", func(_ context.Context, _ struct{}) (string, error) {
return "pong", nil
}),
)
if _, err := agt.Generate(context.Background(), weft.Prompt("hi")); err != nil {
log.Fatal(err)
}
close(events)
fmt.Println("tool starts:", <-done, "dropped:", dropped.Load())
}
Output: tool starts: 1 dropped: 0
func ToolSource ¶
ToolSource replaces the tool set the loop advertises and dispatches against with the given function's return value, fetched fresh exactly once per step (and once per manual Agent.CallTool): the step's advertisement and its dispatch both resolve against that one snapshot, so what the model was shown is exactly what runs. The seam is for registries that change while the agent runs (plugins installed mid-run, MCP servers polled per step); a tool registered mid-step becomes callable on the next step's fetch. The Agent stays immutable: the source is a value; synchronization and freshness of the list belong to the source's owner. A snapshot with a duplicate name fails the run with ErrDuplicateTool and one with a nil entry with ErrNilTool (the runtime analogues of New's duplicate-name panic) rather than silently dropping the second tool or advertising a dereference. A nil function (the default) keeps the static construction-time list — Tool/option registration is then the only source of tools, byte-identical to an agent without a source. Manifest and Agent.Tools still report the static construction-time set: a manifest describes the code, not the registry behind a source.
func TracerProvider ¶ added in v0.2.0
func TracerProvider(tp trace.TracerProvider) Option
TracerProvider sets the OpenTelemetry tracer provider the agent's runs report spans to. Without it, runs use the global provider (otel.GetTracerProvider), which is a no-op until an SDK registers one — so a program that sets up an SDK gets weft's spans with no option at all. Tests and dependency-injected programs pass their own provider here instead of touching the global. Every run reports one invoke_agent span, one chat span per model call, and one execute_tool span per executed tool call, with the GenAI semantic attributes (ADR 0016); no message or tool-argument content is ever put on a span.
Example ¶
ExampleTracerProvider prints the span tree of a two-step run, as an exporter would render it.
package main
import (
"context"
"fmt"
"strings"
"sync"
"testing"
"go.opentelemetry.io/otel/attribute"
"go.opentelemetry.io/otel/codes"
"go.opentelemetry.io/otel/trace"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
// recProvider records every span started through it: name, kind,
// parent, attributes, status, recorded exceptions, and end state.
type recProvider struct {
trace.TracerProvider
mu sync.Mutex
spans []*recSpan
next uint64
}
func newRecProvider() *recProvider { return &recProvider{} }
func (p *recProvider) Tracer(string, ...trace.TracerOption) trace.Tracer {
return &recTracer{p: p}
}
// find returns the one span with the exact name, failing the test when
// it is missing or ambiguous — every assertion here is against a tree
// whose span names are unique per test, so start order never matters.
func (p *recProvider) find(t *testing.T, name string) *recSpan {
t.Helper()
p.mu.Lock()
defer p.mu.Unlock()
var found *recSpan
for _, s := range p.spans {
if s.name == name {
if found != nil {
t.Fatalf("span %q recorded more than once", name)
}
found = s
}
}
if found == nil {
var names []string
for _, s := range p.spans {
names = append(names, s.name)
}
t.Fatalf("span %q not recorded; have %v", name, names)
}
return found
}
type recTracer struct {
trace.Tracer
p *recProvider
}
func (t *recTracer) Start(ctx context.Context, name string, opts ...trace.SpanStartOption) (context.Context, trace.Span) {
cfg := trace.NewSpanStartConfig(opts...)
p := t.p
p.mu.Lock()
p.next++
n := p.next
parent := trace.SpanContextFromContext(ctx)
traceID := parent.TraceID()
if !parent.IsValid() {
var b [16]byte
b[14], b[15] = byte(n>>8), byte(n)
traceID = trace.TraceID(b)
}
var sid [8]byte
sid[7] = byte(n)
sc := trace.NewSpanContext(trace.SpanContextConfig{
TraceID: traceID,
SpanID: trace.SpanID(sid),
TraceFlags: trace.FlagsSampled,
})
s := &recSpan{name: name, kind: cfg.SpanKind(), sc: sc, parent: parent, recording: true}
p.spans = append(p.spans, s)
p.mu.Unlock()
return trace.ContextWithSpan(ctx, s), s
}
type recSpan struct {
trace.Span
mu sync.Mutex
name string
kind trace.SpanKind
sc trace.SpanContext
parent trace.SpanContext
attrs []attribute.KeyValue
status codes.Code
statusMsg string
events []string
ended bool
recording bool
}
func (s *recSpan) End(...trace.SpanEndOption) {
s.mu.Lock()
defer s.mu.Unlock()
s.ended, s.recording = true, false
}
func (s *recSpan) SpanContext() trace.SpanContext { return s.sc }
func (s *recSpan) IsRecording() bool {
s.mu.Lock()
defer s.mu.Unlock()
return s.recording
}
func (s *recSpan) SetAttributes(attrs ...attribute.KeyValue) {
s.mu.Lock()
defer s.mu.Unlock()
s.attrs = append(s.attrs, attrs...)
}
func (s *recSpan) SetStatus(c codes.Code, msg string) {
s.mu.Lock()
defer s.mu.Unlock()
s.status, s.statusMsg = c, msg
}
func (s *recSpan) RecordError(err error, _ ...trace.EventOption) {
s.mu.Lock()
defer s.mu.Unlock()
s.events = append(s.events, "exception: "+err.Error())
}
func (s *recSpan) Name() string {
s.mu.Lock()
defer s.mu.Unlock()
return s.name
}
// attrsMap flattens the span's attributes for exact-set assertions.
func (s *recSpan) attrsMap() map[string]string {
s.mu.Lock()
defer s.mu.Unlock()
out := map[string]string{}
for _, kv := range s.attrs {
out[string(kv.Key)] = kv.Value.String()
}
return out
}
func (s *recSpan) state() (ended bool, status codes.Code, statusMsg string, events []string) {
s.mu.Lock()
defer s.mu.Unlock()
return s.ended, s.status, s.statusMsg, append([]string(nil), s.events...)
}
// spanEcho is the standard tool of the span tests: one string in, one
// string out, nothing async.
var spanEcho = weft.Tool("echo", "Echo the message.", func(_ context.Context, in struct {
Msg string `json:"msg"`
}) (string, error) {
return "echo: " + in.Msg, nil
})
// spanScript is the standard two-step script: one echo call, then a
// final reply. wefttest fixes usage at 10 input / 5 output tokens per
// step, so the token attributes are assertable.
func spanScript() *wefttest.Model {
return wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "echo", Args: `{"msg":"hi"}`}),
wefttest.Say("done"),
)
}
func main() {
tp := newRecProvider()
agt := weft.New(spanScript(), weft.TracerProvider(tp), weft.Name("demo"), spanEcho)
if _, err := agt.Generate(context.Background(), weft.Prompt("hi")); err != nil {
return
}
tp.mu.Lock()
bySpan := map[[8]byte]*recSpan{}
for _, s := range tp.spans {
bySpan[s.sc.SpanID()] = s
}
var print func(s *recSpan, depth int)
print = func(s *recSpan, depth int) {
_, status, _, _ := s.state()
fmt.Printf("%s%s [%v]\n", strings.Repeat(" ", depth), s.name, status)
for _, child := range tp.spans {
if child.parent.SpanID() == s.sc.SpanID() {
print(child, depth+1)
}
}
}
for _, s := range tp.spans {
if !s.parent.IsValid() {
print(s, 0)
}
}
tp.mu.Unlock()
}
Output: invoke_agent demo [Ok] chat script [Ok] execute_tool echo [Ok] chat script [Ok]
func UsageLimit ¶ added in v0.2.0
UsageLimit bounds a run's total token usage — the run's own model calls plus every subagent's (RunResult.Usage). A field left zero is unlimited. The limit is checked after each step, before the loop makes another model call: a step that ends the run — final answer, StopWhen, pending approvals — may overshoot and still succeed, because a budget's job is to stop further spend, not to discard finished work. Exceeding it fails the run with ErrUsageLimit and the partial transcript on RunError.Result. There is no default limit: the right value is workload-specific, and MaxSteps is the default budget. Usage.Total() is not a separate limit — a caller who wants a total sets both fields.
Example ¶
A token budget: exceeded, the run fails with ErrUsageLimit before the next model call; the partial transcript rides on RunError.Result.
package main
import (
"context"
"errors"
"fmt"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
echo := weft.Tool("echo", "", func(_ context.Context, _ struct{}) (string, error) {
return "ok", nil
})
agt := weft.New(wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "echo"}),
wefttest.Say("never reached"),
), echo, weft.UsageLimit(weft.Usage{OutputTokens: 4}))
_, err := agt.Generate(context.Background(), weft.Prompt("again"))
var re *weft.RunError
if errors.As(err, &re) {
fmt.Println(errors.Is(err, weft.ErrUsageLimit), "steps kept:", len(re.Result.Steps))
}
}
Output: true steps kept: 1
func WrapModel ¶ added in v0.2.0
func WrapModel(mw ...ModelMiddleware) Option
WrapModel installs model middleware around the agent's model. The first middleware listed is the outermost — WrapModel(a, b) calls a(b(model)) — and successive WrapModel options append inward. The chain is built once, at New, and sees each step's ModelRequest as the loop built it. The reference set is in package mw: Retry, Fallback, Log, RepairJSON. Nil entries are ignored.
Example ¶
Model middleware wraps the agent's model, chi-style: the first listed is the outermost. The reference set lives in package mw.
package main
import (
"context"
"fmt"
"iter"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
logged := func(next weft.Model) weft.Model {
return loggingModel{next: next}
}
agt := weft.New(wefttest.Script(wefttest.Say("hello")), weft.WrapModel(logged))
res, err := agt.Generate(context.Background(), weft.Prompt("hi"))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Text())
}
type loggingModel struct{ next weft.Model }
func (m loggingModel) Info() weft.ModelInfo { return weft.InfoOf(m.next) }
func (m loggingModel) Stream(ctx context.Context, req weft.ModelRequest) iter.Seq2[weft.ModelEvent, error] {
fmt.Printf("model call: %d messages\n", len(req.Messages))
return m.next.Stream(ctx, req)
}
Output: model call: 1 messages hello
type OutputDecoder ¶ added in v0.3.0
type OutputDecoder[Out any] struct { // contains filtered or unexported fields }
OutputDecoder turns a streaming run's submit_output argument deltas into a filling-in Out — the UI half of structured output: a consumer rendering a form as the model writes it. It is a decoder value, not an event-stream wrapper: feed it the events you already consume and render what comes back.
dec := weft.NewOutputDecoder[Form]()
for ev, err := range run.Events() {
if err != nil { return err }
if p, ok := dec.Feed(ev); ok { render(p) }
}
form, err := dec.Result()
The decoder keys on the submit_output stream identity ToolArgsDelta carries — the tool name and step boundaries — resets its buffer on StepStart and on a fresh submit_output ToolStart, and closes it on the matching ToolFinish. Nested events are ignored: a subagent's structured output is its own decoder's job. Decode is lenient and prefix-shaped: after each delta it attempts the longest closed prefix of the arguments so far (see partial_json.go) and reports it when it changed and decoded — fields not yet present stay zero, a garbage mid-stream prefix keeps the last good partial, and errors surface only at Result, which follows the OutputOf rule (the last submit_output call with a non-error result) and returns ErrNoOutput when none finished. Never model-visible: the decoder reads the stream and writes nothing back.
Cost: every delta rescans the buffered arguments (the close is linear in what has arrived), so a submission's decode cost grows with the square of its size — the price of a fresh partial on every delta. At tool-argument scale (a few KiB) it is noise; a UI feeding very large submissions can trade freshness for linearity by calling Feed less often — the decoder keeps the last good partial across the deltas it skips.
Example ¶
An OutputDecoder renders structured output while it streams: feed it the events you already consume and draw the filling-in form.
package main
import (
"context"
"encoding/json"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
type Form struct {
Name string `json:"name"`
Count int `json:"count"`
}
turn := wefttest.Raw(
weft.ModelToolCallDelta{Index: 0, Name: "submit_output", Args: `{"name":"Ada",`},
weft.ModelToolCallDelta{Index: 0, Name: "submit_output", Args: `"count":3}`},
weft.ModelToolCall{ID: "c1", Name: "submit_output", Args: json.RawMessage(`{"name":"Ada","count":3}`)},
weft.ModelFinish{Reason: weft.StopToolCalls, Usage: weft.Usage{InputTokens: 3, OutputTokens: 2}},
)
run := weft.New(wefttest.Script(turn), weft.Output[Form]()).Stream(context.Background(), weft.Prompt("fill the form"))
dec := weft.NewOutputDecoder[Form]()
for ev, err := range run.Events() {
if err != nil {
log.Fatal(err)
}
if p, ok := dec.Feed(ev); ok {
fmt.Printf("render: name=%q count=%d\n", p.Name, p.Count)
}
}
form, err := dec.Result()
if err != nil {
log.Fatal(err)
}
fmt.Printf("final: name=%q count=%d\n", form.Name, form.Count)
}
Output: render: name="Ada" count=0 render: name="Ada" count=3 final: name="Ada" count=3
func NewOutputDecoder ¶ added in v0.3.0
func NewOutputDecoder[Out any]() *OutputDecoder[Out]
NewOutputDecoder returns a decoder ready to feed a run's events.
func (*OutputDecoder[Out]) Feed ¶ added in v0.3.0
func (d *OutputDecoder[Out]) Feed(ev Event) (partial Out, ok bool)
Feed consumes one run event. It ignores everything except this run's StepStart, the submit_output ToolStart (a whole-call arrival, and the marker of a fresh call), the submit_output ToolArgsDelta stream, and the submit_output ToolFinish; after an args delta it returns the best-effort partial Out and ok == true when the partial changed. A second submit_output call within the step starts the buffer over — Result follows the last call, the OutputOf rule; concatenating two distinct calls' arguments could only ever produce bytes that do not decode.
func (*OutputDecoder[Out]) Result ¶ added in v0.3.0
func (d *OutputDecoder[Out]) Result() (Out, error)
Result decodes the completed submission — the last submit_output call with a non-error result, the OutputOf rule — and returns ErrNoOutput when no valid call finished.
type ParamsOption ¶ added in v0.3.0
ParamsOption is accepted by both New and Stream/Generate: sampling knobs are as per-question as a forced choice (the ThinkingOption shape).
func Params ¶ added in v0.3.0
func Params(p RequestParams) ParamsOption
Params sets the agent's default sampling knobs — Temperature, TopP, MaxTokens, Stop, Seed — for every run's model calls; as a RunOption it overrides that default for one run:
agt := weft.New(m, weft.Params(weft.RequestParams{Temperature: ptr(0.2)}))
agt.Generate(ctx, weft.Params(weft.RequestParams{Temperature: ptr(0.9)}), weft.Prompt(q))
A nil or empty field keeps the adapter's construction-time default for that knob (its Temperature option, and so on); a run-level Params replaces the agent's struct whole, never merging field by field — set every knob the override should carry. A PrepareStep function can edit the request's Params per step ("cold for classification steps, creative for drafting" is a two-line function). Adapters drop knobs their provider lacks, and the provider's own limits apply (google narrows Seed to int32).
Example ¶
Params sets per-step sampling: a PrepareStep function turns the temperature down for the classifying step and back for drafting — one struct, edited per step, no second Model construction.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
p := func(f float64) *float64 { return &f }
classify := weft.Tool("classify", "Classify the request.",
func(_ context.Context, _ struct{}) (string, error) { return "billing", nil })
agt := weft.New(
wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "classify"}),
wefttest.Say("A billing question, answered at temperature 0.9."),
),
weft.Params(weft.RequestParams{Temperature: p(0.9)}),
weft.PrepareStep(func(_ context.Context, step int, req weft.ModelRequest) (weft.ModelRequest, error) {
if step == 0 {
req.Params.Temperature = p(0) // cold for classification
}
return req, nil
}),
classify,
)
res, err := agt.Generate(context.Background(), weft.Prompt("Why did my invoice double?"))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Text())
}
Output: A billing question, answered at temperature 0.9.
type Part ¶
type Part interface {
// contains filtered or unexported methods
}
Part is one content part of a Message. The set of part types is closed: text, tool calls, tool results, reasoning, and files today; approval parts are planned additions that will join this interface.
type PolicyOption ¶ added in v0.2.0
type PolicyOption interface {
Option
ToolOption
}
PolicyOption is accepted by both New and Tool. On an agent it is the default policy for every tool call; on a tool it overrides the agent's default for that tool. The manifest records both levels.
func MaxResultBytes ¶
func MaxResultBytes(n int) PolicyOption
MaxResultBytes sets the maximum size of one tool result's text, in bytes (default 64 KiB). The loop caps longer results — successes, failures, and panics alike — cutting on a rune boundary and appending a marker the model can see, so it knows the output is partial. MaxResultBytes(0) removes the cap; negative values are ignored. On a tool it overrides the agent's cap for that tool alone — a per-tool MaxResultBytes(0) lifts the cap for a tool whose output must arrive whole. Agent.CallTool returns uncapped output: the cap is a run policy, applied by the loop.
func Sequential ¶
func Sequential() PolicyOption
Sequential restricts a step's tool calls to run one at a time, in call order: each tool finishes before the next starts — the safe setting for tools with shared state. Under any parallelism, tools *start* in call order; Sequential additionally serializes their execution. It also sets ModelRequest.SequentialTools, so adapters ask the provider not to emit parallel batches in the first place.
On a tool, Sequential is a barrier: the dispatcher lets every in-flight call of the step finish, runs this call alone, then resumes the step's parallelism for the calls after it. Other tools keep running in parallel; only this one is serialized.
Example ¶
A Sequential tool is a barrier: it runs alone. The step's calls in flight finish first, and the calls after it wait — under any parallelism. Results stay in call order either way.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
write := weft.Tool("write", "Append to the ledger.", func(_ context.Context, _ struct{}) (string, error) {
return "written", nil
}, weft.Sequential())
check := weft.Tool("check", "Verify the ledger.", func(_ context.Context, _ struct{}) (string, error) {
return "verified", nil
})
model := wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "check"}, wefttest.Call{Name: "write"}, wefttest.Call{Name: "check"}),
wefttest.Say("done"),
)
res, err := weft.New(model, write, check, weft.Parallelism(4)).Generate(context.Background(), weft.Prompt("x"))
if err != nil {
log.Fatal(err)
}
for _, r := range res.Steps[0].Results {
fmt.Println(r.Name, "->", r.Content)
}
}
Output: check -> verified write -> written check -> verified
func StrictInput ¶ added in v0.2.0
func StrictInput() PolicyOption
StrictInput rejects tool arguments that carry fields the input struct does not declare, as an ErrInvalidToolInput result naming the field. The default is lenient — undeclared fields are ignored, the encoding/json default — because models routinely add stray keys and a rejection costs a round trip, not accuracy. Use it where an ignored field would be a silent misinterpretation of the call. On the agent it applies to every tool; on a tool to that tool alone. RawTool handlers are unaffected: they receive the raw arguments.
func Timeout ¶ added in v0.2.0
func Timeout(d time.Duration) PolicyOption
Timeout bounds one tool call. On a tool it is that tool's deadline — Timeout(0) on a tool removes the agent's default for that tool alone, the timeout analogue of MaxResultBytes(0); on the agent it is the default for every tool without its own, and non-positive values are ignored there. The handler's ctx carries the deadline; when it expires the loop records an error result ("tool X timed out after 10s") the model sees and moves on, abandoning the handler's goroutine — handlers must honour ctx to release their resources. Without any Timeout only the run's ctx bounds a call.
Example ¶
Per-tool policy: trailing options on Tool override the agent's defaults for that tool alone. Here a slow tool times out into an error result the model sees, and the run carries on.
package main
import (
"context"
"fmt"
"log"
"time"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
slow := weft.Tool("slow", "Takes a while.",
func(ctx context.Context, _ struct{}) (string, error) {
<-ctx.Done() // a well-behaved handler honours the deadline
return "", ctx.Err()
},
weft.Timeout(10*time.Millisecond))
model := wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "slow"}),
wefttest.Say("It did not answer in time."),
)
agt := weft.New(model, slow)
res, err := agt.Generate(context.Background(), weft.Prompt("Try the slow tool."))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Steps[0].Results[0].Content)
fmt.Println(res.Text())
}
Output: tool "slow" timed out after 10ms It did not answer in time.
func WrapTools ¶ added in v0.2.0
func WrapTools(mw ...ToolMiddleware) PolicyOption
WrapTools installs tool middleware. On the agent it wraps every tool call the loop (and Agent.CallTool) dispatches; on a tool it wraps that tool alone, inside the agent's chain: agent middleware → tool middleware → decode → handler. The first middleware listed is the outermost, so WrapTools(a, b) runs a(b(call)), and successive WrapTools options append inward. Middleware sees CallFromContext and may decorate ctx for the handler (see ExampleWrapTools_context); it returns the result text or an error — a *ToolError to give the model a code, an error wrapping ErrApprovalRequired to park the call on RunResult.Pending. The reference set is in package mw: Allow, Audit, MapErrors. Nil entries are ignored.
Example ¶
Tool middleware wraps every call the loop dispatches. A middleware that returns an error produces an error result the model sees; a *weft.ToolError gives it a code.
package main
import (
"context"
"fmt"
"log"
"strings"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
readOnly := func(next weft.ToolCaller) weft.ToolCaller {
return func(ctx context.Context, call weft.ToolCallPart) (string, error) {
if strings.HasPrefix(call.Name, "delete_") {
return "", &weft.ToolError{Code: "DENIED", Message: "this agent is read-only"}
}
return next(ctx, call)
}
}
del := weft.Tool("delete_order", "", func(_ context.Context, _ struct{}) (string, error) { return "deleted", nil })
agt := weft.New(wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "delete_order"}),
wefttest.Say("I cannot do that."),
), del, weft.WrapTools(readOnly))
res, err := agt.Generate(context.Background(), weft.Prompt("delete order 1"))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Steps[0].Results[0].Content)
fmt.Println(res.Text())
}
Output: DENIED: this agent is read-only I cannot do that.
Example (Context) ¶
The context-decoration convention: middleware that has verified something (here, who is calling) adds it to ctx before next, and the handler reads it back through a typed accessor — the same shape as weft.CallFromContext. Decorate only with data the middleware has verified; derive business values in the handler after decode.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
authUser := func(next weft.ToolCaller) weft.ToolCaller {
return func(ctx context.Context, call weft.ToolCallPart) (string, error) {
user, err := verifyUser(ctx) // a session lookup, a token check, ...
if err != nil {
return "", &weft.ToolError{Code: "UNAUTHENTICATED", Message: "sign in first", Err: err}
}
return next(withUser(ctx, user), call)
}
}
myOrders := weft.Tool("my_orders", "List the caller's orders.",
func(ctx context.Context, _ struct{}) (string, error) {
u, ok := userFromContext(ctx)
if !ok {
return "", weft.Errorf("UNAUTHENTICATED", "no user on the call")
}
return "orders for " + u.Name, nil
})
agt := weft.New(wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "my_orders"}),
wefttest.Say("Here they are."),
), myOrders, weft.WrapTools(authUser))
res, err := agt.Generate(context.Background(), weft.Prompt("show my orders"))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Steps[0].Results[0].Content)
}
type exampleUser struct{ Name string }
type userKey struct{}
func withUser(ctx context.Context, u exampleUser) context.Context {
return context.WithValue(ctx, userKey{}, u)
}
func userFromContext(ctx context.Context) (exampleUser, bool) {
u, ok := ctx.Value(userKey{}).(exampleUser)
return u, ok
}
func verifyUser(context.Context) (exampleUser, error) { return exampleUser{Name: "ada"}, nil }
Output: orders for ada
type ReasoningDelta ¶
ReasoningDelta is an increment of provider reasoning, in the order the model produced it relative to TextDelta. Signatures are not streamed; they are on the ReasoningPart of the transcript.
func (ReasoningDelta) MarshalJSON ¶
func (e ReasoningDelta) MarshalJSON() ([]byte, error)
MarshalJSON encodes the event with its "type" discriminator.
type ReasoningPart ¶
type ReasoningPart struct {
Text string `json:"text"`
Signature string `json:"signature,omitempty"`
}
ReasoningPart is one provider reasoning block surfaced by providers that expose it (Anthropic thinking blocks, Gemini thought parts). Weft preserves it in the transcript but does not act on it. Signature is the provider's opaque token for the block (Anthropic rejects thinking sent back without its signature); adapters echo it unchanged. A step yields one ReasoningPart per provider block — a delta carrying a signature closes the block (see ModelReasoningDelta) — placed before the TextPart of the same assistant message, in the order the model produced them.
Example ¶
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
model := wefttest.Script(
wefttest.Think("The user greets; reply in kind.", wefttest.Say("Hello!")),
)
res, err := weft.New(model).Generate(context.Background(), weft.Prompt("Hi."))
if err != nil {
log.Fatal(err)
}
for _, p := range res.Messages[1].Content {
fmt.Printf("%T\n", p)
}
}
Output: weft.ReasoningPart weft.TextPart
func (ReasoningPart) MarshalJSON ¶
func (p ReasoningPart) MarshalJSON() ([]byte, error)
MarshalJSON encodes the part with its "type" discriminator.
type RequestParams ¶ added in v0.3.0
type RequestParams struct {
Temperature *float64
TopP *float64
MaxTokens *int
Stop []string
Seed *int64
}
RequestParams is per-step sampling: the knobs a caller turns between "cold for classification, creative for drafting". Every field is a pointer or slice so the three states stay distinguishable — nil or empty keeps the adapter's construction default (its Temperature, TopP, MaxTokens, Stop, or Seed option), set overrides it for this request alone, and the adapter never replaces a construction value with a zero. A set pointer to 0 is a value (Temperature of exactly 0 is sent), with one provider exception: an anthropic MaxTokens of 0 falls to the adapter's default, because the API requires a positive value. A negative MaxTokens fails the run at the step that carries it — no provider accepts one, and the loop names the bug rather than letting each adapter improvise. A run-level Params option replaces the agent's struct whole, it does not merge field by field; PrepareStep can edit it per step. Adapters drop knobs their provider lacks (Seed on anthropic) under the "adapters document what they drop" rule.
type Role ¶
type Role string
Role is the author of a Message.
There is no system role: the system instruction is agent-level (Instructions) and travels on ModelRequest.System, so a transcript never carries it and adapters never have to merge it.
type Run ¶
type Run struct {
// contains filtered or unexported fields
}
Run is a handle to one streaming execution. Create it with Agent.Stream, then either range over Events (exactly once) or call Wait, which runs the agent to completion and reports the final result.
func (*Run) Close ¶ added in v0.2.0
func (r *Run) Close()
Close releases the run's resources, canceling it if still running. It is for abandoned runs: a run you will consume needs no Close — Events and Wait release everything themselves. Safe to call any number of times, before or after consumption.
func (*Run) Events ¶
Events returns the run's event stream. It is single-use; a second call yields only ErrRunConsumed.
Events arrive in emission order (see ToolStart for the concurrent-tool ordering rule). Delivery is a direct hand-off over an unbuffered channel, and tool events are emitted while holding the step's event-ordering lock — so a slow consumer does not merely receive late: it delays event emission and gates the start of the step's subsequent tools. For latency-sensitive parallel tools, consume promptly (range over Events in a dedicated goroutine that buffers) or use Generate, which needs no consumer.
A failed run delivers its error exactly once as the final element; a successful run ends with RunFinish. Breaking out of the range cancels the run.
type RunError ¶
RunError reports a step-scoped failure: the model stream failed, the context was canceled, or the step budget ran out. Err is the cause — use errors.Is/As on it. Result carries the transcript up to the failure, so partial work is never lost.
type RunFinish ¶
type RunFinish struct {
RunID string `json:"run_id"`
Usage Usage `json:"usage"`
Steps int `json:"steps"`
// Pending mirrors RunResult.Pending: calls the run ended on without
// executing, awaiting Approve/Deny. Such a call had its ToolStart
// and no ToolFinish.
Pending []ToolCallPart `json:"pending,omitempty"`
}
RunFinish is always the final event of a successful run and carries the run's total usage and step count.
func (RunFinish) MarshalJSON ¶
MarshalJSON encodes the event with its "type" discriminator.
type RunOption ¶
type RunOption interface {
// contains filtered or unexported methods
}
RunOption configures a single run.
func Approve ¶ added in v0.2.0
Approve resumes a call left on RunResult.Pending by an earlier run: pass the earlier transcript with Messages and the decision, and the loop executes the call — through the ordinary tool chain, with Call.Approved set — before its next model call. Ids that are not pending are ignored.
func Deny ¶ added in v0.2.0
Deny resolves a pending call without running it: the model sees an error result reading "DENIED: <reason>" and the loop continues. Pending calls given neither Approve nor Deny are denied with the reason "no decision" ("DENIED: no decision").
func Messages ¶
Messages adds existing messages (a session transcript, few-shot examples) to the run's input.
func Resolve ¶ added in v0.3.0
Resolve resumes a pending call with a result computed outside the process — the human-as-tool-executor shape: run the query in prod, paste what happened, and the next model call sees it. The content becomes the call's ToolResultPart verbatim; the handler never runs, no execute_tool span or ToolStart/ToolFinish is emitted (nothing executed), and Call.Approved is never set. MaxResultBytes applies as to any result. Composes with Approve and Deny in one resuming call; the last option for an id wins.
Resolve on a call that is not pending in the resumed transcript is a loud run error at step 0 — a deliberate asymmetry with Approve and Deny, which ignore unknown ids: those are yes/no marks over an id set, while Resolve carries a payload the caller expects the model to see, and dropping it silently is the one thing the error model forbids (ADR 0007's 2026-09-22 amendment).
func ResolveError ¶ added in v0.3.0
ResolveError is Resolve with the result marked as an error: the model sees the content on an error result, the shape a failed execution would have produced.
type RunResult ¶
type RunResult struct {
ID string
// StopReason is the last step's finish reason. StopMaxTokens here
// means the final reply was cut off by the output-token limit — the
// run still succeeds, and the caller decides what truncated text
// means.
StopReason StopReason
Messages []Message
Steps []StepRecord
Usage Usage
// Pending lists the tool calls of the last step that await an
// approval decision (RequireApproval, or middleware returning
// ErrApprovalRequired). The run ended successfully without running
// them and the transcript carries no result for them; resume with
// Messages(res.Messages...) plus Approve/Deny per call.
Pending []ToolCallPart
}
RunResult is the outcome of a completed run: its id, the full transcript (including the input messages), one record per step, and summed usage.
func GenerateAs ¶ added in v0.2.0
GenerateAs runs the agent and returns its structured output, decoded from the last valid submit_output call. The agent must have been built with Output[Out]; a run that ends without a valid submission returns ErrNoOutput alongside the result, so the transcript is still inspectable. Run errors are *RunError as for Generate.
Example ¶
Structured output: Output constrains the final answer to a struct, and GenerateAs returns it decoded. An invalid submission is an ordinary tool error the model repairs; a valid one ends the run.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
type Verdict struct {
Approved bool `json:"approved"`
Reason string `json:"reason" jsonschema:"one sentence"`
}
model := wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "submit_output", Args: `{"approved":true,"reason":"within policy"}`}),
)
agt := weft.New(model, weft.Instructions("Review refund requests."), weft.Output[Verdict]())
v, res, err := weft.GenerateAs[Verdict](context.Background(), agt, weft.Prompt("Refund order 42?"))
if err != nil {
log.Fatal(err)
}
fmt.Println(v.Approved, v.Reason)
fmt.Println("steps:", res.NumSteps())
}
Output: true within policy steps: 1
type RunStart ¶
type RunStart struct {
ID string `json:"id"`
Model ModelInfo `json:"model"`
Agent string `json:"agent,omitempty"`
}
RunStart is always the first event of a run and carries its id and, when reported, the model's identity and the agent's name.
func (RunStart) MarshalJSON ¶
MarshalJSON encodes the event with its "type" discriminator.
type Schema ¶
type Schema struct {
Type string `json:"type,omitempty"`
Format string `json:"format,omitempty"`
Description string `json:"description,omitempty"`
Properties map[string]*Schema `json:"properties,omitempty"`
// AdditionalProperties types a map's values (JSON Schema draft
// 2020-12), so map[string]int stops being "some object". The
// boolean false form is not expressible: reflection always has a
// value type, and hand-written RawTool schemas that need to lock
// properties down must say so in Description.
AdditionalProperties *Schema `json:"additionalProperties,omitempty"`
Required []string `json:"required,omitempty"`
Items *Schema `json:"items,omitempty"`
// contains filtered or unexported fields
}
Schema is the subset of JSON Schema (draft 2020-12) that weft derives from tool input structs. Its shape tracks what the official Go MCP SDK derives via google/jsonschema-go, so tools cross over to MCP without conversion.
func ParseSchema ¶ added in v0.2.0
func ParseSchema(b json.RawMessage) (*Schema, error)
ParseSchema reads a JSON Schema document from outside Go into a *Schema: the structured fields it knows are populated (Type, Properties, Required, … — the manifest and any reader that walks the tree see them), and the document's own bytes are kept and re-emitted by Schema.MarshalJSON, so an enum or a oneOf the Schema type cannot express still reaches the model exactly as written. The top-level type must be an object — providers and MCP both require it, and failing here, at import, beats failing at the first model call.
The structured view is lenient: a keyword whose shape the Schema type cannot hold — a boolean additionalProperties, a type array such as ["string","null"], tuple or boolean items, a non-string description — leaves that field zero (an unconstrained node) and is not an error, because the bytes carry it whole and the view is for readers, not the model. Only the document itself is checked: invalid JSON, trailing data, or a non-object top level return an error naming the problem.
func (*Schema) MarshalJSON ¶ added in v0.2.0
MarshalJSON emits the schema's parsed bytes verbatim when it came from ParseSchema — the manifest's input_schema, the adapters' tool parameters, and any other reader see exactly what the foreign server sent — and the plain struct encoding otherwise (a reflected schema has no bytes to honour, so existing goldens are unchanged).
type StepFinish ¶
type StepFinish struct {
RunID string `json:"run_id"`
Index int `json:"index"`
Reason StopReason `json:"reason"`
Usage Usage `json:"usage"`
Raw string `json:"raw,omitempty"`
}
StepFinish reports that step Index is complete: the model call finished and, when the step requested tools, they have run — it follows the step's ToolFinish events and precedes the stop-condition check (docs/life-of-a-call.md). Raw is the provider's own stop reason when Reason was approximated (see ModelFinish.Raw); empty when the mapping was exact.
func (StepFinish) MarshalJSON ¶
func (e StepFinish) MarshalJSON() ([]byte, error)
MarshalJSON encodes the event with its "type" discriminator.
type StepRecord ¶
type StepRecord struct {
Index int
// StopReason is the mapped reason (stop, tool_calls, max_tokens);
// RawStopReason is the provider's own value when the mapping was
// approximated ("refusal", "content_filter", ...) — see
// ModelFinish.Raw.
StopReason StopReason
RawStopReason string
Usage Usage
Text string
ToolCalls []ToolCallPart
Results []ToolResultPart
// SubagentUsage is the usage of each child run this step started,
// keyed by the parent's call id — including a failed child's partial
// usage. It is already included in RunResult.Usage; Usage above is
// the step's own model call only. Nil when the step ran no subagent.
SubagentUsage map[string]Usage
}
StepRecord captures everything one model step produced: its text, the tool calls it requested, and the results of executing them in call order.
type StepStart ¶
StepStart reports that the model is being called for step Index.
func (StepStart) MarshalJSON ¶
MarshalJSON encodes the event with its "type" discriminator.
type StopCondition ¶
type StopCondition interface {
Stop(steps []StepRecord) bool
}
StopCondition decides, after a step's tool calls have run, whether the run is complete. It sees every step so far; the last element is the step just finished. Returning true ends the run successfully without another model call. The built-ins — HasToolCall, StepCountIs — also implement fmt.Stringer so the manifest (TODO §2.9) can name them; adapt an ordinary function with StopFunc.
func HasToolCall ¶
func HasToolCall(names ...string) StopCondition
HasToolCall stops the run once the step just finished called any of the named tools — the "final answer tool" pattern.
func StepCountIs ¶
func StepCountIs(n int) StopCondition
StepCountIs stops the run after exactly n steps, successfully — unlike MaxSteps, which treats reaching the budget as a failure.
type StopFunc ¶
type StopFunc func(steps []StepRecord) bool
StopFunc adapts an ordinary function to a StopCondition, the http.HandlerFunc shape.
func (StopFunc) Stop ¶
func (f StopFunc) Stop(steps []StepRecord) bool
Stop ends the run when f says so.
type StopReason ¶
type StopReason string
StopReason is why a model step ended.
const ( StopEndTurn StopReason = "stop" StopToolCalls StopReason = "tool_calls" StopMaxTokens StopReason = "max_tokens" )
type TextDelta ¶
TextDelta is an increment of assistant text.
func (TextDelta) MarshalJSON ¶
MarshalJSON encodes the event with its "type" discriminator.
type TextPart ¶
type TextPart struct {
Text string `json:"text"`
}
TextPart is a span of user or assistant text.
func (TextPart) MarshalJSON ¶
MarshalJSON encodes the part with its "type" discriminator.
type ThinkingConfig ¶ added in v0.2.0
type ThinkingConfig struct {
Level ThinkingLevel
Budget int64
}
ThinkingConfig is the per-run reasoning request: a level on the neutral scale, plus a token budget for providers whose depth control is a cap (Anthropic budget_tokens, Gemini thinkingBudget). A level without a Budget leaves the depth to the provider (Anthropic adaptive thinking, Gemini's own level mapping); a Budget without a level pins it. Off wins over a Budget when both are set.
type ThinkingLevel ¶ added in v0.2.0
type ThinkingLevel int
ThinkingLevel is a provider-neutral reasoning-effort scale. The zero value, ThinkUnset, sends nothing and keeps the provider default; every other value asks the adapter to express that depth on the wire in whatever form the provider has — reasoning_effort, a thinking object, a token budget. Adapters map what the provider can express and document what they drop (TODO §5.14).
const ( ThinkUnset ThinkingLevel = iota // provider default; nothing is sent ThinkOff // suppress reasoning where allowed ThinkLow ThinkMedium ThinkHigh )
type ThinkingOption ¶ added in v0.2.0
ThinkingOption is accepted by both New and Stream/Generate: reasoning depth is a per-question concern, not a per-agent one. On an agent it is the default for every run; on a run it overrides that default.
func Thinking ¶ added in v0.2.0
func Thinking(cfg ThinkingConfig) ThinkingOption
Thinking sets the reasoning level for the agent's model calls. As an Option it is every run's default; as a RunOption it overrides that default for one run — the quick-ask shape: fast by default, think on demand, without rebuilding the agent.
agt := weft.New(m, weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkOff}))
agt.Generate(ctx, weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkHigh}), weft.Prompt(q))
The zero Level keeps the provider default; adapters map what the provider can express and document what they drop.
Example ¶
Fast by default, think on demand: the agent option sets every run's default, a run option overrides it for that run alone. The scripted model records what each run asked for.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
model := wefttest.Script(wefttest.Say("ok"), wefttest.Say("ok"))
agt := weft.New(model, weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkOff}))
deep := weft.Thinking(weft.ThinkingConfig{Level: weft.ThinkHigh, Budget: 2048})
if _, err := agt.Generate(context.Background(), deep, weft.Prompt("hard")); err != nil {
log.Fatal(err)
}
if _, err := agt.Generate(context.Background(), weft.Prompt("quick")); err != nil {
log.Fatal(err)
}
for _, req := range model.Requests() {
fmt.Println(req.Thinking == weft.ThinkingConfig{Level: weft.ThinkHigh, Budget: 2048},
req.Thinking.Level == weft.ThinkOff)
}
}
Output: true false false true
type ToolArgsDelta ¶ added in v0.2.0
type ToolArgsDelta struct {
RunID string `json:"run_id"`
Name string `json:"name"`
Args string `json:"args"`
}
ToolArgsDelta reports an increment of a tool call's arguments as the model streams them — the model is "writing" the call, which can take a while for large arguments (generated code, long documents). It is progress only: the call has not been made, and ToolStart still arrives when it executes. Name is the best-known name so far; a provider that streams fragments of several calls interleaves their deltas, distinguished by name where the provider supplies one.
func (ToolArgsDelta) MarshalJSON ¶ added in v0.2.0
func (e ToolArgsDelta) MarshalJSON() ([]byte, error)
MarshalJSON encodes the event with its "type" discriminator.
type ToolCallPart ¶
type ToolCallPart struct {
ID string `json:"id"`
Name string `json:"name"`
Args json.RawMessage `json:"args"`
Signature string `json:"signature,omitempty"`
}
ToolCallPart is a tool invocation requested by the model. Args is the raw JSON the model produced; the loop unmarshals it into the tool's input type before invoking the handler. Signature is the provider's opaque token attached to the call itself (Gemini's thought signatures ride functionCall parts and must return on the same part); empty for providers without one.
func (ToolCallPart) MarshalJSON ¶
func (p ToolCallPart) MarshalJSON() ([]byte, error)
MarshalJSON encodes the part with its "type" discriminator.
type ToolCaller ¶ added in v0.2.0
type ToolCaller func(ctx context.Context, call ToolCallPart) (string, error)
ToolCaller is one link of the tool-call chain: it takes a call and returns the result text the model will see, or an error. The innermost caller decodes the arguments and runs the handler; every ToolMiddleware wraps one.
type ToolChoiceConfig ¶ added in v0.3.0
type ToolChoiceConfig struct {
Mode ToolChoiceMode
Name string
}
ToolChoiceConfig constrains what a step's model call may emit. The zero value is the provider default. Name is required when Mode is ToolChoiceNamed and must be empty under every other mode; the loop fails the run on a mismatch rather than sending a malformed choice (a programming error, not a sentinel condition).
type ToolChoiceMode ¶ added in v0.3.0
type ToolChoiceMode string
ToolChoiceMode selects how the provider must shape a step's tool calls. The zero value, ToolChoiceAuto, keeps the provider default and sends nothing; every other value asks the adapter to express the constraint in the provider's own tool_choice form.
const ( ToolChoiceAuto ToolChoiceMode = "" // provider default; nothing is sent ToolChoiceAny ToolChoiceMode = "any" // some tool must be called ToolChoiceNamed ToolChoiceMode = "tool" // the tool named by Name must be called // ToolChoiceNone forbids tool calls while keeping the catalogue // advertised. It exists for the prompt-cache interplay: removing // tools from the request to stop the model calling them invalidates // the cached prefix (ADR 0013's 2026-09-22 amendment), while none // keeps the bytes and forbids the calls. ToolChoiceNone ToolChoiceMode = "none" )
type ToolChoiceOption ¶ added in v0.3.0
ToolChoiceOption is accepted by both New and Stream/Generate: which tool calls a step must make is a per-question concern as much as a per-agent one (the ThinkingOption shape).
func ToolChoice ¶ added in v0.3.0
func ToolChoice(cfg ToolChoiceConfig) ToolChoiceOption
ToolChoice forces the agent's model calls to include (or forbear from) tool calls. As an Option it is every run's default; as a RunOption it overrides that default for one run:
// a router: the first step must call classify
agt := weft.New(m, classify, weft.ToolChoice(weft.ToolChoiceConfig{Mode: weft.ToolChoiceNamed, Name: "classify"}))
// a final step that must answer in text, cache prefix intact
agt.Generate(ctx, weft.ToolChoice(weft.ToolChoiceConfig{Mode: weft.ToolChoiceNone}), weft.Prompt(q))
Modes: Any (some tool must be called), Named (Name must be — the router and eval-harness shape), None (no call may be made; the catalogue stays advertised, so a prompt-cache prefix on the tool definitions survives), Auto (the zero value: provider default, nothing sent). A PrepareStep function can rewrite the request's ToolChoice per step — force classify on step 0, then auto — the same way it rewrites tools. A non-auto choice the step's tool snapshot cannot satisfy (no tools, a name not advertised, a name under another mode) fails the run with a descriptive error.
Example ¶
ToolChoice forces a step's tool calls — the router shape: classify must be the first call, the rest of the run is unconstrained. One PrepareStep function rewrites the request's ToolChoice per step; the agent-level option would force every step instead.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
classify := weft.Tool("classify", "Classify the request.",
func(_ context.Context, _ struct{}) (string, error) { return "billing", nil })
agt := weft.New(
wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "classify"}),
wefttest.Say("This is a billing question."),
),
weft.PrepareStep(func(_ context.Context, step int, req weft.ModelRequest) (weft.ModelRequest, error) {
if step == 0 {
req.ToolChoice = weft.ToolChoiceConfig{Mode: weft.ToolChoiceNamed, Name: "classify"}
}
return req, nil
}),
classify,
)
res, err := agt.Generate(context.Background(), weft.Prompt("Why did my invoice double?"))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Text())
}
Output: This is a billing question.
type ToolDef ¶
type ToolDef struct {
Name string `json:"name"`
Description string `json:"description,omitempty"`
InputSchema *Schema `json:"input_schema,omitempty"`
OutputSchema *Schema `json:"output_schema,omitempty"`
// contains filtered or unexported fields
}
ToolDef is a named, schema-described tool the model can call. Build one with the generic Tool constructor; the agent loop (and any manual dispatcher) executes it through Invoke.
A ToolDef is mutable until it is registered with New and frozen thereafter: New keeps a deep copy, so mutating the value that was passed in — or a copy returned by Agent.Tools — never reaches the agent or a running run.
func RawTool ¶
func RawTool(name, description string, schema *Schema, fn func(ctx context.Context, args json.RawMessage) (string, error), opts ...ToolOption) *ToolDef
RawTool defines a tool from an explicit schema instead of reflection — for tools defined outside Go source: plugin manifests, MCP remotes, gateways. fn receives the model's raw JSON arguments verbatim: no unmarshalling, no ErrInvalidToolInput — validation belongs to fn, and StrictInput has no effect. A nil schema becomes the empty object schema, so the tool accepts any object input. Panics match Tool: empty name, nil fn. The manifest records no source line (the definition is not in Go source). Trailing options set per-tool policy as for Tool.
func Subagent ¶ added in v0.2.0
func Subagent(name, description string, child *Agent, opts ...ToolOption) *ToolDef
Subagent defines a tool that delegates to another agent. The model calls it with one argument, prompt; the child runs on a fresh transcript holding only that prompt, on the parent call's context, and the tool's result is the child's final text (or, for a child built with Output, the submitted JSON). The child's events arrive in the parent's stream wrapped in Nested; its usage is added to the parent's RunResult.Usage and recorded on StepRecord.SubagentUsage.
A subagent is an ordinary tool: Timeout bounds the child run, MaxResultBytes caps its answer, RequireApproval gates the delegation, Sequential makes it a barrier, WrapTools wraps the delegation once (the child's own seams govern inside it), and the manifest lists it.
A child run that fails is a tool error the parent model sees (SUBAGENT_FAILED), never a parent run error; a child that ends awaiting approval is SUBAGENT_PENDING; a child that is already running above this call is refused with SUBAGENT_CYCLE. Subagent panics if child is nil.
Example ¶
Delegating to another agent: a subagent is a tool whose handler runs another agent on the prompt alone. The child's events arrive wrapped in Nested; its usage rolls into the parent's total.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
researcher := weft.New(wefttest.Script(
wefttest.Say("order 1234 shipped yesterday"),
))
agt := weft.New(wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "research", Args: `{"prompt":"where is order 1234?"}`}),
wefttest.Say("Researched."),
), weft.Subagent("research", "Research a question in depth.", researcher))
res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 1234?"))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Steps[0].Results[0].Content)
fmt.Println("total tokens:", res.Usage.Total())
}
Output: order 1234 shipped yesterday total tokens: 45
func Tool ¶
func Tool[In, Out any](name, description string, fn func(ctx context.Context, in In) (Out, error), opts ...ToolOption) *ToolDef
Tool defines a tool from a plain function. In and Out are inferred from the handler and the input schema is reflected from In's struct tags, so the compiler checks the handler's shape and nothing is written twice:
type RefundInput struct {
OrderID string `json:"order_id" jsonschema:"the order to refund"`
Reason string `json:"reason,omitempty"`
}
weft.Tool("refund_order", "Refund a customer's order",
func(ctx context.Context, in RefundInput) (Receipt, error) {
return billing.Refund(ctx, in.OrderID, in.Reason)
},
weft.Timeout(10*time.Second))
What the model sees as the result is the text of a string Out, and the JSON encoding of any other Out. Inside the handler, CallFromContext reports which call is running. Trailing options set per-tool policy: Timeout, MaxResultBytes, StrictInput.
The handler shape mirrors the official Go MCP SDK's AddTool[In, Out], so a weft tool can be exposed over MCP without an adapter layer.
Tool panics if name is empty, fn is nil, or In is not a struct (or a pointer to one): providers and MCP require an object at the top level of a tool schema, and a scalar there would fail every real call.
Example ¶
A tool is a plain function; the input schema is derived from the struct.
package main
import (
"context"
"encoding/json"
"fmt"
"log"
"github.com/weftgo/weft"
)
func main() {
type WeatherInput struct {
City string `json:"city" jsonschema:"the city to look up"`
Days *int `json:"days,omitempty"`
}
getWeather := weft.Tool("get_weather", "Get a forecast.",
func(_ context.Context, in WeatherInput) (string, error) {
return "sunny in " + in.City, nil
})
b, err := json.MarshalIndent(getWeather.InputSchema, "", " ")
if err != nil {
log.Fatal(err)
}
fmt.Println(getWeather.Name)
fmt.Println(string(b))
}
Output: get_weather { "type": "object", "properties": { "city": { "type": "string", "description": "the city to look up" }, "days": { "type": "integer" } }, "required": [ "city" ] }
func (*ToolDef) Invoke ¶
Invoke runs the tool with raw JSON arguments: it unmarshals into the handler's input type, calls the handler, and returns the result text the model will see (a string output verbatim, anything else JSON-encoded). A decode failure returns an error wrapping ErrInvalidToolInput. Invoke applies the tool's own StrictInput setting; timeouts and result caps are run policy, applied by the loop.
func (*ToolDef) PromptSnippet ¶ added in v0.2.0
PromptSnippet reports the tool's PromptSnippet text; empty when unset.
func (*ToolDef) RequiresApproval ¶ added in v0.2.0
RequiresApproval reports whether the tool was built with RequireApproval. The loop reads it to park the call; AddTools (weft/mcp) refuses such a tool at registration, because MCP has no approval channel and running it unapproved would bypass the gate.
type ToolError ¶ added in v0.2.0
ToolError is a tool failure with a stable code the model can branch on. Code is SCREAMING_SNAKE by convention ("ORDER_NOT_FOUND"); Message is what the model reads; Err is the internal cause — available to tool middleware and audit logs through errors.As/Unwrap, and never shown to the model. A handler returning *ToolError produces the result "<CODE>: <Message>"; the loop renders its own failures with codes too: INVALID_INPUT (arguments that do not decode) and NO_SUCH_TOOL. Codes are not validated. Plain errors keep rendering as err.Error(); install mw.MapErrors to code them centrally.
Example ¶
A *ToolError carries a code the model can branch on and a cause it never sees.
package main
import (
"context"
"errors"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
lookup := weft.Tool("lookup_order", "", func(_ context.Context, in struct {
ID string `json:"id"`
}) (string, error) {
return "", &weft.ToolError{
Code: "ORDER_NOT_FOUND",
Message: "order " + in.ID + " does not exist",
Err: errors.New("pg: no rows in result set"), // for logs and middleware only
}
})
agt := weft.New(wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":"42"}`}),
wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":42}`}),
wefttest.Say("No such order."),
), lookup)
res, err := agt.Generate(context.Background(), weft.Prompt("order 42?"))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Steps[0].Results[0].Content)
fmt.Println(res.Steps[1].Results[0].Content)
}
Output: ORDER_NOT_FOUND: order 42 does not exist INVALID_INPUT: tool "lookup_order": field "id": expected string, got number
func Errorf ¶ added in v0.2.0
Errorf builds a *ToolError with a formatted Message. A %w verb sets Err as fmt.Errorf would — with several %w verbs every cause stays reachable through errors.Is — so the cause is available to middleware while the model sees only the formatted text.
type ToolFinish ¶
type ToolFinish struct {
RunID string `json:"run_id"`
Seq int64 `json:"seq"`
CallID string `json:"call_id"`
Name string `json:"name"`
Content string `json:"content"`
IsError bool `json:"is_error"`
}
ToolFinish reports that a tool invocation completed, successfully or not. Content is the tool's JSON output, or the failure text when IsError is set — the same value the model sees on the matching ToolResultPart, so a UI can render results as they land.
func (ToolFinish) MarshalJSON ¶
func (e ToolFinish) MarshalJSON() ([]byte, error)
MarshalJSON encodes the event with its "type" discriminator.
type ToolMiddleware ¶ added in v0.2.0
type ToolMiddleware func(next ToolCaller) ToolCaller
ToolMiddleware wraps a ToolCaller, the chi shape: it may act before next (deny, decorate ctx, log), after it (map errors, audit), or instead of it. Panic containment stays outside the chain — a panic in middleware is still a tool error result, never a run error.
type ToolOption ¶ added in v0.2.0
type ToolOption interface {
// contains filtered or unexported methods
}
ToolOption configures one tool at definition time. The options that make sense at both levels — MaxResultBytes, Timeout, StrictInput — are PolicyOptions: on the agent they set the default for every tool, on a tool they override that default for the tool alone.
func PromptSnippet ¶ added in v0.2.0
func PromptSnippet(text string) ToolOption
PromptSnippet attaches system-prompt lines to a tool. The loop appends the snippets of the tools it advertises to the agent's Instructions for every model call — one paragraph per tool, in registration order, separated by blank lines — so usage rules live with the tool instead of in a central prompt that drifts as the set grows. Empty snippets add nothing. The manifest records the snippet.
Example ¶
PromptSnippet keeps a tool's usage rules next to the tool; the loop appends them to the instructions of every model call.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
search := weft.Tool("search", "Search the docs.", func(_ context.Context, _ struct{}) (string, error) { return "", nil },
weft.PromptSnippet("Cite the search result you used."))
model := wefttest.Script(wefttest.Say("ok"))
if _, err := weft.New(model, weft.Instructions("You answer questions."), search).Generate(context.Background(), weft.Prompt("hi")); err != nil {
log.Fatal(err)
}
fmt.Println(model.Requests()[0].System)
}
Output: You answer questions. Cite the search result you used.
func RequireApproval ¶ added in v0.2.0
func RequireApproval() ToolOption
RequireApproval marks a tool whose calls never run without a decision. When the model calls it the loop executes the step's other tools, then ends the run successfully with the call on RunResult.Pending (and RunFinish.Pending) and no result in the transcript. Resume with the transcript plus Approve or Deny:
res, _ := agt.Generate(ctx, weft.Prompt("Refund order 42"))
for _, call := range res.Pending { /* ask someone */ }
res, _ = agt.Generate(ctx, weft.Messages(res.Messages...), weft.Approve(call.ID))
Approved calls run through the ordinary chain with Call.Approved set; denied ones (and pending calls given no decision) become error results the model sees. This is a policy and UX seam, not a security boundary: the boundary is the sandbox a tool runs in.
Example ¶
A RequireApproval tool parks its calls: the run ends successfully with them on Pending, and a later run resumes with a decision.
package main
import (
"context"
"fmt"
"log"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
refund := weft.Tool("refund", "Refund an order.", func(_ context.Context, in struct {
Order string `json:"order"`
}) (string, error) {
return "refunded " + in.Order, nil
}, weft.RequireApproval())
agt := weft.New(wefttest.Script(
wefttest.ToolCalls(wefttest.Call{ID: "c1", Name: "refund", Args: `{"order":"42"}`}),
wefttest.Say("Done."),
), refund)
res, err := agt.Generate(context.Background(), weft.Prompt("refund order 42"))
if err != nil {
log.Fatal(err)
}
for _, call := range res.Pending {
fmt.Printf("awaiting approval: %s %s\n", call.Name, call.Args)
}
// Someone decided. Resume with the transcript and the decision.
res, err = agt.Generate(context.Background(), weft.Messages(res.Messages...), weft.Approve("c1"))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Text())
}
Output: awaiting approval: refund {"order":"42"} Done.
func ToolOptions ¶ added in v0.2.0
func ToolOptions(opts ...ToolOption) ToolOption
ToolOptions composes several tool options into one, applied in order — the Tool-defining counterpart of Options, so a package of policy (Timeout, MaxResultBytes, StrictInput, a WrapTools chain) can be named and reused:
productPolicy := weft.ToolOptions(weft.Timeout(5*time.Second), weft.StrictInput())
weft.Tool("lookup", "…", fn, productPolicy)
Nil entries are ignored.
Example ¶
ToolOptions composes tool options into one named value, so a package of per-tool policy travels under one name.
package main
import (
"context"
"fmt"
"log"
"time"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
productPolicy := weft.ToolOptions(
weft.Timeout(5*time.Second),
weft.MaxResultBytes(1024),
)
agt := weft.New(wefttest.Script(
wefttest.ToolCalls(wefttest.Call{Name: "lookup_order", Args: `{"id":"42"}`}),
wefttest.Say("Done."),
), weft.Tool("lookup_order", "Look up an order by id.",
func(_ context.Context, in struct {
ID string `json:"id" jsonschema:"the order id"`
}) (string, error) {
return "order " + in.ID + " shipped", nil
}, productPolicy))
res, err := agt.Generate(context.Background(), weft.Prompt("Where is order 42?"))
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Steps[0].Results[0].Content)
}
Output: order 42 shipped
type ToolResultPart ¶
type ToolResultPart struct {
CallID string `json:"call_id"`
Name string `json:"name"`
Content string `json:"content"`
IsError bool `json:"is_error"`
}
ToolResultPart is the outcome of one tool call, returned to the model as data. Content is the JSON encoding of the tool's output, or the failure message when IsError is set. Tool failures never abort a run; the model sees them and can recover.
func (ToolResultPart) MarshalJSON ¶
func (p ToolResultPart) MarshalJSON() ([]byte, error)
MarshalJSON encodes the part with its "type" discriminator.
type ToolStart ¶
type ToolStart struct {
RunID string `json:"run_id"`
Seq int64 `json:"seq"`
CallID string `json:"call_id"`
Name string `json:"name"`
Args json.RawMessage `json:"args"`
}
ToolStart reports that a tool invocation began. Events from tools running in parallel interleave: pair them by CallID and order by Seq, a per-run counter assigned at emission that totally orders the stream. A call parked by the approval boundary has a ToolStart and no ToolFinish; it is listed on RunFinish.Pending instead.
func (ToolStart) MarshalJSON ¶
MarshalJSON encodes the event with its "type" discriminator.
type Usage ¶
type Usage struct {
InputTokens int64 `json:"input_tokens"`
OutputTokens int64 `json:"output_tokens"`
CachedInputTokens int64 `json:"cached_input_tokens,omitempty"` // ⊆ InputTokens: read from a prompt cache
CacheWriteTokens int64 `json:"cache_write_tokens,omitempty"` // ⊆ InputTokens: written to a prompt cache (anthropic)
ReasoningTokens int64 `json:"reasoning_tokens,omitempty"` // ⊆ OutputTokens
}
Usage is token accounting for one step or one whole run.
The totals are inclusive: CachedInputTokens and CacheWriteTokens are subsets of InputTokens (tokens billed as input — read from or written to a provider prompt cache), ReasoningTokens is a subset of OutputTokens (provider-side reasoning the model burned). The splits are reporting, not budget bases: UsageLimit and Total keep reading the two totals, so a cached-heavy run budgets identically to an uncached one with the same totals. All three are omitempty on the wire — an event from before they existed round-trips unchanged (SchemaVersion stays 1, ADR 0001/0004).
Source Files
¶
Directories
¶
| Path | Synopsis |
|---|---|
|
anthropic
module
|
|
|
core
module
|
|
|
examples
|
|
|
approval
command
Command approval shows the approval boundary end to end, offline: a tool marked RequireApproval parks its call, the run ends with the call on Pending, a person decides, and a second run resumes with the transcript plus the decision.
|
Command approval shows the approval boundary end to end, offline: a tool marked RequireApproval parks its call, the run ends with the call on Pending, a person decides, and a second run resumes with the transcript plus the decision. |
|
getting-started
command
Command getting-started runs a complete weft agent offline: a scripted model (from wefttest), one tool, and the full event stream.
|
Command getting-started runs a complete weft agent offline: a scripted model (from wefttest), one tool, and the full event stream. |
|
google
module
|
|
|
internal
|
|
|
adapterkit
Package adapterkit holds the helpers every first-party adapter needs but no vendor SDK touches: schema rendering, the terminal-error rule, and the FilePart exactly-one guard.
|
Package adapterkit holds the helpers every first-party adapter needs but no vendor SDK touches: schema rendering, the terminal-error rule, and the FilePart exactly-one guard. |
|
jsonclose
Package jsonclose closes truncated JSON documents: the state machine both the OutputDecoder's prefix salvage (partial_json.go) and mw.RepairJSON need.
|
Package jsonclose closes truncated JSON documents: the state machine both the OutputDecoder's prefix salvage (partial_json.go) and mw.RepairJSON need. |
|
jsonconflict
Package jsonconflict holds a fixture type for schema tests that need an embedded field to collide with a struct's own JSON name.
|
Package jsonconflict holds a fixture type for schema tests that need an embedded field to collide with a struct's own JSON name. |
|
mcp
module
|
|
|
Package mw holds the reference middleware for weft's two seams: model middleware (weft.WrapModel) and tool middleware (weft.WrapTools).
|
Package mw holds the reference middleware for weft's two seams: model middleware (weft.WrapModel) and tool middleware (weft.WrapTools). |
|
obsdb
module
|
|
|
clickhouse
module
|
|
|
openai
module
|
|
|
otel
module
|
|
|
runtime
module
|
|
|
store
module
|
|
|
studio
module
|
|
|
cmd
module
|
|
|
thread
module
|
|
|
sqlite
module
|
|
|
Package wefttest provides a scriptable weft.Model, in the spirit of httptest: write the agent's dialogue as a sequence of scripted model steps and test agents offline, deterministically, with no network.
|
Package wefttest provides a scriptable weft.Model, in the spirit of httptest: write the agent's dialogue as a sequence of scripted model steps and test agents offline, deterministically, with no network. |
|
conformance
Package conformance is the provider adapter contract as an executable table.
|
Package conformance is the provider adapter contract as an executable table. |