Documentation
¶
Overview ¶
Package wefttest provides a scriptable weft.Model, in the spirit of httptest: write the agent's dialogue as a sequence of scripted model steps and test agents offline, deterministically, with no network.
Index ¶
- Variables
- func Args(v any) string
- func ConformInfo(m weft.ModelMiddleware) error
- func ConformInfoT(t testing.TB, m weft.ModelMiddleware)
- func Flatten(evs []weft.Event) []weft.Event
- func Golden(t testing.TB, path string, got []byte)
- func Record(t testing.TB, dir string, inner weft.Model) weft.Model
- func Replay(t testing.TB, dir string) weft.Model
- type Call
- type Model
- type Request
- type Turn
Examples ¶
Constants ¶
This section is empty.
Variables ¶
var ErrNoFixture = errors.New("wefttest: no replay fixture for this request")
ErrNoFixture is the stream error Replay yields for a request no fixture answers — a new or changed request since the recording, or a fresh checkout without fixtures. Re-record with Record.
var ErrScriptExhausted = errors.New("wefttest: model script exhausted (agent made more model calls than the script provides)")
ErrScriptExhausted is the stream error a Model yields when the agent makes more model calls than the script provides — almost always a test bug worth failing loudly.
Functions ¶
func Args ¶ added in v0.2.0
Args marshals v as the JSON arguments of a scripted Call:
wefttest.Call{Name: "lookup", Args: wefttest.Args(lookupIn{Order: 42})}
A value that cannot be marshalled is a test-construction bug and panics at the call site naming the type.
Example ¶
Args marshals a typed value into a Call's raw-JSON arguments, so the scripted call and the tool's input struct cannot drift apart.
package main
import (
"context"
"fmt"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
type lookupIn struct {
OrderID int `json:"order_id"`
}
func main() {
model := wefttest.Script(
wefttest.ToolCalls(
wefttest.Call{Name: "lookup_order", Args: wefttest.Args(lookupIn{OrderID: 42})},
),
)
for ev, err := range model.Stream(context.Background(), weft.ModelRequest{}) {
if err != nil {
fmt.Println("stream error:", err)
return
}
if call, ok := ev.(weft.ModelToolCall); ok {
fmt.Printf("%s %s\n", call.Name, call.Args)
}
}
}
Output: lookup_order {"order_id":42}
func ConformInfo ¶ added in v0.2.0
func ConformInfo(m weft.ModelMiddleware) error
ConformInfo reports nil when middleware m forwards the inner model's identity unchanged — the convention ModelMiddleware documents ("implementations should forward Info") — and an error naming the offender otherwise. Use it in a middleware's own test suite to turn the documented convention into a checked fact.
func ConformInfoT ¶ added in v0.2.0
func ConformInfoT(t testing.TB, m weft.ModelMiddleware)
ConformInfoT is ConformInfo as a test helper: it fails t instead of returning an error.
func Flatten ¶ added in v0.2.0
Flatten returns evs with every Nested event replaced by its inner event, recursively — the child's events as the child emitted them — for asserting a nested run's behaviour without unwrapping by hand.
evs := collect(parentStream)
want := []weft.Event{weft.RunStart{}, weft.StepStart{}, ...}
slices.Equal(want, Flatten(evs))
func Golden ¶
Golden compares got against the committed golden file at path. With -update it writes got (creating parent directories); otherwise a mismatch fails the test with a line diff, and a missing file fails with a hint to regenerate. Note the flag goes after the package list (`go test ./... -update`): placed before it, the go tool routes it to the wrong package's binary.
func TestManifest(t *testing.T) {
b, err := weft.Manifest(newAgent())
if err != nil { t.Fatal(err) }
wefttest.Golden(t, "weft.json", b)
}
func Record ¶ added in v0.2.0
Record returns a Model that plays inner and records every request and the stream inner produced for it under dir/<t.Name()>/, one JSON file per request in conversation order, for Replay to answer later. It is the deliberate, local, key-holding half of record/replay: call it from a test only when re-recording, never in CI. The test's directory is replaced, not merged.
The recording switch belongs to the suite that holds the provider key, not to wefttest — Golden's -update rewrites a comparison, while Record spends money. The recommended shape (WEFT_RECORD is the suggested name, so suites converge on one spelling):
func model(t *testing.T) weft.Model {
if os.Getenv("WEFT_RECORD") != "" { // the suite's own switch; wefttest never reads it
return wefttest.Record(t, "testdata/replay", openai.New(os.Getenv("OPENAI_API_KEY"), "gpt-5"))
}
return wefttest.Replay(t, "testdata/replay")
}
Record never checks WEFT_MODEL_REQUESTS: inner does, so a recording made under deny stores the denial's error text and the diff shows it.
func Replay ¶ added in v0.2.0
Replay returns a Model that answers each request from the fixture Record wrote for it under dir/<t.Name()>/, matched by the request's key (messages, tool names, thinking level and the sequential flag — not the system prompt). It makes no request, holds no key, and ignores WEFT_MODEL_REQUESTS like every wefttest model. A request with no fixture fails the stream with ErrNoFixture naming what was wanted; repeated identical requests replay in recorded order.
The fixtures are pretty-printed JSON a reviewer reads in a diff — the request's messages (a changed prompt shows as a changed fixture) and the model's events — so re-recording is the review, exactly as ADR 0013 says for wire fixtures.
Example ¶
Replay answers an agent from what a real model actually said, offline and deterministically. The suite holds the recording switch — Record spends provider money, so the decision belongs to the suite that holds the key, never to a flag wefttest reads. This example shows the shape only: Replay and Record take the *testing.T (for the fixture directory's name), which an example has no equivalent of.
package main
import ()
func main() {
// The suite's model(t), one line at the call site:
//
// func model(t *testing.T) weft.Model {
// if os.Getenv("WEFT_RECORD") != "" { // the suite's own switch
// return wefttest.Record(t, "testdata/replay", myAdapter)
// }
// return wefttest.Replay(t, "testdata/replay")
// }
//
// agt := weft.New(model(t), tools...)
// res, err := agt.Generate(context.Background(), weft.Prompt("Refund order 1234."))
// fmt.Println(res.Text(), err)
//
// A request the recording does not answer fails loudly with
// wefttest.ErrNoFixture naming the directory and the first user
// text — re-record with WEFT_RECORD once, commit the fixtures, and
// every later run replays without a key or a network.
}
Output:
Types ¶
type Call ¶
type Call struct {
Name string
Args string // raw JSON; "{}" when empty
ID string // optional; "call_N" within the turn when empty
}
Call describes one scripted tool invocation.
type Model ¶
type Model struct {
// contains filtered or unexported fields
}
Model is a scripted weft.Model. It plays one Turn per Stream call, in order, and records every ModelRequest it received for assertions.
func (*Model) Info ¶
Info identifies the scripted model, so RunStart.Model and manifests carry a non-zero value in tests.
func (*Model) LastRequest ¶ added in v0.2.0
LastRequest returns the most recent request the agent sent, or the zero Request when none was sent yet — whose matchers report false and empty, so a test that needs a request asserts len(Requests()) > 0 first.
Example ¶
LastRequest is the one-request shorthand for asserting what the agent asked: the advertised tools and the opening prompt, without keeping a handle on every request. A model with no requests yet returns the zero Request, whose matchers report false and empty.
package main
import (
"context"
"fmt"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
type lookupIn struct {
OrderID int `json:"order_id"`
}
var exampleLookup = weft.Tool("lookup_order", "Look up an order by ID.",
func(_ context.Context, in lookupIn) (int, error) {
return in.OrderID, nil
})
func main() {
model := wefttest.Script(wefttest.Say("Order 42 shipped on Tuesday."))
res, err := weft.New(model, exampleLookup).Generate(context.Background(), weft.Prompt("Where is order 42?"))
if err != nil {
fmt.Println(err)
return
}
fmt.Println(res.Text())
req := model.LastRequest()
fmt.Println(req.HasTool("lookup_order"), req.ToolNames(), req.LastText())
}
Output: Order 42 shipped on Tuesday. true [lookup_order] Where is order 42?
func (*Model) Requests ¶
func (m *Model) Requests() []weft.ModelRequest
Requests returns every ModelRequest the agent sent, in order — including a call that exhausted the script (identify it by the ErrScriptExhausted run error) — so assert on system prompts, transcript shape, the tool catalog the model saw, and attempt counts under Retry/Fallback alike.
type Request ¶ added in v0.2.0
type Request struct{ weft.ModelRequest }
Request is one recorded ModelRequest with matchers for the assertions agent tests make most; every field of the request stays reachable through the embedded value.
func (Request) HasTool ¶ added in v0.2.0
HasTool reports whether the request advertised a tool named name.
type Turn ¶
type Turn struct {
// contains filtered or unexported fields
}
Turn is the script for one model call, built by Say, ToolCalls, or Fail.
func MaxTokens ¶
MaxTokens scripts a step that hit the output-token limit: the text so far, then a max_tokens finish. The run records it on RunResult.StopReason (and the StepRecord) rather than failing; truncated tool-call arguments decode as error results the model recovers from.
func Raw ¶ added in v0.2.0
func Raw(events ...weft.ModelEvent) Turn
Raw scripts a step that plays exactly the events given, appending nothing — the way to script a signed reasoning block, a ModelToolCallDelta progress sequence, or a stream that violates the Model contract (no ModelFinish, two of them, an event after the finish) so a test can pin how the loop reports ErrModelContract. Raw still honours ctx like every scripted turn: it is malformed by content, never by ignoring cancellation.
Example ¶
Raw scripts the model's events verbatim — here a signed reasoning block, the shape providers like Anthropic attach a signature to and Think cannot express (it scripts unsigned reasoning only).
package main
import (
"context"
"fmt"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
model := wefttest.Script(
wefttest.Raw(
weft.ModelReasoningDelta{Text: "checking the order", Signature: "sig1"},
weft.ModelTextDelta{Text: "Order 42 arrived broken."},
weft.ModelFinish{Reason: weft.StopEndTurn, Usage: weft.Usage{InputTokens: 10, OutputTokens: 5}},
),
)
for ev, err := range model.Stream(context.Background(), weft.ModelRequest{}) {
if err != nil {
fmt.Println("stream error:", err)
return
}
switch e := ev.(type) {
case weft.ModelReasoningDelta:
fmt.Printf("reasoning %q (signed: %v)\n", e.Text, e.Signature != "")
case weft.ModelTextDelta:
fmt.Printf("text %q\n", e.Text)
case weft.ModelFinish:
fmt.Printf("finish %v\n", e.Reason)
}
}
}
Output: reasoning "checking the order" (signed: true) text "Order 42 arrived broken." finish stop
func Say ¶
Say scripts a completed assistant reply: one text delta, then a normal stop. Usage is fixed at 10 input / 5 output tokens per scripted step, so usage accounting is assertable.
func SayThenFail ¶ added in v0.2.0
SayThenFail scripts a mid-stream failure: one text delta, then the stream errors with err. The loop discards the partial turn and reports err; mw.Retry sees a failure after output.
Example ¶
SayThenFail scripts a mid-stream failure: the caller sees the partial text, then the stream errors — the shape the loop's partial-turn handling and mw.Retry are tested with.
package main
import (
"context"
"errors"
"fmt"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
model := wefttest.Script(
wefttest.SayThenFail("half an answ", errors.New("provider disconnected")),
)
for ev, err := range model.Stream(context.Background(), weft.ModelRequest{}) {
if err != nil {
fmt.Println("stream error:", err)
return
}
if d, ok := ev.(weft.ModelTextDelta); ok {
fmt.Printf("saw %q\n", d.Text)
}
}
}
Output: saw "half an answ" stream error: provider disconnected
func Think ¶
Think prefixes then with a scripted reasoning delta: Think("plan", ToolCalls(...)) is a step that reasons and then requests tools. The reasoning carries no signature; signed blocks are scripted with raw model events.
func ToolCalls ¶
ToolCalls scripts a step that requests tool calls. Call IDs are assigned "call_1", "call_2", ... within the turn unless set explicitly.
func (Turn) WithUsage ¶ added in v0.2.0
WithUsage returns the turn with its ModelFinish.Usage replaced by u. A turn without a finish (Fail, a Raw stream without one) is returned unchanged.
Example ¶
WithUsage rewrites a scripted turn's finish usage, so a UsageLimit overshoot is scripted without arithmetic on the fixed defaults.
package main
import (
"context"
"fmt"
"github.com/weftgo/weft"
"github.com/weftgo/weft/wefttest"
)
func main() {
model := wefttest.Script(
wefttest.Say("done").WithUsage(weft.Usage{InputTokens: 900, OutputTokens: 40}),
)
for ev, err := range model.Stream(context.Background(), weft.ModelRequest{}) {
if err != nil {
fmt.Println("stream error:", err)
return
}
if f, ok := ev.(weft.ModelFinish); ok {
fmt.Printf("usage in=%d out=%d\n", f.Usage.InputTokens, f.Usage.OutputTokens)
}
}
}
Output: usage in=900 out=40
Directories
¶
| Path | Synopsis |
|---|---|
|
Package conformance is the provider adapter contract as an executable table.
|
Package conformance is the provider adapter contract as an executable table. |