wefttest

package
v0.3.2 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 26, 2026 License: MIT Imports: 15 Imported by: 8

Documentation

Overview

Package wefttest provides a scriptable weft.Model, in the spirit of httptest: write the agent's dialogue as a sequence of scripted model steps and test agents offline, deterministically, with no network.

Index

Examples

Constants

This section is empty.

Variables

View Source
var ErrNoFixture = errors.New("wefttest: no replay fixture for this request")

ErrNoFixture is the stream error Replay yields for a request no fixture answers — a new or changed request since the recording, or a fresh checkout without fixtures. Re-record with Record.

View Source
var ErrScriptExhausted = errors.New("wefttest: model script exhausted (agent made more model calls than the script provides)")

ErrScriptExhausted is the stream error a Model yields when the agent makes more model calls than the script provides — almost always a test bug worth failing loudly.

Functions

func Args added in v0.2.0

func Args(v any) string

Args marshals v as the JSON arguments of a scripted Call:

wefttest.Call{Name: "lookup", Args: wefttest.Args(lookupIn{Order: 42})}

A value that cannot be marshalled is a test-construction bug and panics at the call site naming the type.

Example

Args marshals a typed value into a Call's raw-JSON arguments, so the scripted call and the tool's input struct cannot drift apart.

package main

import (
	"context"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

type lookupIn struct {
	OrderID int `json:"order_id"`
}

func main() {
	model := wefttest.Script(
		wefttest.ToolCalls(
			wefttest.Call{Name: "lookup_order", Args: wefttest.Args(lookupIn{OrderID: 42})},
		),
	)
	for ev, err := range model.Stream(context.Background(), weft.ModelRequest{}) {
		if err != nil {
			fmt.Println("stream error:", err)
			return
		}
		if call, ok := ev.(weft.ModelToolCall); ok {
			fmt.Printf("%s %s\n", call.Name, call.Args)
		}
	}
}
Output:
lookup_order {"order_id":42}

func ConformInfo added in v0.2.0

func ConformInfo(m weft.ModelMiddleware) error

ConformInfo reports nil when middleware m forwards the inner model's identity unchanged — the convention ModelMiddleware documents ("implementations should forward Info") — and an error naming the offender otherwise. Use it in a middleware's own test suite to turn the documented convention into a checked fact.

func ConformInfoT added in v0.2.0

func ConformInfoT(t testing.TB, m weft.ModelMiddleware)

ConformInfoT is ConformInfo as a test helper: it fails t instead of returning an error.

func Flatten added in v0.2.0

func Flatten(evs []weft.Event) []weft.Event

Flatten returns evs with every Nested event replaced by its inner event, recursively — the child's events as the child emitted them — for asserting a nested run's behaviour without unwrapping by hand.

evs := collect(parentStream)
want := []weft.Event{weft.RunStart{}, weft.StepStart{}, ...}
slices.Equal(want, Flatten(evs))

func Golden

func Golden(t testing.TB, path string, got []byte)

Golden compares got against the committed golden file at path. With -update it writes got (creating parent directories); otherwise a mismatch fails the test with a line diff, and a missing file fails with a hint to regenerate. Note the flag goes after the package list (`go test ./... -update`): placed before it, the go tool routes it to the wrong package's binary.

func TestManifest(t *testing.T) {
	b, err := weft.Manifest(newAgent())
	if err != nil { t.Fatal(err) }
	wefttest.Golden(t, "weft.json", b)
}

func Record added in v0.2.0

func Record(t testing.TB, dir string, inner weft.Model) weft.Model

Record returns a Model that plays inner and records every request and the stream inner produced for it under dir/<t.Name()>/, one JSON file per request in conversation order, for Replay to answer later. It is the deliberate, local, key-holding half of record/replay: call it from a test only when re-recording, never in CI. The test's directory is replaced, not merged.

The recording switch belongs to the suite that holds the provider key, not to wefttest — Golden's -update rewrites a comparison, while Record spends money. The recommended shape (WEFT_RECORD is the suggested name, so suites converge on one spelling):

func model(t *testing.T) weft.Model {
    if os.Getenv("WEFT_RECORD") != "" { // the suite's own switch; wefttest never reads it
        return wefttest.Record(t, "testdata/replay", openai.New(os.Getenv("OPENAI_API_KEY"), "gpt-5"))
    }
    return wefttest.Replay(t, "testdata/replay")
}

Record never checks WEFT_MODEL_REQUESTS: inner does, so a recording made under deny stores the denial's error text and the diff shows it.

func Replay added in v0.2.0

func Replay(t testing.TB, dir string) weft.Model

Replay returns a Model that answers each request from the fixture Record wrote for it under dir/<t.Name()>/, matched by the request's key (messages, tool names, thinking level and the sequential flag — not the system prompt). It makes no request, holds no key, and ignores WEFT_MODEL_REQUESTS like every wefttest model. A request with no fixture fails the stream with ErrNoFixture naming what was wanted; repeated identical requests replay in recorded order.

The fixtures are pretty-printed JSON a reviewer reads in a diff — the request's messages (a changed prompt shows as a changed fixture) and the model's events — so re-recording is the review, exactly as ADR 0013 says for wire fixtures.

Example

Replay answers an agent from what a real model actually said, offline and deterministically. The suite holds the recording switch — Record spends provider money, so the decision belongs to the suite that holds the key, never to a flag wefttest reads. This example shows the shape only: Replay and Record take the *testing.T (for the fixture directory's name), which an example has no equivalent of.

package main

import ()

func main() {
	// The suite's model(t), one line at the call site:
	//
	//	func model(t *testing.T) weft.Model {
	//	    if os.Getenv("WEFT_RECORD") != "" { // the suite's own switch
	//	        return wefttest.Record(t, "testdata/replay", myAdapter)
	//	    }
	//	    return wefttest.Replay(t, "testdata/replay")
	//	}
	//
	//	agt := weft.New(model(t), tools...)
	//	res, err := agt.Generate(context.Background(), weft.Prompt("Refund order 1234."))
	//	fmt.Println(res.Text(), err)
	//
	// A request the recording does not answer fails loudly with
	// wefttest.ErrNoFixture naming the directory and the first user
	// text — re-record with WEFT_RECORD once, commit the fixtures, and
	// every later run replays without a key or a network.
}

Types

type Call

type Call struct {
	Name string
	Args string // raw JSON; "{}" when empty
	ID   string // optional; "call_N" within the turn when empty
}

Call describes one scripted tool invocation.

type Model

type Model struct {
	// contains filtered or unexported fields
}

Model is a scripted weft.Model. It plays one Turn per Stream call, in order, and records every ModelRequest it received for assertions.

func Script

func Script(turns ...Turn) *Model

Script returns a Model that plays the turns in order.

func (*Model) Info

func (*Model) Info() weft.ModelInfo

Info identifies the scripted model, so RunStart.Model and manifests carry a non-zero value in tests.

func (*Model) LastRequest added in v0.2.0

func (m *Model) LastRequest() Request

LastRequest returns the most recent request the agent sent, or the zero Request when none was sent yet — whose matchers report false and empty, so a test that needs a request asserts len(Requests()) > 0 first.

Example

LastRequest is the one-request shorthand for asserting what the agent asked: the advertised tools and the opening prompt, without keeping a handle on every request. A model with no requests yet returns the zero Request, whose matchers report false and empty.

package main

import (
	"context"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

type lookupIn struct {
	OrderID int `json:"order_id"`
}

var exampleLookup = weft.Tool("lookup_order", "Look up an order by ID.",
	func(_ context.Context, in lookupIn) (int, error) {
		return in.OrderID, nil
	})

func main() {
	model := wefttest.Script(wefttest.Say("Order 42 shipped on Tuesday."))
	res, err := weft.New(model, exampleLookup).Generate(context.Background(), weft.Prompt("Where is order 42?"))
	if err != nil {
		fmt.Println(err)
		return
	}
	fmt.Println(res.Text())
	req := model.LastRequest()
	fmt.Println(req.HasTool("lookup_order"), req.ToolNames(), req.LastText())
}
Output:
Order 42 shipped on Tuesday.
true [lookup_order] Where is order 42?

func (*Model) Requests

func (m *Model) Requests() []weft.ModelRequest

Requests returns every ModelRequest the agent sent, in order — including a call that exhausted the script (identify it by the ErrScriptExhausted run error) — so assert on system prompts, transcript shape, the tool catalog the model saw, and attempt counts under Retry/Fallback alike.

func (*Model) Stream

Stream implements weft.Model.

type Request added in v0.2.0

type Request struct{ weft.ModelRequest }

Request is one recorded ModelRequest with matchers for the assertions agent tests make most; every field of the request stays reachable through the embedded value.

func (Request) HasTool added in v0.2.0

func (r Request) HasTool(name string) bool

HasTool reports whether the request advertised a tool named name.

func (Request) LastText added in v0.2.0

func (r Request) LastText() string

LastText returns the text of the last text part of the request's last message — the user's prompt on the first call, a user follow-up later — and "" when the last message has no text part (a tool message after a tool step).

func (Request) ToolNames added in v0.2.0

func (r Request) ToolNames() []string

ToolNames returns the advertised tool names, sorted.

type Turn

type Turn struct {
	// contains filtered or unexported fields
}

Turn is the script for one model call, built by Say, ToolCalls, or Fail.

func Fail

func Fail(err error) Turn

Fail scripts a model failure: the stream errors immediately with err.

func MaxTokens

func MaxTokens(text string) Turn

MaxTokens scripts a step that hit the output-token limit: the text so far, then a max_tokens finish. The run records it on RunResult.StopReason (and the StepRecord) rather than failing; truncated tool-call arguments decode as error results the model recovers from.

func Raw added in v0.2.0

func Raw(events ...weft.ModelEvent) Turn

Raw scripts a step that plays exactly the events given, appending nothing — the way to script a signed reasoning block, a ModelToolCallDelta progress sequence, or a stream that violates the Model contract (no ModelFinish, two of them, an event after the finish) so a test can pin how the loop reports ErrModelContract. Raw still honours ctx like every scripted turn: it is malformed by content, never by ignoring cancellation.

Example

Raw scripts the model's events verbatim — here a signed reasoning block, the shape providers like Anthropic attach a signature to and Think cannot express (it scripts unsigned reasoning only).

package main

import (
	"context"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	model := wefttest.Script(
		wefttest.Raw(
			weft.ModelReasoningDelta{Text: "checking the order", Signature: "sig1"},
			weft.ModelTextDelta{Text: "Order 42 arrived broken."},
			weft.ModelFinish{Reason: weft.StopEndTurn, Usage: weft.Usage{InputTokens: 10, OutputTokens: 5}},
		),
	)
	for ev, err := range model.Stream(context.Background(), weft.ModelRequest{}) {
		if err != nil {
			fmt.Println("stream error:", err)
			return
		}
		switch e := ev.(type) {
		case weft.ModelReasoningDelta:
			fmt.Printf("reasoning %q (signed: %v)\n", e.Text, e.Signature != "")
		case weft.ModelTextDelta:
			fmt.Printf("text %q\n", e.Text)
		case weft.ModelFinish:
			fmt.Printf("finish %v\n", e.Reason)
		}
	}
}
Output:
reasoning "checking the order" (signed: true)
text "Order 42 arrived broken."
finish stop

func Say

func Say(text string) Turn

Say scripts a completed assistant reply: one text delta, then a normal stop. Usage is fixed at 10 input / 5 output tokens per scripted step, so usage accounting is assertable.

func SayThenFail added in v0.2.0

func SayThenFail(text string, err error) Turn

SayThenFail scripts a mid-stream failure: one text delta, then the stream errors with err. The loop discards the partial turn and reports err; mw.Retry sees a failure after output.

Example

SayThenFail scripts a mid-stream failure: the caller sees the partial text, then the stream errors — the shape the loop's partial-turn handling and mw.Retry are tested with.

package main

import (
	"context"
	"errors"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	model := wefttest.Script(
		wefttest.SayThenFail("half an answ", errors.New("provider disconnected")),
	)
	for ev, err := range model.Stream(context.Background(), weft.ModelRequest{}) {
		if err != nil {
			fmt.Println("stream error:", err)
			return
		}
		if d, ok := ev.(weft.ModelTextDelta); ok {
			fmt.Printf("saw %q\n", d.Text)
		}
	}
}
Output:
saw "half an answ"
stream error: provider disconnected

func Think

func Think(text string, then Turn) Turn

Think prefixes then with a scripted reasoning delta: Think("plan", ToolCalls(...)) is a step that reasons and then requests tools. The reasoning carries no signature; signed blocks are scripted with raw model events.

func ToolCalls

func ToolCalls(calls ...Call) Turn

ToolCalls scripts a step that requests tool calls. Call IDs are assigned "call_1", "call_2", ... within the turn unless set explicitly.

func (Turn) WithUsage added in v0.2.0

func (t Turn) WithUsage(u weft.Usage) Turn

WithUsage returns the turn with its ModelFinish.Usage replaced by u. A turn without a finish (Fail, a Raw stream without one) is returned unchanged.

Example

WithUsage rewrites a scripted turn's finish usage, so a UsageLimit overshoot is scripted without arithmetic on the fixed defaults.

package main

import (
	"context"
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	model := wefttest.Script(
		wefttest.Say("done").WithUsage(weft.Usage{InputTokens: 900, OutputTokens: 40}),
	)
	for ev, err := range model.Stream(context.Background(), weft.ModelRequest{}) {
		if err != nil {
			fmt.Println("stream error:", err)
			return
		}
		if f, ok := ev.(weft.ModelFinish); ok {
			fmt.Printf("usage in=%d out=%d\n", f.Usage.InputTokens, f.Usage.OutputTokens)
		}
	}
}
Output:
usage in=900 out=40

Directories

Path Synopsis
Package conformance is the provider adapter contract as an executable table.
Package conformance is the provider adapter contract as an executable table.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL