agenteval

package module
v0.0.5 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 29, 2026 License: MIT Imports: 17 Imported by: 0

README

agenteval

Evaluation for Go agents over Open Responses: a replay model and recorded tools served from a session, tasks and suites loaded from an fs.FS, a runner that records every run as a session, judges, and scores written back as outcome entries.

go get github.com/ChristopherDavenport/agenteval

What it does

A run is an ordinary agentsession: the runner sends a task through an agentturn configuration with the session recorder attached, so every request, response and tool output is on disk and the exporter renders it as an ATIF trajectory. Judges read that trajectory. Each score is appended to the session as an outcome entry whose target is the run's last entry, whose kind is eval, and whose pass and score are where the format puts them, so a pass rate is computable from the session files alone. The report is a value over those sessions and holds no fact they lack.

suite, _ := agenteval.LoadSuite(os.DirFS("evals"), "smoke")
r := &agenteval.Runner{
	Store:  store,
	Config: func(agenteval.Task) agentturn.Config { return cfg },
	Judges: []agenteval.Judge{judge.Contains("text"), judge.ToolCalled("bash")},
	Cost:   price.Hook(prices),
}
report, _ := r.Run(ctx, suite)
report.WriteJSON(os.Stdout)

A configuration that needs the run's recorder takes ConfigWith instead: a compacting configuration binds compact.WithOnFold there, without which its folds reach no session and the run cannot be replayed, and a layer that annotates the run it is in keeps the same recorder. Answer answers the calls a run left pending when it ended input_required and resumes it, so a product whose policy asks is measured over the whole run rather than the part before its first ask.

r := &agenteval.Runner{
	Store: store,
	ConfigWith: func(t agenteval.Task, rec *session.Recorder) agentturn.Config {
		cfg := product.Config(t)
		cfg.Transform = compact.NewLocal(cfg.Model, compact.WithOnFold(rec.Fold)).Transform
		return cfg
	},
	Answer: engine.Answers,
}

A task file is JSON: an instruction, or prompts for several turns, with expect for the judges and setup and meta for the product. The suite's manifest, the hash of every file loaded, is recorded in each run's env entry, so a result names the suite version it came from.

Replay

replay.NewModel serves a recorded session's model calls in path order, folds included, and in strict mode refuses a request whose hash differs from the recorded one: a hook, a transform or a front that changes what the model would have been sent fails loudly against real traffic. A record that carries no hash for a call is refused rather than served unchecked — at NewModel when a response on the path is unhashed, so a replay that could not have checked what it served says so before it starts instead of passing quietly, and at the call for a fold — unless replay.AllowUnhashed() says to serve it anyway. The one call no mode checks is a fold through the compaction endpoint, which sends no request the format hashes. replay.Tools serves recorded tool outputs by call ID or by name and canonical arguments, so a whole run replays without touching the world.

s, _ := store.Open(ctx, id)
model, _ := replay.NewModel(s, replay.Strict())
cfg.Model = model
cfg.BeforeModelCall = model.BeforeModelCall
cfg.Tools = replay.Tools(s, cfg.Tools, replay.Strict())

Model.BeforeModelCall serves the instructions and the tool list recorded for each call. A product whose layers rebuild those every turn, from a memory store, a skill set or an AGENTS.md, otherwise replays against what the layers say today and diverges at the first call with an error naming two hashes and no layer. Model.Settings is the rest of what each recorded call was made under, for a judge or a check that wants to read it.

Judges

judge.Exact, judge.Contains and judge.ToolCalled are deterministic. judge.Rubric is an agent: the trajectory is its input, the rubric its instructions, and its answer is constrained to a score schema; with a store it records its own session, parented to the one it judged, so a judgement is as replayable as the run.

Comparison and cost

Compare runs a suite under two configurations, pairs the results by task, and names how the two configurations differed, read back from the sessions. The difference it names is the settings the record describes; when those are the same and the two first requests are not, it says so with beyond_settings, because a transform, a hook and an injected item are not settings. A context strategy that folds only later in the run shows as folded, whether each side's path held a compaction entry, when one did and the other did not.

price loads a table of rates and prices a run with ATIF's formula, and the runner and the exporter take the same hook. A result's usage and cost_usd sum every model call on the run's path, the folds a compacting configuration made included, which is what the exported document's final metrics sum, so the two agree on the total. A run with a call the hook cannot price has no cost_usd, where the document reports the sum of the calls it could.

Harbor

The harbor nested module reads a Harbor task directory as a Task (the prompt only; the environment and the verifier stay with Harbor) and a finished trial's reward.json or reward.txt as scores. Task.Meta carries [task] and [metadata] under task. and metadata., because the second table is free-form and invites the first table's words.

Design

The plan is in docs/plans/eval-layer.md and the findings of the design studies that shaped it in docs/feedback.md.

License

MIT.

Documentation

Overview

Package agenteval evaluates Go agents over Open Responses: tasks and suites loaded from an fs.FS, a runner that sends each task through an agentturn configuration and records the run as an agentsession, a judge contract, and scores written back into the session as outcome entries. The replay package serves a recorded session as a model and as tools, so hooks, transforms and fronts are tested against real traffic offline; the judge package holds the deterministic judges and the rubric judge; price costs a run; the harbor nested module reads a Harbor task and a verifier reward.

A run is an ordinary session: the recorder writes it, the exporter renders it, and every score is an outcome entry whose target is the last entry of the run it judges. The report is a value over those sessions and holds no fact they lack.

suite, _ := agenteval.LoadSuite(os.DirFS("evals"), "smoke")
r := &agenteval.Runner{
	Store:  store,
	Config: func(agenteval.Task) agentturn.Config { return cfg },
	Judges: []agenteval.Judge{judge.Contains("text"), judge.ToolCalled("bash")},
}
report, _ := r.Run(ctx, suite)
report.WriteJSON(os.Stdout)

Index

Constants

View Source
const DefaultMaxResumes = 8

DefaultMaxResumes is how many times one prompt is resumed through Runner.Answer when Runner.MaxResumes is zero.

View Source
const SuiteFile = "suite.json"

SuiteFile is the name of the optional file in a suite directory that names the suite. Its content is {"name": "..."}.

View Source
const TaskNS = "agenteval:task"

TaskNS is the namespace of the custom entry the runner writes before each run, whose data is a TaskRecord.

Variables

View Source
var ErrNoPrompts = errors.New("agenteval: session has no user prompts on its path")

ErrNoPrompts is returned by FromSession when the path holds no user message.

View Source
var ErrUnreplayable = errors.New("agenteval: the run cannot be replayed strictly")

ErrUnreplayable is joined onto the result of a run whose session cannot be replayed strictly, because the record does not rebuild every request the run sent.

Functions

func NewOutcome

func NewOutcome(score Score, target, task string) (*agentsession.OutcomeEntry, error)

NewOutcome builds the outcome entry for a score: kind agentsession.OutcomeEval, target the entry named, score, pass and label from the score, and OutcomeDetails carrying the task. The entry is not appended.

func ReadOutcome

ReadOutcome reads a score back out of an outcome entry the runner wrote. It reports false for an entry of another kind or one without OutcomeDetails.

func Summarize

func Summarize(results []Result) map[string]Summary

Summarize computes the per-judge summaries over results.

Types

type Change

type Change struct {
	A any `json:"a"`
	B any `json:"b"`
}

Change is one setting under A and under B.

type Comparison

type Comparison struct {
	Suite    string   `json:"suite"`
	Manifest Manifest `json:"manifest"`
	// A and B are the two reports. They are not written by WriteJSON;
	// the pairs carry both results of every task.
	A, B *Report `json:"-"`
	// Config is the difference between the settings B's runs started
	// under and A's, when every pair differs the same way; Uniform
	// says whether they did. Each pair carries its own. Whether a run
	// folded depends on the task as much as the configuration, so
	// pairs are uniform whatever their Folded says, and Config.Folded
	// says whether any of A's runs folded and any of B's, when those
	// differ.
	Config  ConfigDiff `json:"config"`
	Uniform bool       `json:"uniform"`
	Pairs   []Pair     `json:"pairs"`
	// ByJudge is B's mean minus A's, per judge scored on both sides.
	ByJudge map[string]float64 `json:"by_judge"`
}

Comparison is the same suite run under two configurations and judged the same way, paired by task.

func Compare

func Compare(ctx context.Context, suite *Suite, a, b *Runner) (*Comparison, error)

Compare runs the suite under a and b and pairs the results by task. The runners share nothing but the suite; each records its own sessions in its own store, which may be the same store. The comparison's Config is the difference between the settings the two runs of each task started under, read back from their sessions, so a comparison says how the configurations differed and not only which scored higher.

func (*Comparison) WriteJSON

func (c *Comparison) WriteJSON(w io.Writer) error

WriteJSON writes the comparison as indented JSON with a trailing newline. The two reports are not included; write them separately.

type ConfigDiff

type ConfigDiff struct {
	Model        *Change `json:"model,omitempty"`
	Instructions *Change `json:"instructions,omitempty"`
	Reasoning    *Change `json:"reasoning,omitempty"`
	Text         *Change `json:"text,omitempty"`
	// ToolsAdded names tools B has and A lacks, ToolsRemoved the
	// reverse, and ToolsChanged tools both have under a different
	// definition; each sorted.
	ToolsAdded   []string `json:"tools_added,omitempty"`
	ToolsRemoved []string `json:"tools_removed,omitempty"`
	ToolsChanged []string `json:"tools_changed,omitempty"`
	// Extra holds passthrough members that differ, by wire name; a
	// side that lacks the member has nil.
	Extra map[string]Change `json:"extra,omitempty"`
	// BeyondSettings says the two runs' first calls sent different
	// requests although their settings were the same: a transform, a
	// BeforeTurn or BeforeModelCall hook, or items one side injected.
	// None of those is a setting, so no field above can name it, and a
	// comparison of two configurations that differ only there would
	// otherwise report that nothing differed, which is worse than
	// reporting nothing. It is read from the request hashes the two
	// first responses recorded: two that differ say so, and one hash
	// against none says the other run sent something the record does
	// not rebuild, which is a transform or a hook by definition. It
	// compares the first call of each run, so a transform that
	// changes nothing until later, as a compaction transform does not
	// fold before it has anything to fold, is not visible here;
	// Folded is.
	BeyondSettings bool `json:"beyond_settings,omitempty"`
	// Folded says one run's path held a compaction entry and the
	// other's none, with whether each side folded: a context strategy
	// is not a setting either, and the fold is what it leaves on the
	// path. Two runs that both folded, however often, do not differ
	// here. On a [Comparison] it is over all the runs of each side.
	Folded *Change `json:"folded,omitempty"`
}

ConfigDiff is what differed between two runs' settings: the settings each run's first model call was made under, replayed from its session's config entries.

func DiffSettings

func DiffSettings(a, b agentsession.Settings) ConfigDiff

DiffSettings returns what differs between a and b. It covers the settings the record describes, which is the model, the instructions, the reasoning and text configuration, the tools and the passthrough members, and nothing else: a transform, a hook and an injected item are not settings. Compare reports one of those through ConfigDiff.BeyondSettings.

func (ConfigDiff) Empty

func (d ConfigDiff) Empty() bool

Empty reports whether nothing differed.

type Judge

type Judge interface {
	// Name identifies the judge in scores and outcome entries.
	Name() string
	// Judge scores the trajectory. An error means the judge could not
	// reach a verdict, not that the run failed; the runner records the
	// error on the result and writes no outcome for it.
	Judge(ctx context.Context, t export.Trajectory, task Task) (Score, error)
}

Judge scores a run from its trajectory and the task it ran.

type JudgeFunc

type JudgeFunc struct {
	JudgeName string
	Fn        func(ctx context.Context, t export.Trajectory, task Task) (Score, error)
}

JudgeFunc adapts a function to Judge.

func (JudgeFunc) Judge

func (j JudgeFunc) Judge(ctx context.Context, t export.Trajectory, task Task) (Score, error)

Judge calls Fn and stamps the score with the judge's name.

func (JudgeFunc) Name

func (j JudgeFunc) Name() string

Name returns JudgeName.

type Manifest

type Manifest struct {
	Location string            `json:"location"`
	Files    map[string]string `json:"files"`
}

Manifest says where a suite was loaded from and what it held: the location within the fs.FS and the content hash of every file read, keyed by path, so a result names the suite version it came from.

type OutcomeDetails

type OutcomeDetails struct {
	Task    string          `json:"task"`
	Reason  string          `json:"reason,omitempty"`
	Session string          `json:"session,omitempty"`
	Judge   json.RawMessage `json:"judge,omitempty"`
}

OutcomeDetails is the details member of an outcome entry the runner writes, so a reader gets the task, the reason and the judge's own session at known keys without knowing any judge's private JSON.

type Pair

type Pair struct {
	Task string  `json:"task"`
	A    *Result `json:"a"`
	B    *Result `json:"b"`
	// Delta is B's value minus A's, per judge that scored both.
	Delta map[string]float64 `json:"delta"`
	// Config is the difference between the settings B's run started
	// under and A's.
	Config ConfigDiff `json:"config"`
}

Pair is one task under both configurations.

type Report

type Report struct {
	Suite    string             `json:"suite"`
	Manifest Manifest           `json:"manifest"`
	Results  []Result           `json:"results"`
	ByJudge  map[string]Summary `json:"by_judge"`
}

Report is a runner's results over a suite. ByJudge averages each judge's scores across the tasks it scored, which is not the rule export.PreferScore applies when it chooses between branches of one session: that takes each branch's best score across judges. A report and a preference can therefore name different winners.

func ReadReport

func ReadReport(rd io.Reader) (*Report, error)

ReadReport reads a report written by WriteJSON.

func (*Report) Judges

func (r *Report) Judges() []string

Judges returns the judge names in the report, sorted.

func (*Report) Result

func (r *Report) Result(task string) (*Result, bool)

Result returns the result for a task ID.

func (*Report) WriteJSON

func (r *Report) WriteJSON(w io.Writer) error

WriteJSON writes the report as indented JSON with a trailing newline.

type Result

type Result struct {
	Task Task `json:"task"`
	// SessionID is the session the run was recorded in.
	SessionID string `json:"session_id"`
	// Target is the entry every score of this result targets: the last
	// entry of the run, before any outcome was appended.
	Target string `json:"target,omitempty"`
	// Reason is how the last run of the task ended.
	Reason agentturn.Reason `json:"reason,omitempty"`
	// Runs is how many runs of the loop the task took: one per prompt
	// sent and one per resume.
	Runs int `json:"runs"`
	// Resumes is how many of those runs were resumes through
	// [Runner.Answer].
	Resumes int `json:"resumes,omitempty"`
	// ResumeBound says a prompt was still waiting on input when
	// [Runner.MaxResumes] stopped resuming it, where a task that ends
	// input_required without it stopped because there was no Answer,
	// or it had nothing to say or failed.
	ResumeBound bool `json:"resume_bound,omitempty"`
	// Usage is the sum over the model calls on the run's path: every
	// response, and every fold or branch summary that reported usage.
	// It is the path the exported document's final metrics sum, and the
	// runner prices each call under the model the exporter does, so
	// with the same price hook the two agree on the total whenever
	// every call was priced.
	Usage openresponses.Usage `json:"usage"`
	// CostUSD is the run's cost under Runner.Cost, when every call was
	// priced.
	CostUSD *float64 `json:"cost_usd,omitempty"`
	// Scores are the judges' verdicts, in the runner's judge order. A
	// judge that failed is missing here and named in Err.
	Scores []Score `json:"scores"`
	// Ends are the run ends, in order, for consumers in memory.
	Ends []*agentturn.RunEnd `json:"-"`
	// Err is what went wrong: a run that could not start or ended in
	// error, a store failure, or a judge that could not reach a
	// verdict. A result with an error may still carry scores.
	Err error `json:"-"`
}

Result is one task's run and its scores.

func (Result) MarshalJSON

func (r Result) MarshalJSON() ([]byte, error)

MarshalJSON writes the result with Err as an "error" string.

func (*Result) UnmarshalJSON

func (r *Result) UnmarshalJSON(data []byte) error

UnmarshalJSON reads a result written by MarshalJSON; an "error" string becomes Err.

type Runner

type Runner struct {
	// Store is where each run's session is created. Required.
	Store agentsession.Store
	// Config returns the configuration under test for a task.
	// Required unless ConfigWith is set.
	Config func(Task) agentturn.Config
	// ConfigWith is Config with the recorder that writes the run's
	// session, and wins over Config when set. It is where a
	// configuration binds anything that must reach the record:
	// compact.WithOnFold(rec.Fold), without which a compacting
	// configuration records no fold and its session cannot be replayed
	// strictly, and the recorder itself, which a layer keeps to
	// annotate through Recorder.Annotate the run it is in.
	//
	//	ConfigWith: func(t agenteval.Task, rec *session.Recorder) agentturn.Config {
	//		cfg := product.Config(t)
	//		cfg.Transform = compact.NewLocal(cfg.Model, compact.WithOnFold(rec.Fold)).Transform
	//		return cfg
	//	}
	ConfigWith func(Task, *session.Recorder) agentturn.Config
	// Judges score each run once it has ended. Each score is appended
	// to the run's session as an outcome entry.
	Judges []Judge
	// Answer answers the calls a run left pending when it ended
	// input_required, so an evaluation of a product whose policy asks
	// measures the whole run rather than the part before the first
	// ask. The run is resumed with what it returns and the resume is
	// recorded like any other turn; no answers, or a nil Answer, ends
	// the task there as before. agentpolicy.Engine.Answers satisfies
	// this signature as written.
	Answer func(context.Context, *agentturn.RunEnd) ([]agentturn.Answer, error)
	// MaxResumes bounds how many times one prompt may be resumed
	// through Answer, so an answer source that keeps a call pending
	// cannot loop. Zero means [DefaultMaxResumes].
	MaxResumes int
	// Parallel bounds how many tasks run at once; zero or one means one
	// at a time.
	Parallel int
	// Header, when set, supplies each run's session header: a harness
	// name, a working directory, a fixed ID for a test. Empty fields
	// are filled by the store.
	Header func(Task) agentsession.Header
	// Cost prices one model call, as export.Options.Cost does; see
	// price.Hook. When set, and every call of a run is priced, the
	// result carries the run's cost.
	Cost func(model string, usage openresponses.Usage) (float64, bool)
}

Runner sends each task of a suite through a configuration, records the run as a session and judges it.

func (*Runner) Run

func (r *Runner) Run(ctx context.Context, suite *Suite) (*Report, error)

Run sends every task of the suite through the configuration and returns the report. A task whose run or judging failed has its error on its result; Run itself fails only when it cannot start, or when ctx is done before every task has run, in which case the report holds the results so far.

type Score

type Score struct {
	// Judge is the judge's name; it becomes the outcome entry's label.
	Judge string `json:"judge"`
	// Value is the score on the judge's own scale: comparable across
	// runs of the same judge, not normalised. The deterministic judges
	// use 0 and 1; a Harbor verifier reports whatever its test script
	// wrote. A judge that normalises keeps the raw number in Details.
	Value float64 `json:"value"`
	// Pass is the judge's verdict.
	Pass bool `json:"pass"`
	// Reason says why, in a sentence.
	Reason string `json:"reason,omitempty"`
	// Details is the judge's own record, kept under the outcome's
	// details.judge.
	Details json.RawMessage `json:"details,omitempty"`
	// Session is the ID of the session the judge recorded of its own
	// run, when it is an agent that recorded one.
	Session string `json:"session,omitempty"`
}

Score is one judge's verdict on one run.

type Suite

type Suite struct {
	Name     string
	Tasks    []Task
	Manifest Manifest
}

Suite is a set of tasks loaded together.

func LoadSuite

func LoadSuite(fsys fs.FS, location string) (*Suite, error)

LoadSuite reads a suite from a directory of fsys: every file ending in .json except SuiteFile is one task, read in name order, and SuiteFile, when present, names the suite. Without it the suite is named after the directory. The manifest records every file read. The fs.FS is never written.

type Summary

type Summary struct {
	// Count is how many results the judge scored.
	Count int `json:"count"`
	// Mean is the mean of the judge's values over those results.
	Mean float64 `json:"mean"`
	// Passed is how many of them the judge passed, and PassRate is
	// Passed over Count.
	Passed   int     `json:"passed"`
	PassRate float64 `json:"pass_rate"`
}

Summary is one judge's numbers over the tasks it scored.

type Task

type Task struct {
	// ID names the task within its suite. LoadSuite defaults it to the
	// file name without its extension.
	ID string `json:"id"`
	// Instruction is the common case: one user message.
	Instruction string `json:"instruction,omitempty"`
	// Prompts is the multi-turn case: what the user says, in order,
	// one run of the loop each. It wins over Instruction when set.
	Prompts openresponses.Items `json:"prompts,omitempty"`
	// Setup is product-defined, such as a repository revision to check
	// out before the run. The runner records it and honours none of it.
	Setup map[string]string `json:"setup,omitempty"`
	// Expect is judge-defined: what the deterministic judges compare
	// the run against.
	Expect json.RawMessage `json:"expect,omitempty"`
	// Meta is free-form metadata, recorded with the run.
	Meta map[string]string `json:"meta,omitempty"`
}

Task is one thing to send through a configuration and judge.

func FromSession

func FromSession(s *agentsession.Session) (Task, error)

FromSession turns a recorded run into a task: the user messages on the path to the session's current leaf, in order, become Prompts, so the same conversation can be sent through another configuration. The task is named after the session and Meta records the session and the leaf. Model output, tool outputs, summaries and app-only entries are not prompts and are left out.

func FromSessionAt

func FromSessionAt(s *agentsession.Session, leaf string) (Task, error)

FromSessionAt is FromSession over the path to leaf.

func (Task) Inputs

func (t Task) Inputs() openresponses.Items

Inputs returns what the runner sends: Prompts when set, else one user message carrying Instruction, else nothing.

type TaskRecord

type TaskRecord struct {
	Suite    string            `json:"suite"`
	Location string            `json:"location,omitempty"`
	Task     string            `json:"task"`
	Setup    map[string]string `json:"setup,omitempty"`
	Meta     map[string]string `json:"meta,omitempty"`
}

TaskRecord is the data of a TaskNS custom entry: which suite and task a session ran, and the task's setup and metadata.

Directories

Path Synopsis
Package judge holds the judges: three deterministic ones that read a trajectory against the task's Expect, and Rubric, an agent that reads the trajectory and answers in a fixed schema.
Package judge holds the judges: three deterministic ones that read a trajectory against the task's Expect, and Rubric, an agent that reads the trajectory and answers in a fixed schema.
Package price costs a run.
Package price costs a run.
Package replay serves a recorded session as a model and as tools, so hooks, transforms and fronts are tested against real traffic with nothing behind them.
Package replay serves a recorded session as a model and as tools, so hooks, transforms and fronts are tested against real traffic with nothing behind them.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL