Documentation
¶
Overview ¶
Package agenteval evaluates Go agents over Open Responses: tasks and suites loaded from an fs.FS, a runner that sends each task through an agentturn configuration and records the run as an agentsession, a judge contract, and scores written back into the session as outcome entries. The replay package serves a recorded session as a model and as tools, so hooks, transforms and fronts are tested against real traffic offline; the judge package holds the deterministic judges and the rubric judge; price costs a run; the harbor nested module reads a Harbor task and a verifier reward.
A run is an ordinary session: the recorder writes it, the exporter renders it, and every score is an outcome entry whose target is the last entry of the run it judges. The report is a value over those sessions and holds no fact they lack.
suite, _ := agenteval.LoadSuite(os.DirFS("evals"), "smoke")
r := &agenteval.Runner{
Store: store,
Config: func(agenteval.Task) agentturn.Config { return cfg },
Judges: []agenteval.Judge{judge.Contains("text"), judge.ToolCalled("bash")},
}
report, _ := r.Run(ctx, suite)
report.WriteJSON(os.Stdout)
Index ¶
- Constants
- Variables
- func NewOutcome(score Score, target, task string) (*agentsession.OutcomeEntry, error)
- func ReadOutcome(e *agentsession.OutcomeEntry) (Score, OutcomeDetails, bool)
- func Summarize(results []Result) map[string]Summary
- type Change
- type Comparison
- type ConfigDiff
- type Judge
- type JudgeFunc
- type Manifest
- type OutcomeDetails
- type Pair
- type Report
- type Result
- type Runner
- type Score
- type Suite
- type Summary
- type Task
- type TaskRecord
Constants ¶
const DefaultMaxResumes = 8
DefaultMaxResumes is how many times one prompt is resumed through Runner.Answer when Runner.MaxResumes is zero.
const SuiteFile = "suite.json"
SuiteFile is the name of the optional file in a suite directory that names the suite. Its content is {"name": "..."}.
const TaskNS = "agenteval:task"
TaskNS is the namespace of the custom entry the runner writes before each run, whose data is a TaskRecord.
Variables ¶
var ErrNoPrompts = errors.New("agenteval: session has no user prompts on its path")
ErrNoPrompts is returned by FromSession when the path holds no user message.
var ErrUnreplayable = errors.New("agenteval: the run cannot be replayed strictly")
ErrUnreplayable is joined onto the result of a run whose session cannot be replayed strictly, because the record does not rebuild every request the run sent.
Functions ¶
func NewOutcome ¶
func NewOutcome(score Score, target, task string) (*agentsession.OutcomeEntry, error)
NewOutcome builds the outcome entry for a score: kind agentsession.OutcomeEval, target the entry named, score, pass and label from the score, and OutcomeDetails carrying the task. The entry is not appended.
func ReadOutcome ¶
func ReadOutcome(e *agentsession.OutcomeEntry) (Score, OutcomeDetails, bool)
ReadOutcome reads a score back out of an outcome entry the runner wrote. It reports false for an entry of another kind or one without OutcomeDetails.
Types ¶
type Comparison ¶
type Comparison struct {
Suite string `json:"suite"`
Manifest Manifest `json:"manifest"`
// A and B are the two reports. They are not written by WriteJSON;
// the pairs carry both results of every task.
A, B *Report `json:"-"`
// Config is the difference between the settings B's runs started
// under and A's, when every pair differs the same way; Uniform
// says whether they did. Each pair carries its own. Whether a run
// folded depends on the task as much as the configuration, so
// pairs are uniform whatever their Folded says, and Config.Folded
// says whether any of A's runs folded and any of B's, when those
// differ.
Config ConfigDiff `json:"config"`
Uniform bool `json:"uniform"`
Pairs []Pair `json:"pairs"`
// ByJudge is B's mean minus A's, per judge scored on both sides.
ByJudge map[string]float64 `json:"by_judge"`
}
Comparison is the same suite run under two configurations and judged the same way, paired by task.
func Compare ¶
Compare runs the suite under a and b and pairs the results by task. The runners share nothing but the suite; each records its own sessions in its own store, which may be the same store. The comparison's Config is the difference between the settings the two runs of each task started under, read back from their sessions, so a comparison says how the configurations differed and not only which scored higher.
type ConfigDiff ¶
type ConfigDiff struct {
Model *Change `json:"model,omitempty"`
Instructions *Change `json:"instructions,omitempty"`
Reasoning *Change `json:"reasoning,omitempty"`
Text *Change `json:"text,omitempty"`
// ToolsAdded names tools B has and A lacks, ToolsRemoved the
// reverse, and ToolsChanged tools both have under a different
// definition; each sorted.
ToolsAdded []string `json:"tools_added,omitempty"`
ToolsRemoved []string `json:"tools_removed,omitempty"`
ToolsChanged []string `json:"tools_changed,omitempty"`
// Extra holds passthrough members that differ, by wire name; a
// side that lacks the member has nil.
Extra map[string]Change `json:"extra,omitempty"`
// BeyondSettings says the two runs' first calls sent different
// requests although their settings were the same: a transform, a
// BeforeTurn or BeforeModelCall hook, or items one side injected.
// None of those is a setting, so no field above can name it, and a
// comparison of two configurations that differ only there would
// otherwise report that nothing differed, which is worse than
// reporting nothing. It is read from the request hashes the two
// first responses recorded: two that differ say so, and one hash
// against none says the other run sent something the record does
// not rebuild, which is a transform or a hook by definition. It
// compares the first call of each run, so a transform that
// changes nothing until later, as a compaction transform does not
// fold before it has anything to fold, is not visible here;
// Folded is.
BeyondSettings bool `json:"beyond_settings,omitempty"`
// Folded says one run's path held a compaction entry and the
// other's none, with whether each side folded: a context strategy
// is not a setting either, and the fold is what it leaves on the
// path. Two runs that both folded, however often, do not differ
// here. On a [Comparison] it is over all the runs of each side.
Folded *Change `json:"folded,omitempty"`
}
ConfigDiff is what differed between two runs' settings: the settings each run's first model call was made under, replayed from its session's config entries.
func DiffSettings ¶
func DiffSettings(a, b agentsession.Settings) ConfigDiff
DiffSettings returns what differs between a and b. It covers the settings the record describes, which is the model, the instructions, the reasoning and text configuration, the tools and the passthrough members, and nothing else: a transform, a hook and an injected item are not settings. Compare reports one of those through ConfigDiff.BeyondSettings.
type Judge ¶
type Judge interface {
// Name identifies the judge in scores and outcome entries.
Name() string
// Judge scores the trajectory. An error means the judge could not
// reach a verdict, not that the run failed; the runner records the
// error on the result and writes no outcome for it.
Judge(ctx context.Context, t export.Trajectory, task Task) (Score, error)
}
Judge scores a run from its trajectory and the task it ran.
type JudgeFunc ¶
type JudgeFunc struct {
JudgeName string
Fn func(ctx context.Context, t export.Trajectory, task Task) (Score, error)
}
JudgeFunc adapts a function to Judge.
type Manifest ¶
Manifest says where a suite was loaded from and what it held: the location within the fs.FS and the content hash of every file read, keyed by path, so a result names the suite version it came from.
type OutcomeDetails ¶
type OutcomeDetails struct {
Task string `json:"task"`
Reason string `json:"reason,omitempty"`
Session string `json:"session,omitempty"`
Judge json.RawMessage `json:"judge,omitempty"`
}
OutcomeDetails is the details member of an outcome entry the runner writes, so a reader gets the task, the reason and the judge's own session at known keys without knowing any judge's private JSON.
type Pair ¶
type Pair struct {
Task string `json:"task"`
A *Result `json:"a"`
B *Result `json:"b"`
// Delta is B's value minus A's, per judge that scored both.
Delta map[string]float64 `json:"delta"`
// Config is the difference between the settings B's run started
// under and A's.
Config ConfigDiff `json:"config"`
}
Pair is one task under both configurations.
type Report ¶
type Report struct {
Suite string `json:"suite"`
Manifest Manifest `json:"manifest"`
Results []Result `json:"results"`
ByJudge map[string]Summary `json:"by_judge"`
}
Report is a runner's results over a suite. ByJudge averages each judge's scores across the tasks it scored, which is not the rule export.PreferScore applies when it chooses between branches of one session: that takes each branch's best score across judges. A report and a preference can therefore name different winners.
func ReadReport ¶
ReadReport reads a report written by WriteJSON.
type Result ¶
type Result struct {
Task Task `json:"task"`
// SessionID is the session the run was recorded in.
SessionID string `json:"session_id"`
// Target is the entry every score of this result targets: the last
// entry of the run, before any outcome was appended.
Target string `json:"target,omitempty"`
// Reason is how the last run of the task ended.
Reason agentturn.Reason `json:"reason,omitempty"`
// Runs is how many runs of the loop the task took: one per prompt
// sent and one per resume.
Runs int `json:"runs"`
// Resumes is how many of those runs were resumes through
// [Runner.Answer].
Resumes int `json:"resumes,omitempty"`
// ResumeBound says a prompt was still waiting on input when
// [Runner.MaxResumes] stopped resuming it, where a task that ends
// input_required without it stopped because there was no Answer,
// or it had nothing to say or failed.
ResumeBound bool `json:"resume_bound,omitempty"`
// Usage is the sum over the model calls on the run's path: every
// response, and every fold or branch summary that reported usage.
// It is the path the exported document's final metrics sum, and the
// runner prices each call under the model the exporter does, so
// with the same price hook the two agree on the total whenever
// every call was priced.
Usage openresponses.Usage `json:"usage"`
// CostUSD is the run's cost under Runner.Cost, when every call was
// priced.
CostUSD *float64 `json:"cost_usd,omitempty"`
// Scores are the judges' verdicts, in the runner's judge order. A
// judge that failed is missing here and named in Err.
Scores []Score `json:"scores"`
// Ends are the run ends, in order, for consumers in memory.
Ends []*agentturn.RunEnd `json:"-"`
// Err is what went wrong: a run that could not start or ended in
// error, a store failure, or a judge that could not reach a
// verdict. A result with an error may still carry scores.
Err error `json:"-"`
}
Result is one task's run and its scores.
func (Result) MarshalJSON ¶
MarshalJSON writes the result with Err as an "error" string.
func (*Result) UnmarshalJSON ¶
UnmarshalJSON reads a result written by MarshalJSON; an "error" string becomes Err.
type Runner ¶
type Runner struct {
// Store is where each run's session is created. Required.
Store agentsession.Store
// Config returns the configuration under test for a task.
// Required unless ConfigWith is set.
Config func(Task) agentturn.Config
// ConfigWith is Config with the recorder that writes the run's
// session, and wins over Config when set. It is where a
// configuration binds anything that must reach the record:
// compact.WithOnFold(rec.Fold), without which a compacting
// configuration records no fold and its session cannot be replayed
// strictly, and the recorder itself, which a layer keeps to
// annotate through Recorder.Annotate the run it is in.
//
// ConfigWith: func(t agenteval.Task, rec *session.Recorder) agentturn.Config {
// cfg := product.Config(t)
// cfg.Transform = compact.NewLocal(cfg.Model, compact.WithOnFold(rec.Fold)).Transform
// return cfg
// }
ConfigWith func(Task, *session.Recorder) agentturn.Config
// Judges score each run once it has ended. Each score is appended
// to the run's session as an outcome entry.
Judges []Judge
// Answer answers the calls a run left pending when it ended
// input_required, so an evaluation of a product whose policy asks
// measures the whole run rather than the part before the first
// ask. The run is resumed with what it returns and the resume is
// recorded like any other turn; no answers, or a nil Answer, ends
// the task there as before. agentpolicy.Engine.Answers satisfies
// this signature as written.
Answer func(context.Context, *agentturn.RunEnd) ([]agentturn.Answer, error)
// MaxResumes bounds how many times one prompt may be resumed
// through Answer, so an answer source that keeps a call pending
// cannot loop. Zero means [DefaultMaxResumes].
MaxResumes int
// Parallel bounds how many tasks run at once; zero or one means one
// at a time.
Parallel int
// Header, when set, supplies each run's session header: a harness
// name, a working directory, a fixed ID for a test. Empty fields
// are filled by the store.
Header func(Task) agentsession.Header
// Cost prices one model call, as export.Options.Cost does; see
// price.Hook. When set, and every call of a run is priced, the
// result carries the run's cost.
Cost func(model string, usage openresponses.Usage) (float64, bool)
}
Runner sends each task of a suite through a configuration, records the run as a session and judges it.
func (*Runner) Run ¶
Run sends every task of the suite through the configuration and returns the report. A task whose run or judging failed has its error on its result; Run itself fails only when it cannot start, or when ctx is done before every task has run, in which case the report holds the results so far.
type Score ¶
type Score struct {
// Judge is the judge's name; it becomes the outcome entry's label.
Judge string `json:"judge"`
// Value is the score on the judge's own scale: comparable across
// runs of the same judge, not normalised. The deterministic judges
// use 0 and 1; a Harbor verifier reports whatever its test script
// wrote. A judge that normalises keeps the raw number in Details.
Value float64 `json:"value"`
// Pass is the judge's verdict.
Pass bool `json:"pass"`
// Reason says why, in a sentence.
Reason string `json:"reason,omitempty"`
// Details is the judge's own record, kept under the outcome's
// details.judge.
Details json.RawMessage `json:"details,omitempty"`
// Session is the ID of the session the judge recorded of its own
// run, when it is an agent that recorded one.
Session string `json:"session,omitempty"`
}
Score is one judge's verdict on one run.
type Suite ¶
Suite is a set of tasks loaded together.
func LoadSuite ¶
LoadSuite reads a suite from a directory of fsys: every file ending in .json except SuiteFile is one task, read in name order, and SuiteFile, when present, names the suite. Without it the suite is named after the directory. The manifest records every file read. The fs.FS is never written.
type Summary ¶
type Summary struct {
// Count is how many results the judge scored.
Count int `json:"count"`
// Mean is the mean of the judge's values over those results.
Mean float64 `json:"mean"`
// Passed is how many of them the judge passed, and PassRate is
// Passed over Count.
Passed int `json:"passed"`
PassRate float64 `json:"pass_rate"`
}
Summary is one judge's numbers over the tasks it scored.
type Task ¶
type Task struct {
// ID names the task within its suite. LoadSuite defaults it to the
// file name without its extension.
ID string `json:"id"`
// Instruction is the common case: one user message.
Instruction string `json:"instruction,omitempty"`
// Prompts is the multi-turn case: what the user says, in order,
// one run of the loop each. It wins over Instruction when set.
Prompts openresponses.Items `json:"prompts,omitempty"`
// Setup is product-defined, such as a repository revision to check
// out before the run. The runner records it and honours none of it.
Setup map[string]string `json:"setup,omitempty"`
// Expect is judge-defined: what the deterministic judges compare
// the run against.
Expect json.RawMessage `json:"expect,omitempty"`
// Meta is free-form metadata, recorded with the run.
Meta map[string]string `json:"meta,omitempty"`
}
Task is one thing to send through a configuration and judge.
func FromSession ¶
func FromSession(s *agentsession.Session) (Task, error)
FromSession turns a recorded run into a task: the user messages on the path to the session's current leaf, in order, become Prompts, so the same conversation can be sent through another configuration. The task is named after the session and Meta records the session and the leaf. Model output, tool outputs, summaries and app-only entries are not prompts and are left out.
func FromSessionAt ¶
func FromSessionAt(s *agentsession.Session, leaf string) (Task, error)
FromSessionAt is FromSession over the path to leaf.
func (Task) Inputs ¶
func (t Task) Inputs() openresponses.Items
Inputs returns what the runner sends: Prompts when set, else one user message carrying Instruction, else nothing.
type TaskRecord ¶
type TaskRecord struct {
Suite string `json:"suite"`
Location string `json:"location,omitempty"`
Task string `json:"task"`
Setup map[string]string `json:"setup,omitempty"`
Meta map[string]string `json:"meta,omitempty"`
}
TaskRecord is the data of a TaskNS custom entry: which suite and task a session ran, and the task's setup and metadata.
Directories
¶
| Path | Synopsis |
|---|---|
|
Package judge holds the judges: three deterministic ones that read a trajectory against the task's Expect, and Rubric, an agent that reads the trajectory and answers in a fixed schema.
|
Package judge holds the judges: three deterministic ones that read a trajectory against the task's Expect, and Rubric, an agent that reads the trajectory and answers in a fixed schema. |
|
Package price costs a run.
|
Package price costs a run. |
|
Package replay serves a recorded session as a model and as tools, so hooks, transforms and fronts are tested against real traffic with nothing behind them.
|
Package replay serves a recorded session as a model and as tools, so hooks, transforms and fronts are tested against real traffic with nothing behind them. |