agenteval

package module
v0.0.10 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 2, 2026 License: MIT Imports: 21 Imported by: 0

README

agenteval

Evaluation for Go agents over Open Responses: a replay model and recorded tools served from a session, tasks and suites loaded from an fs.FS, a runner that records every run as a session, judges, and scores written back as outcome entries.

go get github.com/ChristopherDavenport/agenteval

What it does

A run is an ordinary agentsession: the runner sends a task through an agentturn configuration with the session recorder attached, so every request, response and tool output is on disk and the exporter renders it as an ATIF trajectory. Judges read that trajectory. Each score is appended to the session as an outcome entry whose target is the run's last entry, whose kind is eval, and whose pass and score are where the format puts them, so a pass rate is computable from the session files alone. The report is a value over those sessions and holds no fact they lack.

suite, _ := agenteval.LoadSuite(os.DirFS("evals"), "smoke")
r := &agenteval.Runner{
	Store:  store,
	Config: func(agenteval.Task) agentturn.Config { return cfg },
	Judges: []agenteval.Judge{judge.Contains("text"), judge.ToolCalled("bash")},
	Cost:   price.Hook(prices),
}
report, _ := r.Run(ctx, suite)
report.WriteJSON(os.Stdout)

A configuration that needs the run's recorder takes ConfigWith instead: a compacting configuration binds compact.WithOnFold there, without which its folds reach no session and the run cannot be replayed, and a layer that annotates the run it is in keeps the same recorder. A configuration whose construction can fail or holds something to release, such as an agentkit kit, takes Build, which returns an error and a close; SessionOptions opens each run's recorder with options such as session.WithInstructionsParts, so the session records which instructions part changed. A Header that names a Base forks that session, and the agent starts from the context there. Answer answers the calls a run left pending when it ended input_required and resumes it, so a product whose policy asks is measured over the whole run rather than the part before its first ask.

r := &agenteval.Runner{
	Store: store,
	ConfigWith: func(t agenteval.Task, rec *session.Recorder) agentturn.Config {
		cfg := product.Config(t)
		cfg.Transform = compact.NewLocal(cfg.Model, compact.WithOnFold(rec.Fold)).Transform
		return cfg
	},
	Answer: engine.Answers,
}

A task file is JSON: an instruction, or prompts for several turns, with expect for the judges and setup and meta for the product. The suite's manifest, the hash of every file loaded, is recorded in each run's env entry, so a result names the suite version it came from.

Replay

replay.NewModel serves a recorded session's model calls in path order, folds included, and in strict mode refuses a request whose hash differs from the recorded one: a hook, a transform or a front that changes what the model would have been sent fails loudly against real traffic. A record that carries no hash for a call is refused rather than served unchecked — at NewModel when a response on the path is unhashed, so a replay that could not have checked what it served says so before it starts instead of passing quietly, and at the call for a fold — unless replay.AllowUnhashed() says to serve it anyway. The one call no mode checks is a fold through the compaction endpoint, which sends no request the format hashes. replay.Tools serves recorded tool outputs by call ID or by name and canonical arguments, so a whole run replays without touching the world.

s, _ := store.Open(ctx, id)
model, _ := replay.NewModel(s, replay.Strict())
cfg.Model = model
cfg.BeforeModelCall = model.BeforeModelCall
cfg.Tools = replay.Tools(s, cfg.Tools, replay.Strict())

Model.BeforeModelCall serves the instructions and the tool list recorded for each call. A product whose layers rebuild those every turn, from a memory store, a skill set or an AGENTS.md, otherwise replays against what the layers say today and diverges at the first call with an error naming two hashes and no layer. Model.Settings is the rest of what each recorded call was made under, for a judge or a check that wants to read it.

Judges

judge.Exact, judge.Contains and judge.ToolCalled are deterministic. judge.Rubric is an agent: the trajectory is its input, the rubric its instructions, and its answer is constrained to a score schema; with a store it records its own session, parented to the one it judged, so a judgement is as replayable as the run.

Samples and groups

Runner.Samples runs each task several times, each in its own session, with Result.Sample and Task.Meta["sample"] saying which run it was. Runner.GroupJudges score a task's samples together once all of them are judged, as an RL producer computes advantages relative to the group, and each score is an outcome on its sample's session. A group with a sample that was not judged, because its run failed, is not group judged. Each session's task record names the group judges it expects, so a store tells a group that was never scored, because a batch ended before its group step, from one that was.

Comparison and cost

Compare runs a suite under two configurations, pairs the results by task, and names how the two configurations differed, read back from the sessions. The difference it names is the settings the record describes; when those are the same and the two first requests are not, it says so with beyond_settings, because a transform, a hook and an injected item are not settings. A context strategy that folds only later in the run shows as folded, whether each side's path held a compaction entry, when one did and the other did not.

price loads a table of rates and prices a run with ATIF's formula, and the runner and the exporter take the same hook. A result's usage and cost_usd sum every model call on the run's path, the folds a compacting configuration made included, which is what the exported document's final metrics sum, so the two agree on the total, a fold that failed included: its summary calls are counted by both from agentsession v0.0.20 (agentsession#184). A run with a call the hook cannot price has no cost_usd, where the document reports the sum of the calls it could.

Harbor

The harbor nested module reads a Harbor task directory as a Task (the prompt only; the environment and the verifier stay with Harbor) and a finished trial's reward.json or reward.txt as scores. Task.Meta carries [task] and [metadata] under task. and metadata., because the second table is free-form and invites the first table's words.

Design

The plan is in docs/plans/eval-layer.md and the findings of the design studies that shaped it in docs/feedback.md.

License

MIT.

Documentation

Overview

Package agenteval evaluates Go agents over Open Responses: tasks and suites loaded from an fs.FS, a runner that sends each task through an agentturn configuration and records the run as an agentsession, a judge contract, and scores written back into the session as outcome entries. The replay package serves a recorded session as a model and as tools, so hooks, transforms and fronts are tested against real traffic offline; the judge package holds the deterministic judges and the rubric judge; price costs a run; the harbor nested module reads a Harbor task and a verifier reward.

A run is an ordinary session: the recorder writes it, the exporter renders it, and every score is an outcome entry whose target is the last entry of the run it judges. The report is a value over those sessions and holds no fact they lack.

suite, _ := agenteval.LoadSuite(os.DirFS("evals"), "smoke")
r := &agenteval.Runner{
	Store:  store,
	Config: func(agenteval.Task) agentturn.Config { return cfg },
	Judges: []agenteval.Judge{judge.Contains("text"), judge.ToolCalled("bash")},
}
report, _ := r.Run(ctx, suite)
report.WriteJSON(os.Stdout)

Index

Constants

View Source
const (
	UnjudgedA    = "a"
	UnjudgedB    = "b"
	UnjudgedBoth = "both"
)

The values of Pair.Unjudged: which side's result was not judged.

View Source
const DefaultMaxResumes = 8

DefaultMaxResumes is how many times one prompt is resumed through Runner.Answer when Runner.MaxResumes is zero.

View Source
const SampleMeta = "sample"

SampleMeta is the Task.Meta key the runner sets on each sample of a task when Runner.Samples is above one, to the sample's number from "1", and reads to run one sample of a task alone.

View Source
const SuiteFile = "suite.json"

SuiteFile is the name of the optional file in a suite directory that names the suite. Its content is {"name": "..."}.

View Source
const TaskNS = "agenteval:task"

TaskNS is the namespace of the custom entry the runner writes before each run, whose data is a TaskRecord.

Variables

View Source
var ErrNoPrompts = errors.New("agenteval: session has no user prompts on its path")

ErrNoPrompts is returned by FromSession when the path holds no user message.

View Source
var ErrUnreplayable = errors.New("agenteval: the run cannot be replayed strictly")

ErrUnreplayable is joined onto the result of a run whose session cannot be replayed strictly, because the record does not rebuild every request the run sent.

Functions

func NewOutcome

func NewOutcome(score Score, target, task string) (*agentsession.OutcomeEntry, error)

NewOutcome builds the outcome entry for a score: kind agentsession.OutcomeEval, target the entry named, score, pass and label from the score, and OutcomeDetails carrying the task. The entry is not appended.

func ReadOutcome

ReadOutcome reads a score back out of an outcome entry the runner wrote. It reports false for an entry of another kind or one without OutcomeDetails.

func SampleName added in v0.0.10

func SampleName(task string, sample int) string

SampleName is the name the runner gives a sample's session: the task ID for a runner that does not sample (sample 0), and task#n for the nth sample, so a store listing with names tells the samples of a task apart.

func Summarize

func Summarize(results []Result) map[string]Summary

Summarize computes the per-judge summaries over results. It reads the results alone, so a report read back from JSON summarises as the runner's did: a judge is summarised when it scored at least one result, and its Unjudged is the results it did not score.

Types

type Change

type Change struct {
	A any `json:"a"`
	B any `json:"b"`
}

Change is one setting under A and under B.

type Comparison

type Comparison struct {
	Suite    string   `json:"suite"`
	Manifest Manifest `json:"manifest"`
	// A and B are the two reports. They are not written by WriteJSON;
	// the pairs carry both results of every task.
	A, B *Report `json:"-"`
	// Config is the difference between the settings B's runs started
	// under and A's, when every pair differs the same way; Uniform
	// says whether they did. Each pair carries its own. Whether a run
	// folded depends on the task as much as the configuration, so
	// pairs are uniform whatever their Folded and FoldFailed say, and
	// Config.Folded says whether any of A's runs folded and any of
	// B's, when those differ; Config.FoldFailed the same for a fold
	// that failed.
	Config  ConfigDiff `json:"config"`
	Uniform bool       `json:"uniform"`
	Pairs   []Pair     `json:"pairs"`
	// ByJudge is B's mean minus A's, per judge, over the pairs the judge
	// scored on both sides. A pair with an unjudged side, a run that
	// failed under the default of leaving it unjudged, is left out of
	// it and counted on Unjudged: the mean says how the two compared
	// where both answered, and Unjudged how often one did not. An
	// evaluation where a crash counts as a failure, Harbor's included,
	// sets [Runner.JudgeFailedRuns] on both runners, and the failed run
	// then scores 0 here like any other.
	ByJudge map[string]float64 `json:"by_judge"`
	// Unjudged is how many pairs have a side that was not judged; each
	// such pair says which on [Pair.Unjudged].
	Unjudged int `json:"unjudged"`
}

Comparison is the same suite run under two configurations and judged the same way, paired by task.

func Compare

func Compare(ctx context.Context, suite *Suite, a, b *Runner) (*Comparison, error)

Compare runs the suite under a and b and pairs the results by task. The runners share nothing but the suite; each records its own sessions in its own store, which may be the same store. The comparison's Config is the difference between the settings the two runs of each task started under, read back from their sessions, so a comparison says how the configurations differed and not only which scored higher.

A pair whose side was not judged, because its run failed and that runner leaves a failed run unjudged, is marked on Pair.Unjudged and counted on Comparison.Unjudged; its ByJudge covers the pairs both sides answered. A configuration that crashes on the hard tasks and answers the easy ones therefore does not read as the equal of one that answers all of them. For an evaluation where a crash counts as a failure, Harbor's included, which counts a trial with no reward as 0, set Runner.JudgeFailedRuns on both runners.

func (*Comparison) WriteJSON

func (c *Comparison) WriteJSON(w io.Writer) error

WriteJSON writes the comparison as indented JSON with a trailing newline. The two reports are not included; write them separately.

type ConfigDiff

type ConfigDiff struct {
	Model        *Change `json:"model,omitempty"`
	Instructions *Change `json:"instructions,omitempty"`
	Reasoning    *Change `json:"reasoning,omitempty"`
	Text         *Change `json:"text,omitempty"`
	// ToolsAdded names tools B has and A lacks, ToolsRemoved the
	// reverse, and ToolsChanged tools both have under a different
	// definition; each sorted.
	ToolsAdded   []string `json:"tools_added,omitempty"`
	ToolsRemoved []string `json:"tools_removed,omitempty"`
	ToolsChanged []string `json:"tools_changed,omitempty"`
	// Extra holds passthrough members that differ, by wire name; a
	// side that lacks the member has nil.
	Extra map[string]Change `json:"extra,omitempty"`
	// BeyondSettings says the two runs' first calls sent different
	// requests although their settings were the same: a transform, a
	// BeforeTurn or BeforeModelCall hook, or items one side injected.
	// None of those is a setting, so no field above can name it, and a
	// comparison of two configurations that differ only there would
	// otherwise report that nothing differed, which is worse than
	// reporting nothing. It is read from the request hashes the two
	// first responses recorded: two that differ say so, and one hash
	// against none says the other run sent something the record does
	// not rebuild, which is a transform or a hook by definition. It
	// compares the first call of each run, so a transform that
	// changes nothing until later, as a compaction transform does not
	// fold before it has anything to fold, is not visible here;
	// Folded is.
	BeyondSettings bool `json:"beyond_settings,omitempty"`
	// Folded says one run's path held a compaction entry and the
	// other's none, with whether each side folded: a context strategy
	// is not a setting either, and the fold is what it leaves on the
	// path. Two runs that both folded, however often, do not differ
	// here. On a [Comparison] it is over all the runs of each side.
	Folded *Change `json:"folded,omitempty"`
	// FoldFailed is Folded for a fold that failed: one run's path held
	// an agentturn:compaction_failed entry and the other's none. A
	// configuration whose every fold fails, on a summary model too
	// verbose for agentturn to accept its summary, leaves no compaction
	// entry, and without this would compare as one that never tried to
	// compact, though each failed fold paid for its summary calls.
	FoldFailed *Change `json:"fold_failed,omitempty"`
}

ConfigDiff is what differed between two runs' settings: the settings each run's first model call was made under, replayed from its session's config entries.

func DiffSettings

func DiffSettings(a, b agentsession.Settings) ConfigDiff

DiffSettings returns what differs between a and b. It covers the settings the record describes, which is the model, the instructions, the reasoning and text configuration, the tools and the passthrough members, and nothing else: a transform, a hook and an injected item are not settings. Compare reports one of those through ConfigDiff.BeyondSettings.

func (ConfigDiff) Empty

func (d ConfigDiff) Empty() bool

Empty reports whether nothing differed.

type GroupJudge added in v0.0.9

type GroupJudge interface {
	// Name identifies the judge in scores and outcome entries.
	Name() string
	// JudgeGroup returns one score per member, in the group's order.
	// An error, or a count other than the group's, means the judge
	// could not reach a verdict; the runner records the error on
	// every member's result and writes no outcome for it.
	JudgeGroup(ctx context.Context, group []Member, task Task) ([]Score, error)
}

GroupJudge scores the samples of one task together, once each has been judged by the runner's judges: an advantage relative to the group, or a length penalty that applies only when every sample passed. The runner calls it once per task, over its Runner.Samples samples, or the one run when it does not sample.

type GroupJudgeFunc added in v0.0.9

type GroupJudgeFunc struct {
	JudgeName string
	Fn        func(ctx context.Context, group []Member, task Task) ([]Score, error)
}

GroupJudgeFunc adapts a function to GroupJudge.

func (GroupJudgeFunc) JudgeGroup added in v0.0.9

func (j GroupJudgeFunc) JudgeGroup(ctx context.Context, group []Member, task Task) ([]Score, error)

JudgeGroup calls Fn and stamps each score with the judge's name.

func (GroupJudgeFunc) Name added in v0.0.9

func (j GroupJudgeFunc) Name() string

Name returns JudgeName.

type Judge

type Judge interface {
	// Name identifies the judge in scores and outcome entries.
	Name() string
	// Judge scores the trajectory. An error means the judge could not
	// reach a verdict, not that the run failed; the runner records the
	// error on the result and writes no outcome for it.
	Judge(ctx context.Context, t export.Trajectory, task Task) (Score, error)
}

Judge scores a run from its trajectory and the task it ran.

type JudgeFunc

type JudgeFunc struct {
	JudgeName string
	Fn        func(ctx context.Context, t export.Trajectory, task Task) (Score, error)
}

JudgeFunc adapts a function to Judge.

func (JudgeFunc) Judge

func (j JudgeFunc) Judge(ctx context.Context, t export.Trajectory, task Task) (Score, error)

Judge calls Fn and stamps the score with the judge's name.

func (JudgeFunc) Name

func (j JudgeFunc) Name() string

Name returns JudgeName.

type Manifest

type Manifest struct {
	Location string            `json:"location"`
	Files    map[string]string `json:"files"`
}

Manifest says where a suite was loaded from and what it held: the location within the fs.FS and the content hash of every file read, keyed by path, so a result names the suite version it came from.

type Member added in v0.0.9

type Member struct {
	Trajectory export.Trajectory
	Scores     []Score
}

Member is one sample of a group: its trajectory, and what the runner's judges said of it.

type OutcomeDetails

type OutcomeDetails struct {
	Task string `json:"task"`
	// Sample is the sample of the task scored, when the runner sampled
	// it, as [Result.Sample] is.
	Sample  int             `json:"sample,omitempty"`
	Reason  string          `json:"reason,omitempty"`
	Session string          `json:"session,omitempty"`
	Judge   json.RawMessage `json:"judge,omitempty"`
}

OutcomeDetails is the details member of an outcome entry the runner writes, so a reader gets the task, the reason and the judge's own session at known keys without knowing any judge's private JSON.

type Pair

type Pair struct {
	Task string `json:"task"`
	// Sample is the sample of the task both results are, when the
	// runners sampled; the runs are paired by task and sample.
	Sample int     `json:"sample,omitempty"`
	A      *Result `json:"a"`
	B      *Result `json:"b"`
	// Delta is B's value minus A's, per judge that scored both.
	Delta map[string]float64 `json:"delta"`
	// Unjudged is [UnjudgedA], [UnjudgedB] or [UnjudgedBoth] when that
	// side's result was not judged: its run failed and
	// [Runner.JudgeFailedRuns] was off, or it never reached a
	// trajectory. The side has no scores, so Delta holds nothing for
	// the judges, and the pair is counted on [Comparison.Unjudged]
	// rather than in its ByJudge. Empty when both sides were judged.
	Unjudged string `json:"unjudged,omitempty"`
	// Config is the difference between the settings B's run started
	// under and A's.
	Config ConfigDiff `json:"config"`
}

Pair is one task under both configurations.

type Report

type Report struct {
	Suite    string             `json:"suite"`
	Manifest Manifest           `json:"manifest"`
	Results  []Result           `json:"results"`
	ByJudge  map[string]Summary `json:"by_judge"`
}

Report is a runner's results over a suite. ByJudge averages each judge's scores across the tasks it scored, which is not the rule export.PreferScore applies when it chooses between branches of one session: that takes each branch's best score across judges. A report and a preference can therefore name different winners.

func ReadReport

func ReadReport(rd io.Reader) (*Report, error)

ReadReport reads a report written by WriteJSON.

func (*Report) Judges

func (r *Report) Judges() []string

Judges returns the judge names in the report, sorted.

func (*Report) Result

func (r *Report) Result(task string) (*Result, bool)

Result returns the result for a task ID: its first sample, when the runner sampled it.

func (*Report) Sample added in v0.0.9

func (r *Report) Sample(task string, sample int) (*Result, bool)

Sample returns the result of one sample of a task, numbered as Result.Sample is: 0 for a runner that does not sample.

func (*Report) WriteJSON

func (r *Report) WriteJSON(w io.Writer) error

WriteJSON writes the report as indented JSON with a trailing newline.

type Result

type Result struct {
	Task Task `json:"task"`
	// Sample is which run of the task this is, from 1, when
	// [Runner.Samples] is above one; 0 otherwise.
	Sample int `json:"sample,omitempty"`
	// SessionID is the session the run was recorded in.
	SessionID string `json:"session_id"`
	// Target is the entry every score of this result targets: the last
	// entry of the run, before any outcome was appended.
	Target string `json:"target,omitempty"`
	// Reason is how the last run of the task ended.
	Reason agentturn.Reason `json:"reason,omitempty"`
	// Runs is how many runs of the loop the task took: one per prompt
	// sent and one per resume.
	Runs int `json:"runs"`
	// Resumes is how many of those runs were resumes through
	// [Runner.Answer].
	Resumes int `json:"resumes,omitempty"`
	// ResumeBound says a prompt was still waiting on input when
	// [Runner.MaxResumes] stopped resuming it, where a task that ends
	// input_required without it stopped because there was no Answer,
	// or it had nothing to say or failed.
	ResumeBound bool `json:"resume_bound,omitempty"`
	// Usage is the sum over the model calls on the run's path: every
	// response, every fold or branch summary that reported usage, and
	// the summary calls of every fold that failed. It is the path the
	// exported document's final metrics sum, and the runner prices
	// each call under the model the exporter does, so with the same
	// price hook the two agree on the total whenever every call was
	// priced. A failed fold's calls are in the document's from
	// agentsession v0.0.20, which counts a custom entry whose data
	// carries usage (agentsession#184); before it, the document left
	// them out.
	Usage openresponses.Usage `json:"usage"`
	// CostUSD is the run's cost under Runner.Cost, when every call was
	// priced.
	CostUSD *float64 `json:"cost_usd,omitempty"`
	// Scores are the judges' verdicts, in the runner's judge order. A
	// judge that failed is missing here and named in Err; a task whose
	// run failed has none, unless [Runner.JudgeFailedRuns].
	Scores []Score `json:"scores"`
	// Ends are the run ends, in order, for consumers in memory.
	Ends []*agentturn.RunEnd `json:"-"`
	// Err is what went wrong: a run that could not start or ended in
	// error, a store failure, or a judge that could not reach a
	// verdict. A result with an error may still carry scores.
	Err error `json:"-"`
	// contains filtered or unexported fields
}

Result is one task's run and its scores.

func (Result) MarshalJSON

func (r Result) MarshalJSON() ([]byte, error)

MarshalJSON writes the result with Err as an "error" string.

func (*Result) UnmarshalJSON

func (r *Result) UnmarshalJSON(data []byte) error

UnmarshalJSON reads a result written by MarshalJSON; an "error" string becomes Err.

type Runner

type Runner struct {
	// Store is where each run's session is created. Required.
	Store agentsession.Store
	// Config returns the configuration under test for a task.
	// Required unless ConfigWith or Build is set.
	Config func(Task) agentturn.Config
	// ConfigWith is Config with the recorder that writes the run's
	// session, and wins over Config when set. It is where a
	// configuration binds anything that must reach the record:
	// compact.WithOnFold(rec.Fold), without which a compacting
	// configuration records no fold and its session cannot be replayed
	// strictly, and the recorder itself, which a layer keeps to
	// annotate through Recorder.Annotate the run it is in. A task that
	// forks a base (see Header) starts its agent from the context there,
	// which the runner seeds; a compacting configuration seeds its
	// transform the same way with Recorder.CompactOptions, so a fork of a
	// base whose last fold failed does not ask for that summary again.
	//
	//	ConfigWith: func(t agenteval.Task, rec *session.Recorder) agentturn.Config {
	//		cfg := product.Config(t)
	//		opts := append(rec.CompactOptions(), compact.WithOnFold(rec.Fold))
	//		cfg.Transform = compact.NewLocal(cfg.Model, opts...).Transform
	//		return cfg
	//	}
	ConfigWith func(Task, *session.Recorder) agentturn.Config
	// Build is ConfigWith for a configuration whose construction can
	// fail or holds something to release, such as a kit that opens MCP
	// clients, and wins over both when set. Its error ends the task on
	// Result.Err; the close it returns, when not nil, is called once the
	// task is done, after its judges, even when the build failed, and
	// its error is joined there too.
	Build func(ctx context.Context, t Task, rec *session.Recorder) (cfg agentturn.Config, close func() error, err error)
	// SessionOptions, when set, supplies the options each run's
	// recorder is opened with: session.WithInstructionsParts for a
	// configuration that composes its instructions from parts, so the
	// session records which part changed rather than the whole prompt
	// on every config entry. The recorder is opened before the
	// configuration is built, so a function it is given here that needs
	// the configuration, such as one calling a kit's PartsFor, reaches
	// it through a closure over what Build builds a moment later.
	SessionOptions func(Task) []session.Option
	// Judges score each run once it has ended. Each score is appended
	// to the run's session as an outcome entry. A task whose run failed,
	// because the loop or Answer returned an error, is not judged unless
	// JudgeFailedRuns is set: an inference server that went away is not
	// the configuration's answer, and a score of 0 for it would count
	// in the report's mean and, for a product that takes scores as
	// rewards, as a reward.
	Judges []Judge
	// JudgeFailedRuns judges a task whose run failed as one that ended,
	// for an evaluation where a crash counts against the configuration:
	// Harbor's, which counts a trial with no reward as 0. Without it
	// the failed run is counted on [Summary.Unjudged], and a comparison
	// marks its pair on [Pair.Unjudged] rather than scoring it; a
	// comparison under Harbor's rule sets it on both runners.
	JudgeFailedRuns bool
	// Samples is how many times each task is run, each in a session of
	// its own: the group an RL producer scores together. Zero or one
	// means once. Each sample's Task, as every hook and judge sees it,
	// carries Meta["sample"], from "1", so a Header that fixes session
	// IDs can fix one per sample, and its session is named task.ID#n,
	// [SampleName], so a store listing tells the samples apart; a
	// runner that does not sample names the session by the task alone.
	//
	// A task whose Meta["sample"] is set while Samples is above one
	// names the one sample to run, for resuming a batch that was killed
	// or lost a sample to its inference server: the runner runs that
	// sample alone, numbered as it says on Result.Sample, the task
	// record and each outcome's details, and does not run its group
	// step, since one sample is not the group; [Runner.JudgeGroups]
	// finishes the group from the store. A value that is not a number
	// from 1 to Samples is refused.
	Samples int
	// GroupJudges score each task's samples together once all of them
	// are judged, on the goroutine of the last to finish. Each score is
	// appended to its sample's session as an outcome entry, like any
	// other. A group with a sample that was not judged, a run that
	// failed under the default of leaving it unjudged, is not group
	// judged, as an RL producer drops an incomplete group; every other
	// sample's result says so on Err, and [Runner.JudgeGroups] scores
	// the group once the sample has been run again.
	GroupJudges []GroupJudge
	// Answer answers the calls a run left pending when it ended
	// input_required, so an evaluation of a product whose policy asks
	// measures the whole run rather than the part before the first
	// ask. The run is resumed with what it returns and the resume is
	// recorded like any other turn; no answers, or a nil Answer, ends
	// the task there as before. agentpolicy.Engine.Answers satisfies
	// this signature as written.
	Answer func(context.Context, *agentturn.RunEnd) ([]agentturn.Answer, error)
	// MaxResumes bounds how many times one prompt may be resumed
	// through Answer, so an answer source that keeps a call pending
	// cannot loop. Zero means [DefaultMaxResumes].
	MaxResumes int
	// Parallel bounds how many tasks run at once; zero or one means one
	// at a time.
	Parallel int
	// Header, when set, supplies each run's session header: a harness
	// name, a working directory, a fixed ID for a test. Empty fields
	// are filled by the store. A header that names a Base and its
	// ParentSession forks that session, and the agent starts from the
	// context at the base, so a task runs as a continuation of a
	// recorded run and its session replays strictly like any other:
	// through a runner with the same Header under replay.AfterBase(),
	// since the agent starts after the base and never sends the
	// base's requests, or whole, from an agent that sends them.
	Header func(Task) agentsession.Header
	// Cost prices one model call, as export.Options.Cost does; see
	// price.Hook. When set, and every call of a run is priced, the
	// result carries the run's cost.
	Cost func(model string, usage openresponses.Usage) (float64, bool)
}

Runner sends each task of a suite through a configuration, records the run as a session and judges it.

func (*Runner) JudgeGroups added in v0.0.10

func (r *Runner) JudgeGroups(ctx context.Context, suite *Suite) (*Report, error)

JudgeGroups finishes the groups already in the Store: the group step for a task whose batch was killed before it, or whose group was not judged because a sample's run failed and that sample has since been run again through Meta["sample"]. For each task of the suite it takes, per sample from 1 to Samples, the latest session whose run was judged by the runner's Judges and ended, and judges the group with GroupJudges once every sample has one, recording each score on its member's session as Run does. The report's results are the members, in task and sample order, with Task, Sample, SessionID, Target, Reason, Usage and Cost read back from the session, Scores the judges' scores read back followed by the group's, and Err per member; Runs, Resumes and Ends are not read back. A group with a sample that has no such session is not judged, and says so: each member it does have is reported with the missing samples on Err, as Run reports a group it could not judge, and a task with no member at all is one result carrying the error, so a resume that still has samples to run reads which. A group already judged by every group judge is left out, since there is nothing left to do and a second call writes nothing; a group judge every member already holds an outcome from is not run again, so a judge added since scores the groups it has not, and one some member lacks, because that sample was run again after the group was judged, is run over the whole group and appended on every member, a second outcome on those that had one, as scores are appended and never rewritten. A runner that does not sample judges each task's latest session as a group of one.

A sample's sessions are found by name, task#n (see SampleName), through a store that lists names, and by their agenteval:task record otherwise; a session is judged when it holds an outcome from one of the runner's judges for the task and sample, or the runner has none and the session holds a run, and has ended when the entry those outcomes target is a run end whose reason is not error, the runs Run leaves unjudged under the default; with JudgeFailedRuns a judged run counts whatever its reason. Sessions are read without being held where the store is an agentsession.Reader. JudgeGroups fails only when it cannot start or the store fails.

func (*Runner) Run

func (r *Runner) Run(ctx context.Context, suite *Suite) (*Report, error)

Run sends every task of the suite through the configuration and returns the report. A task whose run or judging failed has its error on its result; Run itself fails only when it cannot start, or when ctx is done before every task has run, in which case the report holds the results so far.

type Score

type Score struct {
	// Judge is the judge's name; it becomes the outcome entry's label.
	Judge string `json:"judge"`
	// Value is the score on the judge's own scale: comparable across
	// runs of the same judge, not normalised. The deterministic judges
	// use 0 and 1; a Harbor verifier reports whatever its test script
	// wrote. A judge that normalises keeps the raw number in Details.
	Value float64 `json:"value"`
	// Pass is the judge's verdict.
	Pass bool `json:"pass"`
	// Reason says why, in a sentence.
	Reason string `json:"reason,omitempty"`
	// Details is the judge's own record, kept under the outcome's
	// details.judge.
	Details json.RawMessage `json:"details,omitempty"`
	// Session is the ID of the session the judge recorded of its own
	// run, when it is an agent that recorded one.
	Session string `json:"session,omitempty"`
}

Score is one judge's verdict on one run.

type Suite

type Suite struct {
	Name     string
	Tasks    []Task
	Manifest Manifest
}

Suite is a set of tasks loaded together.

func LoadSuite

func LoadSuite(fsys fs.FS, location string) (*Suite, error)

LoadSuite reads a suite from a directory of fsys: every file ending in .json except SuiteFile is one task, read in name order, and SuiteFile, when present, names the suite. Without it the suite is named after the directory. The manifest records every file read. The fs.FS is never written.

type Summary

type Summary struct {
	// Count is how many results the judge scored. A task whose run
	// failed is not scored unless [Runner.JudgeFailedRuns], so it counts
	// here only then; its result carries the error.
	Count int `json:"count"`
	// Unjudged is how many results the judge did not score, so Count
	// plus Unjudged is the number of results: a run that failed and was
	// left unjudged under the default, or a judge that failed on it;
	// the result's Err says which. A reader that counts a run with no
	// score as 0, as Harbor's mean does, has what it needs to: the mean
	// over every result is Mean times Count over Count plus Unjudged.
	Unjudged int `json:"unjudged"`
	// Mean is the mean of the judge's values over those results.
	Mean float64 `json:"mean"`
	// Passed is how many of them the judge passed, and PassRate is
	// Passed over Count.
	Passed   int     `json:"passed"`
	PassRate float64 `json:"pass_rate"`
}

Summary is one judge's numbers over the tasks it scored.

type Task

type Task struct {
	// ID names the task within its suite. LoadSuite defaults it to the
	// file name without its extension.
	ID string `json:"id"`
	// Instruction is the common case: one user message.
	Instruction string `json:"instruction,omitempty"`
	// Prompts is the multi-turn case: what the user says, in order,
	// one run of the loop each. It wins over Instruction when set.
	Prompts openresponses.Items `json:"prompts,omitempty"`
	// Setup is product-defined, such as a repository revision to check
	// out before the run. The runner records it and honours none of it.
	Setup map[string]string `json:"setup,omitempty"`
	// Expect is judge-defined: what the deterministic judges compare
	// the run against.
	Expect json.RawMessage `json:"expect,omitempty"`
	// Meta is free-form metadata, recorded with the run.
	Meta map[string]string `json:"meta,omitempty"`
}

Task is one thing to send through a configuration and judge.

func FromSession

func FromSession(s *agentsession.Session) (Task, error)

FromSession turns a recorded run into a task: the user messages on the path to the session's current leaf, in order, become Prompts, so the same conversation can be sent through another configuration. The task is named after the session and Meta records the session and the leaf. Model output, tool outputs, summaries and app-only entries are not prompts and are left out.

func FromSessionAt

func FromSessionAt(s *agentsession.Session, leaf string) (Task, error)

FromSessionAt is FromSession over the path to leaf.

func (Task) Inputs

func (t Task) Inputs() openresponses.Items

Inputs returns what the runner sends: Prompts when set, else one user message carrying Instruction, else nothing.

type TaskRecord

type TaskRecord struct {
	Suite    string            `json:"suite"`
	Location string            `json:"location,omitempty"`
	Task     string            `json:"task"`
	Setup    map[string]string `json:"setup,omitempty"`
	Meta     map[string]string `json:"meta,omitempty"`
	// Sample is which of Samples runs of the task this is, from 1,
	// when the runner sampled it, and the session is then named
	// task#sample ([SampleName]). GroupJudges names the group judges
	// the runner was to score the group with, so a reader can tell a
	// group that was never scored, because the batch ended before its
	// group step, by a session that holds no outcome from one of them;
	// [Runner.JudgeGroups] scores such a group from the store.
	Sample      int      `json:"sample,omitempty"`
	Samples     int      `json:"samples,omitempty"`
	GroupJudges []string `json:"group_judges,omitempty"`
}

TaskRecord is the data of a TaskNS custom entry: which suite and task a session ran, and the task's setup and metadata.

Directories

Path Synopsis
Package judge holds the judges: three deterministic ones that read a trajectory against the task's Expect, and Rubric, an agent that reads the trajectory and answers in a fixed schema.
Package judge holds the judges: three deterministic ones that read a trajectory against the task's Expect, and Rubric, an agent that reads the trajectory and answers in a fixed schema.
Package price costs a run.
Package price costs a run.
Package replay serves a recorded session as a model and as tools, so hooks, transforms and fronts are tested against real traffic with nothing behind them.
Package replay serves a recorded session as a model and as tools, so hooks, transforms and fronts are tested against real traffic with nothing behind them.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL