Documentation
¶
Overview ¶
Package eval defines subject-agnostic execution and quality evaluation. Evaluator is the atomic scoring contract: a non-nil error makes its Report unusable. Metric identifies the calculation and measurement semantics; Decision independently identifies the policy behind a categorical verdict. Score, Measurement, Decision, and qualitative feedback are optional outcomes. Changing a threshold changes Decision identity without changing the measured quantity. ProjectionEvaluator adapts aggregate subjects to narrow evaluators.
Independent assessments ¶
Assessment declares an ID before its Evaluator runs. Suite.Run preserves one AssessmentResult for every declaration, including failures, cancellation, and work not started. Under ErrorCollect, a child error does not cancel other assessments. Under ErrorFailFast, pending work stops and completed results remain available. Suite is a collection lifecycle rather than an Evaluator; callers inspect SuiteResult.Err and individual statuses even when Run returns nil. The operation error reports caller cancellation or a fail-fast stop.
CompositeEvaluator is an atomic weighted score calculation. It accepts score-only components by default. An explicit PassPolicy additionally requires decided components. Supporting Report.Details preserve evidence but do not become extra statistical observations. To summarize a child independently, declare it as a separate Assessment with its own ID.
Fixed inputs and generated executions ¶
Dataset owns fixed Case IDs and context identity. Its FixtureID changes when inputs, references, evaluation context, or selected cases change. Generated output is not part of that identity. Subject may be any borrowed immutable aggregate, including an offline input/output/reference sample.
Target.Run performs generation and returns an Execution receipt whose status is independent of quality. A nil Output means no candidate was collected; an existing empty candidate remains present. Trial.Run composes one fixed Case, one Target invocation, and one Suite over TrialSample. It preserves a valid Execution even when assessment fails. A non-nil Target error means no valid execution receipt, so assessments remain not evaluated. Reassessment of a retained execution calls Suite.Run directly without invoking the Target again. The Host owns environment sessions, persistence, repeats, and safe resumption.
Experiments and comparison ¶
Experiment evaluates an offline Dataset through Suite.Run with bounded case concurrency. The default case limit is four and the default Suite limit is one; explicitly increasing both limits multiplies direct assessment calls. NewExperimentReport validates and snapshots already collected CaseResult facts without executing work. ExperimentSummary counts complete cases separately from partial cases, preserves assessment coverage even when no Metric exists, and aggregates only completed, explicitly declared assessments.
ExperimentReport.Compare requires the same FixtureID and set of Case IDs. Numeric differences pair case, AssessmentID, and full Metric identity rather than subtracting means over different successful subsets. Matched and missing observation counts remain explicit. Decision differences additionally require identical policy identity; no significance or repeat-at-k estimate is implied. JSON decoding rejects unknown Report, Decision, and Metric members, including nested details, while metadata and identity parameters retain open JSON values. Owned identities, feedback, policies, and execution reasons must be valid UTF-8 at admission; borrowed generic subjects and outputs keep their own contracts.
Domain vocabularies live outside the kernel: judge supplies generic model-backed evaluation, text owns generated-text metrics, ranking owns provider-neutral ranking metrics, and trajectory owns deterministic Agent execution evaluation. The swebench package exchanges typed prediction and report files with the official harness; it does not reproduce its grader. New domains implement Evaluator directly and do not depend on another domain's vocabulary.
Index ¶
- Constants
- Variables
- type Assessment
- type AssessmentID
- type AssessmentResult
- type AssessmentStatus
- type AssessmentSummary
- type Case
- type CaseID
- type CaseResult
- type Comparison
- type Component
- type CompositeEvaluator
- type CompositeEvaluatorConfig
- type Dataset
- type Decision
- type DecisionDelta
- type Direction
- type Distribution
- type DistributionDelta
- type ErrorPolicy
- type Evaluator
- type EvaluatorFunc
- type Execution
- type ExecutionStatus
- type Experiment
- type ExperimentConfig
- type ExperimentReport
- type ExperimentSummary
- type Metric
- func (m Metric) Direction() Direction
- func (m Metric) MarshalJSON() ([]byte, error)
- func (m Metric) Name() MetricName
- func (m Metric) Namespace() string
- func (m Metric) Parameters() metadata.Map
- func (m Metric) String() string
- func (m Metric) Unit() string
- func (m *Metric) UnmarshalJSON(data []byte) error
- func (m Metric) Validate() error
- type MetricComparison
- type MetricConfig
- type MetricName
- type MetricSummary
- type PassPolicy
- type Projection
- type ProjectionEvaluator
- type Report
- type Score
- type Suite
- type SuiteConfig
- type SuiteResult
- type Target
- type TargetFunc
- type Trial
- type TrialConfig
- type TrialResult
- type TrialSample
- type Verdict
Examples ¶
Constants ¶
const DefaultMaxConcurrency = 4
const DefaultSuiteConcurrency = 1
DefaultSuiteConcurrency keeps nested experiment fan-out sequential unless the Host explicitly budgets concurrent assessments within each case.
const MaxReportDepth = 64
MaxReportDepth bounds recursive detail trees at every public trust boundary.
Variables ¶
var ( ErrInvalidEvaluatorConfig = errors.New("eval: evaluator configuration is invalid") ErrInvalidAssessment = errors.New("eval: invalid assessment") ErrInvalidMetric = errors.New("eval: invalid metric") ErrInvalidScore = errors.New("eval: invalid score") ErrInvalidReport = errors.New("eval: invalid report") ErrInvalidCase = errors.New("eval: invalid case") ErrInvalidDataset = errors.New("eval: invalid dataset") ErrInvalidExperiment = errors.New("eval: invalid experiment") ErrInvalidComparison = errors.New("eval: invalid comparison") ErrNotEvaluated = errors.New("eval: assessment was not evaluated") )
var ErrInvalidExecution = errors.New("eval: invalid execution")
var ErrInvalidTrial = errors.New("eval: invalid trial")
Functions ¶
This section is empty.
Types ¶
type Assessment ¶ added in v0.38.0
type Assessment[T any] struct { ID AssessmentID Evaluator Evaluator[T] }
Assessment binds a declared identity to the one atomic evaluation call.
func (Assessment[T]) Validate ¶ added in v0.38.0
func (a Assessment[T]) Validate() error
type AssessmentID ¶ added in v0.38.0
type AssessmentID string
AssessmentID identifies an evaluation independently of whether it succeeds. The Host keeps this identity stable when comparing the same assessment.
func (AssessmentID) Validate ¶ added in v0.38.0
func (a AssessmentID) Validate() error
type AssessmentResult ¶ added in v0.38.0
type AssessmentResult struct {
ID AssessmentID
Report *Report
Err error
}
AssessmentResult has a Report only after successful evaluation. A failed evaluator's returned Report is discarded under the Evaluator contract. ErrNotEvaluated marks work that never started, preserving any cause through errors.Is. Status is derived from these facts rather than stored separately.
func (AssessmentResult) Status ¶ added in v0.38.0
func (a AssessmentResult) Status() AssessmentStatus
Status is zero when result facts are missing or competing.
func (AssessmentResult) Validate ¶ added in v0.38.0
func (a AssessmentResult) Validate() error
type AssessmentStatus ¶ added in v0.38.0
type AssessmentStatus string
AssessmentStatus describes execution, independently of a quality verdict.
const ( AssessmentCompleted AssessmentStatus = "completed" AssessmentFailed AssessmentStatus = "failed" AssessmentCanceled AssessmentStatus = "canceled" AssessmentNotEvaluated AssessmentStatus = "not_evaluated" )
type AssessmentSummary ¶ added in v0.38.0
type AssessmentSummary struct {
ID AssessmentID
Completed int
Failed int
Canceled int
NotEvaluated int
}
AssessmentSummary preserves declared membership even if no evaluation produced a metric. Its four execution counts partition its case count.
type Case ¶
Case gives a stable identity to one evaluation subject. Subject is borrowed read-only: copying a Case does not copy objects referenced by T.
type CaseResult ¶
type CaseResult struct {
ID CaseID
Metadata metadata.Map
Result SuiteResult
}
CaseResult preserves all assessment outcomes for one fixed case. Result is the sole owner of assessment facts; case completion and verdict are derived.
func (CaseResult) Validate ¶ added in v0.38.0
func (c CaseResult) Validate() error
type Comparison ¶
type Comparison struct {
Baseline ExperimentSummary
Candidate ExperimentSummary
EvaluatedDelta int
ErrorDelta int
Metrics []MetricComparison
}
Comparison preserves execution coverage and paired quality differences. It makes no claim about statistical significance.
type Component ¶
Component assigns score weight and pass criticality to one evaluator. A zero Weight selects 1. Required components must pass independently of the aggregate pass policy and still contribute to the score. To keep a gate out of the quality score, declare it as another assessment in a Suite. Required is only valid when CompositeEvaluatorConfig explicitly selects a pass policy.
type CompositeEvaluator ¶
type CompositeEvaluator[T any] struct { // contains filtered or unexported fields }
CompositeEvaluator combines scored child reports. A configured pass policy additionally requires decided children; it does not change score identity.
func NewCompositeEvaluator ¶
func NewCompositeEvaluator[T any](config CompositeEvaluatorConfig[T]) (*CompositeEvaluator[T], error)
NewCompositeEvaluator copies the component slice; evaluators remain shared.
type CompositeEvaluatorConfig ¶ added in v0.41.0
type CompositeEvaluatorConfig[T any] struct { Components []Component[T] PassPolicy PassPolicy MinimumPassed int MaxConcurrency int }
CompositeEvaluatorConfig defines score aggregation and an optional categorical policy. A zero PassPolicy produces only a score and accepts score-only components. A zero MaxConcurrency selects DefaultMaxConcurrency.
type Dataset ¶
type Dataset[T any] struct { // contains filtered or unexported fields }
Dataset owns an ordered snapshot of case identities and metadata. Its Host-assigned FixtureID must change with inputs, references, or selected cases, but not with observed outputs. Subjects stay borrowed read-only while the Dataset is in use.
func NewDataset ¶
NewDataset snapshots cases and metadata, copying Subject by assignment, and rejects duplicate IDs that would make result correlation ambiguous.
type Decision ¶ added in v0.38.0
type Decision struct {
Policy string `json:"policy"`
Parameters metadata.Map `json:"parameters,omitzero"`
Verdict Verdict `json:"verdict"`
}
Decision records a categorical result and the policy that produced it. Policy and Parameters identify the rule independently of its observed Verdict. A Report without a categorical judgment has no Decision.
func (Decision) MarshalJSON ¶ added in v0.38.0
func (*Decision) UnmarshalJSON ¶ added in v0.38.0
type DecisionDelta ¶ added in v0.38.0
type DecisionDelta struct {
Matched int
BaselineOnly int
CandidateOnly int
Incompatible int
PassedDelta int
FailedDelta int
}
DecisionDelta compares only matched decisions produced by the same policy. Incompatible counts pairs whose policies differ, including threshold changes.
type Direction ¶
type Direction string
Direction describes how a raw measurement relates to quality. It is kept separate from Score, whose direction is always higher-is-better.
type Distribution ¶
type Distribution struct {
Count int
Mean float64
Minimum float64
P10 float64
P50 float64
P90 float64
Maximum float64
}
Distribution summarizes one homogeneous numeric signal. Count distinguishes an absent distribution from a real distribution whose values are all zero.
type DistributionDelta ¶
type DistributionDelta struct {
Present bool
Mean float64
Matched int
BaselineOnly int
CandidateOnly int
}
DistributionDelta summarizes candidate-minus-baseline differences for matched case observations. Unmatched observations are counted, never imputed.
type ErrorPolicy ¶
type ErrorPolicy string
ErrorPolicy controls whether independent failures are collected or stop new scheduling at the Suite or Experiment boundary that owns the policy.
const ( ErrorCollect ErrorPolicy = "collect" ErrorFailFast ErrorPolicy = "fail_fast" )
type Evaluator ¶
type Evaluator[T any] interface { // Evaluate inspects one subject without mutating it and returns a valid, // owned report for the evaluator's metric. Implementations must honor ctx; // a non-nil error means the report must not be consumed. Evaluate(ctx context.Context, subject T) (Report, error) }
type EvaluatorFunc ¶
type Execution ¶ added in v0.38.0
type Execution[O any] struct { Status ExecutionStatus `json:"status"` Output *O `json:"output,omitzero"` Reason string `json:"reason,omitzero"` Metadata metadata.Map `json:"metadata,omitzero"` }
Execution is the target's receipt for one invocation. A nil Output means no candidate; a present empty candidate is not nil. Objects the output references stay borrowed read-only, like Case.Subject.
func (Execution[O]) MarshalJSON ¶ added in v0.38.0
func (*Execution[O]) UnmarshalJSON ¶ added in v0.38.0
type ExecutionStatus ¶ added in v0.38.0
type ExecutionStatus string
ExecutionStatus describes how a target stopped, independently of the quality of its output. A failed or exhausted execution may still produce a candidate that can be graded. Completed never implies a passing assessment.
const ( ExecutionCompleted ExecutionStatus = "completed" ExecutionBudgetExhausted ExecutionStatus = "budget_exhausted" ExecutionCanceled ExecutionStatus = "canceled" ExecutionFailed ExecutionStatus = "failed" )
type Experiment ¶
type Experiment[T any] struct { // contains filtered or unexported fields }
Experiment is an immutable plan for evaluating one Dataset. It owns bounded scheduling and error semantics, but no persistence, artifacts, or product identity.
func NewExperiment ¶
func NewExperiment[T any](config ExperimentConfig[T]) (Experiment[T], error)
func (Experiment[T]) Run ¶
func (e Experiment[T]) Run(ctx context.Context) (ExperimentReport, error)
Example ¶
package main
import (
"context"
"fmt"
"github.com/Tangerg/scope/eval"
)
func main() {
metric, err := eval.NewMetric(eval.MetricConfig{
Namespace: "example",
Name: "non_empty",
})
if err != nil {
panic(err)
}
evaluator := eval.EvaluatorFunc[string](func(_ context.Context, subject string) (eval.Report, error) {
verdict := eval.VerdictFail
if subject != "" {
verdict = eval.VerdictPass
}
return eval.Report{Metric: metric, Decision: &eval.Decision{Policy: "non_empty", Verdict: verdict}}, nil
})
suite, err := eval.NewSuite(eval.SuiteConfig[string]{Assessments: []eval.Assessment[string]{{ID: "non_empty", Evaluator: evaluator}}})
if err != nil {
panic(err)
}
dataset, err := eval.NewDataset("test-fixture",
eval.Case[string]{ID: "first", Subject: "answer"},
)
if err != nil {
panic(err)
}
experiment, err := eval.NewExperiment(eval.ExperimentConfig[string]{
Dataset: dataset, Suite: suite,
})
if err != nil {
panic(err)
}
report, err := experiment.Run(context.Background())
if err != nil {
panic(err)
}
summary := report.Summary()
fmt.Println(summary.Total, summary.Passed)
}
Output: 1 1
type ExperimentConfig ¶
type ExperimentConfig[T any] struct { Dataset Dataset[T] Suite *Suite[T] MaxConcurrency int ErrorPolicy ErrorPolicy }
ExperimentConfig binds fixed case inputs and context to a Suite. MaxConcurrency limits concurrent cases; zero selects DefaultMaxConcurrency. A Suite defaults to one assessment at a time. Explicitly increasing both limits multiplies the maximum number of direct assessment calls.
type ExperimentReport ¶
type ExperimentReport struct {
// contains filtered or unexported fields
}
ExperimentReport owns ordered case facts and the summary derived from them. NewExperimentReport also imports independently generated observations.
func NewExperimentReport ¶ added in v0.38.0
func NewExperimentReport(fixtureID string, results []CaseResult) (ExperimentReport, error)
func (ExperimentReport) Cases ¶
func (e ExperimentReport) Cases() []CaseResult
func (ExperimentReport) Compare ¶
func (e ExperimentReport) Compare(candidate ExperimentReport) (Comparison, error)
Compare requires the same fixture and set of Case IDs, in any order. A changed threshold keeps scores comparable but makes decisions incompatible.
func (ExperimentReport) FixtureID ¶ added in v0.21.0
func (e ExperimentReport) FixtureID() string
func (ExperimentReport) Summary ¶
func (e ExperimentReport) Summary() ExperimentSummary
type ExperimentSummary ¶
type ExperimentSummary struct {
Total int
Evaluated int
Passed int
Failed int
Unjudged int
Errors int
Partial int
Assessments []AssessmentSummary
Metrics []MetricSummary
}
ExperimentSummary counts complete and incomplete cases separately. Errors counts incomplete cases, not quality failures; Partial counts the incomplete cases whose completed assessments still contribute to Metrics.
type Metric ¶
type Metric struct {
// contains filtered or unexported fields
}
Metric is an immutable evaluation identity that can be copied by assignment. Parameters holds structured identity for the calculation. A Report's Decision owns any separate acceptance policy, including score thresholds. Unit and Direction describe optional raw measurements; normalized scores are always unitless and higher-is-better.
func NewMetric ¶
func NewMetric(config MetricConfig) (Metric, error)
NewMetric snapshots parameters and normalizes object order, whitespace, and string escaping. Number spellings retain their exact precision. Later caller mutation cannot change report comparability.
func (Metric) MarshalJSON ¶ added in v0.13.0
func (Metric) Name ¶
func (m Metric) Name() MetricName
func (Metric) Parameters ¶
func (*Metric) UnmarshalJSON ¶ added in v0.13.0
type MetricComparison ¶
type MetricComparison struct {
AssessmentID AssessmentID
Metric Metric
Baseline *MetricSummary
Candidate *MetricSummary
EvaluatedDelta int
ScoreDelta DistributionDelta
MeasurementDelta DistributionDelta
DecisionDelta DecisionDelta
}
MetricComparison relates the same named assessment and calculation. Numeric deltas use paired cases, while the two summaries retain all observations.
type MetricConfig ¶
type MetricSummary ¶
type MetricSummary struct {
AssessmentID AssessmentID
Metric Metric
Evaluated int
Passed int
Failed int
Unjudged int
Scores Distribution
Measurements Distribution
}
MetricSummary contains one top-level observation per successful case and assessment. Supporting Report.Details never become additional samples.
type PassPolicy ¶
type PassPolicy string
PassPolicy controls categorical aggregation independently from score weights.
const ( PassNone PassPolicy = "" PassAll PassPolicy = "all" PassAny PassPolicy = "any" PassAtLeast PassPolicy = "at_least" )
type Projection ¶
Projection narrows an aggregate case to the subject owned by one evaluator.
type ProjectionEvaluator ¶
type ProjectionEvaluator[T, Subject any] struct { // contains filtered or unexported fields }
func NewProjectionEvaluator ¶
func NewProjectionEvaluator[T, Subject any]( evaluator Evaluator[Subject], projection Projection[T, Subject], ) (*ProjectionEvaluator[T, Subject], error)
type Report ¶
type Report struct {
Metric Metric `json:"metric"`
Decision *Decision `json:"decision,omitzero"`
Score *Score `json:"score,omitzero"`
Measurement *float64 `json:"measurement,omitzero"`
Feedback string `json:"feedback,omitzero"`
Metadata metadata.Map `json:"metadata,omitzero"`
Details []Report `json:"details,omitzero"`
}
Report is one evaluation result. Decision, Score, and Measurement are independent and optional, so no evaluation has to invent a threshold or a score. Details are supporting evidence, never additional summary observations.
func (Report) MarshalJSON ¶
func (*Report) UnmarshalJSON ¶
type Score ¶
type Score float64
Score is a normalized quality score in the closed interval [0, 1], where a higher value is always better.
func (Score) Decide ¶ added in v0.38.0
Decide applies the higher-is-better threshold without changing the identity of the calculation that produced the Score.
Example ¶
package main
import (
"fmt"
"github.com/Tangerg/scope/eval"
)
func main() {
score, err := eval.NewScore(0.82)
if err != nil {
panic(err)
}
threshold, err := eval.NewScore(0.8)
if err != nil {
panic(err)
}
decision, err := score.Decide(threshold)
if err != nil {
panic(err)
}
fmt.Println(decision.Verdict, score.Float64())
}
Output: pass 0.82
type Suite ¶ added in v0.38.0
type Suite[T any] struct { // contains filtered or unexported fields }
Suite executes independent assessments and preserves their individual execution outcomes. It is a collection operation, not an atomic Evaluator.
func (*Suite[T]) Run ¶ added in v0.38.0
func (s *Suite[T]) Run(ctx context.Context, subject T) (SuiteResult, error)
Run always returns the results accumulated before failure or cancellation. Under ErrorCollect, individual errors live in Results and do not become the operation error. The caller's cancellation and fail-fast error remain visible.
Example ¶
package main
import (
"context"
"fmt"
"github.com/Tangerg/scope/eval"
)
func main() {
qualityMetric, err := eval.NewMetric(eval.MetricConfig{Namespace: "example", Name: "quality"})
if err != nil {
panic(err)
}
safetyMetric, err := eval.NewMetric(eval.MetricConfig{Namespace: "example", Name: "safety"})
if err != nil {
panic(err)
}
quality := eval.EvaluatorFunc[string](func(context.Context, string) (eval.Report, error) {
score, scoreErr := eval.NewScore(0.9)
return eval.Report{Metric: qualityMetric, Score: &score}, scoreErr
})
safety := eval.EvaluatorFunc[string](func(context.Context, string) (eval.Report, error) {
return eval.Report{Metric: safetyMetric, Decision: &eval.Decision{Policy: "safety", Verdict: eval.VerdictFail}, Feedback: "Answer requires review."}, nil
})
scored, err := eval.NewCompositeEvaluator(eval.CompositeEvaluatorConfig[string]{
Components: []eval.Component[string]{{Evaluator: quality}},
})
if err != nil {
panic(err)
}
// A gate determines acceptance without contributing a score. The Suite
// preserves the quality Composite's score under its own metric identity.
suite, err := eval.NewSuite(eval.SuiteConfig[string]{
Assessments: []eval.Assessment[string]{{ID: "quality", Evaluator: scored}, {ID: "safety", Evaluator: safety}},
})
if err != nil {
panic(err)
}
report, err := suite.Run(context.Background(), "answer")
if err != nil {
panic(err)
}
fmt.Println("acceptance:", report.Verdict())
fmt.Println("quality:", report.Results[0].Report.Score.Float64())
fmt.Println("gate has score:", report.Results[1].Report.Score != nil)
}
Output: acceptance: fail quality: 0.9 gate has score: false
type SuiteConfig ¶
type SuiteConfig[T any] struct { Assessments []Assessment[T] MaxConcurrency int ErrorPolicy ErrorPolicy }
SuiteConfig declares independently identifiable assessments. MaxConcurrency limits direct assessment calls within one Run; zero selects one. ErrorCollect preserves failures without canceling independent assessments.
type SuiteResult ¶ added in v0.38.0
type SuiteResult struct {
Results []AssessmentResult
}
SuiteResult is an ordered set of assessment facts. Verdict and completion are derived, so there is no independently mutable aggregate outcome.
func (SuiteResult) Clone ¶ added in v0.38.0
func (s SuiteResult) Clone() (SuiteResult, error)
func (SuiteResult) Complete ¶ added in v0.38.0
func (s SuiteResult) Complete() bool
func (SuiteResult) Err ¶ added in v0.38.0
func (s SuiteResult) Err() error
func (SuiteResult) Validate ¶ added in v0.38.0
func (s SuiteResult) Validate() error
func (SuiteResult) Verdict ¶ added in v0.38.0
func (s SuiteResult) Verdict() Verdict
Verdict is unspecified until every assessment completes. Successful reports remain individually available even when the suite has no complete verdict.
type Target ¶ added in v0.38.0
Target executes the public input and freezes its candidate before grading; references and hidden tests stay with the assessments. A non-nil error means no usable Execution; known failure, cancellation, or budget exhaustion is an Execution with a stop reason. Implementations honor ctx, finish collecting evidence before returning, and never expose credentials, cleanup handles, or mutable workspaces as candidates.
type TargetFunc ¶ added in v0.38.0
type Trial ¶ added in v0.38.0
type Trial[I, O any] struct { // contains filtered or unexported fields }
Trial composes one execution with independent assessments. Dataset scheduling, retries, persistence, and sandboxes belong to the Host.
Example ¶
package main
import (
"context"
"errors"
"fmt"
"github.com/Tangerg/scope/eval"
)
func main() {
suite, err := eval.NewSuite(eval.SuiteConfig[eval.TrialSample[string, string]]{Assessments: []eval.Assessment[eval.TrialSample[string, string]]{{
ID: "answer", Evaluator: eval.EvaluatorFunc[eval.TrialSample[string, string]](func(_ context.Context, sample eval.TrialSample[string, string]) (eval.Report, error) {
if sample.Execution.Output == nil {
return eval.Report{}, errors.New("answer unavailable")
}
score := eval.Score(0)
if *sample.Execution.Output == "42" {
score = 1
}
metric, err := eval.NewMetric(eval.MetricConfig{Name: "exact_answer"})
if err != nil {
return eval.Report{}, err
}
decision, err := score.Decide(1)
return eval.Report{Metric: metric, Score: &score, Decision: &decision}, err
}),
}}})
if err != nil {
panic(err)
}
trial, err := eval.NewTrial(eval.TrialConfig[string, string]{Suite: suite, Target: eval.TargetFunc[string, string](func(context.Context, string) (eval.Execution[string], error) {
answer := "42"
return eval.Execution[string]{Status: eval.ExecutionCompleted, Output: &answer}, nil
})})
if err != nil {
panic(err)
}
result, err := trial.Run(context.Background(), eval.Case[string]{ID: "question-1", Subject: "What is six times seven?"})
if err != nil {
panic(err)
}
fmt.Println(result.Execution.Status, result.Case.Result.Results[0].Report.Verdict())
}
Output: completed pass
func NewTrial ¶ added in v0.38.0
func NewTrial[I, O any](config TrialConfig[I, O]) (*Trial[I, O], error)
func (*Trial[I, O]) Run ¶ added in v0.38.0
Run returns every accumulated stage result: a valid execution survives failed or canceled grading, and an ExecutionError leaves assessments not evaluated. An execution without output is still graded; assessments decide which evidence they require. The result holds runtime errors and is not a storage schema.
type TrialConfig ¶ added in v0.38.0
type TrialConfig[I, O any] struct { Target Target[I, O] Suite *Suite[TrialSample[I, O]] }
TrialConfig requires both fields; share them only when they support the Host's concurrent calls, because Trial adds no scheduling of its own.
type TrialResult ¶ added in v0.38.0
type TrialResult[O any] struct { Case CaseResult Execution *Execution[O] ExecutionError error }
TrialResult preserves execution and assessment outcomes independently. Case can be included in NewExperimentReport with other cases from the same fixture. ExecutionError is an invocation/protocol error, never a quality judgment.
type TrialSample ¶ added in v0.38.0
TrialSample keeps Execution apart from Case so generated outputs never change fixture identity and can be regraded without invoking the target again.
Source Files
¶
Directories
¶
| Path | Synopsis |
|---|---|
|
Package judge evaluates arbitrary subjects with a structured-output chat model.
|
Package judge evaluates arbitrary subjects with a structured-output chat model. |
|
Package ranking evaluates ranked outputs against graded relevance judgments.
|
Package ranking evaluates ranked outputs against graded relevance judgments. |
|
Package swebench exchanges predictions and reports with the official SWE-bench harness.
|
Package swebench exchanges predictions and reports with the official SWE-bench harness. |
|
Package text evaluates generated text without imposing one shared sample on metrics with different semantic inputs.
|
Package text evaluates generated text without imposing one shared sample on metrics with different semantic inputs. |
|
trajectory
module
|