eval

package module
v0.44.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 7, 2026 License: Apache-2.0 Imports: 14 Imported by: 0

Documentation

Overview

Package eval defines subject-agnostic execution and quality evaluation. Evaluator is the atomic scoring contract: a non-nil error makes its Report unusable. Metric identifies the calculation and measurement semantics; Decision independently identifies the policy behind a categorical verdict. Score, Measurement, Decision, and qualitative feedback are optional outcomes. Changing a threshold changes Decision identity without changing the measured quantity. ProjectionEvaluator adapts aggregate subjects to narrow evaluators.

Independent assessments

Assessment declares an ID before its Evaluator runs. Suite.Run preserves one AssessmentResult for every declaration, including failures, cancellation, and work not started. Under ErrorCollect, a child error does not cancel other assessments. Under ErrorFailFast, pending work stops and completed results remain available. Suite is a collection lifecycle rather than an Evaluator; callers inspect SuiteResult.Err and individual statuses even when Run returns nil. The operation error reports caller cancellation or a fail-fast stop.

CompositeEvaluator is an atomic weighted score calculation. It accepts score-only components by default. An explicit PassPolicy additionally requires decided components. Supporting Report.Details preserve evidence but do not become extra statistical observations. To summarize a child independently, declare it as a separate Assessment with its own ID.

Fixed inputs and generated executions

Dataset owns fixed Case IDs and context identity. Its FixtureID changes when inputs, references, evaluation context, or selected cases change. Generated output is not part of that identity. Subject may be any borrowed immutable aggregate, including an offline input/output/reference sample.

Target.Run performs generation and returns an Execution receipt whose status is independent of quality. A nil Output means no candidate was collected; an existing empty candidate remains present. Trial.Run composes one fixed Case, one Target invocation, and one Suite over TrialSample. It preserves a valid Execution even when assessment fails. A non-nil Target error means no valid execution receipt, so assessments remain not evaluated. Reassessment of a retained execution calls Suite.Run directly without invoking the Target again. The Host owns environment sessions, persistence, repeats, and safe resumption.

Experiments and comparison

Experiment evaluates an offline Dataset through Suite.Run with bounded case concurrency. The default case limit is four and the default Suite limit is one; explicitly increasing both limits multiplies direct assessment calls. NewExperimentReport validates and snapshots already collected CaseResult facts without executing work. ExperimentSummary counts complete cases separately from partial cases, preserves assessment coverage even when no Metric exists, and aggregates only completed, explicitly declared assessments.

ExperimentReport.Compare requires the same FixtureID and set of Case IDs. Numeric differences pair case, AssessmentID, and full Metric identity rather than subtracting means over different successful subsets. Matched and missing observation counts remain explicit. Decision differences additionally require identical policy identity; no significance or repeat-at-k estimate is implied. JSON decoding rejects unknown Report, Decision, and Metric members, including nested details, while metadata and identity parameters retain open JSON values. Owned identities, feedback, policies, and execution reasons must be valid UTF-8 at admission; borrowed generic subjects and outputs keep their own contracts.

Domain vocabularies live outside the kernel: judge supplies generic model-backed evaluation, text owns generated-text metrics, ranking owns provider-neutral ranking metrics, and trajectory owns deterministic Agent execution evaluation. The swebench package exchanges typed prediction and report files with the official harness; it does not reproduce its grader. New domains implement Evaluator directly and do not depend on another domain's vocabulary.

Index

Examples

Constants

View Source
const DefaultMaxConcurrency = 4
View Source
const DefaultSuiteConcurrency = 1

DefaultSuiteConcurrency keeps nested experiment fan-out sequential unless the Host explicitly budgets concurrent assessments within each case.

View Source
const MaxReportDepth = 64

MaxReportDepth bounds recursive detail trees at every public trust boundary.

Variables

View Source
var (
	ErrInvalidEvaluatorConfig = errors.New("eval: evaluator configuration is invalid")
	ErrInvalidAssessment      = errors.New("eval: invalid assessment")
	ErrInvalidMetric          = errors.New("eval: invalid metric")
	ErrInvalidScore           = errors.New("eval: invalid score")
	ErrInvalidReport          = errors.New("eval: invalid report")
	ErrInvalidCase            = errors.New("eval: invalid case")
	ErrInvalidDataset         = errors.New("eval: invalid dataset")
	ErrInvalidExperiment      = errors.New("eval: invalid experiment")
	ErrInvalidComparison      = errors.New("eval: invalid comparison")
	ErrNotEvaluated           = errors.New("eval: assessment was not evaluated")
)
View Source
var ErrInvalidExecution = errors.New("eval: invalid execution")
View Source
var ErrInvalidTrial = errors.New("eval: invalid trial")

Functions

This section is empty.

Types

type Assessment added in v0.38.0

type Assessment[T any] struct {
	ID        AssessmentID
	Evaluator Evaluator[T]
}

Assessment binds a declared identity to the one atomic evaluation call.

func (Assessment[T]) Validate added in v0.38.0

func (a Assessment[T]) Validate() error

type AssessmentID added in v0.38.0

type AssessmentID string

AssessmentID identifies an evaluation independently of whether it succeeds. The Host keeps this identity stable when comparing the same assessment.

func (AssessmentID) Validate added in v0.38.0

func (a AssessmentID) Validate() error

type AssessmentResult added in v0.38.0

type AssessmentResult struct {
	ID     AssessmentID
	Report *Report
	Err    error
}

AssessmentResult has a Report only after successful evaluation. A failed evaluator's returned Report is discarded under the Evaluator contract. ErrNotEvaluated marks work that never started, preserving any cause through errors.Is. Status is derived from these facts rather than stored separately.

func (AssessmentResult) Status added in v0.38.0

Status is zero when result facts are missing or competing.

func (AssessmentResult) Validate added in v0.38.0

func (a AssessmentResult) Validate() error

type AssessmentStatus added in v0.38.0

type AssessmentStatus string

AssessmentStatus describes execution, independently of a quality verdict.

const (
	AssessmentCompleted    AssessmentStatus = "completed"
	AssessmentFailed       AssessmentStatus = "failed"
	AssessmentCanceled     AssessmentStatus = "canceled"
	AssessmentNotEvaluated AssessmentStatus = "not_evaluated"
)

type AssessmentSummary added in v0.38.0

type AssessmentSummary struct {
	ID           AssessmentID
	Completed    int
	Failed       int
	Canceled     int
	NotEvaluated int
}

AssessmentSummary preserves declared membership even if no evaluation produced a metric. Its four execution counts partition its case count.

type Case

type Case[T any] struct {
	ID       CaseID
	Subject  T
	Metadata metadata.Map
}

Case gives a stable identity to one evaluation subject. Subject is borrowed read-only: copying a Case does not copy objects referenced by T.

func (Case[T]) Validate

func (c Case[T]) Validate() error

type CaseID

type CaseID string

CaseID is a stable identity within one Dataset.

func (CaseID) String

func (c CaseID) String() string

func (CaseID) Validate

func (c CaseID) Validate() error

type CaseResult

type CaseResult struct {
	ID       CaseID
	Metadata metadata.Map
	Result   SuiteResult
}

CaseResult preserves all assessment outcomes for one fixed case. Result is the sole owner of assessment facts; case completion and verdict are derived.

func (CaseResult) Validate added in v0.38.0

func (c CaseResult) Validate() error

type Comparison

type Comparison struct {
	Baseline       ExperimentSummary
	Candidate      ExperimentSummary
	EvaluatedDelta int
	ErrorDelta     int
	Metrics        []MetricComparison
}

Comparison preserves execution coverage and paired quality differences. It makes no claim about statistical significance.

type Component

type Component[T any] struct {
	Evaluator Evaluator[T]
	Weight    float64
	Required  bool
}

Component assigns score weight and pass criticality to one evaluator. A zero Weight selects 1. Required components must pass independently of the aggregate pass policy and still contribute to the score. To keep a gate out of the quality score, declare it as another assessment in a Suite. Required is only valid when CompositeEvaluatorConfig explicitly selects a pass policy.

type CompositeEvaluator

type CompositeEvaluator[T any] struct {
	// contains filtered or unexported fields
}

CompositeEvaluator combines scored child reports. A configured pass policy additionally requires decided children; it does not change score identity.

func NewCompositeEvaluator

func NewCompositeEvaluator[T any](config CompositeEvaluatorConfig[T]) (*CompositeEvaluator[T], error)

NewCompositeEvaluator copies the component slice; evaluators remain shared.

func (*CompositeEvaluator[T]) Evaluate

func (c *CompositeEvaluator[T]) Evaluate(ctx context.Context, subject T) (Report, error)

type CompositeEvaluatorConfig added in v0.41.0

type CompositeEvaluatorConfig[T any] struct {
	Components     []Component[T]
	PassPolicy     PassPolicy
	MinimumPassed  int
	MaxConcurrency int
}

CompositeEvaluatorConfig defines score aggregation and an optional categorical policy. A zero PassPolicy produces only a score and accepts score-only components. A zero MaxConcurrency selects DefaultMaxConcurrency.

type Dataset

type Dataset[T any] struct {
	// contains filtered or unexported fields
}

Dataset owns an ordered snapshot of case identities and metadata. Its Host-assigned FixtureID must change with inputs, references, or selected cases, but not with observed outputs. Subjects stay borrowed read-only while the Dataset is in use.

func NewDataset

func NewDataset[T any](fixtureID string, cases ...Case[T]) (Dataset[T], error)

NewDataset snapshots cases and metadata, copying Subject by assignment, and rejects duplicate IDs that would make result correlation ambiguous.

func (Dataset[T]) Cases

func (d Dataset[T]) Cases() []Case[T]

Cases copies cases and metadata in declaration order; Subject may share referenced objects with the Dataset.

func (Dataset[T]) FixtureID added in v0.21.0

func (d Dataset[T]) FixtureID() string

func (Dataset[T]) Len

func (d Dataset[T]) Len() int

type Decision added in v0.38.0

type Decision struct {
	Policy     string       `json:"policy"`
	Parameters metadata.Map `json:"parameters,omitzero"`
	Verdict    Verdict      `json:"verdict"`
}

Decision records a categorical result and the policy that produced it. Policy and Parameters identify the rule independently of its observed Verdict. A Report without a categorical judgment has no Decision.

func (Decision) MarshalJSON added in v0.38.0

func (d Decision) MarshalJSON() ([]byte, error)

func (*Decision) UnmarshalJSON added in v0.38.0

func (d *Decision) UnmarshalJSON(data []byte) error

func (Decision) Validate added in v0.38.0

func (d Decision) Validate() error

type DecisionDelta added in v0.38.0

type DecisionDelta struct {
	Matched       int
	BaselineOnly  int
	CandidateOnly int
	Incompatible  int
	PassedDelta   int
	FailedDelta   int
}

DecisionDelta compares only matched decisions produced by the same policy. Incompatible counts pairs whose policies differ, including threshold changes.

type Direction

type Direction string

Direction describes how a raw measurement relates to quality. It is kept separate from Score, whose direction is always higher-is-better.

const (
	DirectionUnspecified    Direction = ""
	DirectionHigherIsBetter Direction = "higher_is_better"
	DirectionLowerIsBetter  Direction = "lower_is_better"
)

func (Direction) Validate

func (d Direction) Validate() error

type Distribution

type Distribution struct {
	Count   int
	Mean    float64
	Minimum float64
	P10     float64
	P50     float64
	P90     float64
	Maximum float64
}

Distribution summarizes one homogeneous numeric signal. Count distinguishes an absent distribution from a real distribution whose values are all zero.

type DistributionDelta

type DistributionDelta struct {
	Present       bool
	Mean          float64
	Matched       int
	BaselineOnly  int
	CandidateOnly int
}

DistributionDelta summarizes candidate-minus-baseline differences for matched case observations. Unmatched observations are counted, never imputed.

type ErrorPolicy

type ErrorPolicy string

ErrorPolicy controls whether independent failures are collected or stop new scheduling at the Suite or Experiment boundary that owns the policy.

const (
	ErrorCollect  ErrorPolicy = "collect"
	ErrorFailFast ErrorPolicy = "fail_fast"
)

type Evaluator

type Evaluator[T any] interface {
	// Evaluate inspects one subject without mutating it and returns a valid,
	// owned report for the evaluator's metric. Implementations must honor ctx;
	// a non-nil error means the report must not be consumed.
	Evaluate(ctx context.Context, subject T) (Report, error)
}

type EvaluatorFunc

type EvaluatorFunc[T any] func(context.Context, T) (Report, error)

func (EvaluatorFunc[T]) Evaluate

func (e EvaluatorFunc[T]) Evaluate(ctx context.Context, subject T) (Report, error)

type Execution added in v0.38.0

type Execution[O any] struct {
	Status   ExecutionStatus `json:"status"`
	Output   *O              `json:"output,omitzero"`
	Reason   string          `json:"reason,omitzero"`
	Metadata metadata.Map    `json:"metadata,omitzero"`
}

Execution is the target's receipt for one invocation. A nil Output means no candidate; a present empty candidate is not nil. Objects the output references stay borrowed read-only, like Case.Subject.

func (Execution[O]) MarshalJSON added in v0.38.0

func (e Execution[O]) MarshalJSON() ([]byte, error)

func (*Execution[O]) UnmarshalJSON added in v0.38.0

func (e *Execution[O]) UnmarshalJSON(data []byte) error

func (Execution[O]) Validate added in v0.38.0

func (e Execution[O]) Validate() error

type ExecutionStatus added in v0.38.0

type ExecutionStatus string

ExecutionStatus describes how a target stopped, independently of the quality of its output. A failed or exhausted execution may still produce a candidate that can be graded. Completed never implies a passing assessment.

const (
	ExecutionCompleted       ExecutionStatus = "completed"
	ExecutionBudgetExhausted ExecutionStatus = "budget_exhausted"
	ExecutionCanceled        ExecutionStatus = "canceled"
	ExecutionFailed          ExecutionStatus = "failed"
)

type Experiment

type Experiment[T any] struct {
	// contains filtered or unexported fields
}

Experiment is an immutable plan for evaluating one Dataset. It owns bounded scheduling and error semantics, but no persistence, artifacts, or product identity.

func NewExperiment

func NewExperiment[T any](config ExperimentConfig[T]) (Experiment[T], error)

func (Experiment[T]) Run

func (e Experiment[T]) Run(ctx context.Context) (ExperimentReport, error)
Example
package main

import (
	"context"
	"fmt"

	"github.com/Tangerg/scope/eval"
)

func main() {
	metric, err := eval.NewMetric(eval.MetricConfig{
		Namespace: "example",
		Name:      "non_empty",
	})
	if err != nil {
		panic(err)
	}
	evaluator := eval.EvaluatorFunc[string](func(_ context.Context, subject string) (eval.Report, error) {
		verdict := eval.VerdictFail
		if subject != "" {
			verdict = eval.VerdictPass
		}
		return eval.Report{Metric: metric, Decision: &eval.Decision{Policy: "non_empty", Verdict: verdict}}, nil
	})
	suite, err := eval.NewSuite(eval.SuiteConfig[string]{Assessments: []eval.Assessment[string]{{ID: "non_empty", Evaluator: evaluator}}})
	if err != nil {
		panic(err)
	}
	dataset, err := eval.NewDataset("test-fixture",
		eval.Case[string]{ID: "first", Subject: "answer"},
	)
	if err != nil {
		panic(err)
	}
	experiment, err := eval.NewExperiment(eval.ExperimentConfig[string]{
		Dataset: dataset, Suite: suite,
	})
	if err != nil {
		panic(err)
	}
	report, err := experiment.Run(context.Background())
	if err != nil {
		panic(err)
	}
	summary := report.Summary()

	fmt.Println(summary.Total, summary.Passed)
}
Output:
1 1

type ExperimentConfig

type ExperimentConfig[T any] struct {
	Dataset        Dataset[T]
	Suite          *Suite[T]
	MaxConcurrency int
	ErrorPolicy    ErrorPolicy
}

ExperimentConfig binds fixed case inputs and context to a Suite. MaxConcurrency limits concurrent cases; zero selects DefaultMaxConcurrency. A Suite defaults to one assessment at a time. Explicitly increasing both limits multiplies the maximum number of direct assessment calls.

type ExperimentReport

type ExperimentReport struct {
	// contains filtered or unexported fields
}

ExperimentReport owns ordered case facts and the summary derived from them. NewExperimentReport also imports independently generated observations.

func NewExperimentReport added in v0.38.0

func NewExperimentReport(fixtureID string, results []CaseResult) (ExperimentReport, error)

func (ExperimentReport) Cases

func (e ExperimentReport) Cases() []CaseResult

func (ExperimentReport) Compare

func (e ExperimentReport) Compare(candidate ExperimentReport) (Comparison, error)

Compare requires the same fixture and set of Case IDs, in any order. A changed threshold keeps scores comparable but makes decisions incompatible.

func (ExperimentReport) FixtureID added in v0.21.0

func (e ExperimentReport) FixtureID() string

func (ExperimentReport) Summary

func (e ExperimentReport) Summary() ExperimentSummary

type ExperimentSummary

type ExperimentSummary struct {
	Total       int
	Evaluated   int
	Passed      int
	Failed      int
	Unjudged    int
	Errors      int
	Partial     int
	Assessments []AssessmentSummary
	Metrics     []MetricSummary
}

ExperimentSummary counts complete and incomplete cases separately. Errors counts incomplete cases, not quality failures; Partial counts the incomplete cases whose completed assessments still contribute to Metrics.

type Metric

type Metric struct {
	// contains filtered or unexported fields
}

Metric is an immutable evaluation identity that can be copied by assignment. Parameters holds structured identity for the calculation. A Report's Decision owns any separate acceptance policy, including score thresholds. Unit and Direction describe optional raw measurements; normalized scores are always unitless and higher-is-better.

func NewMetric

func NewMetric(config MetricConfig) (Metric, error)

NewMetric snapshots parameters and normalizes object order, whitespace, and string escaping. Number spellings retain their exact precision. Later caller mutation cannot change report comparability.

func (Metric) Direction

func (m Metric) Direction() Direction

func (Metric) MarshalJSON added in v0.13.0

func (m Metric) MarshalJSON() ([]byte, error)

func (Metric) Name

func (m Metric) Name() MetricName

func (Metric) Namespace

func (m Metric) Namespace() string

func (Metric) Parameters

func (m Metric) Parameters() metadata.Map

func (Metric) String

func (m Metric) String() string

func (Metric) Unit

func (m Metric) Unit() string

func (*Metric) UnmarshalJSON added in v0.13.0

func (m *Metric) UnmarshalJSON(data []byte) error

func (Metric) Validate

func (m Metric) Validate() error

type MetricComparison

type MetricComparison struct {
	AssessmentID     AssessmentID
	Metric           Metric
	Baseline         *MetricSummary
	Candidate        *MetricSummary
	EvaluatedDelta   int
	ScoreDelta       DistributionDelta
	MeasurementDelta DistributionDelta
	DecisionDelta    DecisionDelta
}

MetricComparison relates the same named assessment and calculation. Numeric deltas use paired cases, while the two summaries retain all observations.

type MetricConfig

type MetricConfig struct {
	Namespace  string
	Name       MetricName
	Unit       string
	Direction  Direction
	Parameters metadata.Map
}

type MetricName

type MetricName string
const MetricNameComposite MetricName = "composite"

type MetricSummary

type MetricSummary struct {
	AssessmentID AssessmentID
	Metric       Metric
	Evaluated    int
	Passed       int
	Failed       int
	Unjudged     int
	Scores       Distribution
	Measurements Distribution
}

MetricSummary contains one top-level observation per successful case and assessment. Supporting Report.Details never become additional samples.

type PassPolicy

type PassPolicy string

PassPolicy controls categorical aggregation independently from score weights.

const (
	PassNone    PassPolicy = ""
	PassAll     PassPolicy = "all"
	PassAny     PassPolicy = "any"
	PassAtLeast PassPolicy = "at_least"
)

type Projection

type Projection[T, Subject any] func(T) (Subject, error)

Projection narrows an aggregate case to the subject owned by one evaluator.

type ProjectionEvaluator

type ProjectionEvaluator[T, Subject any] struct {
	// contains filtered or unexported fields
}

func NewProjectionEvaluator

func NewProjectionEvaluator[T, Subject any](
	evaluator Evaluator[Subject],
	projection Projection[T, Subject],
) (*ProjectionEvaluator[T, Subject], error)

func (*ProjectionEvaluator[T, Subject]) Evaluate

func (p *ProjectionEvaluator[T, Subject]) Evaluate(ctx context.Context, value T) (Report, error)

type Report

type Report struct {
	Metric      Metric       `json:"metric"`
	Decision    *Decision    `json:"decision,omitzero"`
	Score       *Score       `json:"score,omitzero"`
	Measurement *float64     `json:"measurement,omitzero"`
	Feedback    string       `json:"feedback,omitzero"`
	Metadata    metadata.Map `json:"metadata,omitzero"`
	Details     []Report     `json:"details,omitzero"`
}

Report is one evaluation result. Decision, Score, and Measurement are independent and optional, so no evaluation has to invent a threshold or a score. Details are supporting evidence, never additional summary observations.

func (Report) Clone

func (r Report) Clone() (Report, error)

Clone validates the complete detail tree before allocating its detached copy.

func (Report) MarshalJSON

func (r Report) MarshalJSON() ([]byte, error)

func (*Report) UnmarshalJSON

func (r *Report) UnmarshalJSON(data []byte) error

func (Report) Validate

func (r Report) Validate() error

func (Report) Verdict

func (r Report) Verdict() Verdict

type Score

type Score float64

Score is a normalized quality score in the closed interval [0, 1], where a higher value is always better.

func NewScore

func NewScore(value float64) (Score, error)

func (Score) Decide added in v0.38.0

func (s Score) Decide(threshold Score) (Decision, error)

Decide applies the higher-is-better threshold without changing the identity of the calculation that produced the Score.

Example
package main

import (
	"fmt"

	"github.com/Tangerg/scope/eval"
)

func main() {
	score, err := eval.NewScore(0.82)
	if err != nil {
		panic(err)
	}
	threshold, err := eval.NewScore(0.8)
	if err != nil {
		panic(err)
	}
	decision, err := score.Decide(threshold)
	if err != nil {
		panic(err)
	}

	fmt.Println(decision.Verdict, score.Float64())
}
Output:
pass 0.82

func (Score) Float64

func (s Score) Float64() float64

func (Score) Validate

func (s Score) Validate() error

type Suite added in v0.38.0

type Suite[T any] struct {
	// contains filtered or unexported fields
}

Suite executes independent assessments and preserves their individual execution outcomes. It is a collection operation, not an atomic Evaluator.

func NewSuite added in v0.38.0

func NewSuite[T any](config SuiteConfig[T]) (*Suite[T], error)

func (*Suite[T]) Run added in v0.38.0

func (s *Suite[T]) Run(ctx context.Context, subject T) (SuiteResult, error)

Run always returns the results accumulated before failure or cancellation. Under ErrorCollect, individual errors live in Results and do not become the operation error. The caller's cancellation and fail-fast error remain visible.

Example
package main

import (
	"context"
	"fmt"

	"github.com/Tangerg/scope/eval"
)

func main() {
	qualityMetric, err := eval.NewMetric(eval.MetricConfig{Namespace: "example", Name: "quality"})
	if err != nil {
		panic(err)
	}
	safetyMetric, err := eval.NewMetric(eval.MetricConfig{Namespace: "example", Name: "safety"})
	if err != nil {
		panic(err)
	}
	quality := eval.EvaluatorFunc[string](func(context.Context, string) (eval.Report, error) {
		score, scoreErr := eval.NewScore(0.9)
		return eval.Report{Metric: qualityMetric, Score: &score}, scoreErr
	})
	safety := eval.EvaluatorFunc[string](func(context.Context, string) (eval.Report, error) {
		return eval.Report{Metric: safetyMetric, Decision: &eval.Decision{Policy: "safety", Verdict: eval.VerdictFail}, Feedback: "Answer requires review."}, nil
	})
	scored, err := eval.NewCompositeEvaluator(eval.CompositeEvaluatorConfig[string]{
		Components: []eval.Component[string]{{Evaluator: quality}},
	})
	if err != nil {
		panic(err)
	}
	// A gate determines acceptance without contributing a score. The Suite
	// preserves the quality Composite's score under its own metric identity.
	suite, err := eval.NewSuite(eval.SuiteConfig[string]{
		Assessments: []eval.Assessment[string]{{ID: "quality", Evaluator: scored}, {ID: "safety", Evaluator: safety}},
	})
	if err != nil {
		panic(err)
	}
	report, err := suite.Run(context.Background(), "answer")
	if err != nil {
		panic(err)
	}
	fmt.Println("acceptance:", report.Verdict())
	fmt.Println("quality:", report.Results[0].Report.Score.Float64())
	fmt.Println("gate has score:", report.Results[1].Report.Score != nil)
}
Output:
acceptance: fail
quality: 0.9
gate has score: false

type SuiteConfig

type SuiteConfig[T any] struct {
	Assessments    []Assessment[T]
	MaxConcurrency int
	ErrorPolicy    ErrorPolicy
}

SuiteConfig declares independently identifiable assessments. MaxConcurrency limits direct assessment calls within one Run; zero selects one. ErrorCollect preserves failures without canceling independent assessments.

type SuiteResult added in v0.38.0

type SuiteResult struct {
	Results []AssessmentResult
}

SuiteResult is an ordered set of assessment facts. Verdict and completion are derived, so there is no independently mutable aggregate outcome.

func (SuiteResult) Clone added in v0.38.0

func (s SuiteResult) Clone() (SuiteResult, error)

func (SuiteResult) Complete added in v0.38.0

func (s SuiteResult) Complete() bool

func (SuiteResult) Err added in v0.38.0

func (s SuiteResult) Err() error

func (SuiteResult) Validate added in v0.38.0

func (s SuiteResult) Validate() error

func (SuiteResult) Verdict added in v0.38.0

func (s SuiteResult) Verdict() Verdict

Verdict is unspecified until every assessment completes. Successful reports remain individually available even when the suite has no complete verdict.

type Target added in v0.38.0

type Target[I, O any] interface {
	Run(ctx context.Context, input I) (Execution[O], error)
}

Target executes the public input and freezes its candidate before grading; references and hidden tests stay with the assessments. A non-nil error means no usable Execution; known failure, cancellation, or budget exhaustion is an Execution with a stop reason. Implementations honor ctx, finish collecting evidence before returning, and never expose credentials, cleanup handles, or mutable workspaces as candidates.

type TargetFunc added in v0.38.0

type TargetFunc[I, O any] func(context.Context, I) (Execution[O], error)

func (TargetFunc[I, O]) Run added in v0.38.0

func (t TargetFunc[I, O]) Run(ctx context.Context, input I) (Execution[O], error)

type Trial added in v0.38.0

type Trial[I, O any] struct {
	// contains filtered or unexported fields
}

Trial composes one execution with independent assessments. Dataset scheduling, retries, persistence, and sandboxes belong to the Host.

Example
package main

import (
	"context"
	"errors"
	"fmt"

	"github.com/Tangerg/scope/eval"
)

func main() {
	suite, err := eval.NewSuite(eval.SuiteConfig[eval.TrialSample[string, string]]{Assessments: []eval.Assessment[eval.TrialSample[string, string]]{{
		ID: "answer", Evaluator: eval.EvaluatorFunc[eval.TrialSample[string, string]](func(_ context.Context, sample eval.TrialSample[string, string]) (eval.Report, error) {
			if sample.Execution.Output == nil {
				return eval.Report{}, errors.New("answer unavailable")
			}
			score := eval.Score(0)
			if *sample.Execution.Output == "42" {
				score = 1
			}
			metric, err := eval.NewMetric(eval.MetricConfig{Name: "exact_answer"})
			if err != nil {
				return eval.Report{}, err
			}
			decision, err := score.Decide(1)
			return eval.Report{Metric: metric, Score: &score, Decision: &decision}, err
		}),
	}}})
	if err != nil {
		panic(err)
	}
	trial, err := eval.NewTrial(eval.TrialConfig[string, string]{Suite: suite, Target: eval.TargetFunc[string, string](func(context.Context, string) (eval.Execution[string], error) {
		answer := "42"
		return eval.Execution[string]{Status: eval.ExecutionCompleted, Output: &answer}, nil
	})})
	if err != nil {
		panic(err)
	}
	result, err := trial.Run(context.Background(), eval.Case[string]{ID: "question-1", Subject: "What is six times seven?"})
	if err != nil {
		panic(err)
	}
	fmt.Println(result.Execution.Status, result.Case.Result.Results[0].Report.Verdict())
}
Output:
completed pass

func NewTrial added in v0.38.0

func NewTrial[I, O any](config TrialConfig[I, O]) (*Trial[I, O], error)

func (*Trial[I, O]) Run added in v0.38.0

func (t *Trial[I, O]) Run(ctx context.Context, caseValue Case[I]) (TrialResult[O], error)

Run returns every accumulated stage result: a valid execution survives failed or canceled grading, and an ExecutionError leaves assessments not evaluated. An execution without output is still graded; assessments decide which evidence they require. The result holds runtime errors and is not a storage schema.

type TrialConfig added in v0.38.0

type TrialConfig[I, O any] struct {
	Target Target[I, O]
	Suite  *Suite[TrialSample[I, O]]
}

TrialConfig requires both fields; share them only when they support the Host's concurrent calls, because Trial adds no scheduling of its own.

type TrialResult added in v0.38.0

type TrialResult[O any] struct {
	Case           CaseResult
	Execution      *Execution[O]
	ExecutionError error
}

TrialResult preserves execution and assessment outcomes independently. Case can be included in NewExperimentReport with other cases from the same fixture. ExecutionError is an invocation/protocol error, never a quality judgment.

type TrialSample added in v0.38.0

type TrialSample[I, O any] struct {
	Case      Case[I]
	Execution Execution[O]
}

TrialSample keeps Execution apart from Case so generated outputs never change fixture identity and can be regraded without invoking the target again.

type Verdict

type Verdict string

Verdict is an optional categorical judgment. An unspecified verdict is a valid outcome for measurement-only or qualitative evaluations.

const (
	VerdictUnspecified Verdict = ""
	VerdictPass        Verdict = "pass"
	VerdictFail        Verdict = "fail"
)

func (Verdict) Decided

func (v Verdict) Decided() bool

func (Verdict) Validate

func (v Verdict) Validate() error

Directories

Path Synopsis
Package judge evaluates arbitrary subjects with a structured-output chat model.
Package judge evaluates arbitrary subjects with a structured-output chat model.
Package ranking evaluates ranked outputs against graded relevance judgments.
Package ranking evaluates ranked outputs against graded relevance judgments.
Package swebench exchanges predictions and reports with the official SWE-bench harness.
Package swebench exchanges predictions and reports with the official SWE-bench harness.
Package text evaluates generated text without imposing one shared sample on metrics with different semantic inputs.
Package text evaluates generated text without imposing one shared sample on metrics with different semantic inputs.
trajectory module

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL