ageval

package module
v0.0.0-...-8b226aa Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Jul 30, 2026 License: AGPL-3.0 Imports: 30 Imported by: 0

README

xk6-ageval

ageval is a k6 extension that turns k6 into an evaluation harness for tool-using LLM agents: define test suites of agent tasks in JavaScript, run the agent (simulated, or your real agent CLI) on every k6 iteration — even under load — grade the outcome deterministically and/or with an LLM-as-judge, and gate CI with plain k6 thresholds. Results are standard k6 metrics, so they work with every k6 output and dashboard out of the box.

Status: experimental. The API may change between minor versions.

Quickstart

Build a k6 binary with the extension using xk6:

go install go.k6.io/xk6/cmd/xk6@latest
xk6 build --with github.com/vortegatorres/xk6-ageval@latest

Write an eval as a TestSuite — the suite owns the cases, the agent under test, and the grading; each k6 iteration runs one case attempt:

import { TestSuite, CliAgent } from 'k6/x/ageval';

const suite = new TestSuite({
  id: 'support-agent-core',
  version: '1.0.0',

  // The agent under test: run any agent CLI and parse its stdout…
  agent: new CliAgent({
    command: './myagent',
    parse: (stdout) => JSON.parse(stdout), // or format: 'claude-code' | 'codex'
  }),

  // Optional LLM-as-judge: each case's `expected` doubles as its rubric.
  judge: {
    provider: 'anthropic',
    model: 'claude-haiku-4-5',
    apiKey: __ENV.ANTHROPIC_API_KEY_JUDGE,
    threshold: 0.7,
  },

  cases: [
    {
      id: 'weather-paris',
      input: 'What is the weather in Paris? Use the available tools.',
      expectedTools: [{ name: 'get_weather' }],
      expected: 'States the weather in Paris based on the tool result, concisely.',
      assert: (run) => /paris/i.test(run.output), // optional custom check
    },
  ],
});

export const options = {
  ...suite.options({ attempts: 3, vus: 3 }), // every case attempted 3× (pass@k)
  thresholds: {
    agent_case_pass: ['rate>0.9'],
    agent_tool_correctness: ['rate>0.9'],
  },
};

export default suite.run;
./k6 run eval.test.js

The judge block is optional — without it, grading is purely deterministic (expectedTools, assert) and no LLM API key is needed. See examples/cli for a complete keyless example you can run as-is.

Getting an AgentTestCase

Everything grades an AgentTestCase (input, output, tool calls, usage). Four ways to obtain one:

Producer What it does
new CliAgent({ command, args?, env?, cwd?, format?/parse?, model?, timeoutSeconds? }) Runs a real agent CLI as a subprocess per iteration; trajectory parsed from stdout via a built-in format adapter (claude-code, codex) or a custom parse(stdout) callback.
new SigilAgent({ command, args?, env?, cwd?, model?, timeoutSeconds? }) The same, for an agent instrumented with Sigil. k6 receives what the agent records — generations, tool executions, traces — and rebuilds the trajectory.
new AgentSimulator({ provider, model, apiKey, systemPrompt?, tools?, skills?, maxSteps? }) Simulates an agent: a model/tool loop against a provider (Anthropic) with JS-mocked tools. Good for prompt/tool-schema evals with no agent binary.
new AgentTestCase({ output, toolCalls, input?, expectedTools?, usage?, format?/raw? }) Wraps a trajectory you already recorded (logged run, captured dataset, framework output). No agent runs. Raw payload formats: claude-code, codex, openai, anthropic, a2a.

The result exposes input, output, toolCalls, usage, steps, duration, plus helpers calledTool(), callsOf(), toolSequence(), expectSequence(expected?, opts?), stepReports(), failedSteps(). Each tool call is { name, server, input, output } — server carries the MCP server qualifier when the adapter provides one (e.g. Claude Code's mcp__k6__validate_script → name validate_script, server k6), so assertions can distinguish same-named tools from different servers.

Tool-sequence grading (expectSequence, and expectedTools on suite cases) defaults to a lenient in-order subsequence match that tolerates extra tool calls. To tighten it, pass { mode: 'exact' } and/or { allowOtherCalls: false } — on suite cases via the expectSequence option:

{ id: 'weather', input: '…', expectedTools: [{ name: 'get_weather' }],
  expectSequence: { mode: 'exact' } }

expectedTools: [] asserts the agent called no tools; omitting expectedTools disables tool grading for the case.

Prefer imperative control? The low-level API works without TestSuite:

import { check } from 'k6';
import { CliAgent, judge } from 'k6/x/ageval';

const agent = new CliAgent({ command: './myagent', parse: JSON.parse });

export default function () {
  const run = agent.run({ input: 'Was invoice INV-123 paid?', expectedTools: [{ name: 'get_invoice' }] });
  check(run, { answered: (r) => r.output.length > 0 });
  run.expectSequence();
  judge(run, { name: 'answer_quality', provider: 'anthropic', model: 'claude-haiku-4-5',
               apiKey: __ENV.ANTHROPIC_API_KEY_JUDGE, rubric: 'States whether the invoice is paid.' });
}

Sigil-instrumented agents

Some agents publish their run to Sigil (Grafana's AI observability SDK) instead of printing a parseable trajectory. SigilAgent grades those: k6 stands up a loopback Sigil endpoint, points the child at it, and rebuilds the trajectory from what the agent exports — generations, tool executions, and OTLP trace spans.

import { SigilAgent } from 'k6/x/ageval';

const agent = new SigilAgent({
  command: 'k6',
  args: ['abt', 'ieval', '--model', 'claude-haiku-4-5'],
  name: 'abt',
  model: 'claude-haiku-4-5',
  env: { ANTHROPIC_API_KEY_AGENT: __ENV.ANTHROPIC_API_KEY_AGENT },
  timeoutSeconds: 180,
});

Config matches CliAgent minus format/parse, which are refused because nothing is read from stdout.

Note: Sigil records one generation per model round trip, so a SigilAgent run reports the number of round trips it took as steps. When the agent also records tool executions (as workflow steps or execute_tool trace spans), those are what toolCalls grades — one entry per execution, with what the tool ran with.

Metrics

All metrics are standard k6 types, tagged with agent, model, and (under a suite) suite / case / attempt:

Metric Type Meaning
agent_case_pass Rate Overall per-case pass: AND of all graded criteria (tools, assert, judge)
agent_tool_correctness Rate expectSequence result (expected tool sequence matched)
agent_quality_score Trend LLM-as-judge score, 0..1
agent_judge_pass Rate Judge score ≥ threshold
agent_duration Trend Wall-clock of a full agent run
agent_steps Trend Model round-trips per run
agent_tool_calls Counter Tool calls, tagged by tool name
agent_tokens Counter Agent tokens, tagged by direction (input, output, and — when reported — cache_read, cache_creation)
agent_cost_usd Counter Estimated agent spend, cache-aware (see Models & pricing)
agent_judge_tokens / agent_judge_cost_usd Counter The judge's own spend, tracked separately

Failing judge evals additionally log a parseable ageval_eval {json} line at Error level carrying the rubric, input, and the agent's actual behavior.

Errors are graded, not skipped: under TestSuite, an agent/judge/assert error grades the case failed (agent_case_pass=0) before the iteration aborts. With the low-level API an error aborts the iteration without emitting eval samples — pair your thresholds with k6's own iteration/error metrics there.

Models & pricing

Model names are not validated by the extension — any name is passed through to the provider API, which rejects unknown ones with a clear error. Newly released models work immediately, no extension update needed. The same applies to output caps: maxTokens is passed through and enforced by the API.

Pricing cannot be fetched programmatically (no provider exposes $/MTok via API), so the agent_cost_usd / agent_judge_cost_usd metrics resolve pricing in this order:

  1. An explicit pricing config (USD per million tokens) — works for any model, provider, or negotiated rate. Available on AgentSimulator, CliAgent, SigilAgent, AgentTestCase, judge(), and the suite judge block:

    agent: new CliAgent({
      command: './myagent',
      model: 'some-brand-new-model',
      pricing: { inputPerMTok: 3, outputPerMTok: 15 },
    })
    
  2. A small built-in fallback table for common Anthropic models (best-effort snapshot, dated in anthropic.go).

  3. Neither known → no cost metric is emitted (token metrics still are, so cost stays computable downstream).

Cache tokens are priced at their real rates. Sources that report cache usage (the claude-code adapter, the Anthropic API, recorded runs with usage.cacheReadTokens / usage.cacheCreationTokens) keep cache tokens separate from fresh input; the cost formula prices reads at 0.1× and (5-minute-TTL) writes at 1.25× the input rate — Anthropic's uniform multipliers. If your agent uses the 1-hour cache TTL (2× writes) or another provider's rates, override them explicitly:

pricing: { inputPerMTok: 10, outputPerMTok: 50, cacheReadPerMTok: 1, cacheWritePerMTok: 20 }

Examples

Compatibility

Builds against the go.k6.io/k6/v2 Go module — i.e. k6 v2.x binaries — and Go 1.25+. The extension uses no v2-only k6 APIs; the constraint is simply that /v2 is the current k6 module line xk6 links extensions against (k6 v1.x uses a different module path and can't host a /v2-built extension).

Development

git clone https://github.com/vortegatorres/xk6-ageval && cd xk6-ageval
make build     # k6 binary with the extension built from the local sources
make test      # unit tests (race detector on)
make lint      # golangci-lint
make example   # build + run the keyless CLI suite example end-to-end

Run make help for all targets.

License

AGPL-3.0

Documentation

Overview

Package ageval is a k6 extension for evaluating tool-using LLM agents. Its single data type is the AgentTestCase — holding an agent run's input, output, tool calls and usage. You obtain an AgentTestCase in one of three ways:

  • new AgentTestCase({...}) — wrap a recorded trajectory you already have (a logged production run, a captured dataset, a framework's output, or a raw payload parsed via a `format` adapter); no agent is run;
  • AgentSimulator.run() — an optional producer that simulates an agent by running a model loop against a provider (Anthropic) with mocked tools; or
  • CliAgent.run() — an optional producer that runs a real agent CLI as a subprocess (so the agent runs as part of the k6 test, even under load); or
  • SigilAgent.run() — the same, for an agent that records its run to Sigil instead of printing it: k6 receives the agent's Sigil generations and rebuilds the trajectory from them.

All yield an AgentTestCase that scripts assert on with check()/expectSequence() and an LLM-as-judge, and it emits standard k6 metrics (Trend/Rate/Counter) so results work with every k6 output and threshold with no extra configuration.

TestSuite is the declarative front door on top: a named, versioned set of test cases plus the agent under test and (optionally) how to grade the outcome — `export default suite.run` drives one case attempt per iteration.

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

This section is empty.

Types

type AgentSimulator

type AgentSimulator struct {
	// contains filtered or unexported fields
}

AgentSimulator is the JS-facing agent. It is created via `new AgentSimulator({...})` and exposes a single `run(opts)` method.

func (*AgentSimulator) Run

func (a *AgentSimulator) Run(opts sobek.Value) sobek.Value

Run executes the agent loop synchronously and returns an AgentTestCase. Blocking is intentional and idiomatic (like http.request): we run on the VU goroutine and hold the runtime, so JS mock handlers can be called directly.

run({ input, mocks?, expectedTools?: [{ name, input? }], tags? })

expectedTools is attached to the returned AgentTestCase so expectSequence() can grade against it with no argument.

type AgentTestCase

type AgentTestCase struct {
	Input         string     `js:"input"`
	Output        string     `js:"output"`
	ToolCalls     []ToolCall `js:"toolCalls"`
	ExpectedTools []ToolCall `js:"expectedTools"`
	Usage         RunUsage   `js:"usage"`
	Steps         int        `js:"steps"`
	Duration      float64    `js:"duration"`
	// contains filtered or unexported fields
}

AgentTestCase is the single agent-evaluation data container, exposed to JS. It holds an agent run's input, produced output, recorded tool calls and usage, plus the optional expectedTools to grade against. It is obtained either directly (`new AgentTestCase({...})`) or from a producer (`AgentSimulator.run()`, `CliAgent.run()`, `SigilAgent.run()`), and is the value that check/expectSequence/judge operate on. Exported fields become camelCase properties; exported methods become camelCase methods.

func (*AgentTestCase) CalledTool

func (r *AgentTestCase) CalledTool(name string) bool

CalledTool reports whether a tool with the given name was called at least once.

func (*AgentTestCase) CallsOf

func (r *AgentTestCase) CallsOf(name string) []ToolCall

CallsOf returns every recorded call to the named tool, in order.

func (*AgentTestCase) ExpectSequence

func (r *AgentTestCase) ExpectSequence(expected sobek.Value, opts sobek.Value) bool

ExpectSequence checks the recorded tool calls against an expected sequence and emits the agent_tool_correctness metric (1 on match, 0 otherwise). It returns the boolean result so it can also be used inside check().

expected is an array of `{ name, args? }`; called with no argument it grades against the run's stored expectedTools, and throws when neither is present — an empty expected sequence would match vacuously, and a silently green agent_tool_correctness is worse than an error. opts is `{ mode, allowOtherCalls }`:

  • mode "in-order" (default): expected is an in-order subsequence of the actual calls. With allowOtherCalls=false, no tool outside the expected set may be called.
  • mode "exact": the actual sequence must equal expected one-to-one, in order.

func (*AgentTestCase) FailedSteps

func (r *AgentTestCase) FailedSteps() []ToolCall

FailedSteps returns the step reports whose `success` field is false.

func (*AgentTestCase) StepReports

func (r *AgentTestCase) StepReports() []ToolCall

StepReports returns the calls the agent made to its step-reporting tool.

func (*AgentTestCase) ToolSequence

func (r *AgentTestCase) ToolSequence() []string

ToolSequence returns the ordered list of tool names called during the run.

type CliAgent

type CliAgent struct {
	// contains filtered or unexported fields
}

CliAgent runs a real agent CLI as a subprocess, captures its output, and turns it into an AgentTestCase — so the agent runs as part of a single `k6 run` (no separate capture step). Output is parsed by a built-in `format` (currently "claude-code") or a custom `parse(stdout)` JS callback.

func (*CliAgent) Run

func (a *CliAgent) Run(opts sobek.Value) sobek.Value

Run executes the agent command with the given input, parses its output, and returns an AgentTestCase. Blocking (like http.request) — runs on the VU goroutine.

run({ input, expectedTools?: [{ name, input? }], tags? })

expectedTools is attached to the returned AgentTestCase so expectSequence() can grade against it with no argument.

type ModuleInstance

type ModuleInstance struct {
	// contains filtered or unexported fields
}

ModuleInstance is the per-VU instance of the ageval module.

func (*ModuleInstance) Exports

func (mi *ModuleInstance) Exports() modules.Exports

Exports implements the modules.Instance interface.

type RootModule

type RootModule struct {
	// contains filtered or unexported fields
}

RootModule is the global module instance that creates a ModuleInstance per VU. It also holds the cross-VU TestSuite state: the fallback iteration counter (see suite.go).

func New

func New() *RootModule

New returns a pointer to a new RootModule instance.

func (*RootModule) NewModuleInstance

func (rm *RootModule) NewModuleInstance(vu modules.VU) modules.Instance

NewModuleInstance implements the modules.Module interface and returns a new instance of the module for the given VU. Metrics are registered here, in the init context, where the registry is available.

type RunUsage

type RunUsage struct {
	InputTokens         int64 `js:"inputTokens"`
	OutputTokens        int64 `js:"outputTokens"`
	CacheReadTokens     int64 `js:"cacheReadTokens"`
	CacheCreationTokens int64 `js:"cacheCreationTokens"`
}

RunUsage is the token usage of a run, exposed to JS. Cache tokens are kept separate from fresh input so cost reflects their real (cheaper) rates.

type SigilAgent

type SigilAgent struct {
	// contains filtered or unexported fields
}

SigilAgent runs a real, Sigil-instrumented agent CLI as a subprocess and grades what the agent recorded. k6 stands up a loopback Sigil endpoint, points the child at it, and rebuilds the trajectory from what the child exports: generations, tool executions, and OTLP trace spans.

Use it when the agent under test publishes its run to Sigil; when the agent instead prints a parseable trajectory, CliAgent is the producer to reach for. Sigil records one generation per model round trip, so a SigilAgent run also reports the number of round trips it took as `steps`.

func (*SigilAgent) Run

func (a *SigilAgent) Run(opts sobek.Value) sobek.Value

Run executes the agent command with the given input and returns an AgentTestCase built from the generations the agent recorded. Blocking (like http.request) — runs on the VU goroutine.

run({ input, expectedTools?: [{ name, input? }], tags? })

type TestSuite

type TestSuite struct {
	// contains filtered or unexported fields
}

TestSuite is the declarative front door of ageval: a named, versioned set of test cases, the agent under test, and (optionally) how to grade the outcome. `suite.run` drives one case attempt per k6 iteration — pick the case, run the agent, grade tools/assert/judge — so the script contains no orchestration code.

func (*TestSuite) Options

func (s *TestSuite) Options(opts sobek.Value) sobek.Value

Options returns the k6 options fragment that runs the whole suite: one shared-iterations scenario sized cases×attempts, so under `export default suite.run` every case is attempted `attempts` times (pass@k). Spread it into the script options: `export const options = { ...suite.options({...}), thresholds: {...} }`.

options({ attempts?, vus?, maxDuration? })

func (*TestSuite) Run

func (s *TestSuite) Run() sobek.Value

Run executes one case attempt: it maps the current iteration to a (case, attempt) pair, runs the agent with the case input, grades the outcome (expectedTools → agent_tool_correctness, assert, judge → agent_quality_score/ agent_judge_pass), emits the overall agent_case_pass, and returns the AgentTestCase for optional extra assertions. Designed to be the script's default function: `export default suite.run`.

When the agent run, the judge call, or an assert throws, the case is graded failed (agent_case_pass=0) before the error propagates — so rate thresholds see the failure instead of only ever sampling clean runs.

type ToolCall

type ToolCall struct {
	Name   string         `js:"name"`
	Server string         `js:"server"`
	Input  map[string]any `js:"input"`
	Output string         `js:"output"`
}

ToolCall is one tool invocation recorded during a run, exposed to JS. Server carries the MCP server qualifier when the adapter provides one (e.g. Claude Code's "mcp__k6__validate_script" → Name "validate_script", Server "k6"), so assertions can distinguish same-named tools from different servers while name-based matching stays unqualified.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL