Documentation
¶
Overview ¶
Package ageval is a k6 extension for evaluating tool-using LLM agents. Its single data type is the AgentTestCase — holding an agent run's input, output, tool calls and usage. You obtain an AgentTestCase in one of three ways:
- new AgentTestCase({...}) — wrap a recorded trajectory you already have (a logged production run, a captured dataset, a framework's output, or a raw payload parsed via a `format` adapter); no agent is run;
- AgentSimulator.run() — an optional producer that simulates an agent by running a model loop against a provider (Anthropic) with mocked tools; or
- CliAgent.run() — an optional producer that runs a real agent CLI as a subprocess (so the agent runs as part of the k6 test, even under load); or
- SigilAgent.run() — the same, for an agent that records its run to Sigil instead of printing it: k6 receives the agent's Sigil generations and rebuilds the trajectory from them.
All yield an AgentTestCase that scripts assert on with check()/expectSequence() and an LLM-as-judge, and it emits standard k6 metrics (Trend/Rate/Counter) so results work with every k6 output and threshold with no extra configuration.
TestSuite is the declarative front door on top: a named, versioned set of test cases plus the agent under test and (optionally) how to grade the outcome — `export default suite.run` drives one case attempt per iteration.
Index ¶
- type AgentSimulator
- type AgentTestCase
- func (r *AgentTestCase) CalledTool(name string) bool
- func (r *AgentTestCase) CallsOf(name string) []ToolCall
- func (r *AgentTestCase) ExpectSequence(expected sobek.Value, opts sobek.Value) bool
- func (r *AgentTestCase) FailedSteps() []ToolCall
- func (r *AgentTestCase) StepReports() []ToolCall
- func (r *AgentTestCase) ToolSequence() []string
- type CliAgent
- type ModuleInstance
- type RootModule
- type RunUsage
- type SigilAgent
- type TestSuite
- type ToolCall
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type AgentSimulator ¶
type AgentSimulator struct {
// contains filtered or unexported fields
}
AgentSimulator is the JS-facing agent. It is created via `new AgentSimulator({...})` and exposes a single `run(opts)` method.
func (*AgentSimulator) Run ¶
func (a *AgentSimulator) Run(opts sobek.Value) sobek.Value
Run executes the agent loop synchronously and returns an AgentTestCase. Blocking is intentional and idiomatic (like http.request): we run on the VU goroutine and hold the runtime, so JS mock handlers can be called directly.
run({ input, mocks?, expectedTools?: [{ name, input? }], tags? })
expectedTools is attached to the returned AgentTestCase so expectSequence() can grade against it with no argument.
type AgentTestCase ¶
type AgentTestCase struct {
Input string `js:"input"`
Output string `js:"output"`
ToolCalls []ToolCall `js:"toolCalls"`
ExpectedTools []ToolCall `js:"expectedTools"`
Usage RunUsage `js:"usage"`
Steps int `js:"steps"`
Duration float64 `js:"duration"`
// contains filtered or unexported fields
}
AgentTestCase is the single agent-evaluation data container, exposed to JS. It holds an agent run's input, produced output, recorded tool calls and usage, plus the optional expectedTools to grade against. It is obtained either directly (`new AgentTestCase({...})`) or from a producer (`AgentSimulator.run()`, `CliAgent.run()`, `SigilAgent.run()`), and is the value that check/expectSequence/judge operate on. Exported fields become camelCase properties; exported methods become camelCase methods.
func (*AgentTestCase) CalledTool ¶
func (r *AgentTestCase) CalledTool(name string) bool
CalledTool reports whether a tool with the given name was called at least once.
func (*AgentTestCase) CallsOf ¶
func (r *AgentTestCase) CallsOf(name string) []ToolCall
CallsOf returns every recorded call to the named tool, in order.
func (*AgentTestCase) ExpectSequence ¶
ExpectSequence checks the recorded tool calls against an expected sequence and emits the agent_tool_correctness metric (1 on match, 0 otherwise). It returns the boolean result so it can also be used inside check().
expected is an array of `{ name, args? }`; called with no argument it grades against the run's stored expectedTools, and throws when neither is present — an empty expected sequence would match vacuously, and a silently green agent_tool_correctness is worse than an error. opts is `{ mode, allowOtherCalls }`:
- mode "in-order" (default): expected is an in-order subsequence of the actual calls. With allowOtherCalls=false, no tool outside the expected set may be called.
- mode "exact": the actual sequence must equal expected one-to-one, in order.
func (*AgentTestCase) FailedSteps ¶
func (r *AgentTestCase) FailedSteps() []ToolCall
FailedSteps returns the step reports whose `success` field is false.
func (*AgentTestCase) StepReports ¶
func (r *AgentTestCase) StepReports() []ToolCall
StepReports returns the calls the agent made to its step-reporting tool.
func (*AgentTestCase) ToolSequence ¶
func (r *AgentTestCase) ToolSequence() []string
ToolSequence returns the ordered list of tool names called during the run.
type CliAgent ¶
type CliAgent struct {
// contains filtered or unexported fields
}
CliAgent runs a real agent CLI as a subprocess, captures its output, and turns it into an AgentTestCase — so the agent runs as part of a single `k6 run` (no separate capture step). Output is parsed by a built-in `format` (currently "claude-code") or a custom `parse(stdout)` JS callback.
func (*CliAgent) Run ¶
Run executes the agent command with the given input, parses its output, and returns an AgentTestCase. Blocking (like http.request) — runs on the VU goroutine.
run({ input, expectedTools?: [{ name, input? }], tags? })
expectedTools is attached to the returned AgentTestCase so expectSequence() can grade against it with no argument.
type ModuleInstance ¶
type ModuleInstance struct {
// contains filtered or unexported fields
}
ModuleInstance is the per-VU instance of the ageval module.
func (*ModuleInstance) Exports ¶
func (mi *ModuleInstance) Exports() modules.Exports
Exports implements the modules.Instance interface.
type RootModule ¶
type RootModule struct {
// contains filtered or unexported fields
}
RootModule is the global module instance that creates a ModuleInstance per VU. It also holds the cross-VU TestSuite state: the fallback iteration counter (see suite.go).
func (*RootModule) NewModuleInstance ¶
func (rm *RootModule) NewModuleInstance(vu modules.VU) modules.Instance
NewModuleInstance implements the modules.Module interface and returns a new instance of the module for the given VU. Metrics are registered here, in the init context, where the registry is available.
type RunUsage ¶
type RunUsage struct {
InputTokens int64 `js:"inputTokens"`
OutputTokens int64 `js:"outputTokens"`
CacheReadTokens int64 `js:"cacheReadTokens"`
CacheCreationTokens int64 `js:"cacheCreationTokens"`
}
RunUsage is the token usage of a run, exposed to JS. Cache tokens are kept separate from fresh input so cost reflects their real (cheaper) rates.
type SigilAgent ¶
type SigilAgent struct {
// contains filtered or unexported fields
}
SigilAgent runs a real, Sigil-instrumented agent CLI as a subprocess and grades what the agent recorded. k6 stands up a loopback Sigil endpoint, points the child at it, and rebuilds the trajectory from what the child exports: generations, tool executions, and OTLP trace spans.
Use it when the agent under test publishes its run to Sigil; when the agent instead prints a parseable trajectory, CliAgent is the producer to reach for. Sigil records one generation per model round trip, so a SigilAgent run also reports the number of round trips it took as `steps`.
func (*SigilAgent) Run ¶
func (a *SigilAgent) Run(opts sobek.Value) sobek.Value
Run executes the agent command with the given input and returns an AgentTestCase built from the generations the agent recorded. Blocking (like http.request) — runs on the VU goroutine.
run({ input, expectedTools?: [{ name, input? }], tags? })
type TestSuite ¶
type TestSuite struct {
// contains filtered or unexported fields
}
TestSuite is the declarative front door of ageval: a named, versioned set of test cases, the agent under test, and (optionally) how to grade the outcome. `suite.run` drives one case attempt per k6 iteration — pick the case, run the agent, grade tools/assert/judge — so the script contains no orchestration code.
func (*TestSuite) Options ¶
Options returns the k6 options fragment that runs the whole suite: one shared-iterations scenario sized cases×attempts, so under `export default suite.run` every case is attempted `attempts` times (pass@k). Spread it into the script options: `export const options = { ...suite.options({...}), thresholds: {...} }`.
options({ attempts?, vus?, maxDuration? })
func (*TestSuite) Run ¶
Run executes one case attempt: it maps the current iteration to a (case, attempt) pair, runs the agent with the case input, grades the outcome (expectedTools → agent_tool_correctness, assert, judge → agent_quality_score/ agent_judge_pass), emits the overall agent_case_pass, and returns the AgentTestCase for optional extra assertions. Designed to be the script's default function: `export default suite.run`.
When the agent run, the judge call, or an assert throws, the case is graded failed (agent_case_pass=0) before the error propagates — so rate thresholds see the failure instead of only ever sampling clean runs.
type ToolCall ¶
type ToolCall struct {
Name string `js:"name"`
Server string `js:"server"`
Input map[string]any `js:"input"`
Output string `js:"output"`
}
ToolCall is one tool invocation recorded during a run, exposed to JS. Server carries the MCP server qualifier when the adapter provides one (e.g. Claude Code's "mcp__k6__validate_script" → Name "validate_script", Server "k6"), so assertions can distinguish same-named tools from different servers while name-based matching stays unqualified.