evals

command
v0.2.7 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 20, 2026 License: MIT Imports: 15 Imported by: 0

Documentation

Overview

Command evals runs scripted scenarios against Kram's real agent loop, through a real (or real-ish) configured provider, and checks specific behaviors — the thing unit tests structurally cannot check, since they exercise Kram's own code in isolation while an eval exercises what the model actually does when handed Kram's system prompt and tools.

This directly operationalizes DECISIONS.md's definition of the gap it closes: "a harness that can answer 'did this prompt change make the agent better or worse'". Several scenarios here are regression tests for real bugs found by hand earlier in this project (the silent empty final answer, grep returning binary garbage, memory/skills never being used proactively) — the harness exists so the next regression of the same shape gets caught by `go run ./evals`, not by another hour of manual tmux debugging.

Wiring mirrors cmd/kram: gateway and daemon run in-process on free localhost ports, auto-detected from whichever provider env var/stored credential is available — same autodetection providercatalog backs everywhere else, so an eval run uses exactly the setup a real user has.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL