Documentation
¶
Overview ¶
Package lazycue implements self-healing browser tests written as plain English.
A test is a description string. The system hashes the description to look up a cached DSL test script stored as a JSON file in a .lazycue/ directory next to the tests. If cached, it executes the DSL. If the test passes, it's done. If it fails (or isn't cached), an LLM agent generates or fixes the DSL, and the new version is written back to the cache file.
Index ¶
- Constants
- func CacheFilePath(cacheDir, description string) string
- func DescriptionHash(description string) string
- func FormatSteps(steps []Step) ([]byte, error)
- func GetCachedTest(cacheDir, description string) (*CachedTest, *CacheHit, error)
- func RenderVideo(outPath, name, prompt string, steps []StepResult, pass bool) error
- func RenderVideos(videoDir string, results []*TestResult) error
- func SaveCachedTest(cacheDir, description string, steps []byte, version int, meta *CacheMetadata) error
- func StepSummary(s Step) string
- func VideoAvailable() bool
- func WriteReport(dir string, results []*TestResult) error
- func WriteSummary(path string, results []*TestResult) error
- type AgentConfig
- type AgentMode
- type AgentResult
- type Browser
- func (b *Browser) Close()
- func (b *Browser) Context() context.Context
- func (b *Browser) ExecuteSteps(ctx context.Context, baseURL string, steps []Step) ([]StepResult, error)
- func (b *Browser) Screenshot(ctx context.Context) ([]byte, error)
- func (b *Browser) SetScreenshotSink(sink func(stepIndex int, action string, png []byte))
- type CacheHit
- type CacheMetadata
- type CachedTest
- type Harness
- type HealInfo
- type Options
- type RunMode
- type RunSummary
- type Step
- type StepResult
- type TestResult
- type TestSummary
Constants ¶
const ( ActionWaitVisible = "wait_visible" ActionWaitHidden = "wait_hidden" ActionWaitText = "wait_text" ActionWaitTextGone = "wait_text_gone" ActionFill = "fill" ActionClick = "click" ActionPressKey = "press_key" ActionScreenshot = "screenshot" ActionEval = "eval" ActionAssertVisible = "assert_visible" ActionAssertNotVisible = "assert_not_visible" ActionAssertText = "assert_text" ActionAssertTextContains = "assert_text_contains" ActionAssertAttribute = "assert_attribute" ActionWaitURL = "wait_url" ActionAssertURL = "assert_url" ActionAssertTitle = "assert_title" ActionAssertCount = "assert_count" ActionSleep = "sleep" )
Supported action names.
const CacheBanner = "" /* 217-byte string literal not displayed */
CacheBanner is stored in the "_README" field of every cache file to make it clear the file is machine-managed and should not be hand-edited.
Variables ¶
This section is empty.
Functions ¶
func CacheFilePath ¶
CacheFilePath returns the path to the cache file for a description.
func DescriptionHash ¶
DescriptionHash returns the first 16 hex chars of the SHA-256 hash of a description.
func FormatSteps ¶
FormatSteps serializes steps to indented JSON.
func GetCachedTest ¶
func GetCachedTest(cacheDir, description string) (*CachedTest, *CacheHit, error)
GetCachedTest reads the cached DSL test for a description from cacheDir. Returns (nil, nil, nil) if no cache file exists.
func RenderVideo ¶
func RenderVideo(outPath, name, prompt string, steps []StepResult, pass bool) error
RenderVideo produces an MP4 at outPath that opens with a title card showing the test name and prompt, followed by each captured screenshot with its step instruction overlaid. Steps without a screenshot on disk are skipped. name may be empty (the title card then shows only the prompt).
The whole video is produced by a SINGLE ffmpeg invocation: a filter_complex builds the title card from a color source and each frame from its looped PNG input, then concats them. This avoids spawning O(steps) ffmpeg processes (the old approach rendered a PNG per step plus a final encode), which dominated wall time. Returns an error if ffmpeg/font are unavailable or rendering fails.
func RenderVideos ¶
func RenderVideos(videoDir string, results []*TestResult) error
RenderVideos renders an MP4 for every result that has on-disk screenshots, writing <videoDir>/<DescriptionHash>.mp4 and setting r.VideoPath on success.
Rendering is deliberately kept OUT of the per-test hot path: each MP4 is a non-trivial ffmpeg job, so doing it inline would inflate every test's wall time (and timeout budget). Callers invoke this once, after all tests finish. Renders run concurrently (bounded by GOMAXPROCS) since they're independent.
It is best-effort: if ffmpeg or a font is unavailable it is a no-op (returns nil). Per-video failures are collected and returned joined, but successful videos are still produced and their VideoPath set.
func SaveCachedTest ¶
func SaveCachedTest(cacheDir, description string, steps []byte, version int, meta *CacheMetadata) error
SaveCachedTest writes the cached DSL test for a description to cacheDir as a pretty-printed JSON file, including the managed-by banner.
func StepSummary ¶
StepSummary returns a short human-readable summary of a step, e.g. "navigate /new" or "click #login-button" or "assert_text .title \"Hello\"".
func VideoAvailable ¶
func VideoAvailable() bool
VideoAvailable reports whether the external dependencies needed to render a video (ffmpeg + a usable font) are present.
func WriteReport ¶
func WriteReport(dir string, results []*TestResult) error
WriteReport renders an HTML report for a set of test results, embedding per-step screenshots by relative path. The report is written to dir/index.html. Screenshot paths in results are made relative to dir.
func WriteSummary ¶
func WriteSummary(path string, results []*TestResult) error
WriteSummary writes a RunSummary as JSON to path.
Types ¶
type AgentConfig ¶
type AgentConfig struct {
Mode AgentMode
Description string
PreviousSteps []byte // JSON of previous steps (fix mode)
PreviousError string // Error from previous run (fix mode)
CacheFilePath string // Path to the cached steps JSON for this test (fix mode); lets the agent inspect its git history for prior flakiness
Browser *Browser
BaseURL string
Model string
AnthropicBaseURL string
AnthropicAPIKey string
RepoRoot string
Verbose bool
}
AgentConfig configures an agent run.
type AgentMode ¶
type AgentMode int
AgentMode specifies whether the agent should generate new steps or fix existing ones.
type AgentResult ¶
type AgentResult struct {
Success bool
Error string
StepsJSON []byte
StepResults []StepResult
ScreenshotPath string
InputTokens int
OutputTokens int
// Effort diagnostics, surfaced so a human can see how hard the agent
// worked (e.g. did a heal converge in 1-2 turns, or thrash to the limit?).
Turns int // number of API round trips the agent took
HitMaxTurns bool // true if the agent ran out of turns (maxAgentTurns)
ToolCalls int // total tool_use calls the agent issued (run_steps/screenshot/git_command)
}
AgentResult is the result of an agent run.
func RunAgent ¶
func RunAgent(ctx context.Context, cfg *AgentConfig) (*AgentResult, error)
RunAgent executes the LLM agent loop to generate or fix DSL test steps.
type Browser ¶
type Browser struct {
// contains filtered or unexported fields
}
Browser wraps a headless Chrome instance via chromedp.
func NewBrowser ¶
NewBrowser launches a headless Chrome instance with Pixel 5 viewport (393x851). initScript runs before every page's own scripts when non-empty.
func (*Browser) Close ¶
func (b *Browser) Close()
Close shuts down the browser and waits for the Chrome process to exit.
We use chromedp.Cancel (which closes the browser gracefully and blocks until the process is gone) rather than just cancelling the contexts, so that two Chrome instances never run concurrently during the handoff between one test's browser tearing down and the next test's browser launching. On a small VM that overlap causes CPU contention that can slow a freshly-launched browser's first paint / first SSE frame enough to spuriously trip a wait step.
func (*Browser) ExecuteSteps ¶
func (b *Browser) ExecuteSteps(ctx context.Context, baseURL string, steps []Step) ([]StepResult, error)
ExecuteSteps runs a sequence of DSL steps against the browser. It stops on the first failure and returns results for all attempted steps.
func (*Browser) Screenshot ¶
Screenshot captures a full-page PNG screenshot.
The capture is bounded by a short timeout: it runs after every step (including while the predictable agent is mid-turn on a long `delay:`), and a busy/unresponsive renderer can otherwise leave CaptureScreenshot blocked indefinitely on b.ctx, which has no deadline. A diagnostic screenshot is never worth hanging the whole test, so we cap it and let callers ignore the error.
type CacheMetadata ¶
type CacheMetadata struct {
CreatedAt time.Time `json:"created_at"`
Hostname string `json:"hostname"`
Model string `json:"model"`
InputTokens int `json:"input_tokens"`
OutputTokens int `json:"output_tokens"`
EstimatedCostUSD float64 `json:"estimated_cost_usd"`
CIRun string `json:"ci_run,omitempty"`
GitSHA string `json:"git_sha,omitempty"`
Mode string `json:"mode"`
}
CacheMetadata holds provenance information about a cached test.
type CachedTest ¶
type CachedTest struct {
README string `json:"_README"`
Description string `json:"description"`
Version int `json:"version"`
Steps json.RawMessage `json:"steps"`
Metadata *CacheMetadata `json:"metadata,omitempty"`
}
CachedTest is the JSON wrapper stored in each cache file under .lazycue/.
type Harness ¶
type Harness struct {
// contains filtered or unexported fields
}
Harness holds options for running lazycue tests. Create one as a package-level var and call Test from each test function:
var browser = lazycue.New(lazycue.Options{BaseURL: "http://localhost:3000"})
func TestLogin(t *testing.T) {
browser.Test(t, "Navigate to /login and verify the login form is visible")
}
A Harness accumulates every TestResult it runs, so a TestMain can emit an aggregate report/summary after all tests finish (see Results).
func (*Harness) RenderVideos ¶
RenderVideos renders an MP4 per accumulated result into the harness's VideoDir (no-op if unset). Call it from TestMain after m.Run() and before WriteReport/WriteSummary so the report embeds the videos and the summary records their paths. Best-effort; see the package-level RenderVideos.
func (*Harness) Results ¶
func (h *Harness) Results() []*TestResult
Results returns a copy of every TestResult run through this Harness so far. Use it from TestMain to write an aggregate report/summary, e.g.:
code := m.Run() lazycue.WriteReport(dir, app.Results())
type HealInfo ¶
type HealInfo struct {
TriggerStepSummary string `json:"trigger_step_summary"` // the cached step that failed and triggered the heal (e.g. "wait_visible [data-testid='tool-call-completed']")
TriggerError string `json:"trigger_error"` // the error that step produced (e.g. "timeout after 60s waiting for ...")
TriggerWasTimeout bool `json:"trigger_was_timeout"` // true if the trigger error looks like a transient wait timeout
RetriesBeforeHeal int `json:"retries_before_heal"` // how many fresh-browser retries ran before paying for the heal
AgentTurns int `json:"agent_turns"` // API round trips the heal agent took
AgentToolCalls int `json:"agent_tool_calls"` // tool calls the heal agent issued
AgentHitMaxTurns bool `json:"agent_hit_max_turns"` // true if the agent exhausted its turn budget
StepsChanged bool `json:"steps_changed"` // true if the healed steps actually differ from the cache
CachedStepCount int `json:"cached_step_count"` // number of steps in the cached (pre-heal) test
HealedStepCount int `json:"healed_step_count"` // number of steps the agent emitted
}
HealInfo records why a heal was triggered and how the agent behaved, so a human reviewing CI artifacts can see the cause of a heal and whether the agent did its job (or thrashed). It is only set when Mode == RunModeHealed.
type Options ¶
type Options struct {
BaseURL string // Base URL of the application under test (required)
CacheDir string // Directory holding cache JSON files (default: ".lazycue")
Model string // LLM model (default: "claude-sonnet-4-6")
AnthropicBaseURL string // Anthropic API base URL (default: ANTHROPIC_BASE_URL or https://api.anthropic.com)
AnthropicAPIKey string // Anthropic API key (default: ANTHROPIC_API_KEY)
Verbose bool // Verbose output
ArtifactDir string // If set, write per-step screenshots here and record their paths on StepResults
VideoDir string // If set, render an MP4 per test (prompt title card + captioned screenshots) here, named <DescriptionHash>.mp4
RepoRoot string // Repository root passed to the heal agent's git_command tool (default: `git rev-parse --show-toplevel` from the cwd, falling back to ".")
InitScript string // JavaScript evaluated before every page's own scripts
}
Options configures a lazycue test run.
type RunSummary ¶
type RunSummary struct {
Total int `json:"total"`
Passed int `json:"passed"`
Failed int `json:"failed"`
Cached int `json:"cached"` // ran from cache, no agent
Generated int `json:"generated"` // agent generated fresh (cache miss)
Healed int `json:"healed"` // cached but self-healed
CacheHitPct float64 `json:"cache_hit_pct"`
InputTokens int `json:"input_tokens"`
OutputTokens int `json:"output_tokens"`
EstimatedCostUSD float64 `json:"estimated_cost_usd"`
Tests []TestSummary `json:"tests"`
}
RunSummary is a machine-readable summary of a suite of test runs. It is intended for CI: it makes the cache hit/miss breakdown and agent spend explicit so a pipeline can surface "how much was cached" prominently.
func Summarize ¶
func Summarize(results []*TestResult) RunSummary
Summarize aggregates results into a RunSummary.
type Step ¶
type Step struct {
Action string `json:"action"`
URL string `json:"url,omitempty"`
Selector string `json:"selector,omitempty"`
Value string `json:"value,omitempty"`
Text string `json:"text,omitempty"`
Timeout string `json:"timeout,omitempty"`
Expression string `json:"expression,omitempty"`
Expect string `json:"expect,omitempty"`
Attribute string `json:"attribute,omitempty"`
Key string `json:"key,omitempty"`
Modifiers string `json:"modifiers,omitempty"`
Count int `json:"count,omitempty"`
}
Step represents a single action in a DSL test script.
func ParseSteps ¶
ParseSteps parses a JSON array of steps.
type StepResult ¶
type StepResult struct {
Action string `json:"action"`
Summary string `json:"summary"` // e.g. "click #login-button"
Pass bool `json:"pass"`
Error string `json:"error,omitempty"`
Duration time.Duration `json:"duration"`
Screenshot string `json:"screenshot,omitempty"` // path to PNG captured after this step (if enabled)
Output string `json:"output,omitempty"` // diagnostic output (e.g. the value returned by an eval step)
}
StepResult is the result of executing a single DSL step.
type TestResult ¶
type TestResult struct {
Pass bool
Error string
ScreenshotPath string
Steps []StepResult
CacheVersion int
Description string
Name string // the Go test name (t.Name()), e.g. "TestNewPageSmoke"; "" if unset
Mode RunMode
TotalDuration time.Duration // wall-clock time for the entire test
AgentDuration time.Duration // time spent in the LLM agent (0 if cached)
InputTokens int // total input tokens used by agent (0 if cached)
OutputTokens int // total output tokens used by agent (0 if cached)
EstimatedCost float64 // estimated USD cost
VideoPath string // path to the rendered MP4, if VideoDir was set
Heal *HealInfo // populated only when Mode == RunModeHealed: why/how the heal happened
}
TestResult is the result of running a lazy test.
type TestSummary ¶
type TestSummary struct {
Description string `json:"description"`
Name string `json:"name,omitempty"`
Pass bool `json:"pass"`
Mode string `json:"mode"`
CacheVersion int `json:"cache_version"`
DurationMS int64 `json:"duration_ms"`
AgentMS int64 `json:"agent_ms"`
InputTokens int `json:"input_tokens"`
OutputTokens int `json:"output_tokens"`
EstimatedCost float64 `json:"estimated_cost_usd"`
VideoPath string `json:"video_path,omitempty"`
Error string `json:"error,omitempty"`
Heal *HealInfo `json:"heal,omitempty"` // why/how the test self-healed (only set when Mode == "healed")
}
TestSummary is the per-test slice of a RunSummary.