lazycue

package module
v0.0.0-...-a973f8a Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 3, 2026 License: Apache-2.0 Imports: 26 Imported by: 0

README

LazyCue

We may want Playwright/Selenium/Cypress/etc. tests, but we don't want to maintain them. If we're being honest, we don't want to write them either.

LazyCue is an experiment in "self-healing" browser automation tests. The tests themselves are free form instructions, and an agent, at run time, interprets those instructions. Then, the interpreted instructions are cached, and these cached instructions are re-used over and over again. When that eventually fails, an agent is invoked again to fix them up.

Using this requires that we, as an industry, get more comfortable with CI editing our commits for us. For example, if you have a merge queue, and you submit a change that passes tests but has some trailing whitespace, the Right Answer is for CI to fix the formatting in a new (or amended) commit, and submit that change, even though it isn't quite the same thing as you pushed. LazyCue takes this to the next level: let's not pretend we'd be doing anything but mechanically curing the failing test. Materializing the cache in-repo vs. a separate system (a database, a cache server, elsewhere in Git) is the choice we've made here.

Whence the name?

Some of the products in this space (Playwright, Puppetteer, Stagehand) have theatrical names. So, here we are, with the LLM cueing the browser on what to do. The space is surprisingly dense with names!

The lazy is self-congratulatory.

Prerequisites

LazyCue uses the chromedp package to talk to a Chromium-based browser. On Linux, Headless Shell is good, and you can extract it like so.

sudo mkdir -p /headless-shell && go run github.com/google/go-containerregistry/cmd/crane@latest export chromedp/headless-shell:latest - | sudo tar -x -C /headless-shell

Quick Start

$go run github.com/boldsoftware/shelley/lazycue/cmd/lazycue@latest   --base-url http://xkcd.com/   'Navigate to / and check there is a comic published'
PASS  [generated → v1]  26.015s total, 24.793s agent
  Navigate to / and check theres a comic published
  ✓ navigate /                                           39ms
  ✓ wait_visible #comic                                   2ms
  ✓ assert_visible #comic img                              0s
  ✓ assert_visible #middleContainer                       1ms
  ✓ wait_text Permanent link to this comic:                0s
  ✓ assert_visible a[href*='xkcd.com']                     0s
  ⚡ 26,242 in / 1,032 out tokens  ~$0.094

$go run github.com/boldsoftware/shelley/lazycue/cmd/lazycue@latest   --base-url http://xkcd.com/   'Navigate to / and check there is a comic published'
PASS  [cached v1]  1.317s
  Navigate to / and check theres a comic published
  ✓ navigate /                                           72ms
  ✓ wait_visible #comic                                   1ms
  ✓ assert_visible #comic img                             1ms
  ✓ assert_visible #middleContainer                        0s
  ✓ wait_text Permanent link to this comic:               1ms
  ✓ assert_visible a[href*='xkcd.com']                     0s

Or from Go tests:

var app = lazycue.New(lazycue.Options{BaseURL: "http://localhost:3000"})

func TestHomepage(t *testing.T) {
    app.Test(t, `Navigate to / and verify the page title is "My App". The login button should be visible.`)
}

Workflow of a Run

  1. Hash description → look up .lazycue/<hash>.json next to your tests
  2. If cached: execute DSL steps. If passes → done (fast path, ~1-2s)
  3. If cached but fails mechanically: spawn LLM agent to fix → save new version
  4. If cached but app is genuinely broken: fail the test with explanation
  5. If not cached: spawn LLM agent to generate → save v1
DSL

Cached tests are stored as JSON as arrays of steps, for example:

[
  {"action": "navigate", "url": "/"},
  {"action": "assert_title", "text": "My App"},
  {"action": "wait_visible", "selector": "#login-button", "timeout": "10s"},
  {"action": "click", "selector": "#login-button"},
  {"action": "wait_text", "text": "Welcome back", "timeout": "10s"}
]
Self-Healing vs Genuine Failures

The agent distinguishes between:

  • Mechanical failures: wrong selectors, timing issues, missing waits → self-heals
  • Genuine failures: app doesn't match the description → fails with explanation

The test description is the source of truth. If the description says "title should be X" and the app shows "Y", that's a genuine failure.

Configuration

Flag / Option Default Description
--base-url App URL (required)
--cache-dir .lazycue Directory holding cache JSON files
--artifact-dir Write per-step screenshots + an HTML report (index.html) here
--video-dir LAZYCUE_VIDEO_DIR Render an MP4 per test (prompt title card + captioned screenshots) here
--json Write a machine-readable JSON cache-stats summary here
--model claude-sonnet-4-6 LLM model
--api-url ANTHROPIC_BASE_URL or https://api.anthropic.com Anthropic API base URL
--api-key ANTHROPIC_API_KEY Anthropic API key
--verbose false Verbose output
Options.InitScript JavaScript evaluated before every page's own scripts (Go harness only)

Screenshots & Video Artifacts

LazyCue can capture a screenshot after every executed step and, optionally, stitch them into a short MP4.

  • --artifact-dir DIR writes per-step PNGs into DIR/<prompt-hash>/ (so screenshots are arranged by prompt) plus an HTML report at DIR/index.html.
  • --video-dir DIR (or LAZYCUE_VIDEO_DIR) renders DIR/<prompt-hash>.mp4. The video opens with a title card showing the prompt, then shows each screenshot in order with its step instruction overlaid and a PASS/FAIL badge. Setting --video-dir alone also captures the underlying screenshots (under DIR/<prompt-hash>/), so you don't need --artifact-dir too.

Video rendering shells out to ffmpeg (and drawtext, so a TrueType font is required). The font is auto-detected from common locations; override it with LAZYCUE_FONT=/path/to/font.ttf. If ffmpeg or a font is missing, the test still runs — only the video is skipped.

Rendering happens after the tests finish, not in the per-test path: each MP4 is a non-trivial ffmpeg job, so rendering inline would inflate every test's wall time (and its timeout budget). The Go harness renders in TestMain after m.Run() (call Harness.RenderVideos()); the CLI renders after the run. Each MP4 is produced by a single ffmpeg invocation (a filter_complex that draws the overlays once per frame and concats them), and videos for different tests render concurrently.

lazycue --base-url http://localhost:3000 --video-dir ./lazycue-video \
  'Navigate to /login, sign in, and verify the dashboard loads'
# ... writes ./lazycue-video/<hash>.mp4 and ./lazycue-video/<hash>/step-*.png

How the Cache Works

Cached tests live in a .lazycue/ directory next to your tests. Add them to your repo to take advantage of the caching.

For CI, you'll want to create commits for them as part of your CI, and push those commits along.

Documentation

Overview

Package lazycue implements self-healing browser tests written as plain English.

A test is a description string. The system hashes the description to look up a cached DSL test script stored as a JSON file in a .lazycue/ directory next to the tests. If cached, it executes the DSL. If the test passes, it's done. If it fails (or isn't cached), an LLM agent generates or fixes the DSL, and the new version is written back to the cache file.

Index

Constants

View Source
const (
	ActionNavigate           = "navigate"
	ActionWaitVisible        = "wait_visible"
	ActionWaitHidden         = "wait_hidden"
	ActionWaitText           = "wait_text"
	ActionWaitTextGone       = "wait_text_gone"
	ActionFill               = "fill"
	ActionClick              = "click"
	ActionPressKey           = "press_key"
	ActionScreenshot         = "screenshot"
	ActionEval               = "eval"
	ActionAssertVisible      = "assert_visible"
	ActionAssertNotVisible   = "assert_not_visible"
	ActionAssertText         = "assert_text"
	ActionAssertTextContains = "assert_text_contains"
	ActionAssertAttribute    = "assert_attribute"
	ActionWaitURL            = "wait_url"
	ActionAssertURL          = "assert_url"
	ActionAssertTitle        = "assert_title"
	ActionAssertCount        = "assert_count"
	ActionSleep              = "sleep"
)

Supported action names.

View Source
const CacheBanner = "" /* 217-byte string literal not displayed */

CacheBanner is stored in the "_README" field of every cache file to make it clear the file is machine-managed and should not be hand-edited.

Variables

This section is empty.

Functions

func CacheFilePath

func CacheFilePath(cacheDir, description string) string

CacheFilePath returns the path to the cache file for a description.

func DescriptionHash

func DescriptionHash(description string) string

DescriptionHash returns the first 16 hex chars of the SHA-256 hash of a description.

func FormatSteps

func FormatSteps(steps []Step) ([]byte, error)

FormatSteps serializes steps to indented JSON.

func GetCachedTest

func GetCachedTest(cacheDir, description string) (*CachedTest, *CacheHit, error)

GetCachedTest reads the cached DSL test for a description from cacheDir. Returns (nil, nil, nil) if no cache file exists.

func RenderVideo

func RenderVideo(outPath, name, prompt string, steps []StepResult, pass bool) error

RenderVideo produces an MP4 at outPath that opens with a title card showing the test name and prompt, followed by each captured screenshot with its step instruction overlaid. Steps without a screenshot on disk are skipped. name may be empty (the title card then shows only the prompt).

The whole video is produced by a SINGLE ffmpeg invocation: a filter_complex builds the title card from a color source and each frame from its looped PNG input, then concats them. This avoids spawning O(steps) ffmpeg processes (the old approach rendered a PNG per step plus a final encode), which dominated wall time. Returns an error if ffmpeg/font are unavailable or rendering fails.

func RenderVideos

func RenderVideos(videoDir string, results []*TestResult) error

RenderVideos renders an MP4 for every result that has on-disk screenshots, writing <videoDir>/<DescriptionHash>.mp4 and setting r.VideoPath on success.

Rendering is deliberately kept OUT of the per-test hot path: each MP4 is a non-trivial ffmpeg job, so doing it inline would inflate every test's wall time (and timeout budget). Callers invoke this once, after all tests finish. Renders run concurrently (bounded by GOMAXPROCS) since they're independent.

It is best-effort: if ffmpeg or a font is unavailable it is a no-op (returns nil). Per-video failures are collected and returned joined, but successful videos are still produced and their VideoPath set.

func SaveCachedTest

func SaveCachedTest(cacheDir, description string, steps []byte, version int, meta *CacheMetadata) error

SaveCachedTest writes the cached DSL test for a description to cacheDir as a pretty-printed JSON file, including the managed-by banner.

func StepSummary

func StepSummary(s Step) string

StepSummary returns a short human-readable summary of a step, e.g. "navigate /new" or "click #login-button" or "assert_text .title \"Hello\"".

func VideoAvailable

func VideoAvailable() bool

VideoAvailable reports whether the external dependencies needed to render a video (ffmpeg + a usable font) are present.

func WriteReport

func WriteReport(dir string, results []*TestResult) error

WriteReport renders an HTML report for a set of test results, embedding per-step screenshots by relative path. The report is written to dir/index.html. Screenshot paths in results are made relative to dir.

func WriteSummary

func WriteSummary(path string, results []*TestResult) error

WriteSummary writes a RunSummary as JSON to path.

Types

type AgentConfig

type AgentConfig struct {
	Mode             AgentMode
	Description      string
	PreviousSteps    []byte // JSON of previous steps (fix mode)
	PreviousError    string // Error from previous run (fix mode)
	CacheFilePath    string // Path to the cached steps JSON for this test (fix mode); lets the agent inspect its git history for prior flakiness
	Browser          *Browser
	BaseURL          string
	Model            string
	AnthropicBaseURL string
	AnthropicAPIKey  string
	RepoRoot         string
	Verbose          bool
}

AgentConfig configures an agent run.

type AgentMode

type AgentMode int

AgentMode specifies whether the agent should generate new steps or fix existing ones.

const (
	AgentModeGenerate AgentMode = iota
	AgentModeFix
)

type AgentResult

type AgentResult struct {
	Success        bool
	Error          string
	StepsJSON      []byte
	StepResults    []StepResult
	ScreenshotPath string
	InputTokens    int
	OutputTokens   int
	// Effort diagnostics, surfaced so a human can see how hard the agent
	// worked (e.g. did a heal converge in 1-2 turns, or thrash to the limit?).
	Turns       int  // number of API round trips the agent took
	HitMaxTurns bool // true if the agent ran out of turns (maxAgentTurns)
	ToolCalls   int  // total tool_use calls the agent issued (run_steps/screenshot/git_command)
}

AgentResult is the result of an agent run.

func RunAgent

func RunAgent(ctx context.Context, cfg *AgentConfig) (*AgentResult, error)

RunAgent executes the LLM agent loop to generate or fix DSL test steps.

type Browser

type Browser struct {
	// contains filtered or unexported fields
}

Browser wraps a headless Chrome instance via chromedp.

func NewBrowser

func NewBrowser(parentCtx context.Context, initScript string) (*Browser, error)

NewBrowser launches a headless Chrome instance with Pixel 5 viewport (393x851). initScript runs before every page's own scripts when non-empty.

func (*Browser) Close

func (b *Browser) Close()

Close shuts down the browser and waits for the Chrome process to exit.

We use chromedp.Cancel (which closes the browser gracefully and blocks until the process is gone) rather than just cancelling the contexts, so that two Chrome instances never run concurrently during the handoff between one test's browser tearing down and the next test's browser launching. On a small VM that overlap causes CPU contention that can slow a freshly-launched browser's first paint / first SSE frame enough to spuriously trip a wait step.

func (*Browser) Context

func (b *Browser) Context() context.Context

Context returns the browser's chromedp context.

func (*Browser) ExecuteSteps

func (b *Browser) ExecuteSteps(ctx context.Context, baseURL string, steps []Step) ([]StepResult, error)

ExecuteSteps runs a sequence of DSL steps against the browser. It stops on the first failure and returns results for all attempted steps.

func (*Browser) Screenshot

func (b *Browser) Screenshot(ctx context.Context) ([]byte, error)

Screenshot captures a full-page PNG screenshot.

The capture is bounded by a short timeout: it runs after every step (including while the predictable agent is mid-turn on a long `delay:`), and a busy/unresponsive renderer can otherwise leave CaptureScreenshot blocked indefinitely on b.ctx, which has no deadline. A diagnostic screenshot is never worth hanging the whole test, so we cap it and let callers ignore the error.

func (*Browser) SetScreenshotSink

func (b *Browser) SetScreenshotSink(sink func(stepIndex int, action string, png []byte))

SetScreenshotSink installs a callback invoked after each executed step with a PNG screenshot. Pass nil to disable. Used to produce a visual trace.

type CacheHit

type CacheHit struct {
	Version int
}

CacheHit describes a cache hit.

type CacheMetadata

type CacheMetadata struct {
	CreatedAt        time.Time `json:"created_at"`
	Hostname         string    `json:"hostname"`
	Model            string    `json:"model"`
	InputTokens      int       `json:"input_tokens"`
	OutputTokens     int       `json:"output_tokens"`
	EstimatedCostUSD float64   `json:"estimated_cost_usd"`
	CIRun            string    `json:"ci_run,omitempty"`
	GitSHA           string    `json:"git_sha,omitempty"`
	Mode             string    `json:"mode"`
}

CacheMetadata holds provenance information about a cached test.

type CachedTest

type CachedTest struct {
	README      string          `json:"_README"`
	Description string          `json:"description"`
	Version     int             `json:"version"`
	Steps       json.RawMessage `json:"steps"`
	Metadata    *CacheMetadata  `json:"metadata,omitempty"`
}

CachedTest is the JSON wrapper stored in each cache file under .lazycue/.

type Harness

type Harness struct {
	// contains filtered or unexported fields
}

Harness holds options for running lazycue tests. Create one as a package-level var and call Test from each test function:

var browser = lazycue.New(lazycue.Options{BaseURL: "http://localhost:3000"})

func TestLogin(t *testing.T) {
    browser.Test(t, "Navigate to /login and verify the login form is visible")
}

A Harness accumulates every TestResult it runs, so a TestMain can emit an aggregate report/summary after all tests finish (see Results).

func New

func New(opts Options) *Harness

New creates a Harness with the given options.

func (*Harness) RenderVideos

func (h *Harness) RenderVideos() error

RenderVideos renders an MP4 per accumulated result into the harness's VideoDir (no-op if unset). Call it from TestMain after m.Run() and before WriteReport/WriteSummary so the report embeds the videos and the summary records their paths. Best-effort; see the package-level RenderVideos.

func (*Harness) Results

func (h *Harness) Results() []*TestResult

Results returns a copy of every TestResult run through this Harness so far. Use it from TestMain to write an aggregate report/summary, e.g.:

code := m.Run()
lazycue.WriteReport(dir, app.Results())

func (*Harness) Test

func (h *Harness) Test(t testing.TB, description string)

Test runs a self-healing browser test described in plain English. It calls t.Fatal if the test fails or encounters an error.

type HealInfo

type HealInfo struct {
	TriggerStepSummary string `json:"trigger_step_summary"` // the cached step that failed and triggered the heal (e.g. "wait_visible [data-testid='tool-call-completed']")
	TriggerError       string `json:"trigger_error"`        // the error that step produced (e.g. "timeout after 60s waiting for ...")
	TriggerWasTimeout  bool   `json:"trigger_was_timeout"`  // true if the trigger error looks like a transient wait timeout
	RetriesBeforeHeal  int    `json:"retries_before_heal"`  // how many fresh-browser retries ran before paying for the heal
	AgentTurns         int    `json:"agent_turns"`          // API round trips the heal agent took
	AgentToolCalls     int    `json:"agent_tool_calls"`     // tool calls the heal agent issued
	AgentHitMaxTurns   bool   `json:"agent_hit_max_turns"`  // true if the agent exhausted its turn budget
	StepsChanged       bool   `json:"steps_changed"`        // true if the healed steps actually differ from the cache
	CachedStepCount    int    `json:"cached_step_count"`    // number of steps in the cached (pre-heal) test
	HealedStepCount    int    `json:"healed_step_count"`    // number of steps the agent emitted
}

HealInfo records why a heal was triggered and how the agent behaved, so a human reviewing CI artifacts can see the cause of a heal and whether the agent did its job (or thrashed). It is only set when Mode == RunModeHealed.

func (*HealInfo) Summary

func (h *HealInfo) Summary() string

Summary renders a one-line human-readable explanation of why and how a heal happened, suitable for a test log or CI artifact.

type Options

type Options struct {
	BaseURL          string // Base URL of the application under test (required)
	CacheDir         string // Directory holding cache JSON files (default: ".lazycue")
	Model            string // LLM model (default: "claude-sonnet-4-6")
	AnthropicBaseURL string // Anthropic API base URL (default: ANTHROPIC_BASE_URL or https://api.anthropic.com)
	AnthropicAPIKey  string // Anthropic API key (default: ANTHROPIC_API_KEY)
	Verbose          bool   // Verbose output
	ArtifactDir      string // If set, write per-step screenshots here and record their paths on StepResults
	VideoDir         string // If set, render an MP4 per test (prompt title card + captioned screenshots) here, named <DescriptionHash>.mp4
	RepoRoot         string // Repository root passed to the heal agent's git_command tool (default: `git rev-parse --show-toplevel` from the cwd, falling back to ".")
	InitScript       string // JavaScript evaluated before every page's own scripts
}

Options configures a lazycue test run.

type RunMode

type RunMode string

RunMode describes how the test was resolved.

const (
	RunModeCached    RunMode = "cached"    // test ran from cache
	RunModeGenerated RunMode = "generated" // agent generated fresh
	RunModeHealed    RunMode = "healed"    // agent fixed a cached test
)

type RunSummary

type RunSummary struct {
	Total            int           `json:"total"`
	Passed           int           `json:"passed"`
	Failed           int           `json:"failed"`
	Cached           int           `json:"cached"`    // ran from cache, no agent
	Generated        int           `json:"generated"` // agent generated fresh (cache miss)
	Healed           int           `json:"healed"`    // cached but self-healed
	CacheHitPct      float64       `json:"cache_hit_pct"`
	InputTokens      int           `json:"input_tokens"`
	OutputTokens     int           `json:"output_tokens"`
	EstimatedCostUSD float64       `json:"estimated_cost_usd"`
	Tests            []TestSummary `json:"tests"`
}

RunSummary is a machine-readable summary of a suite of test runs. It is intended for CI: it makes the cache hit/miss breakdown and agent spend explicit so a pipeline can surface "how much was cached" prominently.

func Summarize

func Summarize(results []*TestResult) RunSummary

Summarize aggregates results into a RunSummary.

type Step

type Step struct {
	Action     string `json:"action"`
	URL        string `json:"url,omitempty"`
	Selector   string `json:"selector,omitempty"`
	Value      string `json:"value,omitempty"`
	Text       string `json:"text,omitempty"`
	Timeout    string `json:"timeout,omitempty"`
	Expression string `json:"expression,omitempty"`
	Expect     string `json:"expect,omitempty"`
	Attribute  string `json:"attribute,omitempty"`
	Key        string `json:"key,omitempty"`
	Modifiers  string `json:"modifiers,omitempty"`
	Count      int    `json:"count,omitempty"`
}

Step represents a single action in a DSL test script.

func ParseSteps

func ParseSteps(data []byte) ([]Step, error)

ParseSteps parses a JSON array of steps.

type StepResult

type StepResult struct {
	Action     string        `json:"action"`
	Summary    string        `json:"summary"` // e.g. "click #login-button"
	Pass       bool          `json:"pass"`
	Error      string        `json:"error,omitempty"`
	Duration   time.Duration `json:"duration"`
	Screenshot string        `json:"screenshot,omitempty"` // path to PNG captured after this step (if enabled)
	Output     string        `json:"output,omitempty"`     // diagnostic output (e.g. the value returned by an eval step)
}

StepResult is the result of executing a single DSL step.

type TestResult

type TestResult struct {
	Pass           bool
	Error          string
	ScreenshotPath string
	Steps          []StepResult
	CacheVersion   int
	Description    string
	Name           string // the Go test name (t.Name()), e.g. "TestNewPageSmoke"; "" if unset
	Mode           RunMode
	TotalDuration  time.Duration // wall-clock time for the entire test
	AgentDuration  time.Duration // time spent in the LLM agent (0 if cached)
	InputTokens    int           // total input tokens used by agent (0 if cached)
	OutputTokens   int           // total output tokens used by agent (0 if cached)
	EstimatedCost  float64       // estimated USD cost
	VideoPath      string        // path to the rendered MP4, if VideoDir was set
	Heal           *HealInfo     // populated only when Mode == RunModeHealed: why/how the heal happened
}

TestResult is the result of running a lazy test.

func Run

func Run(ctx context.Context, opts Options, description string) (*TestResult, error)

Run executes a single lazy test described by the given plain-English description.

type TestSummary

type TestSummary struct {
	Description   string    `json:"description"`
	Name          string    `json:"name,omitempty"`
	Pass          bool      `json:"pass"`
	Mode          string    `json:"mode"`
	CacheVersion  int       `json:"cache_version"`
	DurationMS    int64     `json:"duration_ms"`
	AgentMS       int64     `json:"agent_ms"`
	InputTokens   int       `json:"input_tokens"`
	OutputTokens  int       `json:"output_tokens"`
	EstimatedCost float64   `json:"estimated_cost_usd"`
	VideoPath     string    `json:"video_path,omitempty"`
	Error         string    `json:"error,omitempty"`
	Heal          *HealInfo `json:"heal,omitempty"` // why/how the test self-healed (only set when Mode == "healed")
}

TestSummary is the per-test slice of a RunSummary.

Directories

Path Synopsis
cmd
lazycue command
Command lazycue runs a single self-healing browser test described in plain English.
Command lazycue runs a single self-healing browser test described in plain English.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL