flaky

package
v1.13.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Jul 4, 2026 License: GPL-3.0 Imports: 6 Imported by: 0

Documentation

Overview

Package flaky is CommitBrief's deterministic, static flaky-test anti-pattern detector (ADR-0022). Unlike the LLM review path it produces reproducible findings from the diff alone: it scans the added lines of changed test files for high-precision anti-patterns (hard-coded sleeps, unseeded randomness, …) and emits standard render.Finding values — no JSON-schema change, no provider call. Recall is intentionally secondary to precision: a noisy commit-stage gate erodes trust.

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

func Annotate added in v1.12.0

func Annotate(cat *i18n.Catalog, f render.Finding, r RerunResult) render.Finding

Annotate folds a rerun verdict into an existing static flaky finding without touching the JSON schema (locked at v1, ADR-0014): the localized verdict line is appended to the finding's Suggestion, and the finding's Severity is adjusted to reflect empirical truth:

  • flaky → keep the finding; the rerun upgrades it from "suspected" to "confirmed" (text only; severity unchanged).
  • real_failure → this is a genuinely red test, not flakiness; the static anti-pattern note still applies but must not masquerade as a flake, so the verdict line says so plainly.
  • transient → the flake did not reproduce; downgraded to info so a one-off does not block a commit-stage gate, with a line saying it passed every rerun.
  • inconclusive → returned unchanged (no rerun was observed).

Annotate never drops a finding — the static signal is preserved either way; the rerun only refines how it is presented. cat localizes the verdict line to the resolved output language (ADR-0021).

Types

type Detector

type Detector struct {
	// contains filtered or unexported fields
}

Detector turns a parsed diff into deterministic flaky-test findings. The catalog localizes finding text to the resolved output language (ADR-0021); the Severity value stays the English wire vocabulary, as for LLM findings.

func New

func New(cat *i18n.Catalog) *Detector

New returns a Detector that localizes finding text via cat.

func (*Detector) Detect

func (d *Detector) Detect(parsed diff.Diff) []render.Finding

Detect scans the added lines of every changed test file in parsed and returns the matched anti-patterns as findings. The returned slice is non-nil only when there are findings; order follows file → hunk → line.

type Executor added in v1.12.0

type Executor func(ctx context.Context, testID string) (passed bool, err error)

Executor runs a single, already-isolated test once and reports whether it passed. The error channel is reserved for harness failures that are NOT a test result — the runner could not be launched, the context was cancelled, the test binary did not compile. A test that runs and fails its assertions is `passed == false, err == nil`; only an inability to OBSERVE a result is an error. testID is opaque to the orchestrator: it is whatever handle the caller's runner understands (a fully-qualified test name, a node id, …).

Implementations must be safe to call repeatedly for the same testID — each call is one independent rerun.

type RerunResult added in v1.12.0

type RerunResult struct {
	// Verdict is the empirical classification (see the Verdict constants).
	Verdict Verdict
	// Passes and Fails are the observed counts across the reruns that
	// produced a result. Errored attempts (Executor returned a non-nil error)
	// are counted in neither — they are tallied in Errors instead.
	Passes int
	Fails  int
	// Errors is the number of attempts that could not be observed (Executor
	// returned an error). A campaign that errors on every attempt is
	// Inconclusive.
	Errors int
	// Runs is the number of attempts actually made. With early exit it can be
	// fewer than the requested N (we stop as soon as a mixed pass+fail proves
	// flakiness), so callers can report "confirmed after k of N runs".
	Runs int
}

RerunResult is the outcome of a sandbox rerun campaign for one test.

func Rerun added in v1.12.0

func Rerun(ctx context.Context, exec Executor, testID string, runs int) RerunResult

Rerun re-runs a single flagged test up to runs times via exec and classifies the result. It is a pure function of exec's behaviour: given the same exec it returns the same verdict, so it is trivially testable with a fake.

Contract:

  • mixed pass+fail → VerdictFlaky (confirmed)
  • all observed fail → VerdictRealFailure (genuine red — do not quarantine)
  • all observed pass → VerdictTransient (did not reproduce)
  • nothing observed → VerdictInconclusive

Early exit: as soon as BOTH a pass and a fail have been seen the verdict can no longer change (flaky is terminal), so Rerun stops and returns without burning the remaining attempts. Errored attempts never short-circuit the loop — a flickering runner should not be mistaken for a flaky test — but they are recorded so a wholly-unobservable campaign is reported honestly.

runs <= 0 yields an Inconclusive result with no calls to exec. A nil exec is treated the same way (defensive: the seam is unbound), so the orchestrator is safe to call even when the caller never bound a runner.

func (RerunResult) Confirmed added in v1.12.0

func (r RerunResult) Confirmed() bool

Confirmed reports whether the rerun empirically confirmed flakiness. It is the predicate a caller uses to decide whether to keep treating the candidate as flaky after the rerun (true) or to step back from the static finding (a real failure or a non-reproducing transient).

type Verdict added in v1.12.0

type Verdict string

Verdict is the empirical classification produced by re-running a flagged test in isolation. It is intentionally a small closed vocabulary so the rendered/​localized text and any future machine surface stay stable.

const (
	// VerdictFlaky means the reruns produced BOTH a pass and a fail: the test
	// is non-deterministic under identical conditions — confirmed flaky.
	VerdictFlaky Verdict = "flaky"
	// VerdictRealFailure means every rerun failed: the test fails
	// deterministically, so it is a genuine failure, not flakiness. Such a
	// test must NOT be quarantined as flaky — it is correctly red.
	VerdictRealFailure Verdict = "real_failure"
	// VerdictTransient means every rerun passed: the original signal did not
	// reproduce. The flake (if any) was a one-off / already resolved.
	VerdictTransient Verdict = "transient"
	// VerdictInconclusive means no rerun was observed (runs <= 0, or every
	// attempt errored before yielding a result), so no classification can be
	// made. The caller should fall back to the static finding alone.
	VerdictInconclusive Verdict = "inconclusive"
)

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL