README
¶
tp — Task Plan
Spec-to-task lifecycle manager for AI coding agents.
Break specs into atomic, dependency-ordered tasks. Agents execute them with 2 tool calls instead of hundreds.
tp's primary user is the AI coding agent, not the human. Every command, flag, and output is designed for AX (Agent Experience): minimal round-trips, minimal output tokens, deterministic behavior (no prompts), and actionable error hints. The human authors specs and approves releases; the agent plans and the tool executes.
Why
AI agents fail at long tasks. Research shows:
- <15 min tasks: 70%+ success (SWE-bench)
- >50 min tasks: ~23% success (SWE-bench Pro)
- Each tool call costs ~200 tokens of agent context
tp solves this with atomic task decomposition and a 2-call architecture:
tp plan --minimal --json # ONE call: get execution plan
# [agent implements each task, commits each one]
tp done --batch results.ndjson # ONE call: close everything
Token overhead: ~5K (vs ~54K with naive per-task tool calls).
Install
# Homebrew (recommended)
brew tap deligoez/tap
brew install tp
# Go install
go install github.com/deligoez/tp/cmd/tp@latest
# Or build from source
git clone https://github.com/deligoez/tp.git
cd tp && go build -ldflags="-s -w" -o tp ./cmd/tp
# Install Claude Code skill (first time)
npx skills add -g deligoez/tp
# Update skill (after tp updates)
npx skills update -g deligoez/tp
Quick Start
# 1. Create a task file from a spec
tp init spec/my-feature.md
# 2. Add tasks (or use tp import for bulk)
tp add '{"id":"create-model","title":"Create User model","estimate_minutes":8,
"acceptance":"Model exists. Migration runs.","source_sections":["### User Model"],
"source_lines":"15-42","depends_on":[]}'
# 3. Get the execution plan
tp plan --minimal --json
# 4. Implement, commit, and close each task
tp done create-model "User model at app/Models/User.php. Migration runs." --gate-passed --auto-commit
Commands
Primary Workflow
tp plan # Full execution plan (THE primary command)
tp plan --minimal # Minimal: id + acceptance only (~80% fewer tokens)
tp plan --compact # Stripped: no description, source_lines, tags (~40% fewer)
tp plan --from <id> # Start from a specific task onward
tp plan --level 0,1 # Filter by parallelism levels (multi-agent)
tp commit <id> [reason] # Stage + structured commit + record SHA
tp commit <id> --files "*.go" # Selective file staging
tp done <id> <reason> # Close with implicit claim + verification; runs the quality gate
tp done <id> --skip-gate "why" # Skip gate execution, record gate_skipped_reason (needs user approval)
tp done <id> --auto-commit # Stage + commit + close in one call
tp done <id> --auto-commit --files src/engine/*.go # Selective staging + commit + close
tp done <id> --covered-by <id> # Close as covered by another done task
tp done <id> --commit <sha> # Record implementing commit SHA
tp done id1 id2 id3 "reason" # Multi-ID close (shared reason)
tp done --batch file.ndjson # Batch close from NDJSON
Incremental (fallback)
tp next # Resume WIP or claim next ready
tp next --minimal # Minimal output: {id, acceptance} only
tp next --peek # Preview without claiming
Task State
tp claim <id> [id...] # open → wip (batch: multiple IDs)
tp claim --all-ready # Claim all ready tasks at once
tp close <id> <reason> # wip → done (low-level, prefer tp done)
tp reopen <id> # done → open (clears timestamps + SHA)
tp remove <id> # Remove task (--force for dep cleanup)
tp set <id> field=value # Update field (managed fields protected)
tp set --workflow field=value # Update workflow-level fields (convergence params)
tp set --bulk sets.ndjson # Bulk update from NDJSON {id, field, value}
Query
tp list # All tasks
tp list --status open # Filter by status (open/wip/done)
tp list --tag api # Filter by tag
tp list --ids # IDs only
tp list --compact # Minimal fields
tp ready # Tasks with all deps satisfied
tp ready --first # First ready task only
tp ready --count # Count of ready tasks
tp ready --ids # Ready task IDs only
tp show <id> # Full details + spec_excerpt + blocks
tp status # Progress summary (open/wip/done counts)
tp blocked # Tasks waiting on unsatisfied deps
tp graph # Dependency tree
tp graph --tag api # Filter by tag
tp graph --from <id> # Subtree from a task
tp stats # Parallelism analysis
tp report # Per-task duration + estimation accuracy
Spec & Validation
tp lint spec.md # Spec quality + structured element detection
tp review spec.md # Adversarial review prompts (3 personas)
tp review spec.md --perspective code-audit --affected-files src/a.go # Code audit with source files
tp review spec.md --round 2 --findings r1.ndjson # Multi-round with previous findings
tp review spec.md --round 2 --final-round --affected-files src/a.go # Final round: mandatory code read-through
tp review --merge r1.ndjson r2.ndjson -o merged.ndjson # Merge + dedup findings
tp review --resolve findings.ndjson 3 fixed "evidence" # Mark finding as fixed/wontfix/duplicate
tp review --resolve-all findings.ndjson wontfix "reason" # Mark all unresolved findings
tp review --verify spec.md --findings all.ndjson # Lightweight verification (verifier role)
tp review --report r1.ndjson r2.ndjson # Cross-round convergence report
tp review spec.md --diff-from spec-r0.md # Diff-based review (changed sections only)
tp review spec.md --spec-inline # Embed full spec inline (default: reference mode)
tp review --resolve ... --force # Force re-resolve already resolved findings
tp review spec.md --record merged.ndjson # Record a review round (auto-numbered)
tp review spec.md --status --check # Convergence + mechanical checks; exit 0 when converged
tp review spec.md --perspective regression # Standalone regression delta pass
tp review spec.md --no-state # Disable state (pre-0.23.0 manual --round)
tp audit spec.md # Post-implementation: 3 role prompts, verify code matches spec
tp audit spec.md --affected-files src/a.go # Manual file selection
tp audit spec.md --findings review.ndjson # Also verify review findings
tp audit spec.md --record results.ndjson # Record an audit round (non-PASS = finding)
tp audit spec.md --status --check # Audit convergence; exit 0 when converged
tp validate # Task file + section/line coverage + atomicity (--strict)
Data
tp init spec.md # Create empty task file
tp add <json> # Add task (--stdin for piped input)
tp add --bulk tasks.ndjson # Bulk add from NDJSON
tp import file.json # Import + validate (--force to overwrite + relax atomicity)
tp import tasks.json --spec spec/feature.md # Import bare JSON array (auto-wraps)
tp use spec.tasks.json # Set active task file (writes .tp/local.json, git-ignored)
tp use --clear # Clear the active pointer
tp use # Show current active file
Global Flags
--file <path> Explicit task file path
--json Force JSON output (default when piped)
--compact Minimal JSON (~40% smaller); --no-compact forces full
--quiet Suppress info messages; --no-quiet forces info output
--no-color Disable colored output; --color forces color
Task File Discovery
tp finds your .tasks.json automatically:
--fileflag (highest priority)TP_FILEenvironment variable.tp/local.jsonactive pointer (set withtp use)- Legacy
.tp-activemarker (deprecated; removed in v0.25.0) - Auto-detect: scans current directory, then one level of subdirectories
# Set once, use everywhere
export TP_FILE=spec/project.tasks.json
Project Configuration
Multi-spec repos share one workflow policy instead of copying it into every *.tasks.json — so an agent working across specs reads a single source of truth and can't silently drift. A repo-root .tp/ directory holds it:
.tp/config.json(commit to VCS) — shared workflow defaults:quality_gate,review_clean_rounds,audit_clean_rounds,gate_timeout_seconds,review_max_rounds,audit_max_rounds,checks..tp/local.json(git-ignored automatically) — per-checkout state: theactivetask-file pointer (tp use) and CLI flagdefaults..tp/.gitignore— written automatically soconfig.jsonis tracked andlocal.jsonis not.
Discovery walks up from the current directory to the .git boundary to find .tp/ — a single, deterministic anchor the agent never has to disambiguate.
Layered resolution (resolve-at-read)
A task file's workflow block holds only explicit overrides; effective values merge at read time (never materialized, so nothing drifts). Precedence, highest first:
CLI flag > environment > task-file workflow override > .tp/config.json > built-in default
Absent ≠ zero: a field counts as an override only when actually present, so .tp/config.json fills every gap a task file leaves. checks uses replace semantics (the winning layer's array wins whole).
Commands
tp config # Effective configuration as JSON (what the agent will actually run)
tp config --resolved # Annotate each setting with its {value, source} layer
tp config --extract # Hoist policy shared by ALL task files into .tp/config.json
tp config --extract --dry-run # Preview the hoist plan without writing
tp config --extract --force # Merge into an existing .tp/config.json
tp set --workflow --project review_clean_rounds=3 # Edit a project-level workflow field
tp set --local defaults.compact=true # Set a CLI flag default (compact/quiet/no_color)
tp validate --project # Report cross-spec workflow drift (informational; --strict → exit 1)
Negating flags override a defaults entry for a single run: --no-compact, --no-quiet, --color.
Task File Format
Tasks live in a JSON file alongside the spec:
spec/
my-feature.md # spec (source of truth)
my-feature.tasks.json # tasks (derived, git-tracked)
Each task is atomic — one commit, one verb, ≤15 minutes:
{
"id": "create-model",
"title": "Create User model",
"status": "open",
"estimate_minutes": 8,
"acceptance": "Model exists. Migration runs.",
"depends_on": [],
"source_sections": ["### User Model"],
"source_lines": "15-42"
}
source_sections entries use canonical heading form: "## Heading Text" (heading marker prefix +
space + heading text). tp import and tp add are lenient — plain text like "User Model" is
accepted and auto-normalized to canonical form when unambiguous (v0.22.0+). Use the full canonical
form when the same heading text appears at multiple levels (e.g. both ## Setup and ### Setup).
The task file's workflow section supports convergence parameters:
{
"workflow": {
"quality_gate": "go test ./... && golangci-lint run",
"gate_timeout_seconds": 600,
"review_clean_rounds": 2,
"audit_clean_rounds": 2,
"review_max_rounds": 0,
"audit_max_rounds": 0,
"checks": []
}
}
quality_gate— command run automatically attp done/tp close; read-only (author it attp init --quality-gate)gate_timeout_seconds(default: 600, range: 30-3600) — hard timeout for a single gate runreview_clean_rounds/audit_clean_rounds(default: 2, range: 1-10) — consecutive finding-free rounds required before decomposition / after implementationreview_max_rounds/audit_max_rounds(default: 0 = no cap, range: 0-50) — round budget: at the cap while still unconverged,tp review/tp auditprompt generation and--recordrefuse with exit 4 and an escalation hint (raising the cap is a user-approved decision)checks— array of{class, cmd}mechanical detectors, replace semantics (see Mechanical checks & finding class)- Set via
tp set --workflow review_clean_rounds=3(managed fields likequality_gatestay read-only), or during the skill's interview phase
Acceptance Criteria Delimiters
tp parses acceptance criteria using three delimiters:
| Delimiter | Example |
|---|---|
Period + space (. ) |
"Model exists. Migration runs." |
Semicolon + space (; ) |
"Model exists; migration runs" |
Bullet list (\n- ) |
"- Model exists\n- Migration runs" |
JSON arrays are also accepted and joined with \n- on import.
JSON Field Aliases
deps is accepted as shorthand for depends_on:
{"id": "api", "deps": ["model"], ...}
Closure Verification
tp prevents lazy task closure with a deterministic, language-agnostic rule. Every tp done and tp close:
- Evidence lines: for a task with N ≥ 2 acceptance criteria, the reason must contain ≥ N lines each starting with
-at column 0 (indented sub-bullets do not count) — one top-level evidence line per criterion. A single-criterion task accepts any non-empty reason. The error enumerates each parsed criterion. - Forbidden patterns: rejects "deferred", "will be done later", single-word reasons, and "covered by existing" without a path.
# This fails (3 criteria, no evidence lines):
tp done create-model "done"
# This passes (one "- " line per criterion; -- separates the reason from flags):
tp done create-model -- "- User model at app/Models/User.php:18
- migration 0007 applied, schema verified
- go test ./... green"
Automatic quality gate
When workflow.quality_gate is set, tp done/tp close run the command automatically (once per invocation) before closing; a failing gate blocks the close (exit 4) and no task closes. --gate-passed is ignored when a gate is configured (it only records an attestation on gate-less projects). --skip-gate "why" skips execution and records gate_skipped_reason — a user-approved escape hatch, never the agent's own decision.
# --covered-by: task satisfied by another done task (not a deferral)
tp done qa-delegation "test #26 covers this" --covered-by qa-tests
Structured Commits
tp commit generates conventional commit messages with task metadata:
tp commit auth-model "Model and migration created"
feat(auth-model): Create User model
Model and migration created
Task: auth-model
Acceptance: Model exists. Migration runs.
Or commit + close in one call:
tp done auth-model "evidence" --gate-passed --auto-commit
Spec Quality
tp lint detects structured elements (tables, numbered lists, code blocks) for decomposition verification:
tp lint spec.md --json | jq .structured_elements
tp review generates adversarial review prompts that agents feed to sub-agents:
# Default: 3 prompts (implementer, tester, architect)
tp review spec.md --json | jq '.prompts | length'
# → 3
# Perspective-specific review
tp review spec.md --perspective code-audit --affected-files src/a.go --json
# Documentation perspective
tp review spec.md --perspective documentation --json
# Testing perspective
tp review spec.md --perspective testing --json
For multi-round review, use --round and --findings to auto-exclude previously reported issues:
# Round 1: generate prompts, spawn sub-agents, collect findings to findings.ndjson
tp review spec.md --json > review-r1.json
# Round 2: tp auto-injects findings summary, prompts focus on new issues only
tp review spec.md --round 2 --findings findings.ndjson --json > review-r2.json
Code-Aware Review
Inject source files into review prompts to catch state-dependent behaviors that specs miss:
# Inject files into default review — each prompt gets file content + checklist
tp review spec.md --affected-files src/form.vue src/api.ts
# Code audit perspective: C1-C5 checklist (state-dependent behaviors, spec coverage, etc.)
tp review spec.md --perspective code-audit --affected-files src/form.vue
# Final round: force mandatory code read-through to prevent false convergence
tp review spec.md --round 2 --final-round --affected-files src/form.vue
Files are capped at 8000 chars each (50000 total). Prompt budget enforced at 60000 chars total.
Multi-Round Review Workflow
# R1: generate review prompts, spawn sub-agents, collect findings
tp review spec.md # R1: generate review prompts
# Merge findings from multiple sub-agents
tp review --merge r1-*.ndjson -o r1.ndjson # Merge + dedup findings
# Resolve individual findings
tp review --resolve r1.ndjson 3 fixed "evidence" # Mark finding as fixed
# Record the round: tp owns the `.tp-review/` state directory (commit it to VCS), auto-numbers rounds, and
# injects previous findings + the changed-sections diff into R2 automatically
tp review spec.md --record r1.ndjson # record round 1
tp review spec.md # R2: auto diff + findings injected
# Lightweight verification pass
tp review --verify spec.md --findings all.ndjson # Lightweight verification
# Convergence is a recorded fact — loop until this exits 0
tp review spec.md --status --check # converged AND all checks pass?
| Flag | Purpose |
|---|---|
--merge |
Merge and dedup findings from multiple NDJSON files |
--resolve |
Mark a finding as fixed/wontfix/duplicate |
--resolve-all |
Mark all unresolved findings at once |
--verify |
Lightweight verification prompt (single prompt, verifier role) |
--report |
Cross-round convergence report |
--spec-inline |
Embed full spec inline (default: reference by path) |
--diff-from |
Diff-based review (only changed sections inline) |
-o / --output |
Output file path for merge |
--force |
Force re-resolve already resolved findings |
--record <file> |
Record a review round (auto-numbered R; freezes count + clean flag) |
--status / --status --check |
Show convergence state / gate exit 0 on converged + passing checks |
--perspective regression |
Standalone regression pass guarding settled decisions |
--no-state |
Disable state reads/writes (pre-0.23.0 manual --round numbering) |
Mechanical checks & finding class
Review findings may carry an optional class — a kebab-case slug naming a pattern a script could detect across the whole spec. When a class recurs (≥ 2 distinct rounds, or ≥ 5 times in one round), tp review --report and --record surface it under mechanize_candidates, alongside a by_class breakdown. Turn a recurring class into a permanent detector:
tp set --workflow checks='[{"class":"code-citation-drift","cmd":"scripts/check-citations.sh"}]'
tp then runs every registered check at the start of each review round, reports pass/fail under mechanical_checks, and tells reviewers to stop hand-reporting that class. tp review --status --check exits 0 only when the review is converged and every check passes. (checks uses replace semantics — one tp set --workflow call sets the whole array.)
Spec frontmatter (tp: domain & lens)
A spec can open with a YAML frontmatter block; tp reads only the tp: mapping (line numbers stay absolute, and the block is excluded from every parser):
---
tp:
domain: prose # default "software"; only "software" enables software-specific prompts
lens:
all: ["Does any chapter summary leak a plot point ahead of its chapter?"]
implementer: ["Can each section be written without inventing facts not in the outline?"]
tester: []
architect: []
---
- A non-
softwaredomainswaps the three review personas and drops the software-specific questions (error-handling, backward-compatibility, performance) — sotp reviewfits prose, legal, or research specs. lensquestions inject into the matching role (allgoes to every role plus the regression pass).tp lintreports afrontmatterobject and warns on malformed YAML or unknown lens keys (which are ignored).tp auditkeeps its three fixed roles regardless ofdomain.
Lint Checks
tp lint detects structured elements and quality issues:
tp lint spec.md --json | jq '.findings[] | select(.rule)'
| Rule | Severity | What it checks |
|---|---|---|
structured-elements |
info | Tables, numbered lists, code blocks in spec |
acceptance-quality |
warning/info | Removal-only acceptance, vague verbs, short acceptance |
affected-files-scope |
warning | Modify rows in affected files table without scope description |
duplicate-line |
warning | Consecutive identical non-empty lines (edit artifacts) |
numbering-gap |
warning | Gaps in numbered section headings (e.g., 4.1 → 4.3, missing 4.2) |
orphan-list-item |
info | Numbered lists starting at >1 or with gaps (e.g., 1, 3 — missing 2) |
tp validate checks line coverage — verifying that task source_lines cover the entire spec:
tp validate --json | jq .checks.line_coverage
Post-Implementation Audit
tp audit verifies that the spec's requirements actually made it into the code:
# Auto-detect changed files via git diff (zero-config)
tp audit spec.md --json
# Manual file selection
tp audit spec.md --affected-files src/form.vue src/api.ts
# Also verify review findings were addressed
tp audit spec.md --findings findings.ndjson
The command parses the spec's structured elements (table rows, numbered lists), task acceptance criteria, and optionally review findings, then emits one prompt per non-empty role — spec-coverage, security, maintainability-conventions (v0.23.0). Each prompt carries an embedded JSON-array checklist and its per-role affected files; sub-agents return one NDJSON row per checklist item (status ∈ PASS/PARTIAL/FAIL). Record rounds and converge like review:
tp audit spec.md --json | jq '.prompts[].role'
# → "spec-coverage" "security" "maintainability-conventions"
tp audit spec.md --record results.ndjson # non-PASS rows count as findings
tp audit spec.md --status --check # exit 0 only when the audit is converged
Schema break: v0.23.0 audit JSON is incompatible with v0.22.0 —
roleis one of the three above (wasimplementation-auditor), thecategoryfield is removed, andchecklist_items/affected_filesare new. Downstream consumers must update; there is no--legacy-formatflag.
AX (Agent Experience)
tp is designed for AI agents first (AX), not humans (DX):
| Principle | How |
|---|---|
| Minimal tokens | --minimal ~80%, --compact ~40% smaller. 2-call architecture saves ~90% |
| Batch parity | tp claim --all-ready, tp done --batch, tp set --bulk |
| Dependency-aware batch | tp done --batch auto-toposorts by in-batch deps — no manual ordering needed |
| Actionable errors | Every error includes hint field with recovery action |
| Did-you-mean | --covered-by typos suggest similar task IDs |
| Structured commits | tp commit generates conventional commit messages with task metadata |
| Implicit claim | tp done and tp commit auto-claim open tasks |
| WIP resume | tp next returns existing WIP task (crash recovery) |
| Covered-by | Close tasks covered by other tasks without duplicate work |
| Auto-normalize | source_lines accepts "72" (normalized to "72-72") |
| Import flexibility | tp import accepts bare JSON arrays with --spec flag |
| Spec-only review | Review prompts include disclaimer to prevent code-checking |
| Edit hygiene lint | tp lint detects duplicate lines and numbering gaps |
| Estimation calibration | tp add warns when historical estimates are consistently high |
| Duration tracking | tp report shows per-task timing and estimation accuracy |
Claude Code Integration
tp ships with a Claude Code skill via the Agent Skills standard:
# Install skill (first time)
npx skills add -g deligoez/tp
# Update skill (after tp updates)
npx skills update -g deligoez/tp
The skill teaches Claude the 2-call workflow, decomposition rules, NDJSON format, closure verification, and commit conventions automatically.
Research
tp's design is backed by:
| Finding | Source |
|---|---|
| <15 min tasks = 70%+ success | SWE-bench |
| ACI design: 3-4x improvement | SWE-agent, Princeton |
| Planning: 9.85% → 57.58% success | Plan-and-Act |
| 100:1 input-to-output token ratio | Manus |
| 64% token reduction with upfront planning | ReWOO |
See spec/0.1.0.md for the full specification with 22 research references.
License
MIT