README
¶
skill-quality-auditor
A 9-dimension scoring framework for auditing and improving AI skill quality. Combines structural validation with custom scoring across Knowledge Delta, Mindset, Anti-Patterns, Specification Compliance, Progressive Disclosure, Freedom Calibration, Pattern Recognition, Practical Usability, and Eval Validation.
- What it scores & why
- Dimension docs
- Install
- Command Usage
- Flag Shorthands
- Output Formats
- CI Integration
- Repository layout
- Development
What it scores & why
| ID | Dimension | Max | What a low score signals |
|---|---|---|---|
| D1 | Knowledge Delta | 20 | Content restates what the model already knows — no expert uplift |
| D2 | Mindset & Procedures | 15 | Missing mental models or step-by-step guidance the agent needs |
| D3 | Anti-Pattern Coverage | 15 | Common failure modes not called out — agent will repeat them |
| D4 | Specification Compliance | 15 | Frontmatter, structure, or naming deviates from the tile spec |
| D5 | Progressive Disclosure | 15 | Detail is front-loaded; references not used for depth |
| D6 | Freedom Calibration | 15 | Skill is either too prescriptive or too vague for the task |
| D7 | Pattern Recognition | 10 | No trigger conditions — agent won't know when to activate the skill |
| D8 | Practical Usability | 15 | Examples absent or unrealistic; hard to apply in practice |
| D9 | Eval Validation | 20 | No evals — quality claims are unverifiable |
Total: 140 pts. Grade bands:
| Grade | Score |
|---|---|
| A+ | ≥ 133 |
| A | ≥ 126 |
| B+ | ≥ 119 |
| B | ≥ 112 |
| C+ | ≥ 105 |
| C | ≥ 98 |
| D | ≥ 91 |
| F | < 91 |
See cmd/assets/references/quality-thresholds-scoring.md for the full rubric and per-dimension
scoring criteria.
Dimension docs
Each dimension has a dedicated doc with scoring criteria, examples, and academic references:
| Dimension | Doc |
|---|---|
| D1 Knowledge Delta | docs/d1-knowledge-delta.md |
| D2 Mindset & Procedures | docs/d2-mindset-procedures.md |
| D3 Anti-Pattern Coverage | docs/d3-anti-pattern-coverage.md |
| D4 Specification Compliance | docs/d4-specification-compliance.md |
| D5 Progressive Disclosure | docs/d5-progressive-disclosure.md |
| D6 Freedom Calibration | docs/d6-freedom-calibration.md |
| D7 Pattern Recognition | docs/d7-pattern-recognition.md |
| D8 Practical Usability | docs/d8-practical-usability.md |
| D9 Eval Validation | docs/d9-eval-validation.md |
Install
install.sh (Linux / macOS)
curl -fsSL https://raw.githubusercontent.com/pantheon-org/skill-quality-auditor/main/scripts/install.sh | sh
Override install directory or pin a version:
INSTALL_DIR=~/.local/bin curl -fsSL ... | sh
VERSION=v1.2.3 curl -fsSL ... | sh
mise
mise use github:pantheon-org/skill-quality-auditor
Or in mise.toml:
[tools]
"github:pantheon-org/skill-quality-auditor" = "latest"
Go install
go install github.com/pantheon-org/skill-quality-auditor@latest
Updating
| Method | Command |
|---|---|
| install.sh | skill-auditor update |
| mise | mise upgrade skill-auditor |
| Go install | go install github.com/pantheon-org/skill-quality-auditor@latest |
skill-auditor updatealso accepts--check(report without installing) and--version-target vX.Y.Z.
Command Usage
| Stage | Command | What it answers |
|---|---|---|
| Score a skill | evaluate |
What is the overall quality grade and per-dimension breakdown? |
| Score many skills | batch |
How do multiple skills compare, and does any fall below a CI threshold? |
| Find overlap | duplication |
Are any skills too similar to each other? |
| Plan consolidation | aggregate |
How should a family of similar skills be merged? |
| Fix a skill | remediate |
What specific changes would raise this skill's score? |
| Track progress | trend |
Are scores improving or regressing over time? |
| Validate format | validate |
Do artifacts conform to conventions? Does a review report meet spec? |
| Deep analysis | analyze |
What are the keyword signals and structural patterns in this skill? |
| Install skill | init |
How do I install this auditor skill into my agent environment? |
| Self-update | update |
Is a newer release available, and can I install it in place? |
| Housekeeping | prune |
Which old audit snapshots can be removed? |
evaluate
skill-auditor evaluate <skill> [flags]
Flags:
-m, --markdown emit Markdown output (default: JSON)
-s, --store persist result to .context/audits/
-r, --repo-root repo root directory (auto-detected from .git / go.mod if omitted)
<skill> accepts a domain/skill-name key (resolved under <repo-root>/skills/), a directory
containing SKILL.md, or a direct path to SKILL.md.
batch
skill-auditor batch <skill1> [skill2 ...] [flags]
Flags:
-j, --json emit JSON array output
-m, --markdown emit Markdown table output
-s, --store persist each result to .context/audits/
-F, --fail-below exit 1 if any skill scores below this grade (e.g. B+)
-r, --repo-root repo root directory (auto-detected if omitted)
duplication
skill-auditor duplication [skills-dir] [flags]
Flags:
-j, --json emit JSON array of pairs
-m, --markdown emit Markdown output (default)
-s, --store persist report to .context/analysis/
-d, --skills-dir skills directory (default: <repo-root>/skills)
-r, --repo-root repo root directory (auto-detected if omitted)
Pairwise word-level Jaccard similarity across all SKILL.md files. Writes
duplication-report-YYYY-MM-DD.md to .context/analysis/. Exits with code 2 on any
Critical (>35%) pair — suitable as a CI gate.
aggregate
skill-auditor aggregate --family <prefix> [skills-dir] [flags]
Flags:
-j, --json emit JSON output instead of Markdown
-n, --dry-run print plan to stdout without writing to disk
-f, --family skill family prefix to analyse (required, e.g. bdd, typescript)
-d, --skills-dir skills directory (default: <repo-root>/skills)
-r, --repo-root repo root directory (auto-detected if omitted)
Produces a 6-step consolidation plan at .context/analysis/aggregation-plan-<family>-YYYY-MM-DD.md.
remediate
skill-auditor remediate <skill> [flags]
Flags:
-j, --json emit the plan as JSON instead of Markdown
-n, --dry-run print plan to stdout without writing to .context/plans/
-t, --target-score desired total score (default: current + 20, max 140)
-v, --validate validate an existing plan file instead of generating one
-r, --repo-root repo root directory (auto-detected if omitted)
Reads the most recent stored audit for <skill> and generates a schema-compliant remediation
plan at .context/plans/<skill>-remediation-plan.md. Use --validate to check an existing plan
against remediation-plan.schema.json.
trend
skill-auditor trend [flags]
Flags:
-j, --json emit JSON array output
-m, --markdown emit Markdown table output (default)
-s, --store persist report to .context/audits/
-r, --repo-root repo root directory (auto-detected if omitted)
Reads the two most recent stored audits per skill from .context/audits/ and prints a score-delta
table with ↑ / ↓ / — indicators.
validate
skill-auditor validate artifacts [paths...] [flags]
skill-auditor validate review <file> [flags]
Flags (artifacts):
-r, --repo-root repo root directory (auto-detected if omitted)
Flags (review):
-S, --strict-recommended treat recommended fields as errors
-r, --repo-root repo root directory (auto-detected if omitted)
validate artifacts checks SKILL.md line limits, frontmatter name match, asset subdirectory
conventions, script shebangs, and schema file validity. validate review checks a review report
against the embedded requirements spec. Exit code 1 on any error.
analyze
skill-auditor analyze <skill> [flags]
Flags:
-e, --semantic run TF-IDF keyword extraction only
-p, --patterns run rule-based pattern detection only
-P, --pipeline run full pipeline — semantic + patterns + combined report (default)
-j, --json emit JSON output (default)
-m, --markdown emit Markdown output instead of JSON
-s, --store write report to .context/analysis/
-l, --limit int max keywords to include (default 20)
-r, --repo-root repo root directory (auto-detected if omitted)
--semantic extracts TF-IDF top keywords scored against the full skill corpus. --patterns runs
rule-based detectors for required sections, trigger-word frequency, structural conformance, and
anti-pattern signals. Default --pipeline runs both and writes a combined report.
init
skill-auditor init [flags]
Flags:
-n, --dry-run preview what would be created without touching disk
-a, --agent agent(s) to install into (default: auto-detect from current directory)
-g, --global install to global skill directory (~/<agent-path>/) instead of CWD
-I, --interactive choose agents interactively from the full registry list
-m, --method installation method: symlink or copy (default: symlink)
Installs the embedded skill-quality-auditor skill — SKILL.md plus all asset subdirectories
(references/, evals/, schemas/, templates/, requirements/) — into one or more agent skill
directories.
Detection behaviour:
- Default (no flags) — targets the current working directory; auto-detects any agent whose
harness root directory (e.g.
.claude/,.cursor/) already exists under CWD. --global/-g— targets the home directory instead; auto-detects against~.--agent/-a— skips auto-detection and installs into the specified agent(s) explicitly.--interactive/-I— shows a numbered list of all supported agents (*marks detected ones) and lets you choose by number or typeall.
update
skill-auditor update [flags]
Flags:
--check report the latest version without installing
--version-target install a specific version (e.g. v1.2.3)
Fetches the latest release from GitHub and replaces the running binary in-place. Only applicable
when installed via install.sh — Homebrew and mise users should use their own update commands.
prune
skill-auditor prune [flags]
Flags:
-n, --dry-run list audit runs that would be removed without deleting anything
-k, --keep number of audit date-dirs to retain per skill (default 5)
-r, --repo-root repo root directory (auto-detected if omitted)
Removes old date-stamped audit directories from .context/audits/, keeping the N most recent per
skill.
Flag Shorthands
All flags are available via long form and a one-letter shorthand:
| Flag | Short | Commands |
|---|---|---|
--json |
-j |
analyze, batch, duplication, trend, remediate, aggregate |
--markdown |
-m |
evaluate, analyze, batch, duplication, trend |
--store |
-s |
evaluate, analyze, batch, duplication, trend |
--repo-root |
-r |
evaluate, analyze, batch, duplication, trend, remediate, aggregate, prune, validate |
--dry-run |
-n |
aggregate, remediate, prune, init |
--skills-dir |
-d |
aggregate, duplication |
--family |
-f |
aggregate |
--fail-below |
-F |
batch |
--target-score |
-t |
remediate |
--validate |
-v |
remediate |
--keep |
-k |
prune |
--limit |
-l |
analyze |
--strict-recommended |
-S |
validate review |
--global |
-g |
init |
--agent |
-a |
init |
--interactive |
-I |
init |
--method |
-m |
init |
--json and --markdown are mutually exclusive. When neither is passed, commands default to
their natural format (JSON for evaluate, analyze, batch, remediate, aggregate;
Markdown for duplication, trend).
Output Formats
All commands that produce structured data support --json / -j and --markdown / -m.
The two flags are mutually exclusive — pass at most one. When neither is given, each command
uses its own default format (noted in the per-command flag tables above).
JSON output — pass --json to any command:
skill-auditor evaluate skills/my-skill # JSON (default)
skill-auditor evaluate skills/my-skill -m # Markdown
skill-auditor batch skills/skill-a skills/skill-b --json
skill-auditor trend --json
Stored output — pass --store to persist results for later use by remediate and trend:
| Command | Output path |
|---|---|
evaluate --store / batch --store |
.context/audits/<skill>/<date>/ |
duplication |
.context/analysis/duplication-report-YYYY-MM-DD.md |
aggregate |
.context/analysis/aggregation-plan-<family>-YYYY-MM-DD.md |
remediate |
.context/plans/<skill>-remediation-plan.md |
analyze --store |
.context/analysis/pattern-report-<skill>-YYYY-MM-DD.md |
CI Integration
- name: Audit skills
run: |
skill-auditor batch skills/ --fail-below B --store
skill-auditor duplication # exits 2 on Critical pairs
skill-auditor validate artifacts
--fail-below accepts any grade: A+, A, B+, B, C+, C, D, F.
duplication exits with code 2 (not 1) on Critical pairs so it can be distinguished from a
command error in pipeline logic.
Full workflow example:
jobs:
skill-quality:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Install skill-auditor
run: curl -fsSL https://raw.githubusercontent.com/pantheon-org/skill-quality-auditor/main/scripts/install.sh | sh
- name: Batch audit (fail below B)
run: skill-auditor batch skills/ --fail-below B --store
- name: Duplication check
run: skill-auditor duplication
continue-on-error: false # exits 2 on Critical pairs
- name: Artifact validation
run: skill-auditor validate artifacts
Repository layout
go.mod / main.go Go CLI root — build and run from here
cmd/ cobra commands: evaluate, batch, duplication, aggregate,
remediate, trend, validate, prune, analyze, init, update
cmd/assets/ Tessl tile — SKILL.md, tile.json, evals, references,
schemas, templates (single source of truth)
agents/ agent registry (supported environments for init)
docs/ per-dimension documentation with scoring criteria and references
scorer/ D1–D9 dimension scorers
analysis/ TF-IDF keyword extractor + rule-based pattern detectors
duplication/ word-level Jaccard similarity engine
internal/ shared utilities (tokenizer)
reporter/ text/JSON formatters, audit store, report generators
scripts/ install.sh
testdata/ fixture skills for unit tests
Development
go test ./...
go vet ./...
golangci-lint run ./...
shellcheck scripts/install.sh
Pre-commit and pre-push hooks are managed via lefthook:
mise install # installs go, golangci-lint, mdlint, shellcheck, lefthook
lefthook install
Documentation
¶
There is no documentation for this package.
Directories
¶
| Path | Synopsis |
|---|---|
|
Package agents provides the canonical registry of supported agent clients derived from https://github.com/vercel-labs/skills#supported-agents.
|
Package agents provides the canonical registry of supported agent clients derived from https://github.com/vercel-labs/skills#supported-agents. |
|
internal
|
|
|
This file owns aggregation plan formatting: building and rendering the markdown consolidation plan for a skill family.
|
This file owns aggregation plan formatting: building and rendering the markdown consolidation plan for a skill family. |