mdidx

module
v0.0.0-...-7f52054 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: May 19, 2026 License: MIT

README

mdidx

mdidx is a project-oriented semantic search system for Markdown files.

Features

  • Maintains hybrid search using both vectors and BM25 over a corpus of text files
  • Targets ~100MB or less of Markdown documents
  • Discovers projects by walking upward until it finds mdidx.toml
  • Agent-friendly, self-describing CLI with detailed help
  • Emits an installable SKILL.md for agent skill installers
  • Embeddings generated using any OpenAI-compatible APIs
  • Can be run as a watcher to automatically update the index
  • Builds as a single binary, index is a single file

Implementation

  • Written in Go
  • Built on SQLite in WAL mode using FTS5 and sqlite-vec vector search
  • Adds an adaptive IVF-flat coarse vector index for larger corpora while keeping exact full-vector scoring inside probed cells
  • Knows when files are updated and re-indexes incrementally
  • Includes benchmarks to show performance at target scale

Quick start

git clone https://github.com/rbranson/mdidx.git
cd mdidx
make build
./bin/mdidx init
./bin/mdidx index
./bin/mdidx search "deployment checklist"

mdidx uses sqlite-vec through cgo. Use the Makefile for local builds and tests; it enables FTS5 for the bundled SQLite build with CGO_CFLAGS=-DSQLITE_ENABLE_FTS5.

mdidx init creates this default project config:

[index]
paths = ["."]
include = ["**/*.md", "**/*.markdown"]
exclude = [".git/**", "node_modules/**", "vendor/**", ".mdidx/**"]
chunk_target = 1000
chunk_overlap = 150
concurrency = 4

[database]
path = ".mdidx/index.db"

[embedding]
base_url = "https://api.openai.com/v1"
model = "text-embedding-3-small"
batch_size = 64
request_timeout = "1m"
max_retries = 3

[search]
lexical_weight = 0.5
vector_weight = 0.5
candidate_limit = 100

[watch]
first_trigger_delay = "5s"
min_update_interval = "1m"

Set OPENAI_API_KEY for OpenAI-compatible embedding APIs that require authentication. OPENAI_BASE_URL and MDIDX_EMBEDDING_MODEL override the configured endpoint and model.

Commands

mdidx init [--force] [path]
mdidx index [--json] [--concurrency 4] [path...]
mdidx search [--json] [--limit 10] "query"
mdidx watch
mdidx stats [--json]
mdidx doctor [--json]
mdidx config show [--json]
mdidx config paths [--json]
mdidx eval [--json] file.md [distractor.md...]
mdidx version [--json]
mdidx skill emit [--output ~/.codex/skills/mdidx/SKILL.md]

Run mdidx --help or mdidx <command> --help for the full command, flag, config, environment, and output reference.

Installation

Install from the module path:

go install github.com/rbranson/mdidx/cmd/mdidx@latest

Install from a local checkout:

git clone git@github.com:rbranson/mdidx.git
cd mdidx
make install

Build a local binary without installing it:

make build
./bin/mdidx --help
./bin/mdidx version

Build, install, and dist targets stamp version metadata from git when available. Unstamped development builds report dev and unknown values.

Install the agent skill emitted by the binary:

mkdir -p ~/.codex/skills/mdidx
mdidx skill emit --output ~/.codex/skills/mdidx/SKILL.md

Development

Common development targets are collected in the Makefile:

make help          # list targets
make deps          # download and verify modules
make fmt           # format Go source
make vet           # run go vet
make test          # run tests
make check         # fmt, vet, and test
make build         # build ./bin/mdidx
make run ARGS='--help'
make run ARGS='eval README.md --generate-only'
make bench         # benchmark smoke test
make quality-classic # live retrieval quality eval over classic-books-markdown
make integration   # live OpenAI integration test
make clean         # remove build artifacts

Before committing, run:

make check

For a larger local benchmark, set MDIDX_BENCH_MB:

MDIDX_BENCH_MB=100 BENCH_TIME=1x make bench

Retrieval quality evaluation over the classic-books-markdown corpus is opt-in and uses live query embeddings, so it requires a corpus checkout that is already an initialized and indexed mdidx project:

MDIDX_CLASSIC_BOOKS_PATH=~rbranson/src/github.com/mlschmitt/classic-books-markdown OPENAI_API_KEY=... make quality-classic

The eval does not build or update the index. Run mdidx init and mdidx index in the corpus root first. Re-running the eval only searches the existing index, plus optional LLM judging when enabled.

The quality evaluation generates title/author cases from every discovered book, samples passage-level phrase cases from the corpus, and reads thematic cases from internal/core/testdata/classic_quality/themes.json. Set MDIDX_QUALITY_MAX_TITLE_CASES or MDIDX_QUALITY_MAX_PHRASE_CASES to control case volume. Set MDIDX_QUALITY_JUDGE_MODEL to use a chat-completions-compatible LLM as a judge for thematic queries such as "kissing while crying" or "walk through the park".

Live integration tests are opt-in and make real OpenAI-compatible embedding API calls. They require OPENAI_API_KEY and may incur API cost:

OPENAI_API_KEY=... make integration

OPENAI_BASE_URL and MDIDX_EMBEDDING_MODEL are honored when set.

Release-style binaries for common Darwin and Linux targets can be built with:

make dist

Watch mode

mdidx watch uses filesystem notifications through fsnotify, which maps to inotify on Linux and native watchers elsewhere. The first event after idle waits 5 seconds before indexing. After an update runs, later events are coalesced so at most one index update runs per minute by default. The watcher also runs one initial index update when it starts.

Operations

Run a local health check before relying on a project index:

mdidx doctor
mdidx doctor --json

doctor checks project discovery, config validity, index paths, Markdown discovery, SQLite database readiness, embedding configuration, and whether OPENAI_API_KEY is present. It does not make live network calls.

Inspect effective configuration and resolved paths:

mdidx config show --json
mdidx config paths --json

config show includes environment overrides such as OPENAI_BASE_URL and MDIDX_EMBEDDING_MODEL, but never prints API keys. It reports API key presence as a boolean.

mdidx stats --json reports vector_index: "ivf-flat" when a large index has trained coarse vector cells. Smaller indexes search the sqlite-vec table directly to avoid extra indexing overhead.

For automation, use --json success output and parse structured JSON errors from stderr. Commands invoked with --json keep stdout reserved for successful data and emit errors shaped like:

{"error":{"code":"missing_config","message":"...","hint":"...","exit_code":2}}

Large indexing runs can emit progress to stderr:

mdidx index --progress
mdidx index --json --no-progress

Progress never writes to stdout. Use --no-progress for scripts that expect quiet stderr except on errors.

Embedding calls use configurable batching and the OpenAI Go SDK retry policy:

[embedding]
batch_size = 64
request_timeout = "1m"
max_retries = 3

request_timeout bounds each embedding request attempt. max_retries is passed to the OpenAI Go SDK, which retries transient request failures such as rate limits, timeouts, and 5xx responses. Authentication and invalid request errors fail immediately.

Indexing changed files runs concurrently according to index.concurrency. SQLite updates are still written serially after embeddings complete, preserving the project-local database lifecycle.

Retrieval evals

Use mdidx eval to build a repeatable retrieval eval from one or more representative Markdown files:

mdidx eval docs/sample.md
mdidx eval docs/sample.md --json
mdidx eval docs/sample.md docs/nearby-topic.md

The command uses an LLM, gpt-5.4-mini by default, to generate about 100 search queries from the files' indexed sections, saves them under .mdidx/evals/, indexes the eval files, runs each query through mdidx, and then uses the same model only as a relevance judge. Single-file evals search only that source file. Multi-file evals search the project index so the other eval files and any already-indexed Markdown act as distractors. Each pass writes a timestamped run file under .mdidx/evals/runs/ with queries, retrieved snippets, judgments, and aggregate hit@1, hit@3, hit@10, MRR, average judge score, rank-weighted score, nDCG@10, and bucketed metrics.

Generated query sets intentionally mix short keyword searches, medium natural project-search phrases, longer detailed clue searches, and query-type buckets such as exact, paraphrase, ambiguous, narrative, heading-dependent, and confusable cases so retrieval changes are tested against varied user behavior.

Eval query generation, searches, and judge batches run with bounded concurrency by default. Use --concurrency N to tune parallel API calls for your rate limits.

Keep the query set stable while trying retrieval strategy changes, then compare the run JSON files:

mdidx eval docs/sample.md --queries .mdidx/evals/sample.queries.json
mdidx eval docs/sample.md --run-name lexical-heavy --lexical-weight 0.8 --vector-weight 0.2

Use --regenerate to refresh the query set, --generate-only to create the query set without running searches, and --skip-index to evaluate the existing index exactly as-is. Progress is printed to stderr for terminal runs; use --progress to force it for scripted runs or --no-progress to suppress it.

Benchmarks

go test -bench . ./...
MDIDX_BENCH_MB=100 go test -bench BenchmarkIndexAndSearchGeneratedCorpus ./internal/core
MDIDX_CLASSIC_BOOKS_PATH=/path/to/classic-books-markdown OPENAI_API_KEY=... make quality-classic

License

MIT. See LICENSE.

Directories

Path Synopsis
cmd
mdidx command
internal
app
Package app is the CLI boundary for mdidx.
Package app is the CLI boundary for mdidx.
config
Package config owns project discovery and the mdidx.toml contract.
Package config owns project discovery and the mdidx.toml contract.
core
Package core coordinates the user-visible indexing and search workflows.
Package core coordinates the user-visible indexing and search workflows.
embed
Package embed implements the small OpenAI-compatible surface mdidx needs.
Package embed implements the small OpenAI-compatible surface mdidx needs.
eval
Package eval runs repeatable retrieval evaluations.
Package eval runs repeatable retrieval evaluations.
index
Package index discovers Markdown files and converts them into stable chunks.
Package index discovers Markdown files and converts them into stable chunks.
store
Package store contains the SQLite persistence and ranking implementation.
Package store contains the SQLite persistence and ranking implementation.
watch
Package watch keeps a project index fresh without running an index update for every filesystem event.
Package watch keeps a project index fresh without running an index update for every filesystem event.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL