mdidx
mdidx is a project-oriented semantic search system for Markdown files.
Features
- Maintains hybrid search using both vectors and BM25 over a corpus of text files
- Targets ~100MB or less of Markdown documents
- Discovers projects by walking upward until it finds
mdidx.toml
- Agent-friendly, self-describing CLI with detailed help
- Emits an installable
SKILL.md for agent skill installers
- Embeddings generated using any OpenAI-compatible APIs
- Can be run as a watcher to automatically update the index
- Builds as a single binary, index is a single file
Implementation
- Written in Go
- Built on SQLite in WAL mode using FTS5 and sqlite-vec vector search
- Adds an adaptive IVF-flat coarse vector index for larger corpora while keeping
exact full-vector scoring inside probed cells
- Knows when files are updated and re-indexes incrementally
- Includes benchmarks to show performance at target scale
Quick start
git clone https://github.com/rbranson/mdidx.git
cd mdidx
make build
./bin/mdidx init
./bin/mdidx index
./bin/mdidx search "deployment checklist"
mdidx uses sqlite-vec through cgo. Use the Makefile for local builds and tests;
it enables FTS5 for the bundled SQLite build with CGO_CFLAGS=-DSQLITE_ENABLE_FTS5.
mdidx init creates this default project config:
[index]
paths = ["."]
include = ["**/*.md", "**/*.markdown"]
exclude = [".git/**", "node_modules/**", "vendor/**", ".mdidx/**"]
chunk_target = 1000
chunk_overlap = 150
concurrency = 4
[database]
path = ".mdidx/index.db"
[embedding]
base_url = "https://api.openai.com/v1"
model = "text-embedding-3-small"
batch_size = 64
request_timeout = "1m"
max_retries = 3
[search]
lexical_weight = 0.5
vector_weight = 0.5
candidate_limit = 100
[watch]
first_trigger_delay = "5s"
min_update_interval = "1m"
Set OPENAI_API_KEY for OpenAI-compatible embedding APIs that require authentication. OPENAI_BASE_URL and MDIDX_EMBEDDING_MODEL override the configured endpoint and model.
Commands
mdidx init [--force] [path]
mdidx index [--json] [--concurrency 4] [path...]
mdidx search [--json] [--limit 10] "query"
mdidx watch
mdidx stats [--json]
mdidx doctor [--json]
mdidx config show [--json]
mdidx config paths [--json]
mdidx eval [--json] file.md [distractor.md...]
mdidx version [--json]
mdidx skill emit [--output ~/.codex/skills/mdidx/SKILL.md]
Run mdidx --help or mdidx <command> --help for the full command, flag, config, environment, and output reference.
Installation
Install from the module path:
go install github.com/rbranson/mdidx/cmd/mdidx@latest
Install from a local checkout:
git clone git@github.com:rbranson/mdidx.git
cd mdidx
make install
Build a local binary without installing it:
make build
./bin/mdidx --help
./bin/mdidx version
Build, install, and dist targets stamp version metadata from git when available.
Unstamped development builds report dev and unknown values.
Install the agent skill emitted by the binary:
mkdir -p ~/.codex/skills/mdidx
mdidx skill emit --output ~/.codex/skills/mdidx/SKILL.md
Development
Common development targets are collected in the Makefile:
make help # list targets
make deps # download and verify modules
make fmt # format Go source
make vet # run go vet
make test # run tests
make check # fmt, vet, and test
make build # build ./bin/mdidx
make run ARGS='--help'
make run ARGS='eval README.md --generate-only'
make bench # benchmark smoke test
make quality-classic # live retrieval quality eval over classic-books-markdown
make integration # live OpenAI integration test
make clean # remove build artifacts
Before committing, run:
make check
For a larger local benchmark, set MDIDX_BENCH_MB:
MDIDX_BENCH_MB=100 BENCH_TIME=1x make bench
Retrieval quality evaluation over the classic-books-markdown corpus is opt-in
and uses live query embeddings, so it requires a corpus checkout that is already
an initialized and indexed mdidx project:
MDIDX_CLASSIC_BOOKS_PATH=~rbranson/src/github.com/mlschmitt/classic-books-markdown OPENAI_API_KEY=... make quality-classic
The eval does not build or update the index. Run mdidx init and mdidx index
in the corpus root first. Re-running the eval only searches the existing index,
plus optional LLM judging when enabled.
The quality evaluation generates title/author cases from every discovered book,
samples passage-level phrase cases from the corpus, and reads thematic cases
from internal/core/testdata/classic_quality/themes.json. Set
MDIDX_QUALITY_MAX_TITLE_CASES or MDIDX_QUALITY_MAX_PHRASE_CASES to control
case volume. Set MDIDX_QUALITY_JUDGE_MODEL to use a chat-completions-compatible
LLM as a judge for thematic queries such as "kissing while crying" or "walk
through the park".
Live integration tests are opt-in and make real OpenAI-compatible embedding API
calls. They require OPENAI_API_KEY and may incur API cost:
OPENAI_API_KEY=... make integration
OPENAI_BASE_URL and MDIDX_EMBEDDING_MODEL are honored when set.
Release-style binaries for common Darwin and Linux targets can be built with:
make dist
Watch mode
mdidx watch uses filesystem notifications through fsnotify, which maps to inotify on Linux and native watchers elsewhere. The first event after idle waits 5 seconds before indexing. After an update runs, later events are coalesced so at most one index update runs per minute by default.
The watcher also runs one initial index update when it starts.
Operations
Run a local health check before relying on a project index:
mdidx doctor
mdidx doctor --json
doctor checks project discovery, config validity, index paths, Markdown
discovery, SQLite database readiness, embedding configuration, and whether
OPENAI_API_KEY is present. It does not make live network calls.
Inspect effective configuration and resolved paths:
mdidx config show --json
mdidx config paths --json
config show includes environment overrides such as OPENAI_BASE_URL and
MDIDX_EMBEDDING_MODEL, but never prints API keys. It reports API key presence
as a boolean.
mdidx stats --json reports vector_index: "ivf-flat" when a large index has
trained coarse vector cells. Smaller indexes search the sqlite-vec table
directly to avoid extra indexing overhead.
For automation, use --json success output and parse structured JSON errors
from stderr. Commands invoked with --json keep stdout reserved for successful
data and emit errors shaped like:
{"error":{"code":"missing_config","message":"...","hint":"...","exit_code":2}}
Large indexing runs can emit progress to stderr:
mdidx index --progress
mdidx index --json --no-progress
Progress never writes to stdout. Use --no-progress for scripts that expect
quiet stderr except on errors.
Embedding calls use configurable batching and the OpenAI Go SDK retry policy:
[embedding]
batch_size = 64
request_timeout = "1m"
max_retries = 3
request_timeout bounds each embedding request attempt. max_retries is passed
to the OpenAI Go SDK, which retries transient request failures such as rate
limits, timeouts, and 5xx responses. Authentication and invalid request errors
fail immediately.
Indexing changed files runs concurrently according to index.concurrency.
SQLite updates are still written serially after embeddings complete, preserving
the project-local database lifecycle.
Retrieval evals
Use mdidx eval to build a repeatable retrieval eval from one or more
representative Markdown files:
mdidx eval docs/sample.md
mdidx eval docs/sample.md --json
mdidx eval docs/sample.md docs/nearby-topic.md
The command uses an LLM, gpt-5.4-mini by default, to generate about 100 search
queries from the files' indexed sections, saves them under .mdidx/evals/,
indexes the eval files, runs each query through mdidx, and then uses the same
model only as a relevance judge. Single-file evals search only that source file.
Multi-file evals search the project index so the other eval files and any
already-indexed Markdown act as distractors. Each pass writes a timestamped run
file under .mdidx/evals/runs/ with queries, retrieved snippets, judgments, and
aggregate hit@1, hit@3, hit@10, MRR, average judge score,
rank-weighted score, nDCG@10, and bucketed metrics.
Generated query sets intentionally mix short keyword searches, medium natural
project-search phrases, longer detailed clue searches, and query-type buckets
such as exact, paraphrase, ambiguous, narrative, heading-dependent, and
confusable cases so retrieval changes are tested against varied user behavior.
Eval query generation, searches, and judge batches run with bounded concurrency
by default. Use --concurrency N to tune parallel API calls for your rate
limits.
Keep the query set stable while trying retrieval strategy changes, then compare
the run JSON files:
mdidx eval docs/sample.md --queries .mdidx/evals/sample.queries.json
mdidx eval docs/sample.md --run-name lexical-heavy --lexical-weight 0.8 --vector-weight 0.2
Use --regenerate to refresh the query set, --generate-only to create the
query set without running searches, and --skip-index to evaluate the existing
index exactly as-is. Progress is printed to stderr for terminal runs; use
--progress to force it for scripted runs or --no-progress to suppress it.
Benchmarks
go test -bench . ./...
MDIDX_BENCH_MB=100 go test -bench BenchmarkIndexAndSearchGeneratedCorpus ./internal/core
MDIDX_CLASSIC_BOOKS_PATH=/path/to/classic-books-markdown OPENAI_API_KEY=... make quality-classic
License
MIT. See LICENSE.