Lore
A decision-provenance engine for engineering teams. Ask why a decision was
made — and what happened afterwards — and get back evidence where every claim
carries a source URL.
The problem
Teams lose their why. Decisions happen in incident channels, Jira tickets, Notion
pages and PR reviews; then people rotate and six months later nobody can answer
"why did we choose B over A when incident X happened?" The answer exists — scattered
across a trail nobody reads manually. A commit that says only PROJ-4521 carries no
reasoning: the reasoning is in the ticket, and the consequences are in a postmortem
written weeks later.
Codebase RAG tools answer what the code is. Lore answers why it was decided and
what happened next.
How it works
- Ingest — source plugins stream documents from GitHub, GitLab, Notion and Jira
into one normalized
Document model. Every source is optional.
- Link — a resolver turns raw references (ticket keys, URLs, commit SHAs, file
paths) into a typed, directional edge graph: ticket → design doc → PR → review
thread → commit.
- Answer — one pipeline, four seed modes:
resolve anchor → seed → graph walk → hybrid retrieval → rank → EvidenceBundle. Retrieval is BM25 (FTS5) + vector KNN
(sqlite-vec) fused with Reciprocal Rank Fusion.
Three design choices set it apart:
- Code is one anchor among several, not the center. A Jira + Notion workspace with
no repository at all is a first-class configuration;
git blame anchoring is an
optional enrichment.
- Honesty is a feature. Every answer reports its
gaps — "Rate limit the export
endpoint (jira🎫PROJ-4521) stands alone; no linked discussion" — instead of
fabricating a chain. No URL means no evidence, so the node is not returned.
- Every source, model provider and clone reader is a plugin. The engine holds no
vendor name; the ones this binary ships are registered through the same contract a
third party uses. See Plugins.
Two personas, one engine
| Ask-only workspace (no repos) |
Code-anchored workspace |
| Sources: Jira + Notion (± GitHub/GitLab issues and PRs) |
Same, plus local clones registered |
"Why B over A during incident X?" → find_decision |
"Why does auth.go:40-55 exist?" → why |
"What impact did decision A have?" → impact_of |
"How did this file evolve?" → history_of |
| Anchors: query, document, time window |
Additional anchor: code span (blame) |
Nothing in the first column is degraded. find_decision, trace and impact_of run
the full walk — chains, gaps, time anchoring — on a workspace with zero repositories;
the second column is the first plus one extra anchor type. Walk through either one:
ask-only demo · code-anchored demo.
Quickstart
Requirements: Go 1.27+, git, and an OpenAI API key for embeddings — or a local
Ollama daemon instead, with embedder.provider: ollama.
No cgo, no Docker, no database server — SQLite ships inside the binary as pure-Go WASM
(ncruces/go-sqlite3 + sqlite-vec), so the index
is a single portable file and queries work offline after a sync.
go install github.com/setthasit/Lore/cmd/lore@latest # or: git clone … && make bin
export OPENAI_API_KEY=... # embeddings
export LORE_GITHUB_TOKEN=... # fine-grained, read-only PAT; lore.yaml expands it
lore init # writes a commented lore.yaml scaffold and its lore.schema.json
lore source add jira # optional: grow the workspace interactively
lore sync # first run creates ~/.lore/<workspace>.db
lore status # index counts, cursor ages, sync lock
lore ask "why did we pick sqlite?" # prose answer; needs an llm: block
lore init generates the scaffold from the manifests of the plugins this build
registers — a starter source instance whose secret field holds an ${env:VAR}
expansion, so the credential stays in your shell:
workspace: acme
sources:
- use: github
with:
token: ${env:LORE_GITHUB_TOKEN} # read-only PAT for the listed repositories
repos: [] # each entry is "owner/name"; no clone needed
repos: [] # local clones, for blame and history only
embedder:
provider: openai # credentials come from OPENAI_API_KEY
model: text-embedding-3-small # changing the model needs: lore sync --reembed
# llm: # lore ask and --explain answer in prose only with this
# provider: openai
# model: gpt-4o-mini
A secret field holds the credential itself, written either as ${env:VAR} or as a
literal value. A literal is announced once on stderr:
lore: secrets written as literal values in the config: sources[github].with.token; write ${env:VAR} to keep a credential out of the file.
Leave a secret field out and a plugin compiled into the binary falls back to the
variable its manifest suggests — OPENAI_API_KEY for the embedder above. A plugin
installed from outside the binary gets no such fallback.
The embedder can also hold its key itself, beside provider and model:
embedder:
provider: openai
model: text-embedding-3-small
api_key: sk-live-abc # or ${env:OPENAI_API_KEY}
A role binding accepts only the secrets its provider declares. Any other setting, such as
base_url, goes on a providers: entry that the binding names. A binding that names a
declared providers: entry carries no key of its own.
lore source add <plugin> appends another instance of any source plugin this build
registers, asking only for the fields that plugin's manifest declares; two Jira sites
are two instances of one plugin, each with its own id. lore --version prints the
build stamp plus the embedder identity of the workspace. The full lore.yaml
reference lives in 06 — Interfaces & Config.
What an answer looks like
lore ask answers in prose, citing the documents it used — it needs the llm: block
in lore.yaml:
SQLite carries the index because it ships everywhere and needs no server [1].
Postgres with pgvector was the alternative the storage design weighed [2].
**Sources**
1. Index on SQLite, not Postgres — https://github.com/acme/lore/pull/12
2. Storage design — https://notion.so/design/storage
why, trace, impact and history print the evidence itself as a timeline (shape
shown; the values are invented):
provenance of Storage design
anchor: Storage design
https://notion.so/design/storage
2 documents
2025-03-10 Storage design
notion page · 2025-03-10
https://notion.so/design/storage
postgres with pgvector was the alternative
2025-03-12 Index on SQLite, not Postgres
github pr · dev@example.test · 2025-03-12 · follow_up
https://github.com/acme/lore/pull/12
sqlite ships everywhere and needs no server
chains:
notion:page:design/storage → github:pr:acme/lore/pull/12
gaps:
no follow-up evidence after 2025-03-12
Pass --explain to any of those four to get prose instead of the timeline, and --raw
to any query command to get the EvidenceBundle as JSON for scripting. --raw wins
when both are given.
What to expect
Lore is only as good as the trail your team leaves. Where commits name their tickets
and PRs describe their reasoning, chains run four and five hops deep across sources.
Where they do not, the honest outcome is a short chain plus a gaps line — retrieval
still surfaces the unlinked discussion, but nothing invents the missing edge. A first
sync of a large workspace is the slow part (it embeds every chunk); after that, syncs
are incremental and queries are local.
Running the services
Lore is a single binary; the subcommand picks the transport.
lore mcp # MCP over stdio, for a local agent harness
lore serve --http 127.0.0.1:8080 # MCP streamable HTTP at /mcp + lore.v1 gRPC + background sync
serve also runs the sync scheduler, so the index stays fresh while the endpoints are
up. gRPC listens on 127.0.0.1:9090 unless --grpc or server.grpc_addr says
otherwise, and --mtls makes it require a client certificate signed by
server.mtls.client_ca.
A bind address that is not provably loopback is refused unless server.mtls.cert and
server.mtls.key are configured — :8080 and localhost:8080 do not qualify, because
a bare port reaches every interface and a host name is not proof.
Register the stdio server with an MCP host (Claude Code, Cursor, …):
{
"mcpServers": {
"lore": {
"command": "/absolute/path/to/lore",
"args": ["mcp", "--config", "/absolute/path/to/lore.yaml"]
}
}
}
Per-client paths, the streamable-HTTP variant and the troubleshooting table live in the
MCP quickstart.
MCP tools return the structured EvidenceBundle, never prose: the host model is already
an LLM, so it synthesizes in its own context and can immediately call another tool. That
is why no LLM key is needed for MCP usage — only an embedding key.
| MCP tool |
Answers |
find_decision |
"why B over A, around incident X?" — retrieval-seeded, works with zero repos |
why |
"why does auth.go:40-55 look like this?" — blame-seeded |
trace |
everything linked to one commit / PR / ticket / page, in order |
impact_of |
what followed a decision, as a chronological timeline |
history_of |
how one file evolved, commit by commit |
sync_now / sync_status |
trigger a sync round; report cursors, counts and the lock |
The lore.v1 gRPC API is the programmatic surface for anything that is not an agent:
QueryService carries the same five verbs, SyncService adds Trigger, Status and a
server-streaming Watch. Unlike MCP it does synthesize — every request carries an
optional synthesize flag, and leaving it unset means prose alongside the bundle.
CLI surface
| Command |
Purpose |
lore init · lore source add <plugin> |
scaffold and grow lore.yaml |
lore schema |
rewrite the JSON Schema that editors use to complete and check lore.yaml |
lore sync [--source <instance>] [--reembed] |
one sync round; checkpoints per batch, so an interrupted run resumes |
lore status |
index counts, per-source cursor ages, sync lock state |
lore ask <question> |
synthesized prose; --around --source --repo --doc-type --since --until --raw |
lore why <file>:<L1>[-<L2>] ["question"] |
blame-anchored trail; --repo --explain --raw |
lore trace <ref> |
one document's neighborhood; --direction in|out|both --explain --raw |
lore impact <ref | "query"> |
consequences timeline; --question --explain --raw |
lore history <path> |
file timeline; --limit --before pagination; --repo --explain --raw |
lore mcp · lore serve [--http --grpc --mtls] |
MCP stdio · MCP streamable HTTP + lore.v1 gRPC + scheduler |
lore plugin list|search|install|update|verify|remove |
inspect what this build can use; add third-party plugins |
lore build --with <module>@<version> [-o lore] |
compile a binary with third-party plugins linked in |
Every command takes --config (default ./lore.yaml).
Plugins
Sources, model providers and clone readers are the same kind of thing: a plugin with a
manifest, resolved through one registry and configured as instances in lore.yaml. An
official plugin holds no privilege a third-party one lacks. lore plugin list prints
what the running binary can be configured with:
NAME KIND ORIGIN SUMMARY
anthropic provider (complete) builtin Anthropic Claude chat completions
git code builtin Blame and history for one local git clone (read-only)
github source builtin GitHub commits, pull requests, issues and their comments and reviews (read-only)
gitlab source builtin GitLab commits, merge requests, review threads, issues and notes (read-only)
jira source builtin Jira Cloud issues and their comments (read-only)
notion source builtin Notion pages and their block content (read-only)
ollama provider (embed, complete) builtin Local Ollama daemon embeddings and chat completions
openai provider (embed, complete) builtin OpenAI embeddings and chat completions
openai-compatible provider (embed, complete) builtin Embeddings and chat completions from any vendor speaking the OpenAI protocols
embedder: and llm: are role bindings, not vendor switches: each names a provider
instance and a model. openai derives the vector width from the embedding model;
ollama needs embedder.dimensions set to the model's native width (ollama show <model> reports it); openai-compatible reaches Z.AI, OpenRouter, Moonshot, DeepSeek,
Groq, Together, vLLM and LM Studio through a preset: row rather than new code. A
provider name this build neither compiles in nor finds under plugins: is refused at
startup, and a plugin that binds to a role it cannot serve is refused with it.
Third-party plugins load two ways:
- Out of process, no rebuild. Declare it under
plugins: with a coordinate,
lore plugin install, and the host speaks a small NDJSON protocol to the binary — in
any language, as test/fixtures/plugins/pysource.py demonstrates. The artifact is
pinned by SHA-256 in lore.lock, re-verified at every launch, optionally signature-
checked (cosign or minisign) against the pubkey: the declaration names, and started
with an empty environment. In its with: block, your environment reaches it only through
the secrets its manifest declares and the fields it marks expandable.
- Compiled in.
lore build --with github.com/owner/repo@v1.2.3 regenerates the
composition root with that module added and builds a new binary — in-process calls
and compile-time type safety, in exchange for needing a Go toolchain.
Writing one means implementing the sdk contract: 08 — Extensibility
for the Go interfaces and manifests, 09 — Plugin Protocol
for the wire format, 10 — Plugin Distribution for
coordinates, the lockfile and the trust model.
Guides
| Guide |
Contents |
| MCP quickstart |
Claude Code / Cursor config for stdio and streamable HTTP, tool routing, troubleshooting |
| Source setup |
least-privilege credentials for GitHub, GitLab, Notion and Jira |
| Fully local |
Ollama for embeddings and synthesis; what leaves the machine in each mode |
| Ask-only demo |
seeded Jira + Notion workspace, zero repositories, two flagship questions |
| Code-anchored demo |
why on a real OSS repository: commit → PR → issue |
Architecture
Strict unidirectional layering, wired with Uber FX.
Transports never touch the store or a plugin directly — including the MCP path.
flowchart TB
T["Transport — MCP stdio · MCP HTTP · gRPC · CLI"]
S["Service — Query · Why · Trace · Impact · History · Synthesis · SyncOrchestrator · LinkResolver"]
R["Repository — IndexStore (SQLite: FTS5 + sqlite-vec)"]
C["Plugins — source · code · provider instances, built from lore.yaml against a registry"]
X["External plugins — any binary, NDJSON over stdio"]
T --> S
S --> R
S --> C
C --> X
One SQLite file per workspace holds documents, chunks, chunks_fts (BM25),
chunk_vectors (sqlite-vec), edges, pending_refs, cursors, sync_lock and
meta. The index is derived data: sources are ground truth, so deleting it is safe
and the next sync rebuilds it.
Sync is crash-safe by construction: connectors yield Batch{Docs, Cursor}, and the
orchestrator commits the batch then persists that batch's cursor. A single-row lease
with a 15s heartbeat and a 60s TTL takeover keeps a manual lore sync, the background
scheduler and a lore serve daemon from colliding — and a crashed run never wedges the
scheduler.
The design documents in docs/v3/ are the source of truth:
| Doc |
Contents |
| 00 — Design deltas |
every change from the v1 design, with the rationale that drove it |
| 01 — Overview |
problem, concept, differentiators, goals / non-goals |
| 02 — Architecture |
layers, transports, request flows, key decisions |
| 03 — Data Model |
Document / Edge model, schema, chunking, hybrid retrieval |
| 04 — Connectors & Sync |
connector contract, scheduler, lease, link resolver |
| 05 — Query Engine |
pipeline, anchors, tool algorithms, EvidenceBundle |
| 06 — Interfaces & Config |
MCP tools, CLI, gRPC API, lore.yaml reference |
| 07 — Risks & Open Questions |
operating risks with their standing mitigations |
| 08 — Extensibility |
plugin kinds, the sdk contract, manifests, registry, provider roles |
| 09 — Plugin Protocol |
NDJSON wire protocol: ops, streaming, cancellation, errors |
| 10 — Plugin Distribution |
coordinates, lore.lock, signatures, trust model, custom binaries |
Development
make build # go build ./...
make bin # stamped, static binary at bin/lore
make build.matrix # cross-compile linux/darwin/windows × amd64/arm64 with CGO_ENABLED=0
make test # go test ./...
make lint # golangci-lint run — depguard, errcheck, forbidigo, gosec, govet, staticcheck
make gen.mock # go generate ./... — gomock doubles under internal/mocks
make gen.proto # regenerate the lore.v1 stubs from api/proto
make certs.dev # local certificate authority + server/client pairs for mTLS
Tests need no external service: connectors run against httptest fixture servers, the
store against a temp SQLite file, and the end-to-end suite drives the real MCP transports
over a live DI graph — an MCP client session over streamable HTTP on a real socket,
against fixture GitHub/Notion/Jira servers. It includes an ask-only workspace with
zero repositories, which asserts that code-anchored tools refuse with a precondition
error rather than degrading. The plugin host is exercised separately against a real
out-of-process Python plugin (test/fixtures/plugins/pysource.py), which is also made
to exit mid-stream to prove a resumed sync loses no document — skipped where no
python3 is on PATH.
Contributions follow the branch → PR workflow in AGENTS.md: no direct
commits to main, and make build/test/lint green before a PR is opened.
Security posture
- Read-only toward every source. Lore never writes to GitHub, GitLab, Notion or Jira.
- Secrets are held by their own config field, as a literal or as
${env:VAR}.
A literal is announced on stderr at startup. Every resolved secret is scrubbed from
logs and errors — one shorter than 8 characters is named at startup instead — and
never reaches the index; least-privilege tokens are the documented default.
- An external plugin's
with: block reads your environment only where its manifest
allows. Only its declared secrets and the fields it marks expandable expand. Only a
declared secret's value is scrubbed. The plugin is digest-pinned, re-verified at launch,
and started with no inherited environment.
- Off-loopback serving requires TLS, enforced at startup, with mTLS support.
- Private data leaves the machine only toward the configured embedder and, once
llm:
is set, the configured LLM. With provider: ollama on both, nothing leaves at all.
License
Apache License 2.0 — see LICENSE. Contributions are accepted under the same
terms.