README
¶
vecgrep
Local-first semantic code search powered by embeddings.
vecgrep indexes your codebase and enables natural language search using vector embeddings. All processing happens locally via Ollama, ensuring your code never leaves your machine.
Features
- Hybrid Search - Combine semantic (vector) and keyword search for best results
- Three Search Modes - Choose between semantic, keyword, or hybrid search
- Local-First - All embeddings generated locally via Ollama
- OpenAI Support - Optional cloud embeddings via OpenAI API
- Incremental Indexing - Only re-index changed files
- Batch Operations - Efficient bulk indexing with batch inserts
- Language-Aware Chunking - Intelligent code splitting by functions, classes, and blocks
- Rich Filtering - Filter by language, chunk type, directory, file pattern, and line range
- MCP Support - Model Context Protocol server for AI assistant integration
- Web Interface - Browser-based search UI with syntax highlighting
- Similar Code Finder - Find semantically similar code across your codebase
- Search Diagnostics - Explain mode for debugging and optimizing searches
Installation
Prerequisites
- Go 1.25+
- Ollama with an embedding model (default:
nomic-embed-text) - Task (optional, for development)
From Source
git clone https://github.com/abdul-hamid-achik/vecgrep.git
cd vecgrep
task build
# or: go build -o bin/vecgrep ./cmd/vecgrep
Install to GOPATH
task install
# or: go install ./cmd/vecgrep
Quick Start
-
Start Ollama and pull the embedding model:
ollama pull nomic-embed-text -
Initialize vecgrep in your project:
cd /path/to/your/project vecgrep initImportant: Add
.vecgrepto your.gitignorefile:echo ".vecgrep" >> .gitignore -
Index your codebase:
vecgrep index -
Search:
vecgrep search "error handling in HTTP requests"
Using OpenAI (Alternative)
If you prefer cloud embeddings via OpenAI:
-
Set your API key:
export OPENAI_API_KEY=sk-your-key-here -
Configure vecgrep to use OpenAI:
Edit
.vecgrep/config.yaml:embedding: provider: openai model: text-embedding-3-small dimensions: 1536Or set via environment:
export VECGREP_EMBEDDING_PROVIDER=openai export VECGREP_EMBEDDING_MODEL=text-embedding-3-small -
Re-index your codebase:
vecgrep index --full
Usage
Initialize a Project
vecgrep init [--force]
Creates a .vecgrep directory with configuration and database.
Index Files
vecgrep index [paths...] [--full] [--ignore pattern]
Options:
--full- Force full re-index (ignores file hashes)--ignore- Additional patterns to ignore-v, --verbose- Show detailed progress
Search
vecgrep search <query> [options]
Search Modes:
| Mode | Description |
|---|---|
hybrid |
Combines vector similarity with text matching (default) |
semantic |
Pure vector similarity search |
keyword |
Text-based search using pattern matching |
Options:
| Flag | Description |
|---|---|
-n, --limit N |
Maximum results (default: 10) |
-f, --format |
Output format: default, json, compact |
-m, --mode |
Search mode: hybrid, semantic, keyword |
--explain |
Show search diagnostics (index type, nodes visited, duration) |
-l, --lang |
Filter by single language |
--languages |
Filter by multiple languages (comma-separated) |
-t, --type |
Filter by chunk type: function, class, block |
--types |
Filter by multiple chunk types (comma-separated) |
--file |
Filter by file pattern (glob) |
--dir |
Filter by directory prefix |
--lines |
Filter by line range (e.g., 1-100) |
Examples:
# Default hybrid search
vecgrep search "database connection pooling"
# Semantic-only search (vector similarity)
vecgrep search --mode=semantic "error handling patterns"
# Keyword search (text matching)
vecgrep search --mode=keyword "SELECT FROM users"
# Search with diagnostics
vecgrep search --explain "authentication middleware"
# Filter by language
vecgrep search "error handling" -l go -n 5
# Filter by multiple languages
vecgrep search "memory management" --languages=go,rust
# Filter by directory
vecgrep search "config loading" --dir=internal/
# Filter by file pattern
vecgrep search "test helpers" --file="**/*_test.go"
# Filter by chunk types
vecgrep search "handlers" --types=function,method
# Filter by line range
vecgrep search "imports" --lines=1-50
# JSON output for scripting
vecgrep search "API endpoints" --format=json
Web Interface
Start the web server:
vecgrep serve --web
Open http://localhost:8080 in your browser to search with a visual interface.
Options:
-p, --port- Server port (default: 8080)--host- Server host (default: localhost)
MCP Server
Start the MCP server for AI assistant integration:
vecgrep serve --mcp
This runs on stdio for integration with Claude Desktop, Claude Code, etc.
Find Similar Code
vecgrep similar <target> [options]
Find code semantically similar to an existing chunk, file location, or text snippet.
Targets:
42- Chunk ID (numeric)main.go:15- File:line location--text "code"- Inline text snippet
Options:
| Flag | Description |
|---|---|
-n, --limit N |
Maximum results (default: 10) |
-f, --format |
Output format: default, json, compact |
-l, --lang |
Filter by single language |
--languages |
Filter by multiple languages |
-t, --type |
Filter by chunk type |
--types |
Filter by multiple chunk types |
--file |
Filter by file pattern (glob) |
--dir |
Filter by directory prefix |
--lines |
Filter by line range |
--exclude-same-file |
Exclude results from the same file |
-T, --text |
Find similar to text snippet |
Examples:
# Find code similar to chunk ID 42
vecgrep similar 42
# Find code similar to line 50 in search.go
vecgrep similar internal/search/search.go:50
# Find code similar to a text snippet
vecgrep similar --text "func NewSearcher"
# Find similar Go code, excluding same file
vecgrep similar 42 --lang go --exclude-same-file
# Find similar code in specific directory
vecgrep similar --text "error handling" --dir=internal/
Check Status
vecgrep status [options]
Displays index statistics, configuration, and pending changes.
Options:
-f, --format- Output format:default,json
Examples:
vecgrep status # Default text output
vecgrep status --format json # JSON output for scripting
Index Management
Delete a File
Remove a specific file and its chunks from the index:
vecgrep delete <file-path>
Example:
vecgrep delete internal/old_file.go
Clean Database
Remove orphaned data (chunks without files, embeddings without chunks) and optimize:
vecgrep clean
Reset Index
Clear all indexed data (destructive):
vecgrep reset [--force]
Options:
--force- Skip confirmation prompt
Shell Completion
Generate shell completion scripts:
# Bash
vecgrep completion bash > /etc/bash_completion.d/vecgrep
# Zsh
vecgrep completion zsh > "${fpath[1]}/_vecgrep"
# Fish
vecgrep completion fish > ~/.config/fish/completions/vecgrep.fish
Configuration
Configuration is stored in .vecgrep/config.yaml:
Note: The
.vecgrepdirectory contains your local index database and configuration. Add it to.gitignoreto avoid committing it to version control.
embedding:
provider: ollama # or "openai" for cloud embeddings
model: nomic-embed-text # or "text-embedding-3-small" for OpenAI
dimensions: 768 # 1536 for text-embedding-3-small, 3072 for large
ollama_url: http://localhost:11434
openai_api_key: "" # Set via env var OPENAI_API_KEY or VECGREP_OPENAI_API_KEY
openai_base_url: "" # Optional: for Azure OpenAI or custom endpoints
indexing:
chunk_size: 512
chunk_overlap: 64
max_file_size: 1048576
ignore_patterns:
- ".git/**"
- "node_modules/**"
- "vendor/**"
- "*.min.js"
- "*.min.css"
- "*.lock"
search:
default_mode: hybrid # Default search mode: semantic, keyword, or hybrid
vector_weight: 0.7 # Weight for vector similarity in hybrid mode (0-1)
text_weight: 0.3 # Weight for text matching in hybrid mode (0-1)
server:
host: localhost
port: 8080
vector:
veclite:
m: 16 # HNSW max connections per node
ef_construction: 200 # Build quality (higher = better quality, slower build)
ef_search: 100 # Search quality (higher = better recall, slower search)
Vector Backend
vecgrep uses veclite as its vector storage backend with:
- Cosine distance for normalized embedding similarity
- HNSW indexing for fast approximate nearest neighbor search
- Native filtering with support for glob patterns, prefix matching, range queries
- Batch operations for efficient bulk indexing
Configuration Sources
vecgrep loads configuration from multiple sources in priority order:
- Environment variables (
VECGREP_*) - Highest priority - Project root
vecgrep.yamlorvecgrep.yml - XDG-style
.config/vecgrep.yaml - Legacy
.vecgrep/config.yaml - Global defaults
~/.vecgrep/config.yaml - Built-in defaults - Lowest priority
This allows you to set global defaults while overriding per-project settings.
Environment Variables
All environment variables use the VECGREP_ prefix:
| Variable | Description |
|---|---|
VECGREP_EMBEDDING_PROVIDER |
Embedding provider: ollama (default) or openai |
VECGREP_EMBEDDING_MODEL |
Embedding model name |
VECGREP_OLLAMA_URL |
Ollama API URL (default: http://localhost:11434) |
VECGREP_OPENAI_API_KEY |
OpenAI API key (or use OPENAI_API_KEY) |
VECGREP_OPENAI_BASE_URL |
OpenAI base URL (for Azure/custom endpoints) |
VECGREP_HOST |
Server bind address |
VECGREP_PORT |
Server port |
Global Flags
These flags work with all commands:
-c, --config- Custom config file path-v, --verbose- Enable verbose output--version- Show version information
MCP Integration
vecgrep implements the Model Context Protocol for AI assistant integration.
Available Tools
| Tool | Description |
|---|---|
vecgrep_init |
Initialize vecgrep in a directory (creates .vecgrep folder) |
vecgrep_search |
Search with semantic, keyword, or hybrid mode. Supports rich filtering and explain mode. |
vecgrep_index |
Index or re-index files in the project |
vecgrep_status |
Get index statistics (files, chunks, languages) |
vecgrep_similar |
Find code similar to a chunk ID, file:line location, or text snippet |
vecgrep_delete |
Delete a file and its chunks from the index |
vecgrep_clean |
Remove orphaned data and optimize the database |
vecgrep_reset |
Reset the project database (requires confirmation) |
Search Tool Parameters:
| Parameter | Type | Description |
|---|---|---|
query |
string | Search query (required) |
limit |
int | Maximum results |
mode |
string | Search mode: semantic, keyword, or hybrid |
explain |
bool | Return search diagnostics |
language |
string | Filter by single language |
languages |
array | Filter by multiple languages |
chunk_type |
string | Filter by single chunk type |
chunk_types |
array | Filter by multiple chunk types |
file_pattern |
string | Filter by file pattern (glob) |
directory |
string | Filter by directory prefix |
min_line |
int | Filter by minimum start line |
max_line |
int | Filter by maximum start line |
Note: In uninitialized directories, only vecgrep_init is available. After initialization, all tools become available.
Claude Code (CLI)
Add vecgrep as an MCP server:
# Add for all your projects (recommended)
claude mcp add --scope user vecgrep -- vecgrep serve --mcp
# Or add for current project only
claude mcp add --scope local vecgrep -- vecgrep serve --mcp
The MCP server works in any directory. If .vecgrep doesn't exist, use vecgrep_init to initialize it first.
Manage your MCP servers:
claude mcp list # List all servers
claude mcp get vecgrep # Show vecgrep config
claude mcp remove vecgrep # Remove vecgrep
Claude Code (Manual Config)
Add to ~/.claude/settings.json:
{
"mcpServers": {
"vecgrep": {
"command": "vecgrep",
"args": ["serve", "--mcp"],
"cwd": "/path/to/your/project"
}
}
}
Claude Desktop
Add to your Claude Desktop configuration (~/Library/Application Support/Claude/claude_desktop_config.json on macOS):
{
"mcpServers": {
"vecgrep": {
"command": "vecgrep",
"args": ["serve", "--mcp"],
"cwd": "/path/to/your/project"
}
}
}
Note: The cwd should point to a directory with an initialized .vecgrep folder.
Docker
Run vecgrep in a container while using Ollama on your host machine.
Quick Start
# Start Ollama on host (with Metal GPU on macOS)
OLLAMA_METAL=1 OLLAMA_HOST=0.0.0.0 ollama serve
# Run vecgrep container
docker compose up -d
The web interface is available at http://localhost:8080
Configuration
The container connects to Ollama on your host via host.docker.internal:11434.
Volumes:
./.vecgrep:/data/.vecgrep- Persistent index database./:/workspace:ro- Your codebase (read-only)
Index from Container
docker compose exec app vecgrep index /workspace
docker compose exec app vecgrep search "your query"
Upgrading
Breaking Changes in v2.0
Version 2.0 includes significant improvements that require re-indexing:
- Cosine distance replaces Euclidean for better embedding similarity
- Native filtering during search instead of post-filtering
- Batch operations for faster indexing
After upgrading, re-index your codebase:
vecgrep reset --force
vecgrep index
Development
See DEVELOPMENT.md for detailed development workflow.
task doctor # Check your environment
task setup # Install dependencies
task dev # Run with hot reload
task check # Run fmt, lint, test
task build # Build binary
License
MIT License - see LICENSE for details.
Directories
¶
| Path | Synopsis |
|---|---|
|
cmd
|
|
|
vecgrep
command
|
|
|
internal
|
|
|
embed
Package embed provides embedding generation for semantic search.
|
Package embed provides embedding generation for semantic search. |
|
index
Package index provides file indexing and chunking for semantic search.
|
Package index provides file indexing and chunking for semantic search. |
|
mcp
Package mcp implements the MCP server using the official SDK.
|
Package mcp implements the MCP server using the official SDK. |
|
search
Package search provides semantic search functionality.
|
Package search provides semantic search functionality. |
|
web
Package web provides the HTTP server and web UI for vecgrep.
|
Package web provides the HTTP server and web UI for vecgrep. |