omnifeed gives an AI agent the full research loop — search → URLs → content — against self-hosted
SearXNG + crawl4ai,
with a dedicated Reddit engine that returns full comment trees as TOON
(~40% fewer tokens than JSON, lossless) and no Reddit API key.
web_search queries a SearXNG instance (Google/Bing/DDG, Reddit included) and returns ranked URLs with titles and snippets.
fetch_url renders any URL through crawl4ai as clean markdown — and Reddit URLs come back as the full comment tree encoded in TOON.
Why omnifeed
omnifeed
Other Reddit MCPs
Web search → crawl in one self-hosted service
✅ SearXNG + crawl4ai
❌ search-only or crawl-only
Full comment tree (/api/morechildren expansion)
✅ up to 40 rounds (~4k comments)
❌
Token-efficient output
✅ TOON, ~40% smaller than JSON
❌ verbose JSON or truncated bodies
Generic crawl fallback for non-Reddit URLs
✅ via crawl4ai
❌
Front-ends
MCP + Open WebUI + REST
MCP only (most)
Demo
web_search → pick a URL → fetch_url → full Reddit comment tree as TOON. Generate this clip with vhs assets/demo.tape.
Quick start
# No clone needed — fetch the compose file + SearXNG settings, then start:
curl -fsSL https://raw.githubusercontent.com/kinorai/omnifeed/main/docker-compose.yml -o docker-compose.yml
curl -fsSL --create-dirs https://raw.githubusercontent.com/kinorai/omnifeed/main/searxng/settings.yml -o searxng/settings.yml
docker compose up
Starts omnifeed + SearXNG + crawl4ai. Point Open WebUI at http://localhost:8080 with WEB_LOADER_ENGINE=external. (SearXNG is mounted with searxng/settings.yml, which enables the json format the web_search tool needs.)
As an MCP server (Claude Code, Cursor, Windsurf, …)
Stdio — most clients. omnifeed always fetches through crawl4ai (and SearXNG, if you want web_search), so the container your client spawns must be told where they are. The simplest setup joins the network from docker compose up and uses the service names:
Without OMNIFEED_CRAWL4AI_URL the container exits at startup. If crawl4ai/SearXNG run elsewhere, point the URLs there (e.g. http://host.docker.internal:11235/crawl, adding --add-host=host.docker.internal:host-gateway on Linux).
HTTP — remote clients, or the simplest option when the docker compose up stack is already running (no second container, no networking to wire up):
Tools: fetch_url (always) and web_search (only when OMNIFEED_SEARXNG_URL is set). The intended loop is web_search → pick URLs → fetch_url.
/crawl returns [{"page_content": "...", "metadata": {...}}] — already the shape of a LangChain / LlamaIndex Document, so wrapping it as a custom document loader takes only a few lines.
Authentication
The HTTP transports (/crawl, /search, /mcp) share one bearer token. Set OMNIFEED_API_KEY and send Authorization: Bearer <token>:
# Generate a key and print it — clients send it as the bearer token, so save it somewhere.
export OMNIFEED_API_KEY="$(openssl rand -hex 32)"
echo "OMNIFEED_API_KEY=$OMNIFEED_API_KEY"
# `-e OMNIFEED_API_KEY` (no value) forwards the variable from your shell.
docker run -e OMNIFEED_API_KEY \
-e OMNIFEED_CRAWL4AI_URL=http://crawl4ai:11235/crawl \
kinorai/omnifeed
Without a key the proxy refuses to start, so it can't be left open by accident. For a throwaway local run, opt out with OMNIFEED_DEV_NO_AUTH=true (the compose files already do). Stdio MCP needs no token — it inherits the trust of the process that spawned it.
Configuration
Everything is configured with OMNIFEED_-prefixed environment variables. In practice you only ever set three — OMNIFEED_API_KEY, OMNIFEED_CRAWL4AI_URL, and (optionally) OMNIFEED_SEARXNG_URL. The rest have sane defaults.
Variable
Default
Purpose
OMNIFEED_API_KEY
(unset)
Bearer token for /crawl, /search, /mcp. If unset, the proxy refuses to start unless OMNIFEED_DEV_NO_AUTH=true. Stdio MCP is unaffected.
OMNIFEED_CRAWL4AI_URL
(required)
Upstream crawl4ai endpoint. Every engine fetches through it; if empty, the proxy exits at startup.
OMNIFEED_SEARXNG_URL
(unset)
Upstream SearXNG base URL (e.g. http://searxng:8080). When unset, web_search / /search are not exposed. The instance must enable the json format.
OMNIFEED_DEV_NO_AUTH
false
Run the HTTP transports with no auth when no key is set (local/dev only). Ignored if a key is set.
OMNIFEED_LISTEN_ADDR
:8080
HTTP listen address (/crawl, /search)
OMNIFEED_MCP_LISTEN_ADDR
:8081
MCP HTTP/SSE listen address
OMNIFEED_MCP_STDIO
false
Run MCP over stdio (also via --mcp-stdio)
OMNIFEED_METRICS_ADDR
:9090
Prometheus + health listen address
OMNIFEED_CRAWL4AI_TIMEOUT
90s
Per-call timeout to crawl4ai
OMNIFEED_SEARXNG_TIMEOUT
15s
Per-query timeout to SearXNG
OMNIFEED_SEARCH_MAX_RESULTS
25
Hard cap on the search limit argument (1–100)
OMNIFEED_REDDIT_TIMEOUT
4m
Wall-clock cap for a Reddit thread expansion
OMNIFEED_REDDIT_MAX_ROUNDS
3
Default /api/morechildren rounds (max 40 via ?expand=full)
OMNIFEED_REDDIT_FORMAT
toon
Default Reddit output: toon or json
OMNIFEED_REDDIT_FETCH_LIMIT
500
Reddit limit: max comments fetched in the initial tree
OMNIFEED_REDDIT_DEPTH
20
Reddit depth: max nesting depth of the initial tree
OMNIFEED_REDDIT_SORT
top
Reddit sort: one of confidence (=best), top, new, controversial, old, random, qa, live
OMNIFEED_REDDIT_MAX_COMMENTS
0
Hard cap on total comments emitted after expansion (0 = unlimited)
OMNIFEED_REDDIT_MAX_TOP_LEVEL
0
Hard cap on top-level comment threads, replies included (0 = unlimited)
OMNIFEED_REDDIT_KEEP_CREATED
true
Include each comment's created timestamp
OMNIFEED_REDDIT_KEEP_DEPTH
false
Include each comment's depth field
OMNIFEED_MAX_URLS_PER_REQUEST
30
Cap on urls[] length
OMNIFEED_PER_DOMAIN_CONCURRENCY
2
Max concurrent requests to one domain
OMNIFEED_PER_DOMAIN_DELAY
1500ms
Minimum delay between same-domain requests
OMNIFEED_BLOCK_PRIVATE_IPS
true
SSRF protection (keep on in production)
OMNIFEED_LOG_LEVEL
info
debug/info/warn/error
OMNIFEED_LOG_FORMAT
json
json or text
OMNIFEED_ENABLE_PPROF
false
Expose /debug/pprof/* (opt-in)
Controlling Reddit response size
A Reddit thread's comment tree can be huge. The size knobs come in two kinds — it matters which is which:
Upstream Reddit params — forwarded verbatim to Reddit's API, so Reddit owns their behavior: OMNIFEED_REDDIT_FETCH_LIMIT → limit, OMNIFEED_REDDIT_DEPTH → depth, OMNIFEED_REDDIT_SORT → sort. They shape what Reddit sends back (less latency, fewer tokens) but are approximate, and limit/depth bound only the initial fetch. Semantics are Reddit's, not ours — see https://www.reddit.com/dev/api/ → GET [/r/subreddit]/comments/article (limit = "maximum number of comments to return", depth = "maximum depth of subtrees").
omnifeed engine caps — our own, applied after fetch + expansion, so they're exact and independent of Reddit: OMNIFEED_REDDIT_MAX_COMMENTS (truncate the flat comment list) and OMNIFEED_REDDIT_MAX_TOP_LEVEL (keep the first N top-level threads, in sort order, with their replies).
Rule of thumb: reach for the upstream params to fetch less from Reddit; reach for the engine caps when you need a guaranteed ceiling — OMNIFEED_REDDIT_MAX_ROUNDS expansion adds comments on top of limit, so only the caps bound the final total. All five are also per-request on the fetch_url MCP tool (limit, depth, sort, max_comments, max_top_level); a positive value overrides the env default.
API
Two bearer-authenticated JSON endpoints on the loader port (:8080). The full reference — request/response schemas, query params, and error codes — lives in openapi.yaml (paste it into editor.swagger.io or any OpenAPI viewer).
GET /readyz — checks crawl4ai (and SearXNG, when configured); GET /healthz is an alias
GET /metrics — Prometheus (omnifeed_requests_total, omnifeed_request_seconds, omnifeed_reddit_expansion_rounds, omnifeed_search_requests_total, omnifeed_search_request_seconds)
Health and metrics listen on OMNIFEED_METRICS_ADDR (default :9090), separate from the API (:8080) and MCP (:8081) ports.
MCP
JSON-RPC 2.0 at:
stdio when OMNIFEED_MCP_STDIO=true or --mcp-stdio
Streamable HTTP at /mcp (spec 2025-03-26) on OMNIFEED_MCP_LISTEN_ADDR — POST /mcp for one-shot calls, GET /mcp for the SSE stream
GET /mcp/sse is a legacy alias for older dual-endpoint SSE clients; new clients target /mcp
Architecture
Two ports answer two questions. The Searcher port answers query → URLs; the Engine port answers URL → content. MCP tools and REST handlers compose them; transports stay thin.
Reddit's edge 403-blocks non-browser HTTP clients (it fingerprints the TLS/JA3 handshake), so the Reddit engine never calls Reddit directly. It drives crawl4ai's headless Chromium to a www.reddit.com page (which clears the bot wall), then runs a same-origin fetch() of the .json and /api/morechildren endpoints from inside that page. No Reddit auth, cookies, or API key. A per-thread crawl4ai session_id is reused across expansion rounds to keep one warmed context.
Sustained scraping can raise your source IP's risk score. If fetches start returning the block page, slow down, keep expand modest, or route crawl4ai through a residential proxy.
Extending it
New engines (Hacker News, Stack Overflow, …), searchers (Brave, Tavily, …), MCP tools, and transports each plug into one small port without touching the rest. See AGENTS.md → Adding things for the architecture and a step-by-step.
Development
git clone https://github.com/kinorai/omnifeed.git && cd omnifeed
make check # vet + lint + test (what CI runs)
go run ./cmd/omnifeed
See CONTRIBUTING.md for the full workflow, SECURITY.md for vulnerability reporting, and AGENTS.md if you're a coding agent working in this repo.
Star history
Contributing
Star it if it's useful — it helps other AI builders find omnifeed.
New engines, searchers, MCP tools, and transports are all welcome — start with AGENTS.md and CONTRIBUTING.md.
Package antibot recognizes bot-wall / CAPTCHA / challenge pages that upstreams (Cloudflare, Reddit's "network security" wall, PerimeterX, …) return in place of real content — frequently with an HTTP 200, so a crawl that "succeeded" can still be a block.
Package antibot recognizes bot-wall / CAPTCHA / challenge pages that upstreams (Cloudflare, Reddit's "network security" wall, PerimeterX, …) return in place of real content — frequently with an HTTP 200, so a crawl that "succeeded" can still be a block.
Package crawl4ai implements the fallback engine: dispatches generic URLs to an upstream crawl4ai instance and reshapes the response into the canonical Document.
Package crawl4ai implements the fallback engine: dispatches generic URLs to an upstream crawl4ai instance and reshapes the response into the canonical Document.
Package reddit implements the Reddit-specific engine: fetches threads via the public JSON API, expands collapsed reply branches with /api/morechildren, strips fields, drops deleted comments, and emits TOON or JSON.
Package reddit implements the Reddit-specific engine: fetches threads via the public JSON API, expands collapsed reply branches with /api/morechildren, strips fields, drops deleted comments, and emits TOON or JSON.
Package searchapi exposes the web-search use case over plain HTTP/JSON, so non-MCP clients (scripts, RAG pipelines, n8n, ...) can run the same query → result-URLs search the `web_search` MCP tool provides — no MCP client required.
Package searchapi exposes the web-search use case over plain HTTP/JSON, so non-MCP clients (scripts, RAG pipelines, n8n, ...) can run the same query → result-URLs search the `web_search` MCP tool provides — no MCP client required.