searchwire

package module
v0.0.0-...-4cc2adc Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 17, 2026 License: MIT Imports: 25 Imported by: 0

README

searchwire

Zero-config Go metasearch for agent tooling — usable as a Go library or an MCP stdio server.

Searchwire runs ordinary web/text searches for you. Provide a query only — no SearXNG deployment, search host, engine selection, or API key is required for the default path. After searching, callers can fetch a result URL for readable page text (HTML stripped / plain text).

Library

go get github.com/jahrulnr/searchwire
package main

import (
	"context"
	"fmt"
	"log"

	"github.com/jahrulnr/searchwire"
)

func main() {
	searcher := searchwire.New(searchwire.Config{})
	resp, err := searcher.Search(context.Background(), "Go context cancellation")
	if err != nil {
		log.Fatal(err)
	}
	for _, result := range resp.Results {
		fmt.Println(result.Title, result.URL)
	}
	for _, sourceErr := range resp.Errors {
		log.Printf("source %s failed: %s", sourceErr.Source, sourceErr.Error)
	}

	if searcher.CanAnswer() {
		answer, err := searcher.Answer(context.Background(), "What changed in Go recently?")
		if err != nil {
			log.Fatal(err)
		}
		fmt.Println(answer.Provider, answer.Text)
	}

	if len(resp.Results) > 0 {
		page, err := searcher.Fetch(context.Background(), resp.Results[0].URL)
		if err != nil {
			log.Fatal(err)
		}
		fmt.Println(page.Title, page.Truncated, len(page.Text))
	}
}

searchwire.New(searchwire.Config{}) and searchwire.New(searchwire.DefaultConfig()) are equivalent.

Fetch

Searcher.Fetch(ctx, url) / FetchWithLimit(ctx, url, maxBytes) / FetchWithOptions(ctx, url, opts) perform a single HTTP GET:

Field Meaning
URL / FinalURL Requested URL and post-redirect URL
StatusCode / ContentType Response metadata
Title / Text Page title (HTML) and readable text
Links Safe http/https anchors collected from HTML (text + href)
Headers Selected response headers: ETag, Last-Modified, X-RateLimit-Remaining, X-RateLimit-Reset, Retry-After, Content-Range
Redirects Number of redirects followed to reach the final response
Truncated / Bytes Whether the body was capped; bytes read
DescribedBy Absolute URL of the covering llms.txt (rel="describedby" via HTML <link> or HTTP Link: header); empty when absent
MarkdownAlt Absolute URL of a markdown alternate version of this page (rel="alternate" type="text/markdown"); empty when absent
LLMsTxt / LLMsTxtURL Parsed covering llms.txt (auto-attached on 2xx HTML responses; see llms.txt) and its post-redirect URL. nil/empty for non-HTML responses, 4xx/5xx source, or when no file exists.
TempFile Absolute path of a file under the configured temp dir (default os.TempDir()/searchwire, override via Config.TempDir / SEARCHWIRE_TEMP_DIR) holding the full extracted text when Truncated is true. Lets the agent file_read the rest without refetching. Empty when the response was not truncated.

FetchOptions enables conditional GET and Range requests:

Field Header sent Notes
MaxBytes — Per-call byte budget (overrides Config.MaxResponseBytes when smaller)
IfNoneMatch If-None-Match A 304 Not Modified is returned as a FetchResult (not an error)
IfModifiedSince If-Modified-Since Same 304 handling
Range Range A 206 Partial Content is returned as a normal FetchResult
  • Schemes: http / https only.
  • Accepts text/html, application/xhtml+xml, text/plain, text/markdown, text/csv, application/json (and +json vendor types), application/xml, text/xml, application/rss+xml, application/atom+xml (and +xml vendor types), and other text/*.
  • HTML → title + visible text (scripts/styles/nav/footer/aside/form stripped, <pre>/<code> whitespace preserved, links collected, og:title/<h1> title fallbacks). JSON → pretty-printed (invalid JSON returned raw). XML/RSS/Atom → tag-stripped text with newlines preserved. Markdown/CSV/plain text → newlines preserved, inline whitespace collapsed. Non-UTF-8 charsets (declared in Content-Type or <meta charset>) are decoded to UTF-8. Oversized bodies truncate with Truncated: true and a trailing [truncated] marker instead of hard-failing.
  • Non-2xx responses return *HTTPError with RetryAfter (parsed Retry-After seconds, supports integer and HTTP-date forms) and ErrorBody (parsed JSON error envelope when the error body is JSON).
  • Local agent helper only: no JS rendering, no cookie jar, no multi-tenant SSRF suite. Loopback/private hosts are allowed for local-dev agent use.
llms.txt

Searchwire implements the client side of the llmstxt.org v2 proposal — the convention Chrome Lighthouse audits under its Agentic Browsing category:

  • Discovery on every fetch. When a page advertises its LLM-friendly entry points, Fetch reports them resolved to absolute URLs: DescribedBy (the covering llms.txt, from <link rel="describedby"> or the RFC 8288 Link: header) and MarkdownAlt (a markdown version of the fetched page, from <link rel="alternate" type="text/markdown">).
  • Auto-attach on 2xx HTML. When the source fetch is HTML and the server returns 2xx, Fetch probes <dir>/llms.txt first and /llms.txt second, parses the first one that answers, and embeds it as FetchResult.LLMsTxt (with LLMsTxtURL for the post-redirect URL). The probe is best-effort and non-fatal: any failure (404/410 fallthrough, soft-404 SPA, 5xx, network error) leaves the field nil and the source fetch still succeeds. The probe is skipped entirely for non-HTML responses (JSON, plain text, markdown, CSV, XML, RSS, Atom) and for non-2xx source responses — matching the spec that llms.txt only applies to navigable pages.
  • Standalone library call. Searcher.FetchLLMsTxt(ctx, url) is still available for callers that want the parsed *LLMsTxt without going through a full Fetch (e.g. to index a site). It uses the same most-specific-then-root candidate order and 404/410 fallthrough.
  • Parsing. ParseLLMsTxt(content string) parses llms.txt Markdown leniently into a structured LLMsTxt: required Title (H1), optional blockquote Summary, free-form Description, and Sections of [name](url): notes links. A file without the H1 returns an error wrapping ErrInvalidLLMsTxt.
// Library: auto-attached on every successful HTML fetch.
res, err := searcher.Fetch(ctx, "https://example.com/docs/page.html")
if err != nil {
    log.Fatal(err)
}
if res.LLMsTxt != nil {
    for _, section := range res.LLMsTxt.Sections {
        for _, link := range section.Links {
            fmt.Println(section.Name, link.Name, link.URL)
        }
    }
}

// Library: standalone probe without fetching the page.
llms, err := searcher.FetchLLMsTxt(ctx, "https://example.com/docs/page.html")
if errors.Is(err, searchwire.ErrLLMsTxtNotFound) {
    // site publishes no llms.txt — fall back to ordinary fetching
}
for _, section := range llms.LLMsTxt.Sections {
    for _, link := range section.Links {
        fmt.Println(section.Name, link.Name, link.URL)
    }
}

MCP tool note. The previous standalone llms.txt MCP tool was merged into fetch: agents now read the embedded llmsTxt field on the fetch result instead of making a second call. Searcher.FetchLLMsTxt is unchanged for library users. Tool names drop the searchwire_ prefix because the MCP server name already carries it — clients address tools as <server>.<tool> (e.g. searchwire.fetch), so a prefixed name would be doubly redundant.

Searcher.Search(ctx, query) fans the query out to all enabled sources concurrently, deduplicates results by canonical URL (lowercased host, default ports dropped, fragments and utm_* / fbclid / gclid tracking parameters stripped), then ranks the survivors with Reciprocal Rank Fusion (k=60).

type Result struct {
	Title   string
	URL     string // canonical form used for deduplication
	Snippet string // source-provided snippet; GitHub snippets are HTML-stripped and capped at 320 chars
	Sources []string
	Score   float64 // RRF score: higher = more sources agreed on this URL
}

When two results tie on score, earlier-configured sources win (see Built-in sources). Source failures never abort a search while at least one source succeeds — they are reported per-source in Response.Errors.

Per-call limit. By default each search returns Config.Limit results (default 10) and asks each source for that many candidates. To override it without rebuilding the searcher:

resp, err := searcher.SearchWithOptions(ctx, "query", searchwire.SearchOptions{Limit: 25})
// Limit <= 0 falls back to Config.Limit, then to the built-in default of 10.

The same Limit value bounds both what each source is asked for and how many fused results come back.

Defaults
Setting Default Override
Result limit 10 Config.Limit, SearchOptions{Limit} (per call), SEARCHWIRE_LIMIT (adapter)
HTTP timeout 10s Config.Timeout, SEARCHWIRE_TIMEOUT (adapter)
Extracted-text budget per fetch 2 MiB Config.MaxResponseBytes, FetchOptions.MaxBytes / FetchWithLimit (per call), SEARCHWIRE_MAX_BYTES (adapter)
Raw download ceiling 10 MiB Config.MaxRawResponseBytes (safety ceiling, independent of the text budget)

Note: Config.HTTPClient, when provided, replaces the internally constructed client entirely — Config.Timeout applies only to the internal client.

Errors

All sentinel errors are package-level values suitable for errors.Is:

Error Returned by Meaning
ErrEmptyQuery Search Empty or whitespace-only query
ErrEmptyQuestion Answer Empty question
ErrEmptyURL / ErrInvalidURL Fetch Missing or unparseable URL
ErrUnsupportedScheme Fetch Scheme other than http/https
ErrNilSearcher all methods Method called on a nil *Searcher
ErrResponseTooLarge Fetch Body exceeded the raw download ceiling
ErrUnexpectedContentType Fetch Content type outside the accepted allowlist
ErrLLMsTxtNotFound FetchLLMsTxt No llms.txt at the directory or site root (all probes 404/410)
ErrInvalidLLMsTxt FetchLLMsTxt / ParseLLMsTxt File exists but violates the required structure (missing H1)
ErrAnswerUnavailable Answer No answer provider configured
ErrAnswerProviderUnavailable Answer Requested provider not configured
ErrInvalidProviderResponse Answer Provider returned an unusable response

Additionally: every-source failure returns *SearchError (inspect .Failures), and non-2xx responses return *HTTPError (inspect with errors.As; see Fetch above). Error strings from configured providers have their API keys redacted before returning.

MCP Server

Install the stdio adapter:

go install github.com/jahrulnr/searchwire/cmd/searchwire-mcp@latest

Run it as an MCP stdio server (stdout is reserved for the MCP protocol; logs go to stderr):

searchwire-mcp
Tools
Name Input Output
search { "query": string, "limit"?: number } Searchwire response JSON (query, results, errors)
fetch { "url": string, "max_bytes"?: number } Page JSON (url, finalUrl, statusCode, contentType, title, text, links, headers, redirects, truncated, bytes, describedBy, markdownAlt, llmsTxt?, tempFile?)
answer { "question": string, "provider"?: string } Web-grounded answer JSON (question, answer, provider); registered only when an answer-provider key is available

Tool names drop the searchwire_ prefix because the MCP server name already carries it. Clients address tools as <server-name>.<tool> (e.g. searchwire.search); a prefixed tool name would be doubly redundant (searchwire.searchwire_search). The same applies if you mount this server under a different MCP client with its own namespace.

Recommended agent workflow: search → pick URLs → fetch. Snippets alone are incomplete; fetch promising hits for full content. When the fetched page is HTML and the site publishes a covering llms.txt, the parsed file is embedded in fetch's llmsTxt field so the agent can follow LLM-friendly documentation without a second call. When the response is truncated, the full text is also written to a file under SEARCHWIRE_TEMP_DIR (default $TMPDIR/searchwire) and surfaced as tempFile — the agent can read the rest with file_read instead of refetching.

answer is capability-driven: it is absent from MCP tool discovery when none of BRAVE_SEARCH_API_KEY, OPENROUTER_API_KEY, OPENAI_API_KEY, PERPLEXITY_API_KEY, ANTHROPIC_API_KEY, or XAI_API_KEY is available. Restart the MCP process after adding or removing a key so clients receive the updated tool list. When multiple providers are configured, omit provider to use the stable order Brave → OpenRouter → OpenAI → Perplexity → Anthropic → xAI, or pass a provider explicitly.

Brave keys must have access to the Answers plan; a Search-only subscription can still produce a provider billing/authorization error. OpenRouter uses its current openrouter:web_search server tool and defaults to openrouter/auto. OpenAI uses the Responses API web_search tool and defaults to the cost-focused gpt-5.6-luna. Override either model with configuration or its model env var.

Environment
Variable Scope Description
BRAVE_SEARCH_API_KEY library Optional official Brave Search API key; replaces the Brave HTML adapter when set
OPENROUTER_API_KEY library Optional OpenRouter key; enables the openrouter answer provider
OPENROUTER_MODEL library OpenRouter answer model (default openrouter/auto)
OPENAI_API_KEY library Optional OpenAI key; enables the openai answer provider
OPENAI_MODEL library OpenAI answer model (default gpt-5.6-luna)
PERPLEXITY_API_KEY library Optional Perplexity key; enables the perplexity answer provider
PERPLEXITY_PRESET library Perplexity Agent API preset (default low)
ANTHROPIC_API_KEY library Optional Anthropic key; enables the anthropic answer provider
ANTHROPIC_MODEL library Anthropic answer model (default claude-haiku-4-5)
XAI_API_KEY library Optional xAI key; enables the xai answer provider
XAI_MODEL library xAI answer model (default grok-4.5)
GITHUB_TOKEN library Optional GitHub API token (higher rate limits)
SERPER_API_KEY library Optional Serper.dev API key; enables the Serper.dev source
TAVILY_API_KEY library Optional Tavily API key; enables the Tavily source
SEARCHWIRE_LIMIT adapter Default result limit (int; unset = library default)
SEARCHWIRE_TIMEOUT adapter HTTP client timeout as a Go duration string (e.g. 10s; unset = library default)
SEARCHWIRE_MAX_BYTES adapter Default fetch/response byte budget (unset = library default)
SEARCHWIRE_TEMP_DIR adapter Directory used to materialize truncated fetch artifacts (default $TMPDIR/searchwire). The full extracted text is written here when fetch truncates, and surfaced to the agent as the tempFile field so it can read the rest with file_read (or framework equivalents).

Latency: search fans out across 4 sources and may take up to the configured timeout (default 10s). Clients should set appropriate tool timeouts.

Example MCP client config

Generic stdio MCP entry (Cursor, Claude Desktop, NusaShell custom MCP, etc.):

{
  "command": "searchwire-mcp",
  "args": [],
  "env": {
    "GITHUB_TOKEN": ""
  }
}

NusaShell can consume this as an external stdio MCP server via manual config. This repo does not ship a NusaShell plugin bundle.

Built-in sources

Searchwire fans out concurrently to:

  1. Brave Search (HTML by default; official JSON API when configured)
  2. Startpage (HTML)
  3. Wikipedia MediaWiki API (JSON)
  4. GitHub Search API (repositories + issues; falls back to repositories when issues search fails)
  5. Serper.dev (Google Search API; requires SERPER_API_KEY)
  6. Tavily (AI search; requires TAVILY_API_KEY)

Source ordering is also the tie-break order when fused scores are equal.

Unauthenticated GitHub search is limited to about 10 requests per minute across search endpoints.

Configuration

searchwire.Config groups runtime settings and optional integrations:

searcher := searchwire.New(searchwire.Config{
	Limit:      5,
	HTTPClient: httpClient,
	Brave: searchwire.BraveConfig{
		APIKey: "...", // or read BRAVE_SEARCH_API_KEY
	},
	OpenRouter: searchwire.OpenRouterConfig{
		APIKey: "...", // or read OPENROUTER_API_KEY
		Model:  "openrouter/auto",
	},
	OpenAI: searchwire.OpenAIConfig{
		APIKey: "...", // or read OPENAI_API_KEY
		Model:  "gpt-5.6-luna",
	},
	Perplexity: searchwire.PerplexityConfig{
		APIKey: "...", // or read PERPLEXITY_API_KEY
		Preset: "low",
	},
	Anthropic: searchwire.AnthropicConfig{
		APIKey: "...", // or read ANTHROPIC_API_KEY
		Model:  "claude-haiku-4-5",
	},
	XAI: searchwire.XAIConfig{
		APIKey: "...", // or read XAI_API_KEY
		Model:  "grok-4.5",
	},
	GitHub: searchwire.GitHubConfig{
		Token: "ghp_...", // or read GITHUB_TOKEN
	},
	Serper: searchwire.SerperConfig{
		APIKey: "...", // or read SERPER_API_KEY
	},
	Tavily: searchwire.TavilyConfig{
		APIKey: "...", // or read TAVILY_API_KEY
	},
})
Brave
Field Default Description
Enabled true Register the Brave source
APIKey — Official Brave Search API key; overrides env
APIKeyEnv BRAVE_SEARCH_API_KEY Environment variable used for the key

With no key, the zero-configuration Brave HTML adapter remains active. When a key is present, Searchwire replaces it with the official Brave Web Search API under the same brave source name; it does not query both or double-count duplicate results.

The same resolved key enables Searcher.CanAnswer() and Searcher.Answer() using Brave's blocking, single-search Answers endpoint. The initial answer contract returns provider-grounded text; streaming research mode and structured citation extraction are not yet exposed.

LLM answer providers
Provider Key env Model/preset env Default Hosted web tool
OpenRouter OPENROUTER_API_KEY OPENROUTER_MODEL openrouter/auto openrouter:web_search
OpenAI OPENAI_API_KEY OPENAI_MODEL gpt-5.6-luna Responses API web_search
Perplexity PERPLEXITY_API_KEY PERPLEXITY_PRESET low preset Agent API web search + fetch
Anthropic ANTHROPIC_API_KEY ANTHROPIC_MODEL claude-haiku-4-5 Messages API web search
xAI XAI_API_KEY XAI_MODEL grok-4.5 Responses API web search

OpenRouter can route web-grounded requests to supported OpenAI, Anthropic, Google, xAI, and Perplexity model families without adding another Searchwire credential. Set OPENROUTER_MODEL when deterministic model choice or cost is more important than automatic routing.

GitHub
Field Default Description
Enabled true Register the GitHub source
Token — Bearer token; overrides env
TokenEnv GITHUB_TOKEN Env var for token
SearchIssues true Repo+issues (B); false = repos only (A)

When SearchIssues is true and issues search fails, repository results are still returned.

Serper.dev
Field Default Description
Enabled false Register the Serper source
APIKey — Serper.dev API key; overrides env
APIKeyEnv SERPER_API_KEY Env var for the key

Serper.dev exposes a Google Search API that returns organic results as JSON. The source is opt-in: it is only registered when a key resolves and Enabled is not explicitly false. There is no HTML fallback.

Tavily
Field Default Description
Enabled false Register the Tavily source
APIKey — Tavily API key; overrides env
APIKeyEnv TAVILY_API_KEY Env var for the key

Tavily is an AI search API that returns ranked results with content snippets. The source is opt-in: it is only registered when a key resolves and Enabled is not explicitly false. There is no HTML fallback.

Future development

These Config fields exist for upcoming work only. They are ignored by New() today.

Block Planned integration
GoogleConfig Google Programmable Search JSON API (APIKey / GOOGLE_API_KEY, CX / GOOGLE_CX)
CustomSearchConfig Caller-provided search endpoint (URL)

Do not rely on them for behavior until a source adapter lands and tests cover it.

Partial failures

When at least one source succeeds, Search returns a *Response with merged results and any source failures in Response.Errors. When every source fails, Search returns a *SearchError listing each failure.

Limitations

  • HTML adapters can break when source markup changes.
  • Ordinary text/web results only; no images, news, maps, or shopping categories.
  • Fetch is text-oriented (no headless browser / JS execution).
  • No CAPTCHA solving, proxy rotation, or anti-bot bypasses.
  • Not a promise of SearXNG parity or engine compatibility.

Development

gofmt -w *.go cmd/searchwire-mcp/*.go
go mod tidy
go build ./cmd/searchwire-mcp
go install github.com/jahrulnr/searchwire/cmd/searchwire-mcp@latest
go test ./...
go test -race ./...
go vet ./...
SEARCHWIRE_LIVE=1 go test -run TestLiveSearch -v

Identity and license

The metasearch aggregation concept is inspired by systems such as SearXNG, but Searchwire is not a SearXNG client, port, configuration-compatible implementation, or source-code derivative. SearXNG is AGPL; Searchwire is MIT.

This library is pre-v1; APIs may change.

MIT License. See LICENSE.

Documentation

Overview

Package searchwire is a zero-configuration Go metasearch and page-fetch runtime for agent tooling. Callers provide a search query and optional Config; built-in sources fan out concurrently, partial failures are reported in the response, and results are deduplicated and ranked. Built-in sources include Brave Search, Startpage, Wikipedia, GitHub, and the opt-in Serper.dev and Tavily API sources (registered only when their API keys resolve). Searcher.Fetch retrieves one http(s) URL as readable text, with per-call controls for byte budget, conditional GET, and Range requests; it also surfaces llmstxt.org v2 discovery hints (the covering llms.txt file and markdown alternates) when a page advertises them, and on 2xx HTML responses auto-attaches the parsed covering llms.txt (Searcher.LLMsTxt / Searcher.LLMsTxtURL). Truncated responses are also materialized to Config.TempDir (default os.TempDir()/searchwire) and surfaced as Searcher.TempFile so the agent can read the rest with file_read. Searcher.FetchLLMsTxt remains available as a standalone probe for callers that want the parsed llms.txt without fetching the page first. Searcher.SearchWithOptions runs a search with per-call options such as a different result limit than Config.Limit. When a supported provider API key is configured, Searcher.Answer returns web-grounded answers and CanAnswer reports that optional capability. The same library powers the searchwire-mcp stdio adapter under cmd/searchwire-mcp, where the answer tool is registered only when that capability is available at startup.

Index

Constants

View Source
const (
	AnswerProviderBrave      = "brave"
	AnswerProviderOpenRouter = "openrouter"
	AnswerProviderOpenAI     = "openai"
	AnswerProviderPerplexity = "perplexity"
	AnswerProviderAnthropic  = "anthropic"
	AnswerProviderXAI        = "xai"
)

Variables

View Source
var (
	ErrEmptyQuery            = errors.New("searchwire: query is required")
	ErrEmptyQuestion         = errors.New("searchwire: question is required")
	ErrEmptyURL              = errors.New("searchwire: url is required")
	ErrInvalidURL            = errors.New("searchwire: invalid url")
	ErrUnsupportedScheme     = errors.New("searchwire: unsupported url scheme")
	ErrNilSearcher           = errors.New("searchwire: nil searcher")
	ErrResponseTooLarge      = errors.New("searchwire: response exceeds configured limit")
	ErrUnexpectedContentType = errors.New("searchwire: unexpected content type")
	// ErrLLMsTxtNotFound is returned by FetchLLMsTxt when no llms.txt file
	// exists at the enclosing directory or the site root (all probes 404/410).
	ErrLLMsTxtNotFound = errors.New("searchwire: llms.txt not found")
	// ErrInvalidLLMsTxt is returned when llms.txt content exists but violates
	// the required structure (missing H1 heading).
	ErrInvalidLLMsTxt            = errors.New("searchwire: invalid llms.txt")
	ErrAnswerUnavailable         = errors.New("searchwire: answer provider is unavailable")
	ErrAnswerProviderUnavailable = errors.New("searchwire: requested answer provider is unavailable")
	ErrInvalidProviderResponse   = errors.New("searchwire: invalid provider response")
)

Functions

This section is empty.

Types

type Answer

type Answer struct {
	Question string
	Text     string
	Provider string
}

Answer is one web-grounded response from an optional answer provider.

type AnswerOptions

type AnswerOptions struct {
	Provider string
}

AnswerOptions selects an optional configured provider. An empty provider uses the first available provider in Searchwire's stable priority order.

type AnthropicConfig

type AnthropicConfig struct {
	// Enabled explicitly enables or disables the provider. When set to false,
	// env var fallback is skipped. nil means default (enabled if key resolves).
	Enabled *bool
	// APIKey overrides APIKeyEnv (default ANTHROPIC_API_KEY).
	APIKey    string
	APIKeyEnv string
	// Model overrides ModelEnv (default ANTHROPIC_MODEL), then falls back to
	// the cost-focused Claude Haiku 4.5 model.
	Model    string
	ModelEnv string
}

AnthropicConfig controls web-grounded answers through the Messages API.

type BraveConfig

type BraveConfig struct {
	// Enabled registers the Brave source. Default true.
	Enabled *bool

	// APIKey is sent as X-Subscription-Token. When empty, APIKeyEnv is read.
	APIKey string
	// APIKeyEnv names the environment variable for APIKey
	// (default BRAVE_SEARCH_API_KEY).
	APIKeyEnv string
}

BraveConfig controls the built-in Brave source. Without an API key, Searchwire uses Brave's public HTML results. Supplying a key switches the source to the official Brave Search API while preserving the source name.

type Config

type Config struct {
	HTTPClient       HTTPClient
	UserAgent        string
	Limit            int
	MaxResponseBytes int64 // budget for extracted readable text (per fetch)
	// MaxRawResponseBytes is the hard safety ceiling on the raw HTML/plaintext
	// download, independent of the caller text budget.
	MaxRawResponseBytes int64
	Timeout             time.Duration

	Brave  BraveConfig
	Serper SerperConfig
	Tavily TavilyConfig
	GitHub GitHubConfig
	// TempDir overrides the directory used to materialize truncated fetch
	// artifacts. Empty falls back to os.TempDir() + "searchwire". The
	// directory is created on demand. Each artifact is named
	// "<UTC-timestamp>-<sha256-prefix>-<safe-name>.txt" so concurrent calls
	// never collide.
	TempDir string
	// OpenRouter and OpenAI are optional answer providers. They do not add
	// metasearch sources; a resolved API key enables the answer capability.
	OpenRouter OpenRouterConfig
	OpenAI     OpenAIConfig
	Perplexity PerplexityConfig
	Anthropic  AnthropicConfig
	XAI        XAIConfig
	// Google and Custom are future-dev placeholders. They are not read by New()
	// until their source adapters are implemented.
	Google GoogleConfig
	Custom CustomSearchConfig
}

Config configures a Searcher. The zero value uses built-in defaults and zero-configuration sources. Optional integrations read credentials from explicit fields first, then from named environment variables.

func DefaultConfig

func DefaultConfig() Config

DefaultConfig returns the zero-configuration defaults.

type CustomSearchConfig

type CustomSearchConfig struct {
	URL string
}

CustomSearchConfig is a future-dev placeholder for a caller-provided search endpoint. Setting URL has no effect on Search today.

type FetchOptions

type FetchOptions struct {
	// MaxBytes caps the response body. When <= 0, Config.MaxResponseBytes is
	// used (or the library default).
	MaxBytes int64
	// IfNoneMatch is sent as If-None-Match for conditional GET. A 304
	// response is returned as a FetchResult with StatusCode 304 (not an
	// error).
	IfNoneMatch string
	// IfModifiedSince is sent as If-Modified-Since for conditional GET.
	IfModifiedSince string
	// Range is sent as the Range header (e.g. "bytes=10-20"). A 206 response
	// is returned as a normal FetchResult with StatusCode 206.
	Range string
}

FetchOptions controls one fetch call: byte budget, conditional GET headers, and a Range request. Zero values mean "do not send".

type FetchResult

type FetchResult struct {
	URL         string
	FinalURL    string
	StatusCode  int
	ContentType string
	Title       string
	Text        string
	Links       []Link
	// Headers carries selected response headers useful to agents:
	// ETag, Last-Modified, X-RateLimit-Remaining, X-RateLimit-Reset,
	// Retry-After (on 429/503). Empty when absent.
	Headers map[string]string
	// Redirects is the number of redirects followed (0 = direct hit).
	Redirects int
	Truncated bool
	Bytes     int
	// DescribedBy is the absolute URL of the llms.txt file covering this
	// page, discovered via <link rel="describedby"> or an equivalent HTTP
	// Link header (llmstxt.org v2). Empty when the site does not advertise
	// one.
	DescribedBy string
	// MarkdownAlt is the absolute URL of a markdown alternate version of
	// this page (<link rel="alternate" type="text/markdown"> or the HTTP
	// Link header equivalent). Empty when absent.
	MarkdownAlt string
	// LLMsTxt is the parsed llms.txt file covering this page, populated
	// automatically when the source fetch succeeds (2xx), the response
	// content type is HTML, and a covering llms.txt file can be retrieved
	// and parsed. The probe follows the same rules as Searcher.FetchLLMsTxt
	// (most-specific-directory first, then the site root; 404/410 falls
	// through; invalid/soft-404 content is skipped). Non-HTML responses
	// (JSON, plain text, etc.) skip the probe entirely. Any probe failure
	// is non-fatal: callers always get the source result they asked for.
	LLMsTxt *LLMsTxt
	// LLMsTxtURL is the absolute final URL of the llms.txt file that
	// produced LLMsTxt (post-redirect). Empty when LLMsTxt is nil.
	LLMsTxtURL string
	// TempFile is the absolute path of a file under the configured temp
	// directory (Config.TempDir, default os.TempDir()/searchwire) holding
	// the full extracted text when Truncated is true. The file lets the
	// agent read the rest of the response (e.g. via file_read) without
	// having to refetch. Empty when the response was not truncated or when
	// the write failed (the inline Text still carries the truncated
	// payload with the [truncated] marker in that case).
	TempFile string
}

FetchResult is the readable text extracted from one fetched page.

type GitHubConfig

type GitHubConfig struct {
	// Enabled registers the GitHub source. Default true.
	Enabled *bool

	// Token is sent as Bearer auth. When empty, TokenEnv is read.
	Token string
	// TokenEnv names the environment variable for Token (default GITHUB_TOKEN).
	TokenEnv string

	// SearchIssues enables repository+issues mode. When false, only repositories
	// are searched. Default true. When true, repository search still serves as
	// the fallback if issues search fails.
	SearchIssues *bool
}

GitHubConfig controls the built-in GitHub Search API source.

type GoogleConfig

type GoogleConfig struct {
	// APIKey is the Google Custom Search JSON API key (future).
	APIKey string
	// APIKeyEnv names the env var for APIKey (planned default: GOOGLE_API_KEY).
	APIKeyEnv string
	// CX is the programmable search engine ID (future).
	CX string
	// CXEnv names the env var for CX (planned default: GOOGLE_CX).
	CXEnv string
}

GoogleConfig is a future-dev placeholder for Google Programmable Search Engine. Setting these fields has no effect on Search today.

type HTTPClient

type HTTPClient interface {
	Do(*http.Request) (*http.Response, error)
}

HTTPClient is satisfied by *http.Client and lightweight test doubles.

type HTTPError

type HTTPError struct {
	StatusCode int
	Status     string
	Body       string
	// RetryAfter is the parsed Retry-After header value in seconds when the
	// server provided one (e.g. on 429/503). Zero when absent or unparseable.
	RetryAfter int
	// ErrorBody is the structured error envelope parsed from Body when it is
	// valid JSON of the form {"error":...} or {"errors":[...]} or
	// {"message":"..."} or {"detail":"..."}. Nil when Body is not JSON or
	// does not match a known shape.
	ErrorBody map[string]any
}

HTTPError reports a non-2xx HTTP response from a source.

func (*HTTPError) Error

func (e *HTTPError) Error() string
type LLMsLink struct {
	Name string
	URL  string
	// Notes is the optional free-text after the ": " separator.
	Notes string
}

LLMsLink is one "- [name](url): notes" entry from a section's file list.

type LLMsSection

type LLMsSection struct {
	Name  string
	Links []LLMsLink
}

LLMsSection is one H2-delimited group of links.

type LLMsTxt

type LLMsTxt struct {
	// Title is the required H1 project or site name.
	Title string
	// Summary is the blockquote summary following the H1. Empty when absent.
	Summary string
	// Description is free-form prose between the summary and the first H2
	// section, one trimmed line per entry joined with newlines. Empty when
	// absent.
	Description string
	// Sections are the H2-delimited groups in document order.
	Sections []LLMsSection
}

LLMsTxt is a parsed /llms.txt file (llmstxt.org v2 format).

func ParseLLMsTxt

func ParseLLMsTxt(content string) (*LLMsTxt, error)

ParseLLMsTxt parses llms.txt Markdown content per the llmstxt.org v2 format: an optional BOM, a required H1 title, an optional blockquote summary, optional free-form description, then H2-delimited sections whose lists contain "[name](url)" links with optional ": notes". Parsing is lenient: unrecognized lines are ignored rather than rejected. It returns an error wrapping ErrInvalidLLMsTxt only when the required H1 is missing.

type LLMsTxtResult

type LLMsTxtResult struct {
	// URL is the requested llms.txt URL and FinalURL the post-redirect URL.
	URL      string
	FinalURL string
	// StatusCode is the HTTP status of the successful fetch (2xx family).
	StatusCode int
	// Headers carries selected response headers (same set as Fetch) so
	// callers can do conditional GETs on later refreshes.
	Headers map[string]string
	// LLMsTxt is the parsed file content.
	LLMsTxt *LLMsTxt
}

LLMsTxtResult is a fetched and parsed llms.txt file.

type Link struct {
	Text string
	Href string
}

Link is one anchor collected from an HTML page (text + href). Only safe http/https hrefs are collected; javascript:, data:, and protocol-relative URLs are skipped.

type OpenAIConfig

type OpenAIConfig struct {
	// Enabled explicitly enables or disables the provider. When set to false,
	// env var fallback is skipped. nil means default (enabled if key resolves).
	Enabled *bool
	// APIKey overrides APIKeyEnv (default OPENAI_API_KEY).
	APIKey    string
	APIKeyEnv string
	// Model overrides ModelEnv (default OPENAI_MODEL), then falls back to the
	// cost-focused gpt-5.6-luna model.
	Model    string
	ModelEnv string
}

OpenAIConfig controls web-grounded answers through the OpenAI Responses API.

type OpenRouterConfig

type OpenRouterConfig struct {
	// Enabled explicitly enables or disables the provider. When set to false,
	// env var fallback is skipped. nil means default (enabled if key resolves).
	Enabled *bool
	// APIKey overrides APIKeyEnv (default OPENROUTER_API_KEY).
	APIKey    string
	APIKeyEnv string
	// Model overrides ModelEnv (default OPENROUTER_MODEL), then falls back to
	// openrouter/auto.
	Model    string
	ModelEnv string
}

OpenRouterConfig controls web-grounded answers through OpenRouter.

type PerplexityConfig

type PerplexityConfig struct {
	// Enabled explicitly enables or disables the provider. When set to false,
	// env var fallback is skipped. nil means default (enabled if key resolves).
	Enabled *bool
	// APIKey overrides APIKeyEnv (default PERPLEXITY_API_KEY).
	APIKey    string
	APIKeyEnv string
	// Preset overrides PresetEnv (default PERPLEXITY_PRESET), then falls back
	// to Perplexity's cost-focused low preset.
	Preset    string
	PresetEnv string
}

PerplexityConfig controls web-grounded answers through the Agent API.

type Response

type Response struct {
	Query   string
	Results []Result
	Errors  []SourceError
}

Response is the merged output of a metasearch query.

type Result

type Result struct {
	Title   string
	URL     string
	Snippet string
	Sources []string
	Score   float64
}

Result is one ranked web/text hit after deduplication and fusion.

type SearchError

type SearchError struct {
	Failures []SourceError
}

SearchError is returned when every built-in source fails.

func (*SearchError) Error

func (e *SearchError) Error() string

type SearchOptions

type SearchOptions struct {
	// Limit overrides the configured result limit for this call: it caps the
	// fused results returned and bounds each source's requested count.
	// When <= 0, Config.Limit applies.
	Limit int
	// Sources restricts the query to the named registered sources. Names
	// are matched case-insensitively after trimming whitespace; unknown
	// names are ignored. When no registered source matches, the search
	// fails like a fan-out with no sources. Empty means every registered
	// source participates (merged metasearch). Callers can combine this
	// with Searcher.Sources() to implement routing policies (round-robin,
	// random, fixed provider) on top of the searcher.
	Sources []string
}

SearchOptions controls one search call. Zero values keep the configured defaults.

type Searcher

type Searcher struct {
	// contains filtered or unexported fields
}

Searcher fans out to built-in sources, merges duplicates, and ranks results.

func New

func New(cfg Config) *Searcher

New returns a Searcher configured by cfg. Use DefaultConfig() or the zero value for zero-configuration behavior.

func (*Searcher) Answer

func (s *Searcher) Answer(ctx context.Context, question string) (*Answer, error)

Answer returns a web-grounded answer when an answer provider is configured.

func (*Searcher) AnswerWithOptions

func (s *Searcher) AnswerWithOptions(ctx context.Context, question string, options AnswerOptions) (*Answer, error)

AnswerWithOptions returns a web-grounded answer from the selected provider.

func (*Searcher) AvailableAnswerProviders

func (s *Searcher) AvailableAnswerProviders() []string

AvailableAnswerProviders returns configured provider identifiers in stable default-selection order.

func (*Searcher) CanAnswer

func (s *Searcher) CanAnswer() bool

CanAnswer reports whether this Searcher has a configured answer provider.

func (*Searcher) Fetch

func (s *Searcher) Fetch(ctx context.Context, rawURL string) (*FetchResult, error)

Fetch retrieves one http(s) URL and returns readable text. HTML responses are stripped to title + visible text. Plain text is returned as-is. Extracted text — not raw HTML — is capped at Config.MaxResponseBytes; Truncated reports when the returned text was cut to fit. Raw HTML downloads are additionally bounded by an internal hard ceiling (defaultMaxRawResponseBytes). This is a local agent helper, not a hardened proxy.

func (*Searcher) FetchLLMsTxt

func (s *Searcher) FetchLLMsTxt(ctx context.Context, rawURL string) (*LLMsTxtResult, error)

FetchLLMsTxt retrieves and parses the llms.txt file covering rawURL. Per the llmstxt.org v2 proposal a file covers every URL under its path, so the enclosing directory's llms.txt is probed first and the site root second; the most specific file found wins. Directories without their own file typically 404 and fall through to the root probe.

When every candidate 404s/410s, the error wraps ErrLLMsTxtNotFound. A file that exists but lacks the required H1 wraps ErrInvalidLLMsTxt. Other errors (network failures, non-404 HTTP errors, unsupported content types) surface as-is. Conditional GET support matches Fetch via LLMsTxtResult.Headers.

func (*Searcher) FetchWithLimit

func (s *Searcher) FetchWithLimit(ctx context.Context, rawURL string, maxBytes int64) (*FetchResult, error)

FetchWithLimit is Fetch with an optional per-call byte budget for the extracted readable text. When maxBytes <= 0, Config.MaxResponseBytes is used. The budget applies to the returned text, not the raw HTML download, so HTML overhead (markup, scripts, styles) does not consume the caller's budget.

func (*Searcher) FetchWithOptions

func (s *Searcher) FetchWithOptions(ctx context.Context, rawURL string, opts FetchOptions) (*FetchResult, error)

FetchWithOptions is Fetch with conditional-GET / Range / byte-budget controls.

func (*Searcher) Search

func (s *Searcher) Search(ctx context.Context, query string) (*Response, error)

Search runs the query across all built-in sources concurrently. When at least one source succeeds, partial failures are returned in Response.Errors.

func (*Searcher) SearchWithOptions

func (s *Searcher) SearchWithOptions(ctx context.Context, query string, opts SearchOptions) (*Response, error)

SearchWithOptions is Search with per-call options such as a different result limit than the configured default.

func (*Searcher) Sources

func (s *Searcher) Sources() []string

Sources returns the names of the registered search sources in configuration order. Callers use it to implement routing policies (round-robin, random, fixed provider) via SearchOptions.Sources.

type SerperConfig

type SerperConfig struct {
	// Enabled registers the Serper source. Default false.
	Enabled *bool

	// APIKey is sent as X-API-KEY. When empty, APIKeyEnv is read.
	APIKey string
	// APIKeyEnv names the environment variable for APIKey
	// (default SERPER_API_KEY).
	APIKeyEnv string
}

SerperConfig controls the optional Serper.dev (Google Search API) source. Unlike Brave, Serper has no HTML fallback, so the source is only registered when an API key resolves and Enabled is not explicitly false (default false).

type SourceError

type SourceError struct {
	Source string
	Error  string
}

SourceError records one source failure without aborting the whole search.

type TavilyConfig

type TavilyConfig struct {
	// Enabled registers the Tavily source. Default false.
	Enabled *bool

	// APIKey is sent as Bearer auth. When empty, APIKeyEnv is read.
	APIKey string
	// APIKeyEnv names the environment variable for APIKey
	// (default TAVILY_API_KEY).
	APIKeyEnv string
}

TavilyConfig controls the optional Tavily (AI search) source. The source is only registered when an API key resolves and Enabled is not explicitly false (default false).

type XAIConfig

type XAIConfig struct {
	// Enabled explicitly enables or disables the provider. When set to false,
	// env var fallback is skipped. nil means default (enabled if key resolves).
	Enabled *bool
	// APIKey overrides APIKeyEnv (default XAI_API_KEY).
	APIKey    string
	APIKeyEnv string
	// Model overrides ModelEnv (default XAI_MODEL), then falls back to
	// grok-4.5.
	Model    string
	ModelEnv string
}

XAIConfig controls web-grounded answers through the xAI Responses API.

Directories

Path Synopsis
cmd
searchwire-mcp command
Command searchwire-mcp exposes Searchwire as an MCP stdio server.
Command searchwire-mcp exposes Searchwire as an MCP stdio server.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL