Documentation
¶
Overview ¶
Package provider defines the common interface every upstream LLM backend implements, and the normalized streaming event shape the router/server work with regardless of which provider served the request.
Index ¶
Constants ¶
const DefaultTimeout = 120 * time.Second
DefaultTimeout bounds each *phase* of a provider call, not the whole exchange: first the connect/headers wait, then every quiet stretch of the response stream (the watchdog resets on each byte read — see newStreamWatchdog). Its old meaning — an http.Client.Timeout covering the entire request including reading the streamed body — killed any legitimately long generation mid-answer with "stream read: context deadline exceeded ... while reading body", which is exactly what a slow-but-healthy reasoning model produces on a big response. config.Tunables.ProviderTimeout raises or lowers it.
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Anthropic ¶
type Anthropic struct {
// contains filtered or unexported fields
}
Anthropic talks to the native Messages API (x-api-key auth, "system" is a top-level field rather than a message role, streaming uses named SSE events instead of OpenAI's flat "data:" chunks, and tool use/results are content blocks rather than separate message roles).
apiKey is always a real, permanent Anthropic API key — including when the account was connected via the wizard's browser-login flow: that flow (internal/oauthflow.AnthropicAuthorize) exchanges the OAuth token for one of these right away rather than handing back something short-lived, since a raw Claude Pro/Max OAuth token turned out not to be a usable inference credential on its own (see that file's doc comment for what was actually live-verified). So this adapter never needed a second, refreshable credential shape the way internal/provider/openai_responses.go does.
func NewAnthropic ¶
NewAnthropic constructs the Anthropic adapter. baseURL defaults to the public API when empty.
func (*Anthropic) ChatCompletion ¶
func (p *Anthropic) ChatCompletion(ctx context.Context, req openai.ChatCompletionRequest) (<-chan StreamEvent, error)
func (Anthropic) SupportsImages ¶
func (c Anthropic) SupportsImages() bool
func (Anthropic) SupportsTools ¶
func (c Anthropic) SupportsTools() bool
type Gemini ¶
type Gemini struct {
// contains filtered or unexported fields
}
Gemini talks to Google's native streamGenerateContent endpoint (API key as a query param, "contents"/"parts" request shape, "user"/"model" roles instead of "user"/"assistant", and tool calls as functionCall/ functionResponse parts rather than a separate message role).
func NewGemini ¶
NewGemini constructs the Gemini adapter. baseURL defaults to the public API when empty.
func (*Gemini) ChatCompletion ¶
func (p *Gemini) ChatCompletion(ctx context.Context, req openai.ChatCompletionRequest) (<-chan StreamEvent, error)
func (Gemini) SupportsImages ¶
func (c Gemini) SupportsImages() bool
func (Gemini) SupportsTools ¶
func (c Gemini) SupportsTools() bool
type HTTPError ¶
type HTTPError struct {
Provider string
StatusCode int
Status string
// RetryAfter is parsed from the upstream response's Retry-After
// header (seconds form only — the rarer HTTP-date form isn't worth
// the parsing complexity for the providers Kram actually talks to).
// Zero when absent or unparseable; a caller deciding on a Gateway
// Round retry then falls back to its own computed backoff.
RetryAfter time.Duration
// Detail is a bounded upstream error body when the adapter can safely
// capture one. It makes otherwise opaque 400s diagnosable without logging
// request credentials or the full request payload.
Detail string
}
HTTPError wraps a non-2xx upstream response so a caller can report the real status code instead of just a formatted string — this is what lets the router's attempt executor tell an HTTP 500 apart from an HTTP 200 whose *content* got rejected by the ResponseGate (see openai.AttemptInfo.HTTPStatus and DECISIONS.md).
type OpenAICompatible ¶
type OpenAICompatible struct {
// contains filtered or unexported fields
}
OpenAICompatible talks to any backend that already speaks the OpenAI chat-completions wire format: OpenAI itself, OpenRouter, opencode zen, and most other aggregators. Only base URL, API key and an optional pinned model differ between them.
func NewOpenAICompatible ¶
func NewOpenAICompatible(id, baseURL, apiKey, model string, temperature *float64, caps capabilities) *OpenAICompatible
NewOpenAICompatible constructs an adapter for an OpenAI-shaped backend.
func (*OpenAICompatible) ChatCompletion ¶
func (p *OpenAICompatible) ChatCompletion(ctx context.Context, req openai.ChatCompletionRequest) (<-chan StreamEvent, error)
func (*OpenAICompatible) ID ¶
func (p *OpenAICompatible) ID() string
func (*OpenAICompatible) Kind ¶
func (p *OpenAICompatible) Kind() string
func (OpenAICompatible) SupportsImages ¶
func (c OpenAICompatible) SupportsImages() bool
func (OpenAICompatible) SupportsTools ¶
func (c OpenAICompatible) SupportsTools() bool
type OpenAIResponses ¶
type OpenAIResponses struct {
// contains filtered or unexported fields
}
OpenAIResponses talks to the Codex backend a ChatGPT Plus/Pro/Team subscription unlocks via browser login (internal/oauthflow.OpenAIAuthorize) — a different product from the standard OpenAI developer API that OpenAICompatible (openai_compat.go) talks to. It uses OpenAI's Responses wire format ("input" items instead of "messages", typed streaming events instead of flat delta chunks) and only serves a restricted set of Codex-branded models, never the general OpenAI catalog.
This adapter is experimental: its request/response shape follows OpenAI's publicly documented Responses API, but the exact headers this specific ChatGPT-authenticated backend expects beyond "Authorization: Bearer" were not fully confirmed against a real account before shipping (see DECISIONS.md) — treat failures here as a signal to re-check against a live subscription, not necessarily a Kram bug.
func NewOpenAIResponses ¶
func NewOpenAIResponses(id, baseURL string, resolve func(context.Context) (string, error), model string, caps capabilities) *OpenAIResponses
NewOpenAIResponses constructs the adapter. resolve is always required — unlike every other adapter, there is no static-API-key path for this product; a credential only ever comes from a refreshable OAuth token (see internal/credentials.Store.Resolve).
func (*OpenAIResponses) ChatCompletion ¶
func (p *OpenAIResponses) ChatCompletion(ctx context.Context, req openai.ChatCompletionRequest) (<-chan StreamEvent, error)
func (*OpenAIResponses) ID ¶
func (p *OpenAIResponses) ID() string
func (*OpenAIResponses) Kind ¶
func (p *OpenAIResponses) Kind() string
func (OpenAIResponses) SupportsImages ¶
func (c OpenAIResponses) SupportsImages() bool
func (OpenAIResponses) SupportsTools ¶
func (c OpenAIResponses) SupportsTools() bool
type Provider ¶
type Provider interface {
// ID is the stable identifier used in config, logs and telemetry.
ID() string
// Kind is the adapter family (anthropic, gemini, openai-compat) — useful
// for diagnostics, not for routing decisions.
Kind() string
// SupportsImages and SupportsTools reflect the provider's configured
// capabilities (internal/config.ProviderConfig) — callers must check
// these before sending images or tool definitions.
SupportsImages() bool
SupportsTools() bool
// ChatCompletion issues the request upstream and streams back normalized
// events. The channel is always closed by the provider, exactly once,
// after a final event with Done=true or Err set.
ChatCompletion(ctx context.Context, req openai.ChatCompletionRequest) (<-chan StreamEvent, error)
}
Provider is anything that can serve a chat completion request.
func Build ¶
func Build(cfg config.ProviderConfig, resolve func(context.Context) (string, error), timeout time.Duration) (Provider, error)
Build constructs the concrete adapter for a provider config entry. resolve is nil for cfg.AuthMode == "" (today's static-env-var behavior, unchanged) and must be non-nil for cfg.AuthMode == "oauth" — it's called by the adapter on every request to get a currently-valid, possibly-just-refreshed credential (see internal/credentials.Store.Resolve), since gateway providers are built once at startup and held for the gateway's whole lifetime, long past when a short-lived OAuth token would expire. Only openai-responses (the ChatGPT-login-only Codex backend) ever uses this path — Anthropic connected via browser login ends up with a real, permanent API key (see internal/oauthflow/anthropic.go), so it always takes the static branch below like any pasted key. timeout is the per-request HTTP timeout to apply to the built adapter; pass 0 to keep the adapter's DefaultTimeout (the gateway resolves the configured value once and passes it here — see config.Tunables).
type StreamEvent ¶
type StreamEvent struct {
Delta string
// Reasoning is a fragment of a reasoning-capable model's chain-of-
// thought, kept separate from Delta because it is not the model's
// actual answer — a caller must never relay it as assistant message
// content. It exists purely as a liveness signal: a provider that's
// only produced reasoning so far is actively working, not stalled,
// which router.BoundedPeek needs to know (see DECISIONS.md — a
// reasoning phase alone routinely runs past what used to be a fixed
// peek timeout, and was being misread as a dead attempt).
Reasoning string
// ToolCallProgress marks an event that carried tool-call argument
// fragments but no answer content and no reasoning — real evidence
// the provider is actively producing tokens, exactly like Reasoning
// is, but kept as a separate bool rather than folded into it: a
// caller must never mistake tool-call JSON fragments for chain-of-
// thought text. Exists purely as a liveness signal for
// router.BoundedPeek — a provider streaming a long tool call's
// arguments token-by-token, with no leading text, used to look like
// dead silence and get killed as stalled even though it was actively
// working the whole time (every adapter accumulates tool-call
// fragments internally without emitting any StreamEvent at all,
// unless this is set).
ToolCallProgress bool
Done bool
Usage *openai.Usage
ToolCalls []openai.ToolCall
ProviderItems []openai.ProviderItem
Err error
}
StreamEvent is one normalized increment of a chat completion stream. Providers translate their native wire format into a sequence of these. ToolCalls is only set on the final (Done) event — no provider streams partial tool-call deltas as assembled ToolCall values; the daemon's agent loop makes buffered (non-streaming) gateway calls by default and waits for the complete decision before acting on it (see internal/daemon/agent.Config.PreferStreaming for the opt-in exception), so there's nothing further upstream that needs them incrementally either. ToolCallProgress below exists for a narrower reason: not exposing partial arguments, just proving the provider is still alive.