provider

package
v0.7.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 28, 2026 License: MIT Imports: 15 Imported by: 0

Documentation

Overview

Package provider defines the common interface every upstream LLM backend implements, and the normalized streaming event shape the router/server work with regardless of which provider served the request.

Index

Constants

View Source
const DefaultTimeout = 120 * time.Second

DefaultTimeout bounds each *phase* of a provider call, not the whole exchange: first the connect/headers wait, then every quiet stretch of the response stream (the watchdog resets on each byte read — see newStreamWatchdog). Its old meaning — an http.Client.Timeout covering the entire request including reading the streamed body — killed any legitimately long generation mid-answer with "stream read: context deadline exceeded ... while reading body", which is exactly what a slow-but-healthy reasoning model produces on a big response. config.Tunables.ProviderTimeout raises or lowers it.

Variables

This section is empty.

Functions

This section is empty.

Types

type Anthropic

type Anthropic struct {
	// contains filtered or unexported fields
}

Anthropic talks to the native Messages API (x-api-key auth, "system" is a top-level field rather than a message role, streaming uses named SSE events instead of OpenAI's flat "data:" chunks, and tool use/results are content blocks rather than separate message roles).

apiKey is always a real, permanent Anthropic API key — including when the account was connected via the wizard's browser-login flow: that flow (internal/oauthflow.AnthropicAuthorize) exchanges the OAuth token for one of these right away rather than handing back something short-lived, since a raw Claude Pro/Max OAuth token turned out not to be a usable inference credential on its own (see that file's doc comment for what was actually live-verified). So this adapter never needed a second, refreshable credential shape the way internal/provider/openai_responses.go does.

func NewAnthropic

func NewAnthropic(id, baseURL, apiKey, model string, caps capabilities) *Anthropic

NewAnthropic constructs the Anthropic adapter. baseURL defaults to the public API when empty.

func (*Anthropic) ChatCompletion

func (p *Anthropic) ChatCompletion(ctx context.Context, req openai.ChatCompletionRequest) (<-chan StreamEvent, error)

func (*Anthropic) ID

func (p *Anthropic) ID() string

func (*Anthropic) Kind

func (p *Anthropic) Kind() string

func (Anthropic) SupportsImages

func (c Anthropic) SupportsImages() bool

func (Anthropic) SupportsTools

func (c Anthropic) SupportsTools() bool

type Gemini

type Gemini struct {
	// contains filtered or unexported fields
}

Gemini talks to Google's native streamGenerateContent endpoint (API key as a query param, "contents"/"parts" request shape, "user"/"model" roles instead of "user"/"assistant", and tool calls as functionCall/ functionResponse parts rather than a separate message role).

func NewGemini

func NewGemini(id, baseURL, apiKey, model string, caps capabilities) *Gemini

NewGemini constructs the Gemini adapter. baseURL defaults to the public API when empty.

func (*Gemini) ChatCompletion

func (p *Gemini) ChatCompletion(ctx context.Context, req openai.ChatCompletionRequest) (<-chan StreamEvent, error)

func (*Gemini) ID

func (p *Gemini) ID() string

func (*Gemini) Kind

func (p *Gemini) Kind() string

func (Gemini) SupportsImages

func (c Gemini) SupportsImages() bool

func (Gemini) SupportsTools

func (c Gemini) SupportsTools() bool

type HTTPError

type HTTPError struct {
	Provider   string
	StatusCode int
	Status     string
	// RetryAfter is parsed from the upstream response's Retry-After
	// header (seconds form only — the rarer HTTP-date form isn't worth
	// the parsing complexity for the providers Kram actually talks to).
	// Zero when absent or unparseable; a caller deciding on a Gateway
	// Round retry then falls back to its own computed backoff.
	RetryAfter time.Duration
	// Detail is a bounded upstream error body when the adapter can safely
	// capture one. It makes otherwise opaque 400s diagnosable without logging
	// request credentials or the full request payload.
	Detail string
}

HTTPError wraps a non-2xx upstream response so a caller can report the real status code instead of just a formatted string — this is what lets the router's attempt executor tell an HTTP 500 apart from an HTTP 200 whose *content* got rejected by the ResponseGate (see openai.AttemptInfo.HTTPStatus and DECISIONS.md).

func (*HTTPError) Error

func (e *HTTPError) Error() string

type OpenAICompatible

type OpenAICompatible struct {
	// contains filtered or unexported fields
}

OpenAICompatible talks to any backend that already speaks the OpenAI chat-completions wire format: OpenAI itself, OpenRouter, opencode zen, and most other aggregators. Only base URL, API key and an optional pinned model differ between them.

func NewOpenAICompatible

func NewOpenAICompatible(id, baseURL, apiKey, model string, temperature *float64, caps capabilities) *OpenAICompatible

NewOpenAICompatible constructs an adapter for an OpenAI-shaped backend.

func (*OpenAICompatible) ChatCompletion

func (p *OpenAICompatible) ChatCompletion(ctx context.Context, req openai.ChatCompletionRequest) (<-chan StreamEvent, error)

func (*OpenAICompatible) ID

func (p *OpenAICompatible) ID() string

func (*OpenAICompatible) Kind

func (p *OpenAICompatible) Kind() string

func (OpenAICompatible) SupportsImages

func (c OpenAICompatible) SupportsImages() bool

func (OpenAICompatible) SupportsTools

func (c OpenAICompatible) SupportsTools() bool

type OpenAIResponses

type OpenAIResponses struct {
	// contains filtered or unexported fields
}

OpenAIResponses talks to the Codex backend a ChatGPT Plus/Pro/Team subscription unlocks via browser login (internal/oauthflow.OpenAIAuthorize) — a different product from the standard OpenAI developer API that OpenAICompatible (openai_compat.go) talks to. It uses OpenAI's Responses wire format ("input" items instead of "messages", typed streaming events instead of flat delta chunks) and only serves a restricted set of Codex-branded models, never the general OpenAI catalog.

This adapter is experimental: its request/response shape follows OpenAI's publicly documented Responses API, but the exact headers this specific ChatGPT-authenticated backend expects beyond "Authorization: Bearer" were not fully confirmed against a real account before shipping (see DECISIONS.md) — treat failures here as a signal to re-check against a live subscription, not necessarily a Kram bug.

func NewOpenAIResponses

func NewOpenAIResponses(id, baseURL string, resolve func(context.Context) (string, error), model string, caps capabilities) *OpenAIResponses

NewOpenAIResponses constructs the adapter. resolve is always required — unlike every other adapter, there is no static-API-key path for this product; a credential only ever comes from a refreshable OAuth token (see internal/credentials.Store.Resolve).

func (*OpenAIResponses) ChatCompletion

func (p *OpenAIResponses) ChatCompletion(ctx context.Context, req openai.ChatCompletionRequest) (<-chan StreamEvent, error)

func (*OpenAIResponses) ID

func (p *OpenAIResponses) ID() string

func (*OpenAIResponses) Kind

func (p *OpenAIResponses) Kind() string

func (OpenAIResponses) SupportsImages

func (c OpenAIResponses) SupportsImages() bool

func (OpenAIResponses) SupportsTools

func (c OpenAIResponses) SupportsTools() bool

type Provider

type Provider interface {
	// ID is the stable identifier used in config, logs and telemetry.
	ID() string
	// Kind is the adapter family (anthropic, gemini, openai-compat) — useful
	// for diagnostics, not for routing decisions.
	Kind() string
	// SupportsImages and SupportsTools reflect the provider's configured
	// capabilities (internal/config.ProviderConfig) — callers must check
	// these before sending images or tool definitions.
	SupportsImages() bool
	SupportsTools() bool
	// ChatCompletion issues the request upstream and streams back normalized
	// events. The channel is always closed by the provider, exactly once,
	// after a final event with Done=true or Err set.
	ChatCompletion(ctx context.Context, req openai.ChatCompletionRequest) (<-chan StreamEvent, error)
}

Provider is anything that can serve a chat completion request.

func Build

func Build(cfg config.ProviderConfig, resolve func(context.Context) (string, error), timeout time.Duration) (Provider, error)

Build constructs the concrete adapter for a provider config entry. resolve is nil for cfg.AuthMode == "" (today's static-env-var behavior, unchanged) and must be non-nil for cfg.AuthMode == "oauth" — it's called by the adapter on every request to get a currently-valid, possibly-just-refreshed credential (see internal/credentials.Store.Resolve), since gateway providers are built once at startup and held for the gateway's whole lifetime, long past when a short-lived OAuth token would expire. Only openai-responses (the ChatGPT-login-only Codex backend) ever uses this path — Anthropic connected via browser login ends up with a real, permanent API key (see internal/oauthflow/anthropic.go), so it always takes the static branch below like any pasted key. timeout is the per-request HTTP timeout to apply to the built adapter; pass 0 to keep the adapter's DefaultTimeout (the gateway resolves the configured value once and passes it here — see config.Tunables).

type StreamEvent

type StreamEvent struct {
	Delta string
	// Reasoning is a fragment of a reasoning-capable model's chain-of-
	// thought, kept separate from Delta because it is not the model's
	// actual answer — a caller must never relay it as assistant message
	// content. It exists purely as a liveness signal: a provider that's
	// only produced reasoning so far is actively working, not stalled,
	// which router.BoundedPeek needs to know (see DECISIONS.md — a
	// reasoning phase alone routinely runs past what used to be a fixed
	// peek timeout, and was being misread as a dead attempt).
	Reasoning string
	// ToolCallProgress marks an event that carried tool-call argument
	// fragments but no answer content and no reasoning — real evidence
	// the provider is actively producing tokens, exactly like Reasoning
	// is, but kept as a separate bool rather than folded into it: a
	// caller must never mistake tool-call JSON fragments for chain-of-
	// thought text. Exists purely as a liveness signal for
	// router.BoundedPeek — a provider streaming a long tool call's
	// arguments token-by-token, with no leading text, used to look like
	// dead silence and get killed as stalled even though it was actively
	// working the whole time (every adapter accumulates tool-call
	// fragments internally without emitting any StreamEvent at all,
	// unless this is set).
	ToolCallProgress bool
	Done             bool
	Usage            *openai.Usage
	ToolCalls        []openai.ToolCall
	ProviderItems    []openai.ProviderItem
	Err              error
}

StreamEvent is one normalized increment of a chat completion stream. Providers translate their native wire format into a sequence of these. ToolCalls is only set on the final (Done) event — no provider streams partial tool-call deltas as assembled ToolCall values; the daemon's agent loop makes buffered (non-streaming) gateway calls by default and waits for the complete decision before acting on it (see internal/daemon/agent.Config.PreferStreaming for the opt-in exception), so there's nothing further upstream that needs them incrementally either. ToolCallProgress below exists for a narrower reason: not exposing partial arguments, just proving the provider is still alive.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL