Documentation
¶
Overview ¶
Package openai defines the OpenAI-compatible wire format that kram-gateway exposes to clients, regardless of which upstream provider actually serves the request.
Index ¶
- Constants
- type AttemptInfo
- type AttemptOutcome
- type ChatCompletionChoice
- type ChatCompletionChunk
- type ChatCompletionChunkChoice
- type ChatCompletionChunkDelta
- type ChatCompletionRequest
- type ChatCompletionResponse
- type ChatMessage
- type ErrorBody
- type ErrorResponse
- type FailureClass
- type RankedProviderInfo
- type ScoreFactor
- type Tool
- type ToolCall
- type ToolCallFunction
- type ToolFunction
- type Usage
Constants ¶
const RunIDHeader = "X-Kram-Run-Id"
RunIDHeader is an optional HTTP request header a caller may send on POST /v1/chat/completions: an opaque identifier for one agent run (one user turn plus every tool round-trip it causes) — never one that persists across a later, unrelated turn in the same conversation. It's a kram-gateway extension carried out of band as a header rather than a JSON body field, so standard OpenAI-compatible clients that never send it stay fully compatible (see internal/router.RouteContext.RunKey and DECISIONS.md, "Sticky is run-scoped, not session-prefix-scoped").
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type AttemptInfo ¶
type AttemptInfo struct {
Provider string `json:"provider"`
OK bool `json:"ok"`
LatencyMS int64 `json:"latency_ms"`
// Model is the upstream model ID actually requested for this attempt,
// when known — providers can pin a model regardless of what the
// client asked for (config.ProviderConfig.Model), so this can differ
// from the request's own Model field.
Model string `json:"model,omitempty"`
// Outcome classifies how the attempt ended — see AttemptOutcome.
Outcome AttemptOutcome `json:"outcome,omitempty"`
// Reason is a short human-readable explanation, set whenever Outcome
// isn't OutcomeSuccess: an upstream error message, a ResponseGate
// rejection reason, or why a candidate was skipped.
Reason string `json:"reason,omitempty"`
// HTTPStatus is the upstream HTTP status code, when the attempt
// reached one (0 for a transport-level failure that never got a
// response, or for a skipped attempt).
HTTPStatus int `json:"http_status,omitempty"`
// Score is this candidate's ranking score from the strategy that
// produced the routing order, when the strategy scores (nil for
// non-scoring strategies like priority/round-robin). Pointer so "not
// scored" and "scored zero" stay distinguishable.
Score *float64 `json:"score,omitempty"`
// Attempt is this attempt's 1-indexed position in the fallback chain
// actually tried for this request (1 = first candidate tried).
Attempt int `json:"attempt,omitempty"`
// Class is this attempt's FailureClass, set whenever Outcome isn't
// OutcomeSuccess — see Classify and ClassContentRejected's doc
// comment for how a rejection (as opposed to a transport failure)
// gets classified.
Class FailureClass `json:"class,omitempty"`
}
AttemptInfo records one provider attempt made while serving a request — the real fallback trail for a single completion. Combined with Outcome and Reason, this is what lets a caller tell an HTTP 500 apart from an HTTP 200 whose content the ResponseGate rejected — two very different situations that used to look identical (OK=false either way).
OK is kept and still set (OK == Outcome == OutcomeSuccess) purely for backward compatibility: any code that only ever read the old two-field shape keeps working unchanged. New code should read Outcome.
type AttemptOutcome ¶
type AttemptOutcome string
AttemptOutcome classifies how one attempt ended — see AttemptInfo. Deliberately a small, closed set rather than free-form strings: the CLI switches on these to pick a glyph/color, and a routing decision (the gateway/router) is the only thing that should ever produce one.
const ( OutcomeTrying AttemptOutcome = "trying" // in flight — used only by live route.* events, never persisted on a finished attempt OutcomeSuccess AttemptOutcome = "success" // response received and accepted by the ResponseGate OutcomeError AttemptOutcome = "error" // transport/HTTP failure — never reached a response to gate OutcomeRejected AttemptOutcome = "rejected" // response received (HTTP 200 or equivalent) but the ResponseGate/StreamGate rejected it OutcomeSkipped AttemptOutcome = "skipped" // never attempted — filtered out before execution (circuit-open, capability mismatch) )
type ChatCompletionChoice ¶
type ChatCompletionChoice struct {
Index int `json:"index"`
Message ChatMessage `json:"message"`
FinishReason string `json:"finish_reason"`
}
ChatCompletionChoice is one candidate answer in a non-streaming response.
type ChatCompletionChunk ¶
type ChatCompletionChunk struct {
ID string `json:"id"`
Object string `json:"object"`
Created int64 `json:"created"`
Model string `json:"model"`
Choices []ChatCompletionChunkChoice `json:"choices"`
Provider string `json:"provider,omitempty"`
Attempts []AttemptInfo `json:"attempts,omitempty"`
Ranking []RankedProviderInfo `json:"ranking,omitempty"`
Strategy string `json:"strategy,omitempty"`
Usage *Usage `json:"usage,omitempty"`
}
ChatCompletionChunk is one `data: {...}` SSE event for streaming responses. Provider, Attempts and Usage are kram-gateway extensions, set only on the terminal chunk (mirrors ChatCompletionResponse) — used by the daemon's agent loop when a session opts into the streaming path (internal/daemon/agent.Config.PreferStreaming; buffered non-streaming calls, via ChatCompletionResponse instead, are the default — see that field's doc comment for why).
type ChatCompletionChunkChoice ¶
type ChatCompletionChunkChoice struct {
Index int `json:"index"`
Delta ChatCompletionChunkDelta `json:"delta"`
FinishReason *string `json:"finish_reason"`
}
ChatCompletionChunkChoice wraps a delta inside a streaming chunk.
type ChatCompletionChunkDelta ¶
type ChatCompletionChunkDelta struct {
Role string `json:"role,omitempty"`
Content string `json:"content,omitempty"`
ToolCalls []ToolCall `json:"tool_calls,omitempty"`
}
ChatCompletionChunkDelta is the incremental content of one SSE chunk. On the terminal chunk (FinishReason set), ToolCalls carries the fully assembled calls in one go rather than OpenAI's fragmented-by-index deltas — kram-gateway already does that reassembly per provider (internal/provider), and re-fragmenting it for a stream Kram itself is usually the only consumer of would just be extra work with no benefit.
type ChatCompletionRequest ¶
type ChatCompletionRequest struct {
Model string `json:"model"`
Messages []ChatMessage `json:"messages"`
Stream bool `json:"stream"`
MaxTokens *int `json:"max_tokens,omitempty"`
Temperature *float64 `json:"temperature,omitempty"`
Tools []Tool `json:"tools,omitempty"`
}
ChatCompletionRequest is the request body for POST /v1/chat/completions.
type ChatCompletionResponse ¶
type ChatCompletionResponse struct {
ID string `json:"id"`
Object string `json:"object"`
Created int64 `json:"created"`
Model string `json:"model"`
Choices []ChatCompletionChoice `json:"choices"`
Usage Usage `json:"usage"`
// Provider and Attempts are kram-gateway extensions (ignored by
// standard OpenAI clients): which upstream actually served the
// request, and the full fallback trail attempted to get there.
Provider string `json:"provider,omitempty"`
Attempts []AttemptInfo `json:"attempts,omitempty"`
// Ranking is the full candidate ranking a scoring strategy produced
// for this request, if any (nil for non-scoring strategies) — see
// RankedProviderInfo.
Ranking []RankedProviderInfo `json:"ranking,omitempty"`
Strategy string `json:"strategy,omitempty"`
}
ChatCompletionResponse is the non-streaming response body.
type ChatMessage ¶
type ChatMessage struct {
Role string `json:"role"`
Content string `json:"content,omitempty"`
// Images are data: URLs (kram extension, not standard OpenAI wire
// format) — kept separate from Content instead of OpenAI's
// content-parts array so Content can stay a plain string everywhere
// else in the codebase.
Images []string `json:"images,omitempty"`
// ToolCalls is set on an assistant message that is requesting one or
// more tool invocations instead of (or before) answering in text.
ToolCalls []ToolCall `json:"tool_calls,omitempty"`
// ToolCallID and Name identify which call a role:"tool" message answers.
ToolCallID string `json:"tool_call_id,omitempty"`
Name string `json:"name,omitempty"`
}
ChatMessage is a single turn in a chat completion request. Role is one of "system", "user", "assistant", or "tool". An assistant message that wants to call tools sets ToolCalls instead of (or alongside) Content; a "tool" message reports one tool's result and must set ToolCallID.
type ErrorBody ¶
type ErrorBody struct {
Message string `json:"message"`
Type string `json:"type"`
Code string `json:"code,omitempty"`
// Combo, Retryable, RetryAfterMS, Cause and Attempts are kram-gateway
// extensions, populated only when every candidate in a combo failed
// (see internal/server/chat.go's writeGatewayError) — a standard
// OpenAI client ignores unknown fields; internal/daemon/gatewayclient
// decodes them into a typed GatewayError instead of treating this as
// a flat message string, so a caller can actually decide whether
// retrying is worth it instead of guessing from prose. Combo being
// non-empty is the signal this response came from kram-gateway's own
// all-failed path, as opposed to some other OpenAI-compatible
// server's plain error body.
Combo string `json:"combo,omitempty"`
// Retryable and Cause reflect the *last* attempt in the fallback
// trail — the one that actually ended the request — not "any attempt
// was retryable". A combo whose final candidate failed with a 400
// isn't worth retrying even if an earlier candidate hit a 429.
Retryable bool `json:"retryable,omitempty"`
RetryAfterMS int64 `json:"retry_after_ms,omitempty"`
Cause FailureClass `json:"cause,omitempty"`
Attempts []AttemptInfo `json:"attempts,omitempty"`
}
ErrorBody carries the error message and classification.
type ErrorResponse ¶
type ErrorResponse struct {
Error ErrorBody `json:"error"`
}
ErrorResponse is the OpenAI-compatible error envelope.
type FailureClass ¶
type FailureClass string
FailureClass buckets why one provider attempt failed, so a caller can tell "genuinely broken upstream" apart from "Kram sent something the provider correctly rejected" apart from "the caller gave up" — three situations that used to look identical (an error string) and were treated identically (poison the circuit breaker, never worth a retry).
const ( // ClassNetwork is a dial/connection-level failure — no response was // ever reached. ClassNetwork FailureClass = "network" // ClassTimeout is a context deadline or client-side timeout. ClassTimeout FailureClass = "timeout" // ClassServerError is a 5xx — the upstream reached the request but // failed to serve it. ClassServerError FailureClass = "server_error" // ClassRateLimit is a 429 — the upstream is healthy but out of // capacity for this account/window right now. ClassRateLimit FailureClass = "rate_limit" // ClassAuth is a 401/403 — the configured credential is wrong or // revoked. Retrying the same request can never fix this, but the // provider genuinely is unusable until a human fixes the credential, // so this still counts against the circuit breaker (see // CountsAgainstBreaker) even though it's never Retryable. ClassAuth FailureClass = "auth" // ClassInvalidRequest is a 400/404/422 — usually Kram's own malformed // request (a wrong role name, a missing required field, an unknown // model), not evidence the provider itself is unhealthy. See // CountsAgainstBreaker. ClassInvalidRequest FailureClass = "invalid_request" // ClassContentRejected is never produced by Classify — it's assigned // directly by a caller (internal/server/chat.go) for a // ResponseGate/StreamGate rejection or a BoundedPeek timeout, neither // of which is a transport failure Classify has any status/error to // reason about. Kept in this enum so AttemptInfo.Class has one // consistent vocabulary regardless of which layer assigned it. ClassContentRejected FailureClass = "content_rejected" // ClassCancelled is a context.Canceled — the caller gave up (client // disconnected, turn aborted), not evidence the provider is unwell. ClassCancelled FailureClass = "cancelled" // ClassUnknown is anything Classify can't place — treated // conservatively (counts against the breaker, since an unrecognized // failure is not evidence of health either). ClassUnknown FailureClass = "unknown" )
func Classify ¶
func Classify(httpStatus int, err error) FailureClass
Classify derives a FailureClass from a transport-level failure: the upstream HTTP status if one was reached (0 if the request never got a response at all), and the error itself for the no-response case. It has no opinion about ResponseGate/StreamGate rejections — those are not transport failures and must be classified ClassContentRejected directly by the caller, never routed through Classify (see internal/server/chat.go's markFailure/markRejection split).
func (FailureClass) CountsAgainstBreaker ¶
func (c FailureClass) CountsAgainstBreaker() bool
CountsAgainstBreaker reports whether this class is real evidence the provider itself is unhealthy, as opposed to evidence Kram sent it something it correctly rejected (ClassInvalidRequest) or that the caller simply gave up (ClassCancelled). Deliberately broader than Retryable: ClassAuth counts against the breaker (the provider really is unusable right now) even though retrying the same request can never fix it.
func (FailureClass) Retryable ¶
func (c FailureClass) Retryable() bool
Retryable reports whether a Gateway Round retry (a fresh attempt at the whole ranked-candidate list, after a backoff) is ever worth attempting for this class. Never true for a class where retrying the identical request has no chance of a different outcome.
type RankedProviderInfo ¶
type RankedProviderInfo struct {
Provider string `json:"provider"`
Score float64 `json:"score"`
Factors []ScoreFactor `json:"factors,omitempty"`
Reasons []string `json:"reasons,omitempty"` // e.g. "sticky", "last-known-good", "cache-affinity"
}
RankedProviderInfo is one entry in a scoring strategy's full ranking — distinct from AttemptInfo: every ranked candidate appears here whether or not it was actually attempted (fallback can stop before reaching a lower-ranked candidate), which is what lets a UI show "gemini .842, openai .816" as context even when only the top candidate was ever called.
type ScoreFactor ¶
type ScoreFactor struct {
Name string `json:"name"`
Weight float64 `json:"weight"` // 0..1, normalized
Value float64 `json:"value"` // 0..1, this candidate's raw score for the factor
Contribution float64 `json:"contribution"` // weight * value
}
ScoreFactor is one weighted component of a scoring strategy's decision for a single candidate — kept so a UI can show *why* a candidate scored the way it did without ever recomputing the score itself; it only ever renders what the router already decided (see DECISIONS.md).
type Tool ¶
type Tool struct {
Type string `json:"type"` // always "function"
Function ToolFunction `json:"function"`
}
Tool is one function the model may call, described as JSON Schema — the same shape OpenAI's API uses, since every provider adapter already has to translate to/from its own native tool format anyway.
type ToolCall ¶
type ToolCall struct {
ID string `json:"id"`
Type string `json:"type"` // always "function"
Function ToolCallFunction `json:"function"`
// GeminiThoughtSignature is opaque provider-specific state Gemini's
// thinking-enabled models attach to a function call — it must be
// echoed back verbatim on this exact call when the conversation
// continues, or Gemini rejects the request outright ("Function call
// is missing a thought_signature"). Empty for every other provider,
// and for Gemini responses from a non-thinking model. Kept on the
// shared ToolCall (rather than a Gemini-only side channel) because
// tool calls already round-trip through the daemon's session storage
// and back out to whichever provider serves the next turn — the
// value must survive that round-trip intact to be replayed, and
// there's nowhere else in the pipeline that would preserve it.
GeminiThoughtSignature string `json:"gemini_thought_signature,omitempty"`
}
ToolCall is one invocation the model is requesting.
type ToolCallFunction ¶
ToolCallFunction names the tool and carries its arguments as a raw JSON string (not a parsed object) — this matches every provider's actual wire format and avoids losing information to an intermediate schema.
type ToolFunction ¶
type ToolFunction struct {
Name string `json:"name"`
Description string `json:"description,omitempty"`
Parameters json.RawMessage `json:"parameters,omitempty"` // JSON Schema object
}
ToolFunction describes a callable tool's name, purpose, and argument shape.