openai

package
v0.2.7 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 20, 2026 License: MIT Imports: 4 Imported by: 0

Documentation

Overview

Package openai defines the OpenAI-compatible wire format that kram-gateway exposes to clients, regardless of which upstream provider actually serves the request.

Index

Constants

View Source
const RunIDHeader = "X-Kram-Run-Id"

RunIDHeader is an optional HTTP request header a caller may send on POST /v1/chat/completions: an opaque identifier for one agent run (one user turn plus every tool round-trip it causes) — never one that persists across a later, unrelated turn in the same conversation. It's a kram-gateway extension carried out of band as a header rather than a JSON body field, so standard OpenAI-compatible clients that never send it stay fully compatible (see internal/router.RouteContext.RunKey and DECISIONS.md, "Sticky is run-scoped, not session-prefix-scoped").

Variables

This section is empty.

Functions

This section is empty.

Types

type AttemptInfo

type AttemptInfo struct {
	Provider  string `json:"provider"`
	OK        bool   `json:"ok"`
	LatencyMS int64  `json:"latency_ms"`

	// Model is the upstream model ID actually requested for this attempt,
	// when known — providers can pin a model regardless of what the
	// client asked for (config.ProviderConfig.Model), so this can differ
	// from the request's own Model field.
	Model string `json:"model,omitempty"`
	// Outcome classifies how the attempt ended — see AttemptOutcome.
	Outcome AttemptOutcome `json:"outcome,omitempty"`
	// Reason is a short human-readable explanation, set whenever Outcome
	// isn't OutcomeSuccess: an upstream error message, a ResponseGate
	// rejection reason, or why a candidate was skipped.
	Reason string `json:"reason,omitempty"`
	// HTTPStatus is the upstream HTTP status code, when the attempt
	// reached one (0 for a transport-level failure that never got a
	// response, or for a skipped attempt).
	HTTPStatus int `json:"http_status,omitempty"`
	// Score is this candidate's ranking score from the strategy that
	// produced the routing order, when the strategy scores (nil for
	// non-scoring strategies like priority/round-robin). Pointer so "not
	// scored" and "scored zero" stay distinguishable.
	Score *float64 `json:"score,omitempty"`
	// Attempt is this attempt's 1-indexed position in the fallback chain
	// actually tried for this request (1 = first candidate tried).
	Attempt int `json:"attempt,omitempty"`
	// Class is this attempt's FailureClass, set whenever Outcome isn't
	// OutcomeSuccess — see Classify and ClassContentRejected's doc
	// comment for how a rejection (as opposed to a transport failure)
	// gets classified.
	Class FailureClass `json:"class,omitempty"`
}

AttemptInfo records one provider attempt made while serving a request — the real fallback trail for a single completion. Combined with Outcome and Reason, this is what lets a caller tell an HTTP 500 apart from an HTTP 200 whose content the ResponseGate rejected — two very different situations that used to look identical (OK=false either way).

OK is kept and still set (OK == Outcome == OutcomeSuccess) purely for backward compatibility: any code that only ever read the old two-field shape keeps working unchanged. New code should read Outcome.

type AttemptOutcome

type AttemptOutcome string

AttemptOutcome classifies how one attempt ended — see AttemptInfo. Deliberately a small, closed set rather than free-form strings: the CLI switches on these to pick a glyph/color, and a routing decision (the gateway/router) is the only thing that should ever produce one.

const (
	OutcomeTrying   AttemptOutcome = "trying"   // in flight — used only by live route.* events, never persisted on a finished attempt
	OutcomeSuccess  AttemptOutcome = "success"  // response received and accepted by the ResponseGate
	OutcomeError    AttemptOutcome = "error"    // transport/HTTP failure — never reached a response to gate
	OutcomeRejected AttemptOutcome = "rejected" // response received (HTTP 200 or equivalent) but the ResponseGate/StreamGate rejected it
	OutcomeSkipped  AttemptOutcome = "skipped"  // never attempted — filtered out before execution (circuit-open, capability mismatch)
)

type ChatCompletionChoice

type ChatCompletionChoice struct {
	Index        int         `json:"index"`
	Message      ChatMessage `json:"message"`
	FinishReason string      `json:"finish_reason"`
}

ChatCompletionChoice is one candidate answer in a non-streaming response.

type ChatCompletionChunk

type ChatCompletionChunk struct {
	ID       string                      `json:"id"`
	Object   string                      `json:"object"`
	Created  int64                       `json:"created"`
	Model    string                      `json:"model"`
	Choices  []ChatCompletionChunkChoice `json:"choices"`
	Provider string                      `json:"provider,omitempty"`
	Attempts []AttemptInfo               `json:"attempts,omitempty"`
	Ranking  []RankedProviderInfo        `json:"ranking,omitempty"`
	Strategy string                      `json:"strategy,omitempty"`
	Usage    *Usage                      `json:"usage,omitempty"`
}

ChatCompletionChunk is one `data: {...}` SSE event for streaming responses. Provider, Attempts and Usage are kram-gateway extensions, set only on the terminal chunk (mirrors ChatCompletionResponse) — used by the daemon's agent loop when a session opts into the streaming path (internal/daemon/agent.Config.PreferStreaming; buffered non-streaming calls, via ChatCompletionResponse instead, are the default — see that field's doc comment for why).

type ChatCompletionChunkChoice

type ChatCompletionChunkChoice struct {
	Index        int                      `json:"index"`
	Delta        ChatCompletionChunkDelta `json:"delta"`
	FinishReason *string                  `json:"finish_reason"`
}

ChatCompletionChunkChoice wraps a delta inside a streaming chunk.

type ChatCompletionChunkDelta

type ChatCompletionChunkDelta struct {
	Role      string     `json:"role,omitempty"`
	Content   string     `json:"content,omitempty"`
	ToolCalls []ToolCall `json:"tool_calls,omitempty"`
}

ChatCompletionChunkDelta is the incremental content of one SSE chunk. On the terminal chunk (FinishReason set), ToolCalls carries the fully assembled calls in one go rather than OpenAI's fragmented-by-index deltas — kram-gateway already does that reassembly per provider (internal/provider), and re-fragmenting it for a stream Kram itself is usually the only consumer of would just be extra work with no benefit.

type ChatCompletionRequest

type ChatCompletionRequest struct {
	Model       string        `json:"model"`
	Messages    []ChatMessage `json:"messages"`
	Stream      bool          `json:"stream"`
	MaxTokens   *int          `json:"max_tokens,omitempty"`
	Temperature *float64      `json:"temperature,omitempty"`
	Tools       []Tool        `json:"tools,omitempty"`
}

ChatCompletionRequest is the request body for POST /v1/chat/completions.

type ChatCompletionResponse

type ChatCompletionResponse struct {
	ID      string                 `json:"id"`
	Object  string                 `json:"object"`
	Created int64                  `json:"created"`
	Model   string                 `json:"model"`
	Choices []ChatCompletionChoice `json:"choices"`
	Usage   Usage                  `json:"usage"`
	// Provider and Attempts are kram-gateway extensions (ignored by
	// standard OpenAI clients): which upstream actually served the
	// request, and the full fallback trail attempted to get there.
	Provider string        `json:"provider,omitempty"`
	Attempts []AttemptInfo `json:"attempts,omitempty"`
	// Ranking is the full candidate ranking a scoring strategy produced
	// for this request, if any (nil for non-scoring strategies) — see
	// RankedProviderInfo.
	Ranking  []RankedProviderInfo `json:"ranking,omitempty"`
	Strategy string               `json:"strategy,omitempty"`
}

ChatCompletionResponse is the non-streaming response body.

type ChatMessage

type ChatMessage struct {
	Role    string `json:"role"`
	Content string `json:"content,omitempty"`
	// Images are data: URLs (kram extension, not standard OpenAI wire
	// format) — kept separate from Content instead of OpenAI's
	// content-parts array so Content can stay a plain string everywhere
	// else in the codebase.
	Images []string `json:"images,omitempty"`
	// ToolCalls is set on an assistant message that is requesting one or
	// more tool invocations instead of (or before) answering in text.
	ToolCalls []ToolCall `json:"tool_calls,omitempty"`
	// ToolCallID and Name identify which call a role:"tool" message answers.
	ToolCallID string `json:"tool_call_id,omitempty"`
	Name       string `json:"name,omitempty"`
}

ChatMessage is a single turn in a chat completion request. Role is one of "system", "user", "assistant", or "tool". An assistant message that wants to call tools sets ToolCalls instead of (or alongside) Content; a "tool" message reports one tool's result and must set ToolCallID.

type ErrorBody

type ErrorBody struct {
	Message string `json:"message"`
	Type    string `json:"type"`
	Code    string `json:"code,omitempty"`

	// Combo, Retryable, RetryAfterMS, Cause and Attempts are kram-gateway
	// extensions, populated only when every candidate in a combo failed
	// (see internal/server/chat.go's writeGatewayError) — a standard
	// OpenAI client ignores unknown fields; internal/daemon/gatewayclient
	// decodes them into a typed GatewayError instead of treating this as
	// a flat message string, so a caller can actually decide whether
	// retrying is worth it instead of guessing from prose. Combo being
	// non-empty is the signal this response came from kram-gateway's own
	// all-failed path, as opposed to some other OpenAI-compatible
	// server's plain error body.
	Combo string `json:"combo,omitempty"`
	// Retryable and Cause reflect the *last* attempt in the fallback
	// trail — the one that actually ended the request — not "any attempt
	// was retryable". A combo whose final candidate failed with a 400
	// isn't worth retrying even if an earlier candidate hit a 429.
	Retryable    bool          `json:"retryable,omitempty"`
	RetryAfterMS int64         `json:"retry_after_ms,omitempty"`
	Cause        FailureClass  `json:"cause,omitempty"`
	Attempts     []AttemptInfo `json:"attempts,omitempty"`
}

ErrorBody carries the error message and classification.

type ErrorResponse

type ErrorResponse struct {
	Error ErrorBody `json:"error"`
}

ErrorResponse is the OpenAI-compatible error envelope.

type FailureClass

type FailureClass string

FailureClass buckets why one provider attempt failed, so a caller can tell "genuinely broken upstream" apart from "Kram sent something the provider correctly rejected" apart from "the caller gave up" — three situations that used to look identical (an error string) and were treated identically (poison the circuit breaker, never worth a retry).

const (
	// ClassNetwork is a dial/connection-level failure — no response was
	// ever reached.
	ClassNetwork FailureClass = "network"
	// ClassTimeout is a context deadline or client-side timeout.
	ClassTimeout FailureClass = "timeout"
	// ClassServerError is a 5xx — the upstream reached the request but
	// failed to serve it.
	ClassServerError FailureClass = "server_error"
	// ClassRateLimit is a 429 — the upstream is healthy but out of
	// capacity for this account/window right now.
	ClassRateLimit FailureClass = "rate_limit"
	// ClassAuth is a 401/403 — the configured credential is wrong or
	// revoked. Retrying the same request can never fix this, but the
	// provider genuinely is unusable until a human fixes the credential,
	// so this still counts against the circuit breaker (see
	// CountsAgainstBreaker) even though it's never Retryable.
	ClassAuth FailureClass = "auth"
	// ClassInvalidRequest is a 400/404/422 — usually Kram's own malformed
	// request (a wrong role name, a missing required field, an unknown
	// model), not evidence the provider itself is unhealthy. See
	// CountsAgainstBreaker.
	ClassInvalidRequest FailureClass = "invalid_request"
	// ClassContentRejected is never produced by Classify — it's assigned
	// directly by a caller (internal/server/chat.go) for a
	// ResponseGate/StreamGate rejection or a BoundedPeek timeout, neither
	// of which is a transport failure Classify has any status/error to
	// reason about. Kept in this enum so AttemptInfo.Class has one
	// consistent vocabulary regardless of which layer assigned it.
	ClassContentRejected FailureClass = "content_rejected"
	// ClassCancelled is a context.Canceled — the caller gave up (client
	// disconnected, turn aborted), not evidence the provider is unwell.
	ClassCancelled FailureClass = "cancelled"
	// ClassUnknown is anything Classify can't place — treated
	// conservatively (counts against the breaker, since an unrecognized
	// failure is not evidence of health either).
	ClassUnknown FailureClass = "unknown"
)

func Classify

func Classify(httpStatus int, err error) FailureClass

Classify derives a FailureClass from a transport-level failure: the upstream HTTP status if one was reached (0 if the request never got a response at all), and the error itself for the no-response case. It has no opinion about ResponseGate/StreamGate rejections — those are not transport failures and must be classified ClassContentRejected directly by the caller, never routed through Classify (see internal/server/chat.go's markFailure/markRejection split).

func (FailureClass) CountsAgainstBreaker

func (c FailureClass) CountsAgainstBreaker() bool

CountsAgainstBreaker reports whether this class is real evidence the provider itself is unhealthy, as opposed to evidence Kram sent it something it correctly rejected (ClassInvalidRequest) or that the caller simply gave up (ClassCancelled). Deliberately broader than Retryable: ClassAuth counts against the breaker (the provider really is unusable right now) even though retrying the same request can never fix it.

func (FailureClass) Retryable

func (c FailureClass) Retryable() bool

Retryable reports whether a Gateway Round retry (a fresh attempt at the whole ranked-candidate list, after a backoff) is ever worth attempting for this class. Never true for a class where retrying the identical request has no chance of a different outcome.

type RankedProviderInfo

type RankedProviderInfo struct {
	Provider string        `json:"provider"`
	Score    float64       `json:"score"`
	Factors  []ScoreFactor `json:"factors,omitempty"`
	Reasons  []string      `json:"reasons,omitempty"` // e.g. "sticky", "last-known-good", "cache-affinity"
}

RankedProviderInfo is one entry in a scoring strategy's full ranking — distinct from AttemptInfo: every ranked candidate appears here whether or not it was actually attempted (fallback can stop before reaching a lower-ranked candidate), which is what lets a UI show "gemini .842, openai .816" as context even when only the top candidate was ever called.

type ScoreFactor

type ScoreFactor struct {
	Name         string  `json:"name"`
	Weight       float64 `json:"weight"`       // 0..1, normalized
	Value        float64 `json:"value"`        // 0..1, this candidate's raw score for the factor
	Contribution float64 `json:"contribution"` // weight * value
}

ScoreFactor is one weighted component of a scoring strategy's decision for a single candidate — kept so a UI can show *why* a candidate scored the way it did without ever recomputing the score itself; it only ever renders what the router already decided (see DECISIONS.md).

type Tool

type Tool struct {
	Type     string       `json:"type"` // always "function"
	Function ToolFunction `json:"function"`
}

Tool is one function the model may call, described as JSON Schema — the same shape OpenAI's API uses, since every provider adapter already has to translate to/from its own native tool format anyway.

type ToolCall

type ToolCall struct {
	ID       string           `json:"id"`
	Type     string           `json:"type"` // always "function"
	Function ToolCallFunction `json:"function"`
	// GeminiThoughtSignature is opaque provider-specific state Gemini's
	// thinking-enabled models attach to a function call — it must be
	// echoed back verbatim on this exact call when the conversation
	// continues, or Gemini rejects the request outright ("Function call
	// is missing a thought_signature"). Empty for every other provider,
	// and for Gemini responses from a non-thinking model. Kept on the
	// shared ToolCall (rather than a Gemini-only side channel) because
	// tool calls already round-trip through the daemon's session storage
	// and back out to whichever provider serves the next turn — the
	// value must survive that round-trip intact to be replayed, and
	// there's nowhere else in the pipeline that would preserve it.
	GeminiThoughtSignature string `json:"gemini_thought_signature,omitempty"`
}

ToolCall is one invocation the model is requesting.

type ToolCallFunction

type ToolCallFunction struct {
	Name      string `json:"name"`
	Arguments string `json:"arguments"`
}

ToolCallFunction names the tool and carries its arguments as a raw JSON string (not a parsed object) — this matches every provider's actual wire format and avoids losing information to an intermediate schema.

type ToolFunction

type ToolFunction struct {
	Name        string          `json:"name"`
	Description string          `json:"description,omitempty"`
	Parameters  json.RawMessage `json:"parameters,omitempty"` // JSON Schema object
}

ToolFunction describes a callable tool's name, purpose, and argument shape.

type Usage

type Usage struct {
	PromptTokens     int `json:"prompt_tokens"`
	CompletionTokens int `json:"completion_tokens"`
	TotalTokens      int `json:"total_tokens"`
}

Usage reports token accounting for a request.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL