Documentation
¶
Overview ¶
Package llm defines the LLM client interface and shared types.
Index ¶
- Variables
- type APIFormat
- type AnthropicConfig
- type BackendPreset
- type Client
- type ContentBlock
- type Message
- type MockClient
- type OpenAICompatConfig
- type Request
- type Response
- type RetryClient
- type RetryConfig
- type Role
- type StreamHandler
- type StreamingClient
- type SystemSection
- type ToolDefinition
- type ToolResult
- type ToolUse
- type Usage
Constants ¶
This section is empty.
Variables ¶
var ErrOverloaded = errors.New("llm: overloaded")
ErrOverloaded is returned when the backend is temporarily at capacity (HTTP 529). It is retried with a longer delay than ErrRateLimit.
var ErrRateLimit = errors.New("llm: rate limited")
ErrRateLimit is returned when the backend signals a rate-limit condition (HTTP 429).
Functions ¶
This section is empty.
Types ¶
type APIFormat ¶
type APIFormat string
APIFormat identifies the wire protocol used to communicate with a backend.
type AnthropicConfig ¶
type AnthropicConfig struct {
BaseURL string
TokenSource credentials.TokenSource
Model string
MaxTokens int
Timeout time.Duration
Logger *slog.Logger
// PromptCaching enables prompt caching for the system prompt by tagging it
// with cache_control {"type":"ephemeral"}. Reduces input token cost and
// latency on cache hits. Requires the anthropic-beta header.
PromptCaching bool
}
AnthropicConfig holds settings for the Anthropic Messages API backend.
type BackendPreset ¶
BackendPreset is a named, well-known backend configuration. A preset defines default values; individual fields can be overridden in agent.toml.
func LookupPreset ¶
func LookupPreset(name string) (BackendPreset, error)
LookupPreset returns the preset for the given name, or an error if unknown.
type Client ¶
type Client interface {
// Complete sends a request and returns the full response.
Complete(ctx context.Context, req Request) (Response, error)
}
Client is the interface every LLM backend must satisfy.
func NewAnthropicClient ¶
func NewAnthropicClient(cfg AnthropicConfig) (Client, error)
NewAnthropicClient creates a StreamingClient for the Anthropic Messages API.
func NewClientFromConfig ¶
func NewClientFromConfig(cfg config.LLMConfig, credentialsPath string, logger *slog.Logger) (Client, error)
NewClientFromConfig constructs the appropriate Client based on agent config. credentialsPath is the path to credentials.toml; pass "" to use the default.
Resolution order:
- Look up the named preset from cfg.Backend.
- Apply any per-field overrides from cfg (APIFormat, BaseURL).
- Dispatch to the matching Client implementation by APIFormat.
func NewOpenAICompatClient ¶
func NewOpenAICompatClient(cfg OpenAICompatConfig) (Client, error)
NewOpenAICompatClient creates a Client for any OpenAI-compatible API endpoint.
type ContentBlock ¶
type ContentBlock struct {
Text string `json:"text,omitempty"`
Thinking string `json:"thinking,omitempty"`
ToolUse *ToolUse `json:"tool_use,omitempty"`
ToolResult *ToolResult `json:"tool_result,omitempty"`
}
ContentBlock is one item in a message — text, a tool call, a tool result, or model thinking (extended reasoning). Exactly one field is non-zero.
type Message ¶
type Message struct {
Role Role `json:"role"`
Content []ContentBlock `json:"content"`
}
Message is a single turn in the conversation.
func (Message) TextContent ¶
TextContent returns all text blocks concatenated.
func (Message) ThinkingContent ¶
ThinkingContent returns all thinking blocks concatenated.
type MockClient ¶
type MockClient struct {
Responses []Response
Requests []Request // recorded for assertion
// contains filtered or unexported fields
}
MockClient is a deterministic LLM client for unit tests. Each call pops the next response from Responses; if exhausted it returns an error.
func (*MockClient) CompleteStream ¶
func (m *MockClient) CompleteStream(ctx context.Context, req Request, handler StreamHandler) (Response, error)
CompleteStream simulates streaming by calling handler once with the full text of the next canned response, then returning it.
func (*MockClient) Reset ¶
func (m *MockClient) Reset()
Reset clears recorded requests and resets the response pointer.
type OpenAICompatConfig ¶
type OpenAICompatConfig struct {
BaseURL string
TokenSource credentials.TokenSource
Model string
MaxTokens int
Timeout time.Duration
MaxRetries int
Logger *slog.Logger
}
OpenAICompatConfig holds settings for any OpenAI-compatible backend.
type Request ¶
type Request struct {
// SystemPrompt is a single-string system prompt. Used when SystemSections
// is nil. Backends with prompt-caching treat it as one cacheable block.
SystemPrompt string
// SystemSections, when non-nil, replaces SystemPrompt. Each section maps to
// a separate block in the Anthropic system array; sections with
// CacheCheckpoint=true get cache_control markers when the backend has
// prompt caching enabled.
SystemSections []SystemSection
Messages []Message
Tools []ToolDefinition
MaxTokens int
}
Request is the input to a single LLM call.
type Response ¶
type Response struct {
Message Message
StopReason string // "end_turn" | "tool_use" | "max_tokens"
Usage Usage
}
Response is the output of a single LLM call.
func TextResponse ¶
TextResponse is a convenience constructor for a simple text response.
func ToolUseResponse ¶
ToolUseResponse is a convenience constructor for a tool-call response.
type RetryClient ¶
type RetryClient struct {
// contains filtered or unexported fields
}
RetryClient wraps a Client and retries transient errors with exponential backoff and jitter.
Retryable conditions:
- ErrRateLimit (HTTP 429): exponential backoff; honours Retry-After when present
- ErrOverloaded (HTTP 529): separate attempt counter, longer exponential backoff
- Any other error: exponential backoff up to MaxAttempts
Overloaded retries are tracked independently of general retries so that a burst of 529s does not exhaust the general MaxAttempts budget.
Context cancellation and deadline exceeded are never retried.
func NewRetryClient ¶
func NewRetryClient(inner Client, cfg RetryConfig, logger *slog.Logger) *RetryClient
NewRetryClient wraps inner with retry logic using cfg. If logger is nil, slog.Default() is used.
func (*RetryClient) CompleteStream ¶
func (r *RetryClient) CompleteStream(ctx context.Context, req Request, handler StreamHandler) (Response, error)
CompleteStream delegates to the inner client's CompleteStream if it supports streaming, otherwise falls back to Complete without streaming. Streaming responses are not retried once the stream begins.
type RetryConfig ¶
type RetryConfig struct {
// MaxAttempts is the total number of attempts (initial + retries) for general
// and rate-limit errors. Zero or negative values are treated as 3.
MaxAttempts int
// BaseDelay is the initial backoff duration for general and rate-limit errors.
// Zero uses 500ms.
BaseDelay time.Duration
// OverloadedMaxRetries is the number of retries for ErrOverloaded (HTTP 529).
// Zero uses 2.
OverloadedMaxRetries int
// OverloadedBaseDelay is the starting delay between overloaded retries.
// Zero uses 30s. Delay grows exponentially: BaseDelay * 2^(retry-1).
OverloadedBaseDelay time.Duration
}
RetryConfig controls the retry behaviour of RetryClient.
type StreamHandler ¶
type StreamHandler func(delta string)
StreamHandler is called for each text delta during streaming. delta is a non-empty fragment of the assistant's response text.
type StreamingClient ¶
type StreamingClient interface {
Client
// CompleteStream sends a request and calls handler for each text delta as
// tokens arrive. It returns the full Response once the stream ends.
CompleteStream(ctx context.Context, req Request, handler StreamHandler) (Response, error)
}
StreamingClient is an optional extension of Client that supports incremental token streaming. Backends that support streaming implement both interfaces.
type SystemSection ¶
SystemSection is one piece of the system prompt. When CacheCheckpoint is true, the backend places a prompt-cache breakpoint after this section (Anthropic: cache_control {"type":"ephemeral"}).
type ToolDefinition ¶
type ToolDefinition struct {
Name string
Description string
InputSchema []byte // JSON Schema object
}
ToolDefinition describes a tool the model may call.
type ToolResult ¶
type ToolResult struct {
ToolUseID string `json:"tool_use_id"`
Content string `json:"content"`
IsError bool `json:"is_error,omitempty"`
}
ToolResult is the response to a prior ToolUse.