llm

package
v0.0.0-...-1112840 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: May 1, 2026 License: MIT Imports: 18 Imported by: 0

Documentation

Overview

Package llm defines the LLM client interface and shared types.

Index

Constants

This section is empty.

Variables

View Source
var ErrOverloaded = errors.New("llm: overloaded")

ErrOverloaded is returned when the backend is temporarily at capacity (HTTP 529). It is retried with a longer delay than ErrRateLimit.

View Source
var ErrRateLimit = errors.New("llm: rate limited")

ErrRateLimit is returned when the backend signals a rate-limit condition (HTTP 429).

Functions

This section is empty.

Types

type APIFormat

type APIFormat string

APIFormat identifies the wire protocol used to communicate with a backend.

const (
	APIFormatOpenAI    APIFormat = "openai"
	APIFormatAnthropic APIFormat = "anthropic"
)

type AnthropicConfig

type AnthropicConfig struct {
	BaseURL     string
	TokenSource credentials.TokenSource
	Model       string
	MaxTokens   int
	Timeout     time.Duration
	Logger      *slog.Logger

	// PromptCaching enables prompt caching for the system prompt by tagging it
	// with cache_control {"type":"ephemeral"}. Reduces input token cost and
	// latency on cache hits. Requires the anthropic-beta header.
	PromptCaching bool
}

AnthropicConfig holds settings for the Anthropic Messages API backend.

type BackendPreset

type BackendPreset struct {
	BaseURL   string
	APIFormat APIFormat
}

BackendPreset is a named, well-known backend configuration. A preset defines default values; individual fields can be overridden in agent.toml.

func LookupPreset

func LookupPreset(name string) (BackendPreset, error)

LookupPreset returns the preset for the given name, or an error if unknown.

type Client

type Client interface {
	// Complete sends a request and returns the full response.
	Complete(ctx context.Context, req Request) (Response, error)
}

Client is the interface every LLM backend must satisfy.

func NewAnthropicClient

func NewAnthropicClient(cfg AnthropicConfig) (Client, error)

NewAnthropicClient creates a StreamingClient for the Anthropic Messages API.

func NewClientFromConfig

func NewClientFromConfig(cfg config.LLMConfig, credentialsPath string, logger *slog.Logger) (Client, error)

NewClientFromConfig constructs the appropriate Client based on agent config. credentialsPath is the path to credentials.toml; pass "" to use the default.

Resolution order:

  1. Look up the named preset from cfg.Backend.
  2. Apply any per-field overrides from cfg (APIFormat, BaseURL).
  3. Dispatch to the matching Client implementation by APIFormat.

func NewOpenAICompatClient

func NewOpenAICompatClient(cfg OpenAICompatConfig) (Client, error)

NewOpenAICompatClient creates a Client for any OpenAI-compatible API endpoint.

type ContentBlock

type ContentBlock struct {
	Text       string      `json:"text,omitempty"`
	Thinking   string      `json:"thinking,omitempty"`
	ToolUse    *ToolUse    `json:"tool_use,omitempty"`
	ToolResult *ToolResult `json:"tool_result,omitempty"`
}

ContentBlock is one item in a message — text, a tool call, a tool result, or model thinking (extended reasoning). Exactly one field is non-zero.

type Message

type Message struct {
	Role    Role           `json:"role"`
	Content []ContentBlock `json:"content"`
}

Message is a single turn in the conversation.

func (Message) TextContent

func (m Message) TextContent() string

TextContent returns all text blocks concatenated.

func (Message) ThinkingContent

func (m Message) ThinkingContent() string

ThinkingContent returns all thinking blocks concatenated.

func (Message) ToolUses

func (m Message) ToolUses() []ToolUse

ToolUses returns all tool-call blocks in this message.

type MockClient

type MockClient struct {
	Responses []Response
	Requests  []Request // recorded for assertion
	// contains filtered or unexported fields
}

MockClient is a deterministic LLM client for unit tests. Each call pops the next response from Responses; if exhausted it returns an error.

func (*MockClient) Complete

func (m *MockClient) Complete(_ context.Context, req Request) (Response, error)

Complete returns the next canned response.

func (*MockClient) CompleteStream

func (m *MockClient) CompleteStream(ctx context.Context, req Request, handler StreamHandler) (Response, error)

CompleteStream simulates streaming by calling handler once with the full text of the next canned response, then returning it.

func (*MockClient) Reset

func (m *MockClient) Reset()

Reset clears recorded requests and resets the response pointer.

type OpenAICompatConfig

type OpenAICompatConfig struct {
	BaseURL     string
	TokenSource credentials.TokenSource
	Model       string
	MaxTokens   int
	Timeout     time.Duration
	MaxRetries  int
	Logger      *slog.Logger
}

OpenAICompatConfig holds settings for any OpenAI-compatible backend.

type Request

type Request struct {
	// SystemPrompt is a single-string system prompt. Used when SystemSections
	// is nil. Backends with prompt-caching treat it as one cacheable block.
	SystemPrompt string

	// SystemSections, when non-nil, replaces SystemPrompt. Each section maps to
	// a separate block in the Anthropic system array; sections with
	// CacheCheckpoint=true get cache_control markers when the backend has
	// prompt caching enabled.
	SystemSections []SystemSection

	Messages  []Message
	Tools     []ToolDefinition
	MaxTokens int
}

Request is the input to a single LLM call.

type Response

type Response struct {
	Message    Message
	StopReason string // "end_turn" | "tool_use" | "max_tokens"
	Usage      Usage
}

Response is the output of a single LLM call.

func TextResponse

func TextResponse(text string) Response

TextResponse is a convenience constructor for a simple text response.

func ToolUseResponse

func ToolUseResponse(id, name string, input []byte) Response

ToolUseResponse is a convenience constructor for a tool-call response.

type RetryClient

type RetryClient struct {
	// contains filtered or unexported fields
}

RetryClient wraps a Client and retries transient errors with exponential backoff and jitter.

Retryable conditions:

  • ErrRateLimit (HTTP 429): exponential backoff; honours Retry-After when present
  • ErrOverloaded (HTTP 529): separate attempt counter, longer exponential backoff
  • Any other error: exponential backoff up to MaxAttempts

Overloaded retries are tracked independently of general retries so that a burst of 529s does not exhaust the general MaxAttempts budget.

Context cancellation and deadline exceeded are never retried.

func NewRetryClient

func NewRetryClient(inner Client, cfg RetryConfig, logger *slog.Logger) *RetryClient

NewRetryClient wraps inner with retry logic using cfg. If logger is nil, slog.Default() is used.

func (*RetryClient) Complete

func (r *RetryClient) Complete(ctx context.Context, req Request) (Response, error)

Complete calls the wrapped client, retrying on transient failures.

func (*RetryClient) CompleteStream

func (r *RetryClient) CompleteStream(ctx context.Context, req Request, handler StreamHandler) (Response, error)

CompleteStream delegates to the inner client's CompleteStream if it supports streaming, otherwise falls back to Complete without streaming. Streaming responses are not retried once the stream begins.

type RetryConfig

type RetryConfig struct {
	// MaxAttempts is the total number of attempts (initial + retries) for general
	// and rate-limit errors. Zero or negative values are treated as 3.
	MaxAttempts int
	// BaseDelay is the initial backoff duration for general and rate-limit errors.
	// Zero uses 500ms.
	BaseDelay time.Duration

	// OverloadedMaxRetries is the number of retries for ErrOverloaded (HTTP 529).
	// Zero uses 2.
	OverloadedMaxRetries int
	// OverloadedBaseDelay is the starting delay between overloaded retries.
	// Zero uses 30s. Delay grows exponentially: BaseDelay * 2^(retry-1).
	OverloadedBaseDelay time.Duration
}

RetryConfig controls the retry behaviour of RetryClient.

type Role

type Role string

Role represents the speaker of a conversation turn.

const (
	RoleUser      Role = "user"
	RoleAssistant Role = "assistant"
	RoleTool      Role = "tool"
)

type StreamHandler

type StreamHandler func(delta string)

StreamHandler is called for each text delta during streaming. delta is a non-empty fragment of the assistant's response text.

type StreamingClient

type StreamingClient interface {
	Client
	// CompleteStream sends a request and calls handler for each text delta as
	// tokens arrive. It returns the full Response once the stream ends.
	CompleteStream(ctx context.Context, req Request, handler StreamHandler) (Response, error)
}

StreamingClient is an optional extension of Client that supports incremental token streaming. Backends that support streaming implement both interfaces.

type SystemSection

type SystemSection struct {
	Content         string
	CacheCheckpoint bool
}

SystemSection is one piece of the system prompt. When CacheCheckpoint is true, the backend places a prompt-cache breakpoint after this section (Anthropic: cache_control {"type":"ephemeral"}).

type ToolDefinition

type ToolDefinition struct {
	Name        string
	Description string
	InputSchema []byte // JSON Schema object
}

ToolDefinition describes a tool the model may call.

type ToolResult

type ToolResult struct {
	ToolUseID string `json:"tool_use_id"`
	Content   string `json:"content"`
	IsError   bool   `json:"is_error,omitempty"`
}

ToolResult is the response to a prior ToolUse.

type ToolUse

type ToolUse struct {
	ID    string `json:"id"`
	Name  string `json:"name"`
	Input []byte `json:"input"` // raw JSON
}

ToolUse is a tool call requested by the model.

type Usage

type Usage struct {
	InputTokens  int
	OutputTokens int
}

Usage holds token counts for a single call.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL