llm

package
v0.4.2 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 24, 2026 License: MIT Imports: 15 Imported by: 0

Documentation

Overview

Package llm talks to language models behind one interface, Provider (Chat and Stream), with the message and tool types they share.

NewProvider builds the generic providers: "openai", "gemini", "anthropic", "mistral", "ollama", "llamacpp" and "openai-compat" (any OpenAI-compatible server, BaseURL required), each wrapped with WithRetry. A provider written outside this package implements Provider; one that speaks the OpenAI chat/completions format can start from NewOpenAICompat.

Index

Examples

Constants

View Source
const (
	ProviderOpenAI       ProviderType = "openai"
	ProviderGemini       ProviderType = "gemini"
	ProviderOllama       ProviderType = "ollama"
	ProviderAnthropic    ProviderType = "anthropic"
	ProviderMistral      ProviderType = "mistral"
	ProviderLlamaCpp     ProviderType = "llamacpp"      // llama.cpp local server
	ProviderOpenAICompat ProviderType = "openai-compat" // any OpenAI-compatible server; BaseURL required

	MediaAudio MediaType = "audio"
	MediaImage MediaType = "image"
)

Variables

This section is empty.

Functions

func ListLlamaCppModels

func ListLlamaCppModels() []string

ListLlamaCppModels queries the local llama.cpp server for the list of loaded models. Returns nil if the server is unreachable.

func ListOllamaModels

func ListOllamaModels() []string

ListOllamaModels queries the local Ollama server for available models. Returns nil if the server is unreachable - Ollama is optional.

Types

type Config

type Config struct {
	Provider         ProviderType
	Model            string
	APIKey           string // empty = the provider's environment variable
	BaseURL          string // empty = the provider's default URL; required by "openai-compat"
	OllamaNumCtx     int    // Ollama context window, 0 = Ollama's default
	OllamaNumPredict int    // Ollama max generated tokens, 0 = MaxTokens
	// Temperature, when set, is sent as the sampling temperature. Some
	// models refuse it (recent Anthropic ones, OpenAI reasoning models).
	Temperature *float64
	// MaxTokens caps the tokens generated by one model call; 0 leaves the
	// provider's default (32000 for Anthropic, which requires a value).
	MaxTokens int
	// PromptCache marks the prompt for caching where the provider needs to
	// be told (Anthropic: the system prompt and the tools, and the
	// conversation so far). OpenAI and Gemini cache on their own.
	PromptCache bool
	// ResponseSchema, when set, is a JSON Schema the model's answers must
	// follow. Not every model takes it together with tools.
	ResponseSchema json.RawMessage
	// ExtraBody adds raw fields to every request body of the providers that
	// speak OpenAI's format (openai, mistral, llamacpp, openai-compat), after
	// the ones above, which it can override: top_p, presence_penalty,
	// chat_template_kwargs... The other providers ignore it.
	ExtraBody map[string]any
	// Timeout is the longest one model call may take, streamed answer
	// included; 0 keeps the provider's default: 600 s for the OpenAI format,
	// 300 for Anthropic and Ollama, 120 for Gemini. A local model that
	// thinks at length needs more.
	Timeout time.Duration
	// DropReasoning leaves what the model thought out of the history sent
	// back, for the providers that speak OpenAI's format; by default it
	// goes back as reasoning_content, which is what llama.cpp reads.
	DropReasoning bool
}

Config describes a provider for NewProvider.

type FunctionCall

type FunctionCall struct {
	Name      string          `json:"name"`
	Arguments json.RawMessage `json:"arguments"`
}

FunctionCall represents the function name and arguments of a tool call.

type FunctionDef

type FunctionDef struct {
	Name        string          `json:"name"`
	Description string          `json:"description"`
	Parameters  ToolParams      `json:"parameters"`
	Schema      json.RawMessage `json:"-"`
}

FunctionDef defines the name, description and parameters of a tool function. Schema, when set, is a raw JSON Schema of the arguments that replaces Parameters for every provider. A provider written outside this package should send JSONSchema(), or the encoded FunctionDef, rather than read Parameters, which is empty when only Schema is set.

func (FunctionDef) JSONSchema added in v0.2.0

func (d FunctionDef) JSONSchema() json.RawMessage

JSONSchema returns the schema of the arguments: Schema when set, Parameters encoded otherwise.

func (FunctionDef) MarshalJSON added in v0.2.0

func (d FunctionDef) MarshalJSON() ([]byte, error)

MarshalJSON writes JSONSchema under "parameters", the OpenAI shape.

func (*FunctionDef) UnmarshalJSON added in v0.2.0

func (d *FunctionDef) UnmarshalJSON(b []byte) error

UnmarshalJSON reads what MarshalJSON writes: "parameters" is kept whole in Schema, and fills Parameters as far as it fits.

type Media

type Media struct {
	Type     MediaType `json:"type"`
	MimeType string    `json:"mime_type"` // "audio/wav", "image/jpeg", etc.
	Data     []byte    `json:"data"`      // raw bytes (base64 in JSON)
}

Media represents binary content such as audio or an image attached to a message.

type MediaType

type MediaType string

MediaType identifies the kind of media attached to a message.

type Message

type Message struct {
	Role      string     `json:"role"` // user, assistant, system, tool
	Content   string     `json:"content"`
	ToolCalls []ToolCall `json:"tool_calls,omitempty"`
	ToolID    string     `json:"tool_call_id,omitempty"` // for tool responses
	Name      string     `json:"name,omitempty"`         // tool name for responses

	Media []Media `json:"media,omitempty"` // for messages containing media

	// Reasoning is what the model thought before it answered, as a
	// server that speaks OpenAI's format streams it apart
	// (reasoning_content). It goes back with the history, since a
	// template such as Qwen's renders it for the turn under way, unless
	// the provider is told to drop it.
	Reasoning string `json:"reasoning_content,omitempty"`

	// Usage is set by a provider on the message it returns; nil when
	// the provider does not report it. It is not part of the JSON.
	Usage *Usage `json:"-"`
}

Message is a universal message compatible with all providers.

type OpenAICompatConfig

type OpenAICompatConfig struct {
	APIKey  string // sent as "Authorization: Bearer <APIKey>" when not empty
	BaseURL string // up to and including the version, e.g. http://localhost:8000/v1
	Model   string
	Name    string        // returned by Name and used in error messages
	Timeout time.Duration // whole-request timeout, 600s when zero
	// ExtraBody adds raw fields to every request body (tool_choice,
	// temperature, reasoning_effort...). tool_choice is left out of
	// requests that carry no tools.
	ExtraBody map[string]any
	// ExtraBodyFunc, when set, is called for every request and its fields
	// are added after ExtraBody's, for what changes between two requests.
	ExtraBodyFunc func() map[string]any
	ExtraHeaders  map[string]string // extra HTTP headers on every request
	// AudioFormat selects the audio media encoding: "input_audio"
	// (OpenAI convention, expected by llama.cpp) or "audio_url" (vLLM/Gemma
	// recipes). Empty defaults to "audio_url".
	AudioFormat string
	// StreamUsage asks for the token usage at the end of a stream
	// (stream_options.include_usage). Some servers refuse the field.
	StreamUsage bool
	// OnReasoning, when set, receives what the model thinks before it
	// answers, as the server streams it apart from the answer
	// (reasoning_content, or reasoning). It is also kept in the message's
	// Reasoning, never in its Content.
	OnReasoning func(chunk string)
	// DropReasoning sends the history back without the reasoning_content
	// of its assistant messages.
	DropReasoning bool
}

OpenAICompatConfig describes an OpenAI-compatible chat/completions endpoint for NewOpenAICompat.

type Provider

type Provider interface {
	Chat(ctx context.Context, messages []Message, tools []Tool) (*Message, error)
	Stream(ctx context.Context, messages []Message, tools []Tool, onChunk func(string) error) (*Message, error)
	Name() string
	ModelName() string
}

Provider is the interface for all LLMs.

func NewOpenAICompat

func NewOpenAICompat(cfg OpenAICompatConfig) Provider

NewOpenAICompat returns a provider for an OpenAI-compatible endpoint. It does not retry: wrap it with WithRetry to get the behaviour of NewProvider.

Example

A provider for an OpenAI-compatible server that needs its own settings, with the same retries as NewProvider.

package main

import (
	"fmt"

	"github.com/ThiraSoft/agentkit/llm"
)

func main() {
	p := llm.WithRetry(llm.NewOpenAICompat(llm.OpenAICompatConfig{
		BaseURL:   "http://localhost:8000/v1",
		Model:     "local-model",
		Name:      "local",
		ExtraBody: map[string]any{"temperature": 0.2},
	}))
	fmt.Println(p.Name(), p.ModelName())
}
Output:
local local-model

func NewProvider

func NewProvider(cfg Config) (Provider, error)

NewProvider builds the provider described by cfg, wrapped with WithRetry. APIKey and BaseURL, when set, take precedence over the environment variable and the default URL of the provider. "openai-compat" requires BaseURL. An unknown provider is an error.

Example
package main

import (
	"fmt"
	"log"

	"github.com/ThiraSoft/agentkit/llm"
)

func main() {
	p, err := llm.NewProvider(llm.Config{
		Provider: llm.ProviderOpenAICompat,
		Model:    "local-model",
		BaseURL:  "http://localhost:8080/v1",
	})
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(p.Name(), p.ModelName())
}
Output:
openai-compat local-model

func UnwrapProvider

func UnwrapProvider(p Provider) Provider

UnwrapProvider removes the decorators (retry) to get back the concrete provider, for instance to type-assert a provider you wrote yourself.

func WithRetry

func WithRetry(p Provider) Provider

WithRetry wraps p so that transient errors (rate limit, overload, network drops) are retried with exponential backoff, up to 3 attempts. A stream is retried only while no chunk has been emitted, so shown text is never duplicated. WithRetry returns nil for nil and p itself when p is already wrapped. NewProvider applies it to every provider it builds; use it on a provider built with NewOpenAICompat or written outside this package.

type ProviderType

type ProviderType string

ProviderType identifies an LLM provider implementation.

type StreamError added in v0.3.1

type StreamError struct {
	Provider string // the provider's name
	Type     string // the server's kind of error, such as server_error
	Message  string
}

StreamError is an error the server sent down a stream it had already opened, in place of the rest of the answer. Nothing the model said is in it, so a client can show it apart from the answer rather than as part of it.

func (*StreamError) Error added in v0.3.1

func (e *StreamError) Error() string

type Tool

type Tool struct {
	Type     string      `json:"type"` // "function"
	Function FunctionDef `json:"function"`
}

Tool describes a function tool that an LLM can call.

type ToolArg

type ToolArg struct {
	Type        string `json:"type"` // JSON Schema type: "string", "number", "object", etc.
	Description string `json:"description"`
	Example     string `json:"example,omitempty"`
	Default     string `json:"default,omitempty"`
	Items       any    `json:"items,omitempty"` // for arrays
}

ToolArg describes a single argument in a tool schema.

type ToolCall

type ToolCall struct {
	ID               string       `json:"id,omitempty"`
	Type             string       `json:"type,omitempty"` // "function"
	Function         FunctionCall `json:"function"`
	ThoughtSignature string       `json:"thoughtSignature,omitempty"` // for Gemini
}

ToolCall represents a tool invocation requested by the model.

type ToolParams

type ToolParams struct {
	Type       string         `json:"type"` // "object"
	Properties ToolProperties `json:"properties"`
	Required   []string       `json:"required,omitempty"`
}

ToolParams defines the parameters schema for a tool.

type ToolProperties

type ToolProperties map[string]ToolArg

ToolProperties maps argument names to their schema definitions.

type Usage added in v0.2.0

type Usage struct {
	InputTokens      int
	OutputTokens     int
	CacheReadTokens  int // prompt tokens read from the provider's cache
	CacheWriteTokens int // prompt tokens written to it (Anthropic)
}

Usage counts the tokens of one model call as the provider reports them; a count the provider does not report stays zero. InputTokens counts the whole prompt, the part read from or written to a cache included.

func (Usage) Add added in v0.2.0

func (u Usage) Add(v Usage) Usage

Add returns u plus v, count by count.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL