go-litellm

module
v1.5.5 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 5, 2026 License: Apache-2.0

README

go-litellm

A Go client library for interacting with the LiteLLM API. This package provides a simple, type-safe, and developer-friendly way to perform completions, audio transcription, image-based queries, structured JSON output, embeddings, token counting, tool operations, model browsing, and more.


Installation

go get github.com/andrejsstepanovs/go-litellm@latest

Quick Start

Initialize Client
package main

import (
    "context"
    "net/url"
    "time"
    "log"

    "github.com/andrejsstepanovs/go-litellm/client"
    "github.com/andrejsstepanovs/go-litellm/conf/connections/litellm"
)

func main() {
    baseURL, _ := url.Parse("http://localhost:4000")
    conn := litellm.Connection{
        URL: *baseURL,
        Targets: litellm.Targets{
            System: litellm.Target{Timeout: time.Second},
            LLM:    litellm.Target{Timeout: time.Minute * 2},
            MCP:    litellm.Target{Timeout: time.Minute * 5},
        },
    }

    cfg := client.Config{
        APIKey:      "sk-1234",
        Temperature: 0.7,
    }

    ai, err := client.New(cfg, conn)
    if err != nil {
        log.Fatal(err)
    }

    _ = ai // ready to use
}

Examples

1. Simple Chat Completion
model, _ := ai.Model(ctx, "claude-4")
messages := request.Messages{request.UserMessageSimple("What is the capital of France?")}
req := request.NewCompletionRequest(model, messages, nil, nil, 1)
resp, _ := ai.Completion(ctx, req)
fmt.Println(resp.String()) // The capital of France is Paris.
2. Audio Transcription (Speech-to-Text)
model, _ := ai.Model(ctx, "whisper-1")
extraParams := map[string]any{}
//extraParams["punctuate"] = true
//extraParams["smart_format"] = true
res, _ := ai.SpeechToText(ctx, model, "path/to/audio.oga", extraParams)
fmt.Println(res.Text)
3. Image Analysis / Captioning
model, _ := ai.Model(ctx, "gpt-4o-mini")
messages := request.Messages{
    request.UserMessageImage("Describe this image", "https://example.com/image.jpg"),
}
req := request.NewCompletionRequest(model, messages, nil, nil, 1)
resp, _ := ai.Completion(ctx, req)
fmt.Println(resp.String())
4. Structured JSON Output (Strict Schema)
type City struct {
    CityName        string `json:"city_name"`
    PopulationCount int    `json:"population_count"`
}
type ListOfCities struct {
    Cities []City `json:"cities"`
}

schema := request.JSONSchema{
    Name: "list_of_cities",
    Schema: map[string]interface{}{
        "type": "object",
        "properties": map[string]interface{}{
            "cities": map[string]interface{}{
                "type": "array",
                "items": map[string]interface{}{
                    "type": "object",
                    "properties": map[string]interface{}{
                        "city_name": map[string]interface{}{"type": "string"},
                        "population_count": map[string]interface{}{"type": "integer"},
                    },
                    "required": []string{"city_name", "population_count"},
                },
            },
        },
        "required": []string{"cities"},
    },
    Strict: true,
}

model, _ := ai.Model(ctx, "claude-4")
messages := request.Messages{request.UserMessageSimple("List the 3 largest cities")}
req := request.NewCompletionRequest(model, messages, nil, nil, 0.2)
req.SetJSONSchema(schema)
resp, _ := ai.Completion(ctx, req)

var cities ListOfCities
json.Unmarshal(resp.Bytes(), &cities)
fmt.Printf("%+v\n", cities)
5. Token Count Calculation
model, _ := ai.Model(ctx, "claude-4")
messages := request.Messages{request.UserMessageSimple("Hello")}
req := &request.TokenCounterRequest{
    Model:    model.ModelId,
    Messages: messages,
}
count, _ := ai.TokenCounter(ctx, req)
fmt.Println("Total tokens:", count.TotalTokens)
6. Cache Controls

go-litellm provides both manual and automatic ways to instruct upstream providers (like Anthropic and Gemini) to cache prompt blocks.

Manual Caching

Use CachePoint() on a simple message, or .Cache() on specific blocks for granular control:

model, _ := ai.Model(ctx, "claude-4")

messages := request.Messages{
    request.SystemMessageSimple("Reusable context with a default cache point").CachePoint(),
    request.SystemMessage(request.MessageContents{
        request.MessageContent{
            Type: "text",
            Text: "You are an AI assistant tasked with analyzing legal documents.",
        },
        request.MessageContent{
            Type: "text",
            Text: longLegalAgreement,
        }.Cache(request.CacheControlEphemeral, request.CacheTTL("1h")),
    }),
    request.UserMessageSimple("What are the key terms and conditions?"),
}

fmt.Println("Cache points:", messages.CacheControlCount())

req := request.NewCompletionRequest(model, messages, nil, nil, 1)
resp, _ := ai.Completion(ctx, req)
fmt.Println(resp.String())
Automatic Caching

If you are using LiteLLM proxy, you can use the SetCacheControlInjectionPoints feature to tell LiteLLM to automatically append ephemeral cache markers to specific message roles without modifying messages manually:

model, _ := ai.Model(ctx, "claude-4")

messages := request.Messages{
    request.SystemMessageSimple("Reusable context"),
    request.UserMessageSimple("What are the key terms?"),
}

req := request.NewCompletionRequest(model, messages, nil, nil, 1)
// Automatically inserts {"type": "ephemeral"} on system and user messages
req.SetCacheControlInjectionPoints([]string{"system", "user"})

resp, _ := ai.Completion(ctx, req)
Cache Metrics

When utilizing caching, you can retrieve the tokens used for cache hits and creations directly from the response usage metrics:

fmt.Printf("Cache Read Tokens: %d\n", resp.Usage.CacheReadTokens())
fmt.Printf("Cache Creation Tokens: %d\n", resp.Usage.CacheCreationTokens())
7. List Available Tools
tools, _ := ai.Tools(ctx)
for _, tool := range tools {
    fmt.Printf("%s - %s\n", tool.Name, tool.Description)
}
8. Call a Tool
tool := common.ToolCallFunction{
    Name: "current_time",
    Arguments: map[string]string{"timezone": "Europe/Riga"},
}
res, _ := ai.ToolCall(ctx, tool)
fmt.Println(res.String())
9. Browse Available Models
models, _ := ai.Models(ctx)
for _, m := range models {
    fmt.Println(m.ID, m.OwnedBy)
}
10. Tool-Aware Conversation Example

This example demonstrates how to:

  1. Maintain a conversation history.
  2. Send the list of available tools so the AI knows what it can call.
  3. When AI requests a tool call, execute it, append the result to the history, and send it back.
  4. Repeat until AI no longer requests tools.
package main

import (
    "context"
    "fmt"
    "log"
    "net/url"
    "time"

    "github.com/andrejsstepanovs/go-litellm/client"
    "github.com/andrejsstepanovs/go-litellm/conf/connections/litellm"
    "github.com/andrejsstepanovs/go-litellm/mcp"
    "github.com/andrejsstepanovs/go-litellm/models"
    "github.com/andrejsstepanovs/go-litellm/request"
    "github.com/andrejsstepanovs/go-litellm/response"
)

func main() {
    ctx := context.Background()

    // Create LiteLLM client
    conn := litellm.Connection{
        URL: *mustParseURL("http://localhost:4000"),
        Targets: litellm.Targets{
            System: litellm.Target{Timeout: time.Second},
            LLM:    litellm.Target{Timeout: time.Minute},
            MCP:    litellm.Target{Timeout: time.Minute},
        },
    }
    cfg := client.Config{APIKey: "sk-1234", Temperature: 0.7}
    ai, err := client.New(cfg, conn)
    if err != nil {
        log.Fatal(err)
    }

    // Pick a model and list tools
    model, _ := ai.Model(ctx, "claude-4")
    tools, _ := ai.Tools(ctx)

    // Initial conversation
    messages := request.Messages{
        request.UserMessageSimple("What's the current time in Riga?"),
    }

    finalResp := runToolAwareConversation(ctx, ai, model, tools, messages, 5)
    fmt.Println("Final Answer:", finalResp.String())
}

func runToolAwareConversation(ctx context.Context, ai *client.Litellm, model models.ModelMeta, tools mcp.AvailableTools, messages request.Messages, maxIter int) response.Response {
    if maxIter <= 0 {
        return response.Response{}
    }

    req := request.NewCompletionRequest(model, messages, tools.ToLLMCallTools(), nil, 0.7)
    resp, err := ai.Completion(ctx, req)
    if err != nil {
        log.Fatal("Completion error:", err)
    }

    if resp.Choice().FinishReason == response.FINISH_REASON_TOOL {
        for _, toolCall := range resp.Choice().Message.ToolCalls.SortASC() {
            toolResp, err := ai.ToolCall(ctx, toolCall.Function)
            if err != nil {
                log.Fatal("Tool call error:", err)
            }
            for _, tr := range toolResp {
                messages = append(messages, request.ToolCallMessage(toolCall, tr))
            }
        }
        return runToolAwareConversation(ctx, ai, model, tools, messages, maxIter-1)
    }

    return resp
}

func mustParseURL(s string) *url.URL {
    u, err := url.Parse(s)
    if err != nil {
        log.Fatal(err)
    }
    return u
}

11. Custom Headers

Configure headers once on the client and they're automatically attached to every outgoing request (chat completions, embeddings, token counting, model listing, audio transcription/speech, tool calls, etc.) — mirroring litellm's Python extra_headers kwarg:

cfg := client.Config{
    APIKey:      "sk-1234",
    Temperature: 0.7,
    ExtraHeaders: map[string]string{
        "X-App-Name": "YourAppName",
        "X-User-Id":  "user_123",
    },
}

ai, err := client.New(cfg, conn)

Keys and values are trimmed of surrounding whitespace before use. When creating the client via client.New, Config.Validate() rejects any header whose key or value is empty (or whitespace-only) after trimming.


12. Reasoning Responses

Some reasoning models emit the model's final answer in a reasoning field rather than the standard content slot. Two flavours are encountered in practice:

Field Used by Shape
reasoning_content DeepSeek-style models served via LiteLLM string
reasoning OpenRouter-style models (e.g. openai/gpt-oss-20b) string, object, or array

response.ResponseMessage exposes both:

type ResponseMessage struct {
    Content          string           `json:"content"`
    ReasoningContent string           `json:"reasoning_content"`
    Reasoning        json.RawMessage  `json:"reasoning,omitempty"`
    Role             string           `json:"role"`
    ToolCalls        common.ToolCalls `json:"tool_calls,omitempty"`
}

Reasoning is kept as a json.RawMessage because OpenRouter may serialise it as either a plain string or a structured object/array of reasoning details. Use the ReasoningString() helper to read it as a string regardless of the shape:

msg := resp.Message()

switch {
case msg.Content != "":
    fmt.Println("content:", msg.Content)
case msg.ReasoningContent != "":
    fmt.Println("reasoning_content:", msg.ReasoningContent)
case msg.ReasoningString() != "":
    // gpt-oss-20b on OpenRouter falls here.
    fmt.Println("reasoning:", msg.ReasoningString())
}

ReasoningString():

  • Returns the raw string when reasoning is a JSON string.
  • Re-emits the JSON when reasoning is an object or array of structured reasoning details (e.g. OpenRouter's reasoning_details), so callers can still inspect or extract fields like text.
  • Returns "" for an empty payload, a nil receiver, or any other "empty" value.

ResponseMessage.IsEmpty() is also aware of the new field: a message that carries only reasoning is no longer considered empty.

The existing Response.ReasoningString() helper continues to return only ReasoningContent to preserve backwards compatibility — use Response.Message().ReasoningString() for the OpenRouter-style field.


Supported Endpoints

  • /models – list available models
  • /v2/model/info – detailed model info
  • /model_group/info – fetch model metadata
  • /utils/token_counter – count tokens for a given request
  • /v1/embeddings – generate embeddings
  • /audio/transcriptions – speech-to-text
  • /audio/speech – text-to-speech
  • /mcp-rest/tools/list – list tools
  • /mcp-rest/tools/call – invoke tools
  • /chat/completions – chat completions with support for text, images, strict schemas, and tool calling integration

Contributing

Contributions are welcome! Feel free to open an issue or submit a PR.

License

Apache 2.0

Directories

Path Synopsis

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL