tooladapter

package module
v1.0.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 13, 2025 License: Apache-2.0 Imports: 13 Imported by: 0

README

OpenAI Tool Adapter

OpenAI Tool Adapter

Go Version Coverage Go Report Card

A high-performance Go package that enables seamless function calling for Large Language Models that lack native tool support. By transforming OpenAI-style tool requests into prompt-based instructions and intelligently parsing responses back into structured tool calls, it creates a transparent compatibility layer that maintains the familiar OpenAI API while supporting any instruction-following model.

✨ Key Features

  • 🔄 OpenAI SDK Compatible - Uses official github.com/openai/openai-go types for seamless integration
  • ⚡ Streaming & Unary Support - Full compatibility with both standard and streaming chat completions
  • 🎯 Robust Response Parsing - Finite state machine JSON extraction reliably handles diverse LLM response formats
  • 💬 Multi-turn Tool Conversations - Automatic handling of ToolMessage results in conversation history
  • 🚀 High-Performance Processing - Optimized prompt generation and JSON parsing for minimal latency
  • 🔍 Complete Observability - Type-safe metrics and structured logging without vendor lock-in
  • 🛡️ Production Ready - Battle-tested with comprehensive error handling and edge case coverage

🚀 Quick Start

Installation
go get github.com/juburr/openai-tool-adapter
go get github.com/openai/openai-go
Basic Usage

[!IMPORTANT]
Only use this adapter for models that lack native function calling support. Determine your model's capabilities independently and apply the adapter only when necessary.

package main

import (
    "context"
    "fmt"
    "log"
    
    "github.com/juburr/openai-tool-adapter"
    "github.com/openai/openai-go"
    "github.com/openai/openai-go/option"
)

func main() {
    // Initialize the tool adapter
    adapter := tooladapter.New()
    
    // Create OpenAI client
    client := openai.NewClient(option.WithAPIKey("your-api-key"))
    
    // Define your tools (standard OpenAI format)
    tools := []openai.ChatCompletionToolParam{
        {
            Type: "function",
            Function: openai.FunctionDefinitionParam{
                Name:        "get_weather",
                Description: openai.String("Get current weather for a location"),
                Parameters: openai.FunctionParameters{
                    "type": "object",
                    "properties": map[string]interface{}{
                        "location": map[string]interface{}{
                            "type":        "string",
                            "description": "City name",
                        },
                    },
                    "required": []string{"location"},
                },
            },
        },
    }
    
    // Create your request (standard OpenAI format)
    request := openai.ChatCompletionNewParams{
        Model: openai.ChatModelGPT4o, // Any model without native tool support
        Messages: []openai.ChatCompletionMessageParamUnion{
            openai.UserMessage("What's the weather like in San Francisco?"),
        },
        Tools: tools,
    }
    
    // Transform request for non-native models
    transformedRequest, err := adapter.TransformCompletionsRequest(request)
    if err != nil {
        log.Fatal(err)
    }
    
    // Send to your LLM service
    ctx := context.Background()
    response, err := client.Chat.Completions.New(ctx, transformedRequest)
    if err != nil {
        log.Fatal(err)
    }
    
    // Transform response back to OpenAI format
    finalResponse, err := adapter.TransformCompletionsResponse(response)
    if err != nil {
        log.Fatal(err)
    }
    
    // Use response with standard OpenAI tool calling logic
    if len(finalResponse.Choices) > 0 && len(finalResponse.Choices[0].Message.ToolCalls) > 0 {
        for _, toolCall := range finalResponse.Choices[0].Message.ToolCalls {
            fmt.Printf("Function: %s\nArguments: %s\n", 
                toolCall.Function.Name, 
                toolCall.Function.Arguments)
        }
    }
}
Multi-turn Conversations with Tool Results

The adapter automatically handles tool results in multi-turn conversations. When you include ToolMessage types in your conversation history, they are extracted and converted into natural language prompts that the model can understand:

// Multi-turn conversation with tool results
request := openai.ChatCompletionNewParams{
    Model: "gpt-4", 
    Messages: []openai.ChatCompletionMessageParamUnion{
        openai.UserMessage("What's the weather in San Francisco?"),
        // Assistant responds with tool calls (from previous interaction)
        openai.AssistantMessage("I'll check the weather for you.", 
            openai.ToolCall{
                ID: "call_123", 
                Type: "function", 
                Function: openai.Function{
                    Name: "get_weather", 
                    Arguments: `{"location": "San Francisco"}`,
                },
            },
        ),
        // Tool result from executing the tool
        openai.ToolMessage("The weather in San Francisco is 72°F and sunny.", "call_123"),
        // New user request
        openai.UserMessage("Can you format that into a nice summary?"),
    },
    Tools: tools, // Optional: can be omitted if no new tools needed
}

transformedRequest, err := adapter.TransformCompletionsRequest(request)
// Tool results are automatically converted to natural language context

The adapter handles four different scenarios:

  1. No tools or results: Request passes through unchanged
  2. Tools only: Original behavior - tools injected into system prompt
  3. Tool results only: Results converted to natural language context (useful for final iterations)
  4. Both tools and results: Tool definitions + previous results both included in prompt
Configuration Options
// Configure with multiple options including tool processing policies
adapter := tooladapter.New(
    tooladapter.WithCustomPromptTemplate(template),
    tooladapter.WithLogger(logger),
    tooladapter.WithMetricsCallback(callback),
    tooladapter.WithSystemMessageSupport(true), // Enable for models that support system messages
    tooladapter.WithToolPolicy(tooladapter.ToolStopOnFirst), // Control tool processing
    tooladapter.WithToolMaxCalls(5), // Limit tool calls for safety
)

// Use pre-configured option sets
adapter := tooladapter.New(tooladapter.WithLogLevel(slog.LevelInfo))
Streaming Support
// Create streaming request
stream := client.Chat.Completions.NewStreaming(ctx, transformedRequest)

// Wrap with adapter
adaptedStream := adapter.TransformStreamingResponse(stream)
defer adaptedStream.Close()

// Process stream with real-time tool call detection
for adaptedStream.Next() {
    chunk := adaptedStream.Current()
    // Handle tool calls and content as they arrive
}

📖 Documentation

Core Documentation
Advanced Topics

🔧 Configuration Reference

Available Options
Option Description Use Case
WithCustomPromptTemplate(string) Override default tool prompt template Custom instruction formatting
WithLogger(*slog.Logger) Set custom structured logger Production logging integration
WithLogLevel(slog.Level) Set logging level with default handler Simple log level control
WithMetricsCallback(func) Enable metrics collection Performance monitoring
WithSystemMessageSupport(bool) Enable/disable system message support Model-specific message role handling
WithToolCollectWindow(time.Duration) Set collection timeout window Time-based tool collection limits
WithToolPolicy(ToolPolicy) Control tool processing behavior Latency vs completeness trade-offs
WithToolMaxCalls(int) Limit maximum tool calls processed Safety and resource management
WithToolCollectMaxBytes(int) Limit maximum bytes during tool collection Memory safety and resource protection
WithCancelUpstreamOnStop(bool) Cancel upstream context when stopping Resource conservation in streaming
WithStreamingToolBufferSize(int) Set maximum streaming buffer size Control memory usage during streaming tool parsing
WithPromptBufferReuseLimit(int) Set buffer pool reuse threshold Memory management in high-throughput environments
WithStreamingEarlyDetection(int) Enable early tool call detection in streaming Prevent preface text emission when tool calls follow
Pre-configured Option Sets
Option Set Configuration Best For
WithLogger() Custom logger (JSON or text) Any environment
WithLogLevel() Level-only using default handler Simple setups

📊 Observability

The adapter provides comprehensive observability without vendor lock-in:

// Complete observability setup
adapter := tooladapter.New(
    // Structured logging (integrates with any log system)
    tooladapter.WithLogger(slog.New(slog.NewJSONHandler(os.Stdout, nil))),
    
    // Type-safe metrics (works with any monitoring platform)
    tooladapter.WithMetricsCallback(func(data tooladapter.MetricEventData) {
        switch eventData := data.(type) {
        case tooladapter.ToolTransformationData:
            // Send to Prometheus, DataDog, New Relic, etc.
            yourMetrics.RecordTransformation(eventData)
            
        case tooladapter.FunctionCallDetectionData:
            // High-precision timing with nanosecond accuracy
            yourMetrics.RecordProcessingTime(eventData.Performance.ProcessingDuration)
        }
    }),
)

⚡ Performance

The OpenAI Tool Adapter delivers excellent performance across both transformation entry points. Expect microsecond-level transformations with very few memory allocations. Performance remains highly predictable regardless of complexity, making it suitable for production workloads.

Benchmark Data

Date: Aug 13, 2025
Release: v1.0.0
Processor: AMD Ryzen 9 9950X3D

🏎️ Request Transformations (tool injection into prompts)

  • Tiny (1 tool): 1.75 μs/op, 2.99 KB/op, 32 allocs/op
  • Small (5 tools): 8.38 μs/op, 11.69 KB/op, 167 allocs/op
  • Medium (20 tools): 33.51 μs/op, 44.49 KB/op, 689 allocs/op
  • Large (50 tools): 81.58 μs/op, 108.90 KB/op, 1686 allocs/op

🐎 Response Transformations (function call detection & parsing)

  • Tiny (no tool calls): 0.24 μs/op, 0.19 KB/op, 1 alloc/op
  • Small (single tool call): 2.58 μs/op, 3.35 KB/op, 32 allocs/op
  • Medium (multiple tool calls): 6.33 μs/op, 6.18 KB/op, 46 allocs/op
  • Large (many tool calls): 21.62 μs/op, 23.40 KB/op, 80 allocs/op
  • Large (complex tool call): 7.76 μs/op, 9.17 KB/op, 36 allocs/op
  • Large (mixed content): 4.90 μs/op, 7.55 KB/op, 33 allocs/op
Running Benchmarks

Run performance benchmarks to validate performance on your hardware:

# Run all benchmarks
go test -bench=. -benchmem ./...

# JSON performance benchmarks
go test -bench=JSON ./...

# Parser-specific benchmarks
go test -bench=Parser -benchmem ./...

Performance Characteristics:

  • Linear scaling - Performance scales proportionally with content size
  • Zero-allocation operations - Core string processing has no memory overhead
  • Sub-microsecond parsing - Most operations complete in microseconds or faster
  • Predictable behavior - Consistent allocation patterns across scenarios

🧪 Testing

# Run all tests
go test ./...

# Run with coverage
go test -race -coverprofile=coverage.out ./...
go tool cover -html=coverage.out

# Test with benchmarking
go test -bench=. ./...

# Run fuzzing tests
go test -fuzz=FuzzJSONExtractor -fuzztime=30s
go test -fuzz=FuzzValidateFunctionName -fuzztime=30s
go test -fuzz=FuzzTransformCompletionsResponse -fuzztime=30s

# Run end-to-end integration tests
# Requires a running vLLM instance with Gemma 3
cd e2e && go test -tags e2e -v .

The test suite includes:

  • High code coverage with comprehensive edge case testing
  • Fuzz testing for JSON parsing, function validation, and response transformation
  • Production scenario testing including resource exhaustion and malicious input handling
  • Concurrency stress testing with race condition detection
  • Integration testing for real-world usage patterns

🤝 Contributing

Contributions are welcome! Please:

  1. Fork the repository
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Commit your changes (git commit -m 'feat: new amazing feature description')
  4. Push to the branch (git push origin feature/amazing-feature)
  5. Open a Pull Request

Commit prefixes should use the Conventional Commits standard: feat:, fix:, docs:, style:, refactor:, perf:, test:, build:, ci:, revert, or chore:. These popular extensions are also okay: deps:, sec:, infra:, release:, and wip:.

📄 License

This project is licensed under the Apache 2 License - see the LICENSE file for details.

🙏 Acknowledgments

  • Philipp Schmid for pioneering the prompt-based function calling technique
  • Google's Gemma team for validating and documenting the approach
  • The OpenAI team for their excellent Go SDK
  • The open-source community for continuous improvements and feedback
  • Claude Code for being my copilot on this journey and allowing me to throw this package together so quickly

⭐ Star this repository if it helps you build better LLM applications!

Documentation

Overview

Package tooladapter provides OpenAI tool compatibility for Large Language Models that lack native function calling support. It transforms OpenAI-style tool requests into prompt-based format and parses model responses back into structured tool calls.

CONCURRENCY SUMMARY:

  • Adapter: Thread-safe, can be shared across goroutines
  • StreamAdapter: NOT thread-safe, single-consumer design
  • Parser functions: Thread-safe, stateless operations

Index

Constants

View Source
const (
	MaxFunctionNameLength = 64
	MaxPrefixLength       = 64
)

Function name validation constants.

View Source
const (
	// DefaultPromptTemplate provides a robust, concise template that works across LLM families.
	// It emphasizes immediate, JSON-only tool calls when appropriate, and natural language otherwise.
	DefaultPromptTemplate = `` /* 690-byte string literal not displayed */

)

Variables

This section is empty.

Functions

func ApplyOptions

func ApplyOptions(adapter *Adapter, opts []Option)

ApplyOptions applies a slice of options to an adapter, handling errors gracefully. This helper function makes it easier to apply multiple options while handling validation errors properly.

func ExtractFunctionCalls

func ExtractFunctionCalls(candidates []string) []functionCall

ExtractFunctionCalls preserves the previous API by returning only the parsed calls. It will return either a slice parsed from an array or a single-element slice from an object.

func ExtractFunctionCallsDetailed

func ExtractFunctionCallsDetailed(candidates []string) ([]functionCall, bool)

ExtractFunctionCalls attempts to parse function calls from JSON candidates.

VALIDATION STRATEGY: This function provides comprehensive validation through a two-stage process: 1. JSON Structure Validation: DisallowUnknownFields() ensures only "name" and "parameters" fields are present 2. Content Validation: ValidateFunctionCall() ensures required fields are present and valid

The validation handles all edge cases: - Empty names: {"name": "", "parameters": null} -> rejected by ValidateFunctionName - Missing names: {"parameters": null} -> JSON unmarshals to empty string, rejected - Null names: {"name": null, "parameters": null} -> JSON unmarshals to empty string, rejected - Whitespace-only names: {"name": " ", "parameters": null} -> rejected by character validation - Extra fields: {"name": "func", "parameters": null, "extra": "field"} -> rejected by DisallowUnknownFields

This multi-layered approach ensures only valid OpenAI-compatible function calls are extracted. ExtractFunctionCallsDetailed attempts to parse function calls and returns whether the matched JSON was an array (true) or a single object (false). Returns nil, false when no match.

func HasCompleteJSON

func HasCompleteJSON(content string) bool

HasCompleteJSON checks if the given text contains at least one valid function call.

func ValidateFunctionCall

func ValidateFunctionCall(call functionCall) bool

ValidateFunctionCall checks if a parsed object represents a valid function call.

func ValidateFunctionCallArray

func ValidateFunctionCallArray(calls []functionCall) bool

ValidateFunctionCallArray checks if a parsed array contains valid function calls.

func ValidateFunctionName

func ValidateFunctionName(name string) error

ValidateFunctionName validates function names manually for performance. This function is thread-safe and can be called concurrently.

Types

type Adapter

type Adapter struct {
	// contains filtered or unexported fields
}

Adapter translates standard tool-call requests into a prompt-based format.

THREAD SAFETY: Adapter instances are safe for concurrent use by multiple goroutines. All public methods can be called concurrently without external synchronization.

Concurrency design:

  • All fields are immutable after construction (set once during New())
  • sync.Pool handles concurrent buffer access internally
  • slog.Logger is thread-safe
  • Metrics callbacks should be implemented as thread-safe by users
  • No shared mutable state between method calls

Usage patterns:

  • Single adapter instance can handle requests from multiple goroutines
  • Each method call is independent and stateless
  • StreamAdapter instances are NOT thread-safe (single-consumer design)

func New

func New(opts ...Option) *Adapter

New creates a new tool adapter with optional configurations

func (*Adapter) GenerateToolCallID

func (a *Adapter) GenerateToolCallID() string

GenerateToolCallID generates a unique ID for a tool call using UUIDv7. UUIDv7 provides the performance benefits of timestamp-based generation while maintaining full RFC 4122 compliance and battle-tested collision resistance.

THREAD SAFETY: This method is safe for concurrent use by multiple goroutines. Each call generates a cryptographically unique ID with proper collision resistance.

Benefits of UUIDv7 over UUIDv4: - Timestamp-based prefix enables natural sorting and better database performance - Better concurrent performance than gofrs/uuid implementation - Still cryptographically secure with proper collision resistance - Maintains standard UUID format that OpenAI expects - Includes timestamp extraction methods for debugging/analytics

Performance: ~270ns/op vs ~210ns/op for UUIDv4, but provides ordering benefits.

func (*Adapter) TransformCompletionsRequest

func (a *Adapter) TransformCompletionsRequest(req openai.ChatCompletionNewParams) (openai.ChatCompletionNewParams, error)

TransformCompletionsRequest modifies a chat completion request to inject tool definitions. This is the backward-compatible version that uses context.Background(). For production use with timeouts and cancellation, use TransformCompletionsRequestWithContext.

func (*Adapter) TransformCompletionsRequestWithContext

func (a *Adapter) TransformCompletionsRequestWithContext(ctx context.Context, req openai.ChatCompletionNewParams) (openai.ChatCompletionNewParams, error)

TransformCompletionsRequestWithContext modifies a chat completion request to inject tool definitions and process tool results with context support for cancellation and timeouts.

func (*Adapter) TransformCompletionsResponse

func (a *Adapter) TransformCompletionsResponse(resp openai.ChatCompletion) (openai.ChatCompletion, error)

TransformCompletionsResponse processes LLM responses to extract and format tool calls. This is the backward-compatible version that uses context.Background(). For production use with timeouts and cancellation, use TransformCompletionsResponseWithContext.

func (*Adapter) TransformCompletionsResponseWithContext

func (a *Adapter) TransformCompletionsResponseWithContext(ctx context.Context, resp openai.ChatCompletion) (openai.ChatCompletion, error)

TransformCompletionsResponseWithContext processes LLM responses to extract and format tool calls with context support for cancellation and timeouts. This function now processes ALL choices in the response, not just the first one.

func (*Adapter) TransformStreamingResponse

func (a *Adapter) TransformStreamingResponse(stream ChatCompletionStreamInterface) *StreamAdapter

TransformStreamingResponse creates a stream adapter that processes tool calls. This is the backward-compatible version that uses context.Background(). For production use with timeouts and cancellation, use TransformStreamingResponseWithContext.

func (*Adapter) TransformStreamingResponseWithContext

func (a *Adapter) TransformStreamingResponseWithContext(ctx context.Context, stream ChatCompletionStreamInterface) *StreamAdapter

TransformStreamingResponseWithContext creates a stream adapter that processes tool calls with context support for cancellation and timeouts.

type ChatCompletionStreamInterface

type ChatCompletionStreamInterface interface {
	Next() bool
	Current() openai.ChatCompletionChunk
	Err() error
	Close() error
}

ChatCompletionStreamInterface represents the streaming interface returned by OpenAI SDK This matches the interface returned by client.Chat.Completions.NewStreaming()

type FunctionCallDetectionData

type FunctionCallDetectionData struct {
	// FunctionCount is the number of function calls detected and parsed
	FunctionCount int `json:"function_count"`

	// FunctionNames lists the names of all functions called
	FunctionNames []string `json:"function_names"`

	// ContentLength is the length of the original response content in characters
	ContentLength int `json:"content_length"`

	// JSONCandidates is the number of potential JSON blocks found in the content
	JSONCandidates int `json:"json_candidates"`

	// Streaming indicates whether this detection occurred in streaming mode
	Streaming bool `json:"streaming"`

	// Performance contains timing and resource metrics for this detection
	Performance PerformanceMetrics `json:"performance"`
}

FunctionCallDetectionData contains metrics about function call parsing. This event is emitted when the adapter extracts function calls from LLM responses and converts them back to OpenAI-compatible format.

func (FunctionCallDetectionData) EventType

func (d FunctionCallDetectionData) EventType() MetricEvent

type JSONCandidate

type JSONCandidate struct {
	Content []rune
	Start   int
	End     int
}

JSONCandidate represents a potential JSON block found in the text. Content is a slice of the original input, avoiding allocations.

type JSONExtractor

type JSONExtractor struct {
	// contains filtered or unexported fields
}

JSONExtractor uses a state machine to reliably extract JSON objects and arrays.

func NewJSONExtractor

func NewJSONExtractor(input string) *JSONExtractor

NewJSONExtractor creates a new JSON extractor for the given input text.

func (*JSONExtractor) ExtractJSONBlocks

func (je *JSONExtractor) ExtractJSONBlocks() []string

ExtractJSONBlocks finds all potential JSON objects and arrays in the input text. It uses a single-pass parser for efficiency.

type MetricEvent

type MetricEvent string

MetricEvent represents the type of metric event being emitted. Each event corresponds to a significant operation within the tool adapter.

const (
	// MetricEventToolTransformation fires when tools are injected into system prompt.
	// This event indicates that the adapter has successfully converted OpenAI-style
	// tool definitions into a prompt-based format for models without native tool support.
	MetricEventToolTransformation MetricEvent = "tool_transformation"

	// MetricEventFunctionCallDetection fires when function calls are parsed from responses.
	// This event indicates that the adapter has successfully extracted and converted
	// function calls from LLM response text back into OpenAI-compatible tool calls.
	MetricEventFunctionCallDetection MetricEvent = "function_call_detection"
)

type MetricEventData

type MetricEventData interface {
	EventType() MetricEvent
}

MetricEventData is implemented by all metric event data structures. This interface enables type-safe handling of different event types while maintaining a clean callback signature.

type Option

type Option func(*Adapter)

Option is a function that configures the Adapter. This functional options pattern provides several key benefits: 1. Backwards compatibility - new options don't break existing code 2. Optional parameters - users only specify what they want to change 3. Self-documenting - option names clearly indicate their purpose 4. Validation - each option can validate its input independently

func DefaultOptions

func DefaultOptions() []Option

DefaultOptions returns a set of sensible default options. This is useful for applications that want to start with good defaults and then customize specific aspects.

func WithCancelUpstreamOnStop

func WithCancelUpstreamOnStop(cancel bool) Option

WithCancelUpstreamOnStop controls whether the upstream stream is cancelled when stopping tool collection in streaming mode.

This option only applies to streaming mode with ToolStopOnFirst or ToolCollectThenStop policies. When true, the adapter will cancel the upstream context to stop further content generation.

Default: true

func WithCustomPromptTemplate

func WithCustomPromptTemplate(template string) Option

WithCustomPromptTemplate overrides the default system prompt template. This allows customization for specific LLM families or use cases.

The template must contain exactly one %s placeholder where tool definitions will be inserted. The function validates this requirement.

func WithLogLevel

func WithLogLevel(level slog.Level) Option

WithLogLevel sets the logging level for the default logger. This is a convenience option when you want to use slog.Default() but control the level. For production use, consider using WithLogger with a properly configured logger.

func WithLogger

func WithLogger(logger *slog.Logger) Option

WithLogger sets a custom slog.Logger for the adapter. This enables structured logging for operational observability in production.

Logging strategy: - INFO: Operational events (tool transformations, function call detection) - DEBUG: Detailed information (performance metrics) - WARN: Unexpected but recoverable situations - ERROR: Actual errors that affect functionality

If no logger is provided, a no-op logger is used to avoid breaking existing code.

func WithMetricsCallback

func WithMetricsCallback(callback func(MetricEventData)) Option

WithMetricsCallback sets a callback function that receives metric events. This enables integration with monitoring systems like Prometheus, DataDog, or custom metrics collection.

The callback receives typed event data that can be safely type-switched to access specific metrics. All event data includes performance metrics for operational monitoring.

Example usage:

adapter := tooladapter.New(
    tooladapter.WithMetricsCallback(func(data tooladapter.MetricEventData) {
        switch eventData := data.(type) {
        case tooladapter.ToolTransformationData:
            // Handle tool transformation metrics
            myMetrics.ToolTransformations.Inc()
        case tooladapter.FunctionCallDetectionData:
            // Handle function call detection metrics
            myMetrics.FunctionCalls.Add(float64(eventData.FunctionCount))
            myMetrics.ProcessingDuration.Observe(float64(eventData.Performance.ProcessingDurationMs))
        }
    }),
)

The callback is called synchronously during adapter operations, so it should be fast to avoid impacting request processing performance. For expensive operations like database writes, consider using a background goroutine or message queue.

IMPORTANT: The adapter includes panic recovery for metrics callbacks. If your callback panics, the panic will be caught, logged, and the adapter will continue normal operation. This ensures that metrics collection failures never impact core functionality. However, you should still implement proper error handling in your callbacks for best practices.

func WithPromptBufferReuseLimit

func WithPromptBufferReuseLimit(thresholdBytes int) Option

WithPromptBufferReuseLimit sets the maximum size of prompt generation buffers that will be returned to the buffer pool for reuse. Larger buffers are discarded to prevent the buffer pool from growing unbounded when processing very large tool definitions.

This affects the internal buffer pool used for building tool prompts during request transformation. Buffers that exceed this threshold are garbage collected rather than being pooled for reuse.

Use cases:

  • Increase for applications with consistently large tool schemas
  • Decrease for memory-sensitive environments
  • Set very low for testing pool behavior

Default: 64KB (64 * 1024 bytes)

func WithStreamingEarlyDetection

func WithStreamingEarlyDetection(lookAheadChars int) Option

WithStreamingEarlyDetection enables early tool call detection in streaming responses by looking ahead within the first N characters of content for tool call patterns. This improves buffering heuristics when models emit explanatory text before JSON.

The adapter will search for tool call patterns like {"name": or [{"name": within the specified character limit. This helps catch tool calls that start after some preface text, improving content suppression for ToolStopOnFirst/ToolCollectThenStop.

Recommended values:

  • 80-100 characters: Good balance of recall vs false positives
  • 120 characters: More generous, catches longer prefaces
  • 0 (default): Disabled, uses only immediate JSON detection

Use cases:

  • Enable when models frequently add explanatory text before tool calls
  • Keep disabled for maximum performance and minimal false positives
  • Use with ToolStopOnFirst or ToolCollectThenStop for content suppression

Note: ToolAllowMixed policy streams content regardless, so this mainly benefits policies that suppress content when tool calls are detected.

func WithStreamingToolBufferSize

func WithStreamingToolBufferSize(limitBytes int) Option

WithStreamingToolBufferSize sets the maximum amount of content that can be buffered while parsing streaming responses for tool calls. This prevents memory exhaustion during streaming by limiting how much text is held in memory while searching for complete JSON tool call structures.

When this limit is exceeded during streaming, the buffered content is processed as regular text rather than continuing to search for tool calls.

Use cases:

  • Increase for models that generate very large tool calls
  • Decrease for memory-constrained environments
  • Set very low for testing buffer overflow behavior

Default: 10MB (10 * 1024 * 1024 bytes)

func WithSystemMessageSupport

func WithSystemMessageSupport(supported bool) Option

WithNoSystemInstructionRole sets which role to use when no system message is present. Default is false to support models that ignore or lack a system role (e.g., Gemma 3), but you should set this to true if your model supports or requires a system message.

func WithToolCollectMaxBytes

func WithToolCollectMaxBytes(maxBytes int) Option

WithToolCollectMaxBytes sets the maximum number of bytes to collect during JSON tool call processing as a safety cap.

This prevents memory exhaustion from malformed or malicious responses. Set to 0 for no limit (not recommended for production).

Default: 65536 (64KB) - provides DoS protection while allowing legitimate use cases

func WithToolCollectWindow

func WithToolCollectWindow(duration time.Duration) Option

WithToolCollectWindow sets the maximum time to wait for additional tools when using ToolCollectThenStop policy in streaming mode.

If set to 0, uses structure-only batching (no timer). This option is ignored for non-streaming mode and other policies. An upper bound is unnecessary, as overly high durations are equivalent in behavior to both non-streaming mode and the ToolDrainAll policy.

Default: 200ms

func WithToolMaxCalls

func WithToolMaxCalls(maxCalls int) Option

WithToolMaxCalls sets the maximum number of tool calls to collect across both streaming and non-streaming modes.

This provides a safety cap to prevent excessive tool call processing. Set to 0 for no limit (not recommended for production).

Default: 8

func WithToolPolicy

func WithToolPolicy(policy ToolPolicy) Option

WithToolPolicy sets the tool processing policy for the adapter. This controls how tool calls are detected, collected, and emitted.

Available policies:

  • ToolStopOnFirst: Stop on first tool call (lowest latency, safest)
  • ToolCollectThenStop: Collect tools until array closes or limits reached
  • ToolDrainAll: Read entire response and collect all tools
  • ToolAllowMixed: Allow both text content and tools to be emitted

Default: ToolStopOnFirst

type ParseState

type ParseState int

ParseState represents the current state of the JSON parser's state machine.

const (
	StateInObject ParseState = iota // Inside a JSON object
	StateInArray                    // Inside a JSON array
	StateInString                   // Inside a string literal
	StateInEscape                   // Processing an escape sequence
)

type PerformanceMetrics

type PerformanceMetrics struct {
	// ProcessingDuration is the total time spent processing the operation
	// Uses time.Duration for nanosecond precision - callers can convert as needed
	ProcessingDuration time.Duration `json:"processing_duration"`

	// MemoryAllocatedBytes tracks memory allocation during the operation (when available)
	MemoryAllocatedBytes int64 `json:"memory_allocated_bytes,omitempty"`

	// SubOperations provides timing breakdowns for complex operations
	// Keys might include: "prompt_generation", "json_parsing", "validation", etc.
	// Uses time.Duration for precise timing measurements
	// Note: This map is created fresh for each metric event and is never modified
	// after creation, making it safe for concurrent read access.
	SubOperations map[string]time.Duration `json:"sub_operations,omitempty"`
}

PerformanceMetrics contains timing and resource usage information. These metrics are included with most events to provide operational visibility into the adapter's performance characteristics.

Thread Safety: PerformanceMetrics instances are immutable after creation. The SubOperations map is created fresh for each metric event and is never modified after the metric is emitted. This makes it safe for concurrent access by metric callbacks, even if they spawn goroutines.

type StreamAdapter

type StreamAdapter struct {
	// contains filtered or unexported fields
}

StreamAdapter wraps an OpenAI streaming response to intercept and transform tool calls. It implements the same interface as the original stream while processing tool calls.

THREAD SAFETY: StreamAdapter instances are NOT thread-safe and designed for single-consumer use. Each StreamAdapter should be used by only one goroutine, following the same pattern as the underlying OpenAI streaming response.

Concurrency design:

  • Internal mutex (mu) protects all mutable fields during method calls
  • Methods are safe to call sequentially from a single goroutine
  • Multiple goroutines should NOT call methods on the same StreamAdapter instance
  • Context cancellation is thread-safe and can be triggered from any goroutine
  • Close() can be safely called concurrently with other operations

Usage pattern:

go func() {
    stream := adapter.TransformStreamingResponseWithContext(ctx, sourceStream)
    defer stream.Close()
    for stream.Next() {  // Single goroutine only
        chunk := stream.Current()
        // process chunk
    }
}()

func (*StreamAdapter) Close

func (s *StreamAdapter) Close() error

Close closes the underlying stream and cancels the context.

func (*StreamAdapter) Current

Current returns the current chunk in the stream.

func (*StreamAdapter) Err

func (s *StreamAdapter) Err() error

Err returns any error from the stream.

func (*StreamAdapter) Next

func (s *StreamAdapter) Next() bool

type ToolPolicy

type ToolPolicy int

ToolPolicy defines how tool calls are handled during response processing. Different policies provide various trade-offs between latency, completeness, and behavior.

const (
	// ToolStopOnFirst stops processing on the first valid tool call (safest + lowest latency).
	// Emits only the first tool call detected and ignores subsequent content/calls.
	ToolStopOnFirst ToolPolicy = iota

	// ToolCollectThenStop halts content emission but collects tools until array closes,
	// timeout, or limits are reached (short batching window or structure-only).
	ToolCollectThenStop

	// ToolDrainAll reads to end of stream and collects all tool calls before emitting
	// (read to EOS, collect everything).
	ToolDrainAll

	// ToolAllowMixed streams both text content and tools together without suppression
	// (stream text and tools together).
	ToolAllowMixed
)

func (ToolPolicy) String

func (tp ToolPolicy) String() string

String returns a human-readable string representation of the ToolPolicy.

type ToolTransformationData

type ToolTransformationData struct {
	// ToolCount is the number of tools being transformed
	ToolCount int `json:"tool_count"`

	// ToolNames lists the names of all tools being transformed
	ToolNames []string `json:"tool_names"`

	// PromptLength is the length of the generated system prompt in characters
	PromptLength int `json:"prompt_length"`

	// Performance contains timing and resource metrics for this transformation
	Performance PerformanceMetrics `json:"performance"`
}

ToolTransformationData contains metrics about tool-to-prompt transformations. This event is emitted when the adapter converts OpenAI tool definitions into system prompts for models without native tool support.

func (ToolTransformationData) EventType

func (d ToolTransformationData) EventType() MetricEvent

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL