Documentation
¶
Index ¶
- Constants
- Variables
- func AllowedCacheTypes() []string
- func ApplyOpenAIForwardHeaders(req *http.Request, headers []ForwardHeader)
- func CapNumCtxToEmbeddingModelMax(modelDir string, numCtx int) int
- func HTTPErrorMessage(err error) string
- func HTTPStatusCode(err error) int
- func IsROCMHost() bool
- func ModelMaxPositionEmbeddings(modelDir string) int
- func NewHTTPStatusError(status int, body string) error
- func NormalizeCacheType(value string) (string, error)
- func NormalizeNGPULayers(requested int) (int, error)
- func OpenAIModelUsesThinkingTypeDisabled(modelName string) bool
- func PrefersNativeAnthropicMessages(eng Engine) bool
- func ROCMSingleEngineMode() bool
- func ROCMUnifiedMemoryMode() bool
- func ResolveEmbeddingPooling(modelName string) string
- func ResolveNGPULayers(requested int) int
- func ResolveNumCtx(modelDir string, requested int) int
- func ResolveNumCtxWithModelMax(modelDir string, requested int, configured bool) int
- func ResolveNumCtxWithModelSetting(modelDir string, requested, modelSetting int, configured bool) int
- func ResolveNumParallel(requested int) int
- func SupportsNativeToolStreaming(eng Engine) bool
- func UseModelMaxCtxByDefault(configured bool) bool
- type AnthropicMessagesProxier
- type CacheUsage
- type CacheUsageCollector
- type ChatCompletionProxier
- type ContentPart
- type ConvertProgressFunc
- type EmbeddingsProxier
- type Engine
- func LoadEmbeddingEngineWithProgress(modelDir string, lm *model.LocalModel, progress ConvertProgressFunc, ...) (Engine, error)
- func LoadEngine(modelDir string, lm *model.LocalModel) (Engine, error)
- func LoadEngineWithProgress(modelDir string, lm *model.LocalModel, progress ConvertProgressFunc, ...) (Engine, error)
- func LoadEngineWithSpeculativeProgress(modelDir string, lm *model.LocalModel, progress ConvertProgressFunc, ...) (Engine, error)
- func NewOpenAICompatibleEngine(baseURL, modelName, token string) Engine
- func NewOpenAICompatibleEngineWithHeaders(baseURL, modelName, token string, headers []ForwardHeader) Engine
- func NewOpenAIEngine(baseURL, modelName, token string) Engine
- func NewRemoteEngine(baseURL, modelName string, numCtx, numParallel, nGPULayers int, ...) Engine
- type ForwardHeader
- type ImageURL
- type Message
- type NativeToolStreamer
- type Options
- type Session
- type SpeculativeConfig
- type TokenCallback
Constants ¶
const ROCMSingleEngineEnv = "CSGHUB_LITE_ROCM_SINGLE_ENGINE"
const ROCMUnifiedMemoryEnv = "CSGHUB_LITE_ROCM_UNIFIED_MEMORY"
Variables ¶
var ErrUnsupportedFormat = errors.New("unsupported model format for inference")
Functions ¶
func AllowedCacheTypes ¶
func AllowedCacheTypes() []string
AllowedCacheTypes returns the llama-server KV cache dtypes accepted by csghub-lite.
func ApplyOpenAIForwardHeaders ¶ added in v0.9.43
func ApplyOpenAIForwardHeaders(req *http.Request, headers []ForwardHeader)
ApplyOpenAIForwardHeaders adds configured compatibility headers to an upstream request. Host is handled through Request.Host because net/http does not send Header["Host"] as the authority.
func CapNumCtxToEmbeddingModelMax ¶ added in v0.9.77
CapNumCtxToEmbeddingModelMax caps the context to what an embedding model can actually consume. It differs from the chat cap in honouring limits below 1024: 512-token embedding models are the norm (bge, gte, and most sentence encoders), and llama-server allocates the context per slot, so a 512-token model given the 8192 default reserves sixteen times the KV cache it can ever use -- multiplied again by the slot count.
A suspiciously small limit is ignored rather than trusted, since it would leave the model unable to embed anything useful.
func HTTPErrorMessage ¶
func HTTPStatusCode ¶
func IsROCMHost ¶ added in v0.9.12
func IsROCMHost() bool
func NewHTTPStatusError ¶
func NormalizeCacheType ¶
NormalizeCacheType returns a lower-case llama-server cache dtype or "" when unset.
func NormalizeNGPULayers ¶
NormalizeNGPULayers accepts -1 (unset) or any non-negative llama-server --n-gpu-layers value.
func OpenAIModelUsesThinkingTypeDisabled ¶ added in v0.9.64
OpenAIModelUsesThinkingTypeDisabled reports whether the model family disables reasoning via the "thinking": {"type": "disabled"} request field. Callers that build requests directly (such as the evaluation judge) use it to send the field only to models that accept it; strict OpenAI-compatible endpoints reject unrecognized arguments like "thinking" with a 400.
func PrefersNativeAnthropicMessages ¶ added in v0.9.51
PrefersNativeAnthropicMessages reports whether an engine represents a third-party compatible API where an incoming Messages request should be preserved before falling back to Chat Completions.
func ROCMSingleEngineMode ¶ added in v0.9.12
func ROCMSingleEngineMode() bool
func ROCMUnifiedMemoryMode ¶ added in v0.9.27
func ROCMUnifiedMemoryMode() bool
ROCMUnifiedMemoryMode reports whether llama-server should run with GGML_CUDA_ENABLE_UNIFIED_MEMORY=1. On AMD APUs the ROCm backend misreports free device memory (llama.cpp reads /proc/meminfo MemAvailable), so its fit feature keeps full GPU offload while the real hipMalloc is capped by the small VRAM carve-out + GTT and fails. hipMallocManaged (unified memory) allocates from system RAM instead, which is what upstream recommends for APUs (see ggml-org/llama.cpp#18159). Discrete ROCm GPUs are left alone because unified memory hurts their performance.
func ResolveEmbeddingPooling ¶
ResolveEmbeddingPooling returns the llama-server pooling strategy for common embedding model families. The env override is intentionally kept as an escape hatch because GGUF metadata and model cards occasionally disagree.
func ResolveNGPULayers ¶
ResolveNGPULayers returns the effective llama-server GPU layer offload count. Explicit requests win; otherwise the flag stays unset (-1) so llama-server's built-in fit feature can auto-adjust GPU offload to the free device memory. Passing an explicit -ngl disables that auto-fit, which broke large models on small-VRAM hosts (e.g. AMD APUs) when we always forced 9999.
func ResolveNumCtx ¶
ResolveNumCtx returns the effective llama-server context window (--ctx-size / -c). Explicit requests win, then CSGHUB_LITE_LLAMA_NUM_CTX capped at the model maximum, then the optional model maximum, then a conservative model-aware fallback.
func ResolveNumCtxWithModelMax ¶ added in v0.9.54
ResolveNumCtxWithModelMax resolves the context window using the persisted model-maximum default. CSGHUB_LITE_LLAMA_USE_MODEL_MAX_CTX overrides that default when the environment variable is explicitly set.
func ResolveNumCtxWithModelSetting ¶ added in v0.9.67
func ResolveNumCtxWithModelSetting(modelDir string, requested, modelSetting int, configured bool) int
ResolveNumCtxWithModelSetting resolves the context window with a per-model setting slotted between the request and the global default. The order is request > per-model setting > global setting, and a model without a setting resolves exactly as it did before.
func ResolveNumParallel ¶
ResolveNumParallel returns the effective number of parallel slots for llama-server. Explicit requests win, then CSGHUB_LITE_LLAMA_NUM_PARALLEL, then defaultLlamaParallel.
func SupportsNativeToolStreaming ¶ added in v0.9.22
SupportsNativeToolStreaming reports whether eng can stream tool-call responses directly from its backend.
func UseModelMaxCtxByDefault ¶ added in v0.9.54
UseModelMaxCtxByDefault returns the effective model-maximum default. An explicitly set environment variable has precedence over persisted config.
Types ¶
type AnthropicMessagesProxier ¶ added in v0.9.51
type AnthropicMessagesProxier interface {
AnthropicMessages(ctx context.Context, reqBody map[string]interface{}, headers http.Header) (*http.Response, error)
}
AnthropicMessagesProxier exposes direct access to an Anthropic-compatible /v1/messages API. Request headers are supplied separately so protocol version and beta feature headers can be preserved without forwarding local authentication credentials.
type CacheUsage ¶ added in v0.9.40
type CacheUsage struct {
ReadInputTokens int64
CreationInputTokens int64
EligibleInputTokens int64
}
CacheUsage contains prompt-cache metrics aggregated across inference calls made while handling one API request.
type CacheUsageCollector ¶ added in v0.9.40
type CacheUsageCollector struct {
// contains filtered or unexported fields
}
CacheUsageCollector safely aggregates cache usage for concurrent inference work associated with one request.
func WithCacheUsageCollector ¶ added in v0.9.40
func WithCacheUsageCollector(ctx context.Context) (context.Context, *CacheUsageCollector)
WithCacheUsageCollector attaches a request-scoped cache usage collector.
func (*CacheUsageCollector) Snapshot ¶ added in v0.9.40
func (c *CacheUsageCollector) Snapshot() CacheUsage
Snapshot returns the cache usage recorded so far.
type ChatCompletionProxier ¶
type ChatCompletionProxier interface {
ChatCompletion(ctx context.Context, reqBody map[string]interface{}) (*http.Response, error)
}
ChatCompletionProxier exposes direct access to the underlying OpenAI-compatible /v1/chat/completions API for advanced use cases such as native Ollama tool-calling compatibility.
type ContentPart ¶
type ContentPart struct {
Type string `json:"type"`
Text string `json:"text,omitempty"`
ImageURL *ImageURL `json:"image_url,omitempty"`
}
ContentPart represents one part of a multimodal message.
type ConvertProgressFunc ¶
ConvertProgressFunc receives conversion progress updates. If nil, conversion progress is not reported.
type EmbeddingsProxier ¶
type EmbeddingsProxier interface {
Embeddings(ctx context.Context, reqBody map[string]interface{}) (*http.Response, error)
}
EmbeddingsProxier exposes direct access to the underlying OpenAI-compatible /v1/embeddings API.
type Engine ¶
type Engine interface {
// Generate produces text from a prompt, calling onToken for each generated token.
Generate(ctx context.Context, prompt string, opts Options, onToken TokenCallback) (string, error)
// Chat produces a response from a conversation history.
Chat(ctx context.Context, messages []Message, opts Options, onToken TokenCallback) (string, error)
// Close releases the model resources.
Close() error
// ModelName returns the loaded model identifier.
ModelName() string
}
Engine is the interface for model inference backends.
func LoadEmbeddingEngineWithProgress ¶
func LoadEmbeddingEngineWithProgress(modelDir string, lm *model.LocalModel, progress ConvertProgressFunc, verbose bool, numCtx, numParallel, nGPULayers int, cacheTypeK, cacheTypeV, dtype string) (Engine, error)
LoadEmbeddingEngineWithProgress is like LoadEngineWithProgress but starts llama-server in embedding mode for OpenAI-compatible /v1/embeddings.
func LoadEngine ¶
func LoadEngine(modelDir string, lm *model.LocalModel) (Engine, error)
LoadEngine loads a model and returns an Engine for inference. If the model is SafeTensors, it auto-converts to GGUF first. By default, llama-server output is not mirrored to stderr, but it is still captured for diagnostics and appended to the llama-server log file.
func LoadEngineWithProgress ¶
func LoadEngineWithProgress(modelDir string, lm *model.LocalModel, progress ConvertProgressFunc, verbose bool, numCtx, numParallel, nGPULayers int, cacheTypeK, cacheTypeV, dtype string) (Engine, error)
LoadEngineWithProgress is like LoadEngine but accepts a progress callback for SafeTensors → GGUF conversion. When verbose is true, llama-server output is printed to stderr.
func LoadEngineWithSpeculativeProgress ¶ added in v0.9.37
func LoadEngineWithSpeculativeProgress(modelDir string, lm *model.LocalModel, progress ConvertProgressFunc, verbose bool, numCtx, numParallel, nGPULayers int, cacheTypeK, cacheTypeV, dtype string, speculative SpeculativeConfig) (Engine, error)
func NewOpenAICompatibleEngineWithHeaders ¶ added in v0.9.43
func NewOpenAICompatibleEngineWithHeaders(baseURL, modelName, token string, headers []ForwardHeader) Engine
func NewOpenAIEngine ¶
type ForwardHeader ¶ added in v0.9.43
type Message ¶
type Message struct {
Role string `json:"role"`
Content interface{} `json:"content"`
ReasoningContent string `json:"reasoning_content,omitempty"`
}
Message represents a chat message. Content can be a string for text-only, or an array of content parts for multimodal (e.g., image + text) messages.
type NativeToolStreamer ¶ added in v0.9.22
type NativeToolStreamer interface {
SupportsNativeToolStreaming() bool
}
NativeToolStreamer marks engines whose backend natively emits OpenAI-compatible streaming tool-call deltas, so tool requests can be proxied with stream enabled instead of being aggregated and normalized locally after the full completion.
type Options ¶
type Options struct {
Temperature float64
TopP float64
TopK int
MaxTokens int
Seed int
NumCtx int
Stop []string
// DisableThinking forces routing-style requests to skip provider thinking
// modes (Qwen enable_thinking=false, GLM/Kimi/DeepSeek thinking.type=disabled).
DisableThinking bool
}
Options controls generation parameters.
func DefaultOptions ¶
func DefaultOptions() Options
DefaultOptions returns sensible defaults. MaxTokens follows Ollama and llama.cpp semantics: -1 means no explicit generation cap.
type Session ¶
type Session struct {
// contains filtered or unexported fields
}
Session holds a conversation with context.
func NewSession ¶
NewSession creates a new chat session with the given engine.
func (*Session) SetSystemPrompt ¶
SetSystemPrompt sets or replaces the system prompt.
type SpeculativeConfig ¶ added in v0.9.37
type SpeculativeConfig struct {
Types []string
DraftModel string
DraftNMax int
DraftNMin int
DraftPMin *float64
}
func NormalizeSpeculativeConfig ¶ added in v0.9.37
func NormalizeSpeculativeConfig(config SpeculativeConfig) (SpeculativeConfig, error)
func (SpeculativeConfig) Enabled ¶ added in v0.9.37
func (config SpeculativeConfig) Enabled() bool
func (SpeculativeConfig) Key ¶ added in v0.9.37
func (config SpeculativeConfig) Key() string
type TokenCallback ¶
type TokenCallback func(token string)
TokenCallback is called for each generated token during streaming.