openai

package module
v0.3.8 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 29, 2026 License: MIT Imports: 19 Imported by: 0

Documentation

Overview

Package openai is the weft adapter for the OpenAI Chat Completions API and OpenAI-compatible servers (gateways, local models), wrapping the official openai-go SDK. weft never owns the HTTP: every request goes through the SDK's client, and the SDK's stream type is the contract.

Mapping highlights (ADR 0013 has the full tables):

  • Text deltas pass through as ModelTextDelta; tool-call argument fragments also surface live as ModelToolCallDelta progress while the assembled call is buffered by delta index and emitted whole as a ModelToolCall before ModelFinish — the core's uniform-streaming rent.
  • Stop reasons map to weft's three; anything unmapped keeps its raw value on ModelFinish.Raw.
  • A RoleTool message fans out to one OpenAI tool message per result; assistant ReasoningParts are dropped (Chat Completions has no reasoning input).
  • SequentialTools sets parallel_tool_calls: false.
  • ModelRequest.Thinking maps by dialect: reasoning_effort for the official API and unrecognized hosts; a thinking object injected into the request body (per-request SDK middleware) for the known gateway hosts — z.ai, bigmodel.cn, moonshot.ai/cn — whose param the typed params cannot carry. Dialect(...) overrides detection; Off has no effort form and sends nothing there (ADR 0013 amendment 9 has the full table).

Retry stance: transport retries (429, 5xx, net.Error) belong to the SDK via MaxRetries; the weft loop never retries a model call, and logic retries are model-seam middleware.

The Responses API (reasoning items, built-in tools) is not wrapped here; it would be an openai/responses sub-package.

Index

Examples

Constants

This section is empty.

Variables

This section is empty.

Functions

func Model

func Model(name string, opts ...Option) weft.Model

Model returns a weft.Model backed by the OpenAI Chat Completions API (or any compatible server, via BaseURL). A Model is immutable and safe for concurrent runs.

Example

Constructing a model performs no I/O; a key ($OPENAI_API_KEY) is read only when a run starts. Swap providers by swapping this one line.

package main

import (
	"context"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/openai"
)

func main() {
	m := openai.Model("gpt-4o-mini")
	echo := weft.Tool("echo", "Echo a message.", func(ctx context.Context, in struct {
		Message string `json:"message" jsonschema:"the text to repeat"`
	}) (string, error) {
		return in.Message, nil
	})
	agt := weft.New(m, echo)
	_ = agt // a run is agt.Generate(ctx, weft.Prompt("Echo: hello"))
}

Types

type Option

type Option interface {
	// contains filtered or unexported methods
}

Option configures the adapter at construction, the same functional style as the core. The zero configuration reads $OPENAI_API_KEY (and $OPENAI_BASE_URL); the SDK's transport default applies (2 retries on 429/5xx/connection errors).

func APIKey

func APIKey(k string) Option

APIKey sets the API key. Default: $OPENAI_API_KEY.

func BaseURL

func BaseURL(u string) Option

BaseURL points the adapter at an OpenAI-compatible server — gateways, local models. Also read from $OPENAI_BASE_URL when the option is not given. The provider name stays "openai": a base URL host is not a provider identity and would leak into telemetry cardinality.

func Client

func Client(c *openai.Client) Option

Client uses an already-configured SDK client (Azure endpoints, custom transports, test doubles); it overrides BaseURL, APIKey, and MaxRetries. The WEFT_MODEL_REQUESTS kill switch guards egress from clients the adapter builds from credentials; an injected client's destinations are the caller's responsibility — which is why it stays reachable under deny (ADR 0013's kill-switch clause).

func Dialect added in v0.2.0

func Dialect(d ThinkingDialect) Option

Dialect pins the thinking wire form instead of detecting it from the base URL. It matters for custom gateways behind hosts the adapter does not recognize and for clients whose endpoint it cannot inspect.

Example

The thinking wire form is detected from the base URL (reasoning_effort for the official API and unrecognized hosts, a thinking object for the known gateways); Dialect pins it when detection can't know — a custom gateway behind an unrecognized host, or a client whose endpoint the adapter cannot inspect.

package main

import (
	"github.com/weftgo/weft/openai"
)

func main() {
	m := openai.Model("kimi-k2.7",
		openai.BaseURL("https://gw.internal/v1"),
		openai.Dialect(openai.DialectObject))
	_ = m // runs now send thinking:{"type":...} per weft.Thinking
}

func ExtraBody added in v0.3.0

func ExtraBody(fields map[string]any) Option

ExtraBody adds fields to every request's JSON body — the generic valve for vendor knobs weft has no option for (the public version of the gateway thinking injection). Deep-merged into the body weft built: nested maps merge recursively, every other value replaces, and **your key wins on conflict** — the escape hatch is you taking responsibility for bytes weft did not choose, and the default-bytes tests do not cover what it sends. Construction-time only, and a snapshot: the values are deep-copied when the option applies, so mutating the map you passed afterwards never reaches the Model (safe for concurrent runs). It applies to the requests the adapter makes, including through an injected Client(c).

func ExtraHeaders added in v0.3.0

func ExtraHeaders(h http.Header) Option

ExtraHeaders adds HTTP headers to every request, verbatim. A header the SDK itself sets (Authorization, Content-Type) is yours not to clobber — the option does not check. Construction-time only, and a snapshot: the slices are copied when the option applies, so mutating the header values you passed afterwards never reaches the Model.

func IdleTimeout

func IdleTimeout(d time.Duration) Option

IdleTimeout is the maximum gap between two stream chunks before the call fails wrapping weft.ErrStreamIdle (default 60s; zero disables it). The wait for response headers is the first gap: it covers the SDK's transport retries and their sleeps too (a 429 whose retry-after the SDK honours sleeps inside it), so a retry sequence longer than the timeout fails ErrStreamIdle with the provider's error discarded — hand retries to mw.Retry with MaxRetries(0), or raise the timeout. The ctx deadline stays the hard limit on the whole call — a slow but actively streaming response is never killed.

func MaxRetries

func MaxRetries(n int) Option

MaxRetries sets the SDK's transport retry count (429/5xx/connection errors only; the SDK's own default is 2 when the option is absent). MaxRetries(0) switches the SDK's retries off — the pairing for mw.Retry, which then owns every retry and reads the provider's retry-after itself instead of the SDK sleeping on it under the idle timer (see IdleTimeout). A negative n is 0. The weft loop never retries a model call; logic retries are model-seam middleware (TODO §4.1).

func MaxTokens

func MaxTokens(n int) Option

MaxTokens caps a step's output tokens (max_completion_tokens). Zero keeps the provider default.

func Seed added in v0.3.0

func Seed(s int64) Option

Seed sets the sampling seed for deterministic-ish runs — a hint the provider treats best-effort, not a contract. Not sent unless the option is given; a per-request weft.RequestParams.Seed overrides it for one call.

func Stop added in v0.3.0

func Stop(seqs ...string) Option

Stop sets stop sequences the API stops on; not sent unless the option is given. A per-request weft.RequestParams.Stop overrides it for one call.

func Temperature

func Temperature(t float64) Option

Temperature sets the sampling temperature; it is not sent unless the option is given. A per-request weft.RequestParams.Temperature overrides it for one call.

func TopP added in v0.3.0

func TopP(p float64) Option

TopP sets nucleus sampling; it is not sent unless the option is given. A per-request weft.RequestParams.TopP overrides it for one call.

type ThinkingDialect added in v0.2.0

type ThinkingDialect int

ThinkingDialect selects how weft's ThinkingConfig reaches an OpenAI-compatible server: the official API takes reasoning_effort, while gateways such as z.ai and Moonshot take a thinking object the SDK's typed params have no field for. DialectAuto (the default) detects from the base URL; the explicit dialects override detection.

const (
	DialectAuto ThinkingDialect = iota
	// DialectEffort sends reasoning_effort ("low"|"medium"|"high") —
	// the official Chat Completions knob.
	DialectEffort
	// DialectObject injects thinking:{"type":"enabled"|"disabled"}
	// into the request body — the z.ai and Moonshot shape.
	DialectObject
	// DialectNone never sends thinking parameters, for strict
	// compatible servers that reject unknown fields.
	DialectNone
)

Directories

Path Synopsis
Command example runs a two-step agent conversation against the real OpenAI API (or any OPENAI_BASE_URL server).
Command example runs a two-step agent conversation against the real OpenAI API (or any OPENAI_BASE_URL server).

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL