runtime

package
v0.14.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 11, 2026 License: MIT Imports: 32 Imported by: 0

Documentation

Overview

Package runtime is the playground's in-app side (WEFT-PLAYGROUND.md §10.2, [D6]): it registers the agents your code built with the Studio the app is already observed by, receives experiment commands over the runtime link, and executes them as real runs of those agents — your tools, your model keys, your process. Studio never runs your agent; it asks, this package answers.

The whole integration is one deferred call:

defer runtime.Install(
    runtime.Studio(url, token),            // or runtime.Local(srv) in setup A
    runtime.Agents(support, billing),      // the agents a runtime exposes
    runtime.Models(map[string]core.Model{  // allowed alternates, by display name
        "glm-5.3-flash": glmFlash,
    }),
    runtime.Limits(runtime.Budget{MaxTokensPerExperiment: 200_000, MaxRunsPerExperiment: 60}),
    runtime.AllowSideEffects("send_email"), // real only under side_effects "allow"
    runtime.Threads(store),                // thread.Storage; nil = ephemeral only
)()

Safety (WEFT-PLAYGROUND.md §6, non-negotiable)

Without Install no link opens, and even then only when WEFT_ENV=dev or Enabled(true) is set — production binaries expose nothing unless they opt in explicitly. Commands can only narrow: tools the agent registered can be turned off, never added (OnlyTools); models come from the Models allow-list or the app's own ModelResolver; MaxSteps and Parallelism only lower; park_on only adds parking; sampling params and the tool choice are neutral (a named choice must name a tool the command keeps on and does not park). Side effects never re-fire silently (§6 rule 3): a tool counts as "never" unless its code vouched core.Replay(core.ReplaySafe) — a vouched tool runs in every mode. A tool the runtime opted in with AllowSideEffects runs for real only when the command asks for side_effects "allow"; in the other modes it is a side effect like any other. In side_effects "substitute" (the default), a parked call that matches a recorded call of the source (same tool, same arguments as JSON; repeated calls in the order the source made them) is answered with the recorded result — the runtime acts as ADR 0007's resolver, the handler never runs; a miss stays parked for a human (the panel's continue / skip / resolve). "park" answers nothing from the record: every such call waits at the boundary. "allow" runs the opted-in tools for real, and is refused unless every tool the command leaves on is opted in or vouched ReplaySafe. A parked run resumes once each of its parked calls has a decision, under all of them; a decision naming a call that is not parked is rejected. The debugger's breakpoints (§8.3) park their tools on every run this package starts, and steer (§8.4) delivers into a run it holds: an ephemeral run's steering queue, or a fork turn's session as a thread steer under that turn's run options — a steer thread defers to a follow-up turn keeps the park rule; the app's own turns are never breakable or steerable from here (D7, PQ7).

The rule is default-deny (core.ParkAllExcept): a run lists the tools that may execute — the vouched-safe ones, an Output agent's submit_output, and under "allow" the opted-in names — and every other call parks, matched by name against each step's own tool set. So a tool that reaches the run only through core.ToolSource parks like any unannotated tool, and the rule follows a core.Subagent delegation into the child run: the child's own unvouched tools park there. One limit: a child's parked call is not the panel's to decide — the delegating call reads SUBAGENT_PENDING (ADR 0014) and the parent run carries on, the side effect never having fired. Names are matched in parent and child alike, so opt a name in only if every tool of that name down the delegation may run.

Everything a command carries is re-validated here, whatever Studio checked: unknown agents, tools, models, modes and options are rejected, never defaulted; the source run id, from_step and the transcript edits are checked against the transcript this runtime resolved itself. The link holds bounded state — at most 256 commands admitted and 16 runs executing at once, the newest 4096 command ids for at-most-once, 128 parked runs, 64 forks — and stopping it cancels the runs it started.

The engines: "live" runs the agent's own model (or a Models alternate); "scripted" (§5.5) answers each model call with the source run's recorded turn at zero tokens — its own core.Model over the messages records, keyed like wefttest's fixtures and missing loudly ("no recorded turn") when the input changed, never silently answering a prompt experiment. Thread modes: "ephemeral" (nothing written to thread storage) and "fork" (§5.4) — the source session opens read-side, Fork copies it to a new session with lineage (thread mints the run ids, stamps weft.session.forked_from), and the command's input becomes the fork's next turn under the same shaping; the source's approval grants are revoked in the fork (the app user's standing consent does not decide an experiment's parked call); a later fork command naming the fork's latest turn continues it in place, one naming an earlier turn forks from that turn. A fork's parked call is the fork session's own approval boundary: a decision is recorded in the fork and resumes it as the fork's next turn (the record is never consulted in a fork — every such call parks).

Runs this package starts are ordinary weft runs: they flow through the weft/otel pipeline to every destination, carrying weft.playground = true, weft.playground.command, weft.experiment.id and weft.forked_from, and never weft.session.id in ephemeral mode (an experiment is not a turn of the session). Per-experiment budget caps (Limits) are counted from each command's own usage; a breach rejects the next command of that experiment and never touches the app's own runs.

Where it sits

A module of its own because it imports thread (fork mode, transcript reads) and carries an executor, a registry and budget accounting — none of which belongs in the exporter-wiring module weft/otel (review §1.8). Arrows point down only: this package imports root, thread, obsdb, otel and studio (studio for *studio.Server alone, in runtime.Local); nothing imports it back.

Index

Examples

Constants

This section is empty.

Variables

This section is empty.

Functions

func Install

func Install(opts ...Option) (shutdown func())

Install registers the agents with a Studio and starts the runtime link: it dials out (never listens), so it works behind NAT. It never panics and never fails the program — a runtime that cannot be built (nothing enabled, no endpoint, no named agents) logs a WARN and opens nothing. The returned function stops the link — the command stream ends, the runs the link started are canceled, and it waits (bounded) for both; call it on exit, like otel.Install's. Calling it again is harmless.

The link registers on connect and after every reconnect, receives commands over an SSE stream (resuming with Last-Event-ID), acks every command before executing it (at-most-once: a repeated command id is ignored), executes it as a run of the named agent, and reports the run's end with a finished ack. The run's content flows back through the normal OTel pipeline, not through the link.

Example

ExampleInstall is the whole integration: one deferred call in main. Enabled(false) keeps the example inert — in a real dev binary the link opens because WEFT_ENV=dev (or Enabled(true)).

package main

import (
	"fmt"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/runtime"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	support := weft.New(
		wefttest.Script(wefttest.Say("Your order shipped yesterday.")),
		weft.Name("acme-support"),
	)
	glmFlash := wefttest.Script(wefttest.Say("…"))
	shutdown := runtime.Install(
		runtime.Studio("http://127.0.0.1:7331", ""),
		runtime.Agents(support),
		runtime.Models(map[string]weft.Model{"glm-5.3-flash": glmFlash}),
		runtime.Limits(runtime.Budget{MaxTokensPerExperiment: 200_000, MaxRunsPerExperiment: 60}),
		runtime.AllowSideEffects("send_email"), // for real only under side_effects "allow"
		runtime.Enabled(false),                 // dev-only by default: WEFT_ENV=dev
	)
	defer shutdown()
	fmt.Println("link:", "closed")
}
Output:
link: closed

Types

type Budget

type Budget struct {
	MaxTokensPerExperiment int64 // input + output tokens
	MaxRunsPerExperiment   int64
}

Budget caps one experiment, counted per experiment_id from the usage of the commands this runtime ran (§6 rule 6). Zero fields are no caps. A breach rejects the *next* command of that experiment (budget_exceeded) — never a run in flight, never the app's own runs.

type Option

type Option func(*config)

Option configures Install.

func Agents

func Agents(agents ...*core.Agent) Option

Agents registers the agents this runtime exposes. Only registered agents are playable (WEFT-PLAYGROUND.md §5.2). An agent without a name (core.Name) cannot register — its manifest would be nameless — and is skipped with a WARN.

func AllowSideEffects

func AllowSideEffects(tools ...string) Option

AllowSideEffects opts tools in to really running in playground runs whose command asks for side_effects "allow" (WEFT-PLAYGROUND §5.1, §6 rule 3): under "allow" a tool named here is never parked and never substituted — its handler executes whenever the experiment's model calls it. In "substitute" (the default) and "park" an opted-in tool is a side effect like any other: a call matching one the source run recorded is answered with the recorded result, any other call parks at the approval boundary ("park" answers nothing from the record). A tool marked core.Replay(core.ReplaySafe) is not a side effect and runs in every mode, opted in or not. A command that asks for "allow" is refused unless every tool it leaves on is named here or vouched ReplaySafe — "allow" runs this list for real, it does not widen it.

Name a tool here only when re-running it is harmless. Names match wherever the run reaches: a tool a core.ToolSource supplies under that name, and a tool of that name in a core.Subagent child (the child is under the same rule — its other tools park).

func Enabled

func Enabled(on bool) Option

Enabled opens (or forbids) the link explicitly. The default — no Enabled option — is on only when WEFT_ENV=dev: production binaries expose the runtime link only by explicit opt-in (§6 rule 1).

func Limits

func Limits(b Budget) Option

Limits sets this runtime's per-experiment caps. Without it, a runtime accepts unlimited playground runs within the agent's own budgets (MaxSteps, UsageLimit).

func Local

func Local(srv *studio.Server) Option

Local takes the embedded Studio server (setup A): the link talks to srv's handler in-process, no socket. srv is the studio.New(...) server the app mounts; nil is ignored. In-process is not exempt from srv's own token gate: a server built with studio.Token needs the token here as well — add Studio("", token).

func ModelResolver added in v0.13.0

func ModelResolver(resolve func(ctx context.Context, name string) (core.Model, error)) Option

ModelResolver lets a command name a model the runtime did not list in Models — "try this on claude-haiku-4-5" without pre-registering every model. A command's model override that is neither the agent's own model name nor a Models name is handed to resolve while the command is validated, before its ack: a model back runs the command (the run's RunStart.Model and weft.override.model name it); an error rejects the command with "model <name>: <err.Error()>".

resolve is the app's code, so the playground still only narrows: Studio proposes a name, the app decides whether it exists — build the client from the app's own credentials, refuse names it does not support. The error text is shown to whoever ran the experiment, so it must carry no key, URL or secret; the app controls the message. resolve may be called concurrently and once per command; cache clients if building one is costly. Each call is bounded: its ctx is canceled after 10 s and the command is rejected with "model <name>: resolver timed out" (the command runs before its ack, and an ack Studio waits on too long is marked lost); a resolver that ignores ctx is abandoned at the bound and its late result dropped — honour ctx, or the abandoned call keeps running. Registration reports the flag (resolver: true), and Studio accepts an unlisted name only then. A nil resolve is ignored.

Example

ExampleModelResolver lets the playground try a model the app never listed in runtime.Models: a command's provider-qualified name reaches the resolver before its ack, and the app decides whether it exists — it builds the client from its own credentials (here a scripted model stands in for anthropic.New) and refuses everything else. The error text is the rejected command's reason, so it names no key or URL.

package main

import (
	"context"
	"errors"
	"fmt"
	"strings"

	"github.com/weftgo/weft"
	"github.com/weftgo/weft/runtime"
	"github.com/weftgo/weft/wefttest"
)

func main() {
	support := weft.New(wefttest.Script(wefttest.Say("ok")), weft.Name("acme-support"))
	resolve := func(ctx context.Context, name string) (weft.Model, error) {
		provider, model, ok := strings.Cut(name, "/")
		if !ok || provider != "anthropic" || !strings.HasPrefix(model, "claude-") {
			return nil, errors.New("this app serves anthropic/claude-* models only")
		}
		return wefttest.Script(wefttest.Say("…")), nil // anthropic.New(anthropic.Model(model)) in a real app
	}
	shutdown := runtime.Install(
		runtime.Agents(support),
		runtime.ModelResolver(resolve),
		runtime.Enabled(false), // dev-only by default: WEFT_ENV=dev
	)
	defer shutdown()
	_, err := resolve(context.Background(), "openai/gpt-9")
	fmt.Println(err)
}
Output:
this app serves anthropic/claude-* models only

func Models

func Models(models map[string]core.Model) Option

Models declares the model alternates a command may switch to, by display name — the allow-list the playground's model picker reads and the resolver that turns a command's model string back into a core.Model (§5.2). A name missing here is refused: the playground cannot add a model, only choose among the ones the code registered.

func Studio

func Studio(url, token string) Option

Studio dials the Studio at url with token (setups B and C: the local binary or the hosted one). Empty token is the dev mode of setup B. Without a Studio or Local option, Install falls back to the pipeline's own Studio destination — otel.StudioEndpoint() — then, when that is empty, to the discovery file a running `weft studio` / `weft dev` wrote (unless WEFT_STUDIO_URL is set or WEFT_DISCOVERY=off; a stale file is never trusted), and opens nothing when there is none.

func Threads

func Threads(store thread.Storage) Option

Threads sets the thread.Storage transcript reads resolve through first (WEFT-PLAYGROUND.md §10.3): readers never lock, so an open session is readable. Without it the runtime is ephemeral-only and resolves transcripts from its local obsdb or from Studio. nil is ignored.

Directories

Path Synopsis
examples
local command
Command local is setup A's playground in one process: an embedded Studio on loopback, the app's own pipeline exporting into it, one agent with a scripted model and two tools, and the runtime link — WEFT-PLAYGROUND.md §7's P0 slice, drivable with curl.
Command local is setup A's playground in one process: an embedded Studio on loopback, the app's own pipeline exporting into it, one agent with a scripted model and two tools, and the runtime link — WEFT-PLAYGROUND.md §7's P0 slice, drivable with curl.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL