agent

package
v0.3.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 24, 2026 License: MIT Imports: 20 Imported by: 0

Documentation

Overview

Package agent is Kram's tool-calling loop: the piece that actually makes the daemon useful rather than a plain chat relay. Each turn it sends the session's history (and tool definitions) to the gateway; if the model asks to call tools, they run and their results feed back in, looping until the model answers in plain text, reaches a real stagnation guard, or exhausts the segmented emergency budget.

Design choices are grounded in patterns that recur across production agent loops (opencode/Crush, Hermes Agent — see the research notes in this session): tool execution waits for a complete (non-streaming) model response rather than interleaving with token streaming — every project that documents this decouples the two; the iteration budget has a soft landing (a warning, then one forced "final answer" call) rather than a hard cutoff, matching Hermes's approach; and context compaction is capped at a handful of attempts per run to guard against the "model re-executes its own summary" infinite loop documented in both opencode and Crush (see internal/daemon/compaction).

Index

Constants

This section is empty.

Variables

View Source
var ErrContextOverflow = errors.New("context kept overflowing after repeated compaction attempts")

ErrContextOverflow is returned when compaction couldn't bring the conversation back under budget within MaxCompactionsPerRun attempts — a deliberate hard failure instead of an infinite compact/retry loop.

View Source
var ErrNotFound = errors.New("session not found")

ErrNotFound is returned when a session ID doesn't exist.

Functions

This section is empty.

Types

type Config

type Config struct {
	// Model selects which gateway combo this session's calls go to.
	Model string
	// MaxTurns is the number of model calls in one automatic continuation
	// segment (tool round-trips included). Reaching it no longer ends a real
	// task by itself; MaxSegmentsPerRun controls the emergency total.
	MaxTurns int
	// MaxSegmentsPerRun lets a productive long task continue automatically
	// across MaxTurns-sized segments. Default 4 (50 * 4 = 200 calls). Segment
	// boundaries are ephemeral structured UI checkpoints, while the final
	// boundary still gets the existing tool-free soft landing.
	MaxSegmentsPerRun int
	// MaxCompactionsPerRun caps consecutive compaction attempts within a
	// single Run before giving up with ErrContextOverflow.
	MaxCompactionsPerRun int
	// MaxContextTokens is the effective-history budget before compaction
	// triggers (see internal/daemon/compaction).
	MaxContextTokens int
	// Workspace is the project root — used to load AGENTS.md/CLAUDE.md as
	// persistent project context, injected into every turn.
	Workspace string
	// MaxGatewayRounds bounds how many times callModelWithRetry retries a
	// whole gateway call (a fresh ranked-candidate pass) after a
	// retryable GatewayError, before giving up — see retry.go. Runs
	// entirely inside one iteration of runLoop's turn loop, so retrying
	// never consumes MaxTurns: no new logical decision by the model
	// happened, just another attempt at the same one. Default 3.
	MaxGatewayRounds int
	// PreferStreaming opts a session back into the streaming gateway
	// call path (see streamCall) instead of the buffered default (see
	// bufferedCall). Streaming commits to one candidate the moment
	// router.BoundedPeek sees a meaningful first signal — if that
	// candidate then fails mid-stream, the whole turn fails with it,
	// since HTTP headers are already sent and no further fallback is
	// possible. The buffered path doesn't have that problem: kram-
	// gateway's own non-streaming branch already tries every ranked
	// candidate to completion before writing anything back (see
	// internal/server/chat.go), so it's the default. False (buffered) is
	// what almost every caller wants; this exists as an escape hatch,
	// not a recommendation.
	PreferStreaming bool
	// ToolOrder curates the generated Tools overview's presentation
	// order (see compileToolsOverview) — it never changes which tools
	// are offered, only where each one is listed. nil means today's
	// plain alphabetical order. When set, it must contain
	// tools.ToolOrderRest exactly once, marking where every unlisted
	// tool is inserted (still alphabetical); New validates this,
	// including that every named tool actually exists in the registry,
	// and fails loudly rather than silently dropping a typo.
	ToolOrder []string
	// SystemPromptOverride, when non-empty, replaces the "base"
	// PromptPart's content (identity/workflow/skills/memory/delegation/
	// asking/writing-code/output/safety — see systemPrompt) with this
	// text instead of Kram's own. Every other preamble part — the
	// generated tools overview, background-job guidance, project
	// context, memory — is unaffected and still assembled normally
	// around it; this deliberately can't suppress those, since the
	// tools overview exists specifically so a tool can't go silently
	// unmentioned (see the VisibleTools() fix), and an override that
	// could also drop that would reopen the exact bug that fix closed.
	// Empty (the default) means today's systemPrompt(workspace) output,
	// unchanged. Sourcing this from a file or a CLI flag is left to the
	// caller — this field takes the resolved text directly, matching
	// how ToolOrder is a resolved value too, not a file path.
	SystemPromptOverride string
}

Config tunes the loop's limits. All fields have sane defaults applied by New if left zero.

type ContextCategory

type ContextCategory struct {
	Name   string `json:"name"`
	Tokens int    `json:"tokens"`
}

ContextCategory is one real, distinct contributor to a session's context window usage — not a placeholder. MCP and skill discovery are exposed through registered tool definitions, so their cost is included in the tool category alongside conversation content and compaction summaries.

type ContextUsage

type ContextUsage struct {
	Budget     int               `json:"budget"`
	Used       int               `json:"used"`
	Free       int               `json:"free"`
	Categories []ContextCategory `json:"categories"`
}

ContextUsage is a session's current context-window breakdown, estimated the same way internal/daemon/compaction decides when to compact — same chars/4 approximation, so this panel and the compaction trigger never disagree with each other.

type Event

type Event struct {
	Kind       EventKind
	Content    string   // EventDelta
	Reasoning  string   // EventReasoning
	ToolName   string   // EventToolStart, EventToolResult
	ToolArgs   string   // EventToolStart
	ToolResult string   // EventToolResult
	ToolOK     bool     // EventToolResult
	ProcessID  string   // EventToolResult when run_background started a process
	Notice     string   // EventNotice
	QuestionID string   // EventQuestion
	Question   string   // EventQuestion
	Options    []string // EventQuestion

	ApprovalID      string // EventApproval
	ApprovalTool    string // EventApproval
	ApprovalSubject string // EventApproval

	RouteCall *RouteCall // EventRouteDone
	Segment   int        // EventSegment, one-based
	Segments  int        // EventSegment, configured maximum
}

Event is one thing that happened during Run, emitted live via the OnEvent callback so a caller (the daemon's HTTP layer, ultimately the CLI) can show it as it happens instead of only after the whole turn completes.

type EventFunc

type EventFunc func(Event)

EventFunc receives live events during Run. A nil EventFunc is valid — Run behaves the same either way, just without the live callback.

type EventKind

type EventKind string

EventKind identifies what an Event carries.

const (
	// EventDelta is a text fragment as the model generates it — the
	// live-streaming case. Only fires for genuine content; tool-call
	// turns produce no deltas (the model isn't "saying" anything, it's
	// deciding what to call).
	EventDelta EventKind = "delta"
	// EventReasoning is a fragment of a reasoning-capable model's chain-
	// of-thought, mirroring gatewayclient.StreamDelta.Reasoning — kept as
	// a distinct kind from EventDelta, never folded into Content, so a
	// caller can't accidentally treat it as (or concatenate it into) the
	// model's actual answer. Best-effort: most providers in Kram's
	// fallback chain never send it, and only streamCall (reached when
	// Config.PreferStreaming opts a session in) can observe it at all —
	// bufferedCall's single non-streaming gateway call has no
	// per-fragment signal of any kind to relay.
	EventReasoning EventKind = "reasoning"
	// EventToolStart fires right before a tool call executes.
	EventToolStart EventKind = "tool_start"
	// EventToolResult fires right after, success or failure either way.
	EventToolResult EventKind = "tool_result"
	// EventNotice carries an out-of-band notice (image capability
	// fallback, a compaction just happened) as soon as it's known, rather
	// than only at the very end.
	EventNotice EventKind = "notice"
	// EventQuestion fires when the ask_question tool pauses the turn —
	// the caller shows Question/Options and is expected to answer via
	// Service.AnswerQuestion(QuestionID, ...) on a separate call, since the
	// turn (and this SSE stream) stays blocked until it does.
	EventQuestion EventKind = "question"
	// EventApproval fires when the permission policy marks a tool call
	// Ask — distinct from EventQuestion: the model isn't uncertain here,
	// it knows exactly what it wants to do, but policy requires the
	// user's sign-off first. The caller shows ApprovalTool/ApprovalSubject
	// and is expected to answer via Service.AnswerApproval(ApprovalID,
	// "once"|"always"|"deny") on a separate call, same blocking shape as
	// EventQuestion.
	EventApproval EventKind = "approval"
	// EventRouteStart fires right before a model call goes out — the
	// earliest point a live route bar can show "routing…" instead of
	// staying blank until the whole call finishes. Real per-attempt
	// progress (which candidate is being tried right now, mid-fallback)
	// isn't observable here: the gateway's fallback loop happens inside
	// one HTTP round-trip the daemon only sees the result of, so the next
	// signal is EventRouteDone — see DECISIONS.md, "Live route events are
	// per model call, not per attempt."
	EventRouteStart EventKind = "route_start"
	// EventRouteDone fires once a model call's full routing story is
	// known: every attempt made, the winner, and (for a scoring strategy)
	// the full candidate ranking — see RouteCall.
	EventRouteDone EventKind = "route_done"
	// EventHeartbeat is a periodic, payload-free liveness signal emitted
	// while a buffered (non-streaming) gateway call is in flight — see
	// Service.bufferedCall. No per-candidate progress is observable on
	// this path (same reason EventRouteStart's doc comment gives: the
	// gateway's fallback loop happens inside one HTTP round-trip the
	// daemon only sees the result of), so this exists purely to keep a
	// caller's own "still working" liveness clock fresh during a
	// multi-candidate or multi-round wait that can legitimately run
	// longer than a single provider call would. Deliberately a distinct
	// kind from EventNotice, which renders as a visible transcript line —
	// a heartbeat firing every few seconds would spam that; a caller with
	// no special handling for this kind is expected to silently ignore
	// it (which, for anything that treats "any event" as a liveness
	// signal, is already exactly the desired behavior).
	EventHeartbeat EventKind = "heartbeat"
	// EventSegment marks an automatic continuation budget boundary. Unlike a
	// notice it is ephemeral operational state: the TUI folds it into the live
	// activity line instead of appending transcript noise below the animation.
	EventSegment EventKind = "segment"
)

type PromptPart

type PromptPart struct {
	ID        string
	Placement PromptPlacement
	Refresh   RefreshPolicy
	Source    string
	Content   string
}

PromptPart is one named, addressable unit of what the model is told — the first piece of Kram's Instruction IR. Source names where the content came from (builtin/AGENTS.md/memory/runtime), for a future prompt-inspection view — not consumed by anything yet in v1.

type PromptPlacement

type PromptPlacement int

PromptPlacement is where a part lands relative to conversation history — kept as real data on the part, not just encoded in which compiler function produced it, so a future inspector (or a later phase adding more post-history reminders) has one place to look.

const (
	PlacementPreamble    PromptPlacement = iota // before history
	PlacementPostHistory                        // after history
)

type RefreshPolicy

type RefreshPolicy int

RefreshPolicy documents which of the three real cadences a part follows — not three arbitrary "stability tiers", but the three that actually exist in runLoop today: a value fixed for the Service's lifetime, one fixed for a single run (frozen once per user turn, for provider prefix-cache reasons — see recentMemoryMessage's doc comment), or one re-evaluated on every internal tool-loop iteration within that run. Nothing conditions behavior on this in v1; it exists so the already-real distinction between systemPrompt (Static), AGENTS.md (Iteration — read fresh every loop pass), and memory (Run — frozen once per turn) is inspectable instead of implicit in append order, and so a future Model/Agent Profile phase has a real vocabulary to extend instead of starting from scratch.

const (
	RefreshStatic    RefreshPolicy = iota // fixed for the Service's lifetime
	RefreshRun                            // fixed once per runLoop call (per user turn)
	RefreshIteration                      // re-evaluated on every internal tool-loop pass
)

type RouteCall

type RouteCall struct {
	Index    int                         `json:"index"` // 1-indexed position within the run
	Combo    string                      `json:"combo"`
	Strategy string                      `json:"strategy"`
	Attempts []openai.AttemptInfo        `json:"attempts"`
	Ranking  []openai.RankedProviderInfo `json:"ranking,omitempty"`
	Winner   string                      `json:"winner,omitempty"`
}

RouteCall is one model call's full routing story — every attempt made while serving it (success, error, or gate-rejected), not just the one that eventually won. Ranking (when the combo's strategy scores) is the gateway's own full candidate ranking, unchanged — nothing here is recomputed, only carried through.

type RouteTrace

type RouteTrace struct {
	Combo    string      `json:"combo"`
	Strategy string      `json:"strategy"`
	Calls    []RouteCall `json:"calls"`
}

RouteTrace accumulates every model call's routing story across one user turn (one Service.Run). This exists specifically to fix a real bug: the loop used to do `result.Attempts = callResult.Attempts` on every iteration, so a turn with several tool round-trips silently lost every model call's fallback trail except the very last one — a multi-step task like "read a file, grep, edit, run tests, answer" would only ever show the routing story for "answer," discarding four other model calls' worth of real attempts. RunResult.Attempts is kept as-is (the last call's trail, for callers that only want the immediate footer view); RouteTrace is the new, complete picture.

type RunResult

type RunResult struct {
	Message      store.Message
	ToolActivity []ToolActivity
	Attempts     []openai.AttemptInfo // fallback trail of the final (deciding) gateway call — kept for the simple footer view
	// RouteTrace is the full picture Attempts alone can't show: every
	// model call this run made (a turn can be several, across tool
	// round-trips), each with its own complete fallback trail — see
	// route.go for why this exists and what bug it fixes.
	RouteTrace  RouteTrace
	Usage       openai.Usage // summed across every gateway call this turn
	Compactions int
	ImageNotice string // set if images were attached but the combo can't accept them
}

RunResult is everything a caller gets back from one user turn — which may have involved any number of tool round-trips and compactions under the hood.

type Service

type Service struct {
	// contains filtered or unexported fields
}

Service runs the agent loop for a workspace.

func New

func New(st *store.Store, gw *gatewayclient.Client, tr *tools.Registry, cfg Config) (*Service, error)

New builds an agent Service. New builds a Service, or fails if cfg.ToolOrder is malformed or names a tool tr doesn't actually have registered — the "fail loud instead of a typo silently vanishing" guarantee the tool-order feature exists to provide (see tools.ValidateToolOrder / tools.UnknownToolOrderNames).

func (*Service) AnswerApproval

func (s *Service) AnswerApproval(id, decision string) bool

AnswerApproval delivers the user's decision ("once", "always", or "deny") to the pending approval waiting on id, if any is still pending. Returns false if id is unknown (already answered, timed out, or never existed) or decision isn't one of the three valid values, so the caller (the daemon's HTTP handler) can report a clear error instead of silently no-opping.

func (*Service) AnswerQuestion

func (s *Service) AnswerQuestion(id, ans string) bool

AnswerQuestion delivers ans to the ask_question call waiting on id, if any is still pending. Returns false if id is unknown — already answered, timed out, or never existed — so the caller (the daemon's HTTP handler) can report a clear 404 instead of silently no-opping.

func (*Service) BackgroundProcessOutput added in v0.2.8

func (s *Service) BackgroundProcessOutput(id string, cursor *int64) (tools.BackgroundProcessOutput, bool)

func (*Service) BackgroundProcesses added in v0.2.8

func (s *Service) BackgroundProcesses() []tools.BackgroundProcessInfo

BackgroundProcesses and BackgroundProcessOutput are read-only pass-throughs for the local TUI. Keeping the server dependent on Service, rather than on a second Registry reference, preserves the daemon's construction boundary.

func (*Service) ContextUsage

func (s *Service) ContextUsage(ctx context.Context, sessionID string) (ContextUsage, error)

ContextUsage reports what's actually consuming this session's context budget right now.

func (*Service) ReplaceDisabledTools added in v0.2.5

func (s *Service) ReplaceDisabledTools(names []string)

ReplaceDisabledTools applies a persisted tools/skills profile to this live daemon so the first post-setup session cannot observe startup-time settings.

func (*Service) Run

func (s *Service) Run(ctx context.Context, sessionID, userContent string, images []string, onEvent EventFunc) (RunResult, error)

Run handles one user message end to end: persist it, run the tool loop until the model produces a final text answer (or the budget runs out), and return that answer plus everything that happened along the way. onEvent (may be nil) receives a live play-by-play — text deltas as the model generates them, tool start/result, and notices — as they happen rather than only in the returned RunResult once everything is done.

func (*Service) RunTask

func (s *Service) RunTask(ctx context.Context, goal, taskContext, model string, depth int) (string, error)

RunTask implements tools.Delegator: runs goal (plus optional context) to completion in a brand-new session, isolated from every other conversation — the subagent sees only what's passed here, not the parent's history, matching Hermes Agent's "spawn a junior engineer" model rather than a shared-context delegation. model, if empty, falls back to the parent's own combo. depth is the nesting level the *child* will run at (the caller — delegate_task — passes its own depth+1); it's threaded through runLoop so a grandchild's own delegate_task call sees the right depth and gets blocked once maxSpawnDepth is reached.

func (*Service) Skills

func (s *Service) Skills() []tools.Skill

Skills passes through the registry's discovered-skills listing for the same endpoint.

func (*Service) Tools

func (s *Service) Tools() []tools.ToolInfo

Tools passes through the registry's full tool listing (enabled or not) for the daemon's GET /tools endpoint.

type ToolActivity

type ToolActivity struct {
	Name      string `json:"name"`
	Args      string `json:"args"`
	Result    string `json:"result"`
	OK        bool   `json:"ok"`
	ProcessID string `json:"process_id,omitempty"`
}

ToolActivity records one tool call the loop made, for callers (the CLI) that want to show what the agent actually did, not just its final answer.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL