compaction

package
v0.7.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 28, 2026 License: MIT Imports: 6 Imported by: 0

Documentation

Overview

Package compaction keeps a long session's history within a model's context budget using the tiered strategy every production agent loop we looked at converges on (opencode, Crush): cheap structural pruning first, full LLM summarization only as a last resort.

Two known failure modes shaped this design and are worth naming explicitly, since both opencode (github.com/sst/opencode issues #27924, #15533) and Crush (charmbracelet/crush issue #2551) hit them in production: a compaction summary that reads like an instruction gets re-executed by the model as a new task, which promptly overflows again and re-triggers compaction — an infinite loop. Compact() guards against this by wrapping the summary in an unambiguous "reference only" framing and storing it as a system message rather than a user turn; the agent loop additionally caps consecutive compactions within a single Run.

Index

Constants

View Source
const CompactionMarkerName = "kram:compaction_summary"

CompactionMarkerName tags a system message as a compaction summary rather than a normal turn — see store.Message.Name.

View Source
const (

	// DefaultMaxTokens is the effective-history budget before compaction
	// kicks in. Deliberately conservative — most models have far more
	// headroom, but a smaller budget means summaries happen sooner and
	// cheaper rather than in one enormous, expensive pass.
	DefaultMaxTokens = 60_000
)

Variables

This section is empty.

Functions

func Compact

func Compact(ctx context.Context, gw *gatewayclient.Client, model string, effective []store.Message) (store.Message, error)

Compact summarizes the effective history via one gateway call and returns a new compaction-marker message ready to append to the store. It does not append it itself — the caller decides that, so it can be done atomically alongside whatever triggered compaction.

func EffectiveHistory

func EffectiveHistory(all []store.Message) []store.Message

EffectiveHistory returns what the model should actually see: everything from the most recent compaction marker onward (the marker itself becomes the lead system message), or the full history if compaction has never run for this session.

func EmergencyPrune added in v0.6.0

func EmergencyPrune(effective []store.Message, maxTokens int) []store.Message

EmergencyPrune is the fallback for when Compact's own summary call fails — the model being unreachable is precisely when compaction tends to be needed most, and failing the whole turn over it turned a transient summarizer error into a dead session. It keeps the newest whole user-turns that fit within maxTokens (plus the leading compaction-summary marker when present), dropping older turns entirely. Cutting only at user-message boundaries keeps every assistant tool_calls message next to its tool results — a split pair is a protocol error several providers hard-reject. Never returns fewer messages than the most recent user turn, even if that alone still exceeds the budget: sending an oversized last turn at least lets the provider say so, where sending nothing guarantees failure.

func EstimateTokens

func EstimateTokens(msgs []store.Message) int

EstimateTokens is a rough size estimate for a slice of messages.

func NeedsCompaction

func NeedsCompaction(effective []store.Message, maxTokens int) bool

NeedsCompaction reports whether the effective history is over budget.

func PruneForModel

func PruneForModel(effective []store.Message) []store.Message

PruneForModel is the cheap, structural first pass: old tool-result content past the protected tail is replaced with a placeholder. This never touches the database — it's applied only to the in-memory copy sent to the model, so the CLI/audit trail always shows the real result.

Types

This section is empty.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL