Documentation
¶
Overview ¶
Package compaction keeps a long session's history within a model's context budget using the tiered strategy every production agent loop we looked at converges on (opencode, Crush): cheap structural pruning first, full LLM summarization only as a last resort.
Two known failure modes shaped this design and are worth naming explicitly, since both opencode (github.com/sst/opencode issues #27924, #15533) and Crush (charmbracelet/crush issue #2551) hit them in production: a compaction summary that reads like an instruction gets re-executed by the model as a new task, which promptly overflows again and re-triggers compaction — an infinite loop. Compact() guards against this by wrapping the summary in an unambiguous "reference only" framing and storing it as a system message rather than a user turn; the agent loop additionally caps consecutive compactions within a single Run.
Index ¶
- Constants
- func Compact(ctx context.Context, gw *gatewayclient.Client, model string, ...) (store.Message, error)
- func EffectiveHistory(all []store.Message) []store.Message
- func EmergencyPrune(effective []store.Message, maxTokens int) []store.Message
- func EstimateTokens(msgs []store.Message) int
- func NeedsCompaction(effective []store.Message, maxTokens int) bool
- func PruneForModel(effective []store.Message) []store.Message
Constants ¶
const CompactionMarkerName = "kram:compaction_summary"
CompactionMarkerName tags a system message as a compaction summary rather than a normal turn — see store.Message.Name.
const ( // DefaultMaxTokens is the effective-history budget before compaction // kicks in. Deliberately conservative — most models have far more // headroom, but a smaller budget means summaries happen sooner and // cheaper rather than in one enormous, expensive pass. DefaultMaxTokens = 60_000 )
Variables ¶
This section is empty.
Functions ¶
func Compact ¶
func Compact(ctx context.Context, gw *gatewayclient.Client, model string, effective []store.Message) (store.Message, error)
Compact summarizes the effective history via one gateway call and returns a new compaction-marker message ready to append to the store. It does not append it itself — the caller decides that, so it can be done atomically alongside whatever triggered compaction.
func EffectiveHistory ¶
EffectiveHistory returns what the model should actually see: everything from the most recent compaction marker onward (the marker itself becomes the lead system message), or the full history if compaction has never run for this session.
func EmergencyPrune ¶ added in v0.6.0
EmergencyPrune is the fallback for when Compact's own summary call fails — the model being unreachable is precisely when compaction tends to be needed most, and failing the whole turn over it turned a transient summarizer error into a dead session. It keeps the newest whole user-turns that fit within maxTokens (plus the leading compaction-summary marker when present), dropping older turns entirely. Cutting only at user-message boundaries keeps every assistant tool_calls message next to its tool results — a split pair is a protocol error several providers hard-reject. Never returns fewer messages than the most recent user turn, even if that alone still exceeds the budget: sending an oversized last turn at least lets the provider say so, where sending nothing guarantees failure.
func EstimateTokens ¶
EstimateTokens is a rough size estimate for a slice of messages.
func NeedsCompaction ¶
NeedsCompaction reports whether the effective history is over budget.
func PruneForModel ¶
PruneForModel is the cheap, structural first pass: old tool-result content past the protected tail is replaced with a placeholder. This never touches the database — it's applied only to the in-memory copy sent to the model, so the CLI/audit trail always shows the real result.
Types ¶
This section is empty.