Documentation
¶
Overview ¶
Package compact is the reference compaction Transform for agentturn: when a transcript grows past a token budget, the older part is folded into a shorter form and the recent part is kept verbatim.
Two folds are provided behind the same Transform. New uses the model's own compaction endpoint and splices the returned compaction item in; NewLocal asks an ordinary model call for a summary and splices that in as a message, for the many servers that do not implement compaction.
c := compact.New(model, compact.WithBudget(120_000))
cfg.Transform = c.Transform
l := compact.NewLocal(model, compact.WithBudget(120_000), compact.WithModel("gpt-5-mini"))
cfg.Transform = l.Transform
A Transform runs before every model call and shapes that call only; the loop's working transcript is never replaced. The transform therefore remembers what it compacted: as long as the transcript still begins with the prefix it folded, the next call reuses the result and only folds again when the kept part outgrows the budget. When it folds again, the previous result is folded with the new prefix, so what an earlier fold kept is summarised again rather than lost. A front that wants to persist the result registers WithOnFold, which reports every fold, successful or not, with the index at which the transcript was split; Transform.Last returns the latest summary. WithPin keeps chosen items of the folded prefix verbatim after the summary, for context a harness injected that must not become whatever the summary made of it. WithRequest edits the summary request NewLocal sends, so it can carry the reasoning setting the agent's own requests do.
Index ¶
- Constants
- Variables
- func Estimate(items openresponses.Items) int
- func PrefixHash(items agentturn.Transcript) string
- func SummaryMessage(summary string) openresponses.Item
- type Compactor
- type Fold
- type Option
- func WithBackOff(enabled bool) Option
- func WithBudget(tokens int) Option
- func WithEstimator(fn func(openresponses.Items) int) Option
- func WithFailedFold(split int, prefixHash string, tokens int) Option
- func WithFilter(fn func(agentturn.Transcript) agentturn.Transcript) Option
- func WithKeepLast(n int) Option
- func WithMinFold(tokens int) Option
- func WithModel(name string) Option
- func WithOnFold(fn func(context.Context, Fold) error) Option
- func WithPin(fn func(openresponses.Item) bool) Option
- func WithRequest(fn func(*openresponses.Request)) Option
- func WithSummaryItem(fn func(summary string) openresponses.Item) Option
- func WithSummaryPrompt(prompt string) Option
- type Transform
Constants ¶
const DefaultBudget = 100_000
DefaultBudget is the token budget when none is set.
const DefaultKeepLast = 4
DefaultKeepLast is the number of recent items never compacted when none is set.
const DefaultSummaryMaxOutputTokens = 8192
DefaultSummaryMaxOutputTokens caps the MaxOutputTokens NewLocal sets on a summary request: half the budget, and no more than this.
const DefaultSummaryPrompt = `` /* 426-byte string literal not displayed */
DefaultSummaryPrompt is the instruction NewLocal sends with the items to fold when WithSummaryPrompt is not given.
Variables ¶
var ErrSummaryIncomplete = errors.New("compact: summary response is incomplete")
ErrSummaryIncomplete is the error of a NewLocal fold whose summary response the server ended incomplete twice, such as one cut at MaxOutputTokens each time: a summary cut short has lost what it had not yet said, so it is never applied. It is wrapped with the server's reason when it gave one. Like ErrSummaryTooLarge, such a fold is reported to WithOnFold with it and the transcript is sent unfolded: Transform returns no error.
var ErrSummaryNoText = errors.New("compact: summary response has no text")
ErrSummaryNoText is the error of a NewLocal fold whose summary response had no text twice, such as one that ended in a function call or in reasoning alone each time. Like ErrSummaryTooLarge, such a fold is reported to WithOnFold with it and the transcript is sent unfolded: Transform returns no error, so a model that will not summarise a prefix does not fail every turn of the conversation.
var ErrSummaryTooLarge = errors.New("compact: summary is larger than what it folds")
ErrSummaryTooLarge is the error of a NewLocal fold whose summary, asked twice, was each time estimated at no fewer tokens than the items it was to replace. Such a fold is reported to WithOnFold with this error, wrapped with the two estimates, and the transcript is sent unfolded: Transform returns no error, since a summary that would not shrink the request leaves the request as it was rather than failing the turn.
After any of the three errors the transform backs off: see Transform.Transform.
Functions ¶
func Estimate ¶
func Estimate(items openresponses.Items) int
Estimate is the default token estimator: the JSON size of the items divided by four.
func PrefixHash ¶ added in v0.0.15
func PrefixHash(items agentturn.Transcript) string
PrefixHash identifies a prefix of a transcript, the way the transform recognises a transcript it folded or failed to fold: the SHA-256, in lowercase hex, of the JSON array encoding/json makes of the items, each as openresponses marshals it. It is empty, which never matches, when the items do not marshal. It is what Fold.PrefixHash reports and WithFailedFold takes, so a session record keeps it; a change to an item's encoding in a later openresponses only makes a remembered prefix not match, and the transcript is folded as usual.
func SummaryMessage ¶ added in v0.0.4
func SummaryMessage(summary string) openresponses.Item
SummaryMessage is the default WithSummaryItem: a user message that opens with "Summary of the conversation so far:" and carries the summary.
Types ¶
type Compactor ¶
type Compactor interface {
Compact(ctx context.Context, req openresponses.CompactRequest) (*openresponses.CompactResponse, error)
}
Compactor is the compaction side of an openresponses.Adapter.
type Fold ¶ added in v0.0.5
type Fold struct {
// Split is the index into that transcript at which the kept tail
// starts: the items before it were folded, and the items from it on
// were sent verbatim after Output.
Split int
// First is the item at Split, the first one kept, nil when the fold
// kept none or failed. A recorder that names the entry holding it finds it by
// this item rather than by Split, which counts items of the
// transcript the transform was given: after another transform in a
// chain, not the agent's.
First openresponses.Item
// Output stands in for the folded items on the request: the
// compaction endpoint's output for [New], the summary message for
// [NewLocal]. nil when the fold failed. The request the transform
// returns is Output, then Pinned, then the items from Split on, so
// the two members name different things and a reader that wants
// what was sent joins them in that order.
Output openresponses.Items
// Summary is the item a recorder writes as the compaction summary:
// the compaction item from [New], the summary message from
// [NewLocal]. nil when the fold failed.
Summary openresponses.Item
// TokensBefore is the estimate that triggered the fold.
TokensBefore int
// Usage is what the fold's model calls reported, when they did. For
// a fold that asked more than once, it is the sum over every call,
// whether the fold then succeeded or failed: what the fold cost,
// not what its last call did.
Usage *openresponses.Usage
// ResponseID is the ID of the response the fold's model call
// produced, from either endpoint, so a recorder can tie the fold to
// a call the server made and a replay can recognise the fold's call
// among the run's. For a fold that asked more than once, it is the
// last call's.
ResponseID string
// OutputTypes are the item types of the last call's output, in
// order, such as [reasoning function_call] for a summary that ended
// in a call: enough to tell a model that answered with something
// other than text from a stream that ended early. nil when the call
// produced no response.
OutputTypes []string
// Attempts is the number of model calls the fold made: one, or two
// for a [NewLocal] fold that asked again.
Attempts int
// Pinned are the items of the folded prefix that [WithPin] kept, in
// their order. They follow Output on the request and are not part
// of it.
Pinned openresponses.Items
// Request is the request [NewLocal] sent for the fold's last
// attempt, as [WithRequest] left it: the items being folded and the
// summary prompt. A failed fold carries it too. Its input is no
// path's context, so a hash of it never rebuilds from a stored
// path; a recorder that keeps it must mark it as the fold's own
// call. nil for [New], whose compaction request is not a Request.
Request *openresponses.Request
// PrefixHash is the [PrefixHash] of the transcript's first Split
// items when the fold failed and the transform backs off from that
// prefix, as [Transform.Transform] describes, and empty otherwise:
// with Split and TokensBefore, what [WithFailedFold] takes to back
// off in another process.
PrefixHash string
// Err is set when the fold failed; Transform returns it, save for
// [ErrSummaryTooLarge], [ErrSummaryIncomplete] and
// [ErrSummaryNoText], after which Transform sends the transcript
// unfolded and returns no error. A fold cut off by an abort carries
// the context error. A failed fold still reports what its calls
// did: Usage, ResponseID, OutputTypes, Attempts and Request, as far
// as the calls got.
Err error
}
Fold describes one compaction attempt on the transcript passed to Transform.Transform.
type Option ¶
type Option func(*Transform)
Option configures a Transform.
func WithBackOff ¶ added in v0.0.16
WithBackOff turns the back-off after a failed fold off when enabled is false: a prefix whose fold failed with an unfolded send is asked about again on the next call over budget, as before v0.0.13, and WithFailedFold seeds nothing. The failed fold is still reported to WithOnFold with its prefix hash. It is for replaying a recording made across restarts before v0.0.15, when the back-off lived in the transform's memory alone and a host that restarted asked again, or by a host that resumed without agentturn/session's CompactOptions: a replay in one process would otherwise back off where the recording asked. The default is on.
func WithBudget ¶
WithBudget sets the estimated token count above which the transcript is compacted.
func WithEstimator ¶
func WithEstimator(fn func(openresponses.Items) int) Option
WithEstimator replaces the token estimator. The default divides the JSON size of the items by four.
func WithFailedFold ¶ added in v0.0.15
WithFailedFold seeds the transform with a failed fold it backs off from, as though it had made it: the fold reported with Fold.Split split, Fold.PrefixHash prefixHash and Fold.TokensBefore tokens. A process that restarts, or resumes a conversation in another, passes the last such fold its record holds, so it does not ask again for a summary that already failed; agentturn/session reads it from the record. A transcript that does not begin with that prefix ignores it, and an empty prefixHash or a negative split seeds nothing. The next fold the transform backs off from replaces it, as in one process.
func WithFilter ¶
func WithFilter(fn func(agentturn.Transcript) agentturn.Transcript) Option
WithFilter sets the filter applied to the items sent to be folded, so app-only items never reach the server. The default is agentturn.DefaultFilter; pass the same function as Config.Filter.
func WithKeepLast ¶
WithKeepLast sets how many recent items are always kept verbatim. The cut never separates a function call from its output.
func WithMinFold ¶ added in v0.0.15
WithMinFold sets the estimated token count below which the part of the transcript a fold would replace is left as it is: the call goes over budget, nothing is asked and nothing is reported to WithOnFold, as for a call within the budget. The part weighed is what the fold would send to be folded, the previous fold's output included. It is for a transcript over budget because of its kept tail, such as a large tool output among the last WithKeepLast items, whose prefix is too small for any summary to shrink.
For NewLocal the default is the larger of an eighth of the budget and twice the estimate of the summary item WithSummaryItem makes from no text, the summary's wrapper and as much again for its text, both as the other options leave them. A prefix that small saves little even when summarised in a word, which is not worth a summary call, and a few short items are likely to be refused as no smaller than their summary after a second call, which is a failed fold. For New the default is zero, since the size of a compaction item is the server's. Zero folds whatever is over budget.
func WithOnFold ¶ added in v0.0.5
WithOnFold adds a function called after every fold attempt, from the goroutine that called Transform and before Transform returns, with the fold that was applied or the failure. Each WithOnFold adds one more, and they are called in the order they were given, so a recorder and a product's own counter both hear every fold. An error one returns ends the calling: the ones after it are not called for that fold. An error it returns fails the Transform and so the turn, the way a subscriber's error fails a run, so a recorder that could not write the fold stops the run rather than letting the record drift from the request. A fold discarded because another caller folded the same transcript meanwhile is not reported. A recorder registers here; see agentturn/session.
func WithPin ¶ added in v0.0.7
func WithPin(fn func(openresponses.Item) bool) Option
WithPin keeps the items fn reports through a fold: whatever part of the folded prefix they were in, they follow the summary in the request, in their order, and the fold that summarised them summarised them too, so nothing is lost if the pin is later dropped. It is for context a harness injected and must not lose to a summary: a stream rule's reminder, a policy notice, an instruction the user gave once. WithKeepLast keeps a window at the end; this keeps a member of the part that is folded.
A pinned item is on the request and not in the transcript, so a session recorder names it in the compaction entry, whose pinned member the context algorithm places after the summary. The calls after such a fold therefore keep their request hashes, and a session resumed from the record still carries the pinned items.
func WithRequest ¶ added in v0.0.12
func WithRequest(fn func(*openresponses.Request)) Option
WithRequest sets a function that edits the request NewLocal sends for a summary before it is sent, after the transform has set its model, input and store. It is how the summary is asked the way the agent asks everything else: a thinking model left at its server's default reasoning may think through the whole fold and end with a function call copied from the transcript, so a caller whose own requests set Reasoning sets the same here, and may set MaxOutputTokens, Temperature or any other field. A replay tells the fold's call from a turn by its lack of tools and instructions, so leave those empty. The transform sets MaxOutputTokens to half the budget, at most DefaultSummaryMaxOutputTokens, before fn runs, so a summary that runs away is cut by the server; fn may change or clear it. fn runs once per attempt on a fresh request; the edited request is the one reported as Fold.Request. New ignores it.
func WithSummaryItem ¶ added in v0.0.4
func WithSummaryItem(fn func(summary string) openresponses.Item) Option
WithSummaryItem sets how NewLocal turns the summary text into the item that stands in for the folded prefix. The default is a user message opening with "Summary of the conversation so far:", which every server accepts on input. A server that accepts compaction items on input could be given one here instead. New ignores it.
func WithSummaryPrompt ¶ added in v0.0.4
WithSummaryPrompt replaces DefaultSummaryPrompt for NewLocal. It is sent as the last user message of the summary request, after the items to fold. New ignores it.
type Transform ¶
type Transform struct {
// contains filtered or unexported fields
}
Transform compacts transcripts. Its Transform method is the value for agentturn.Config.Transform. It is safe for concurrent use and does not hold its lock across the model call, but it remembers one compacted prefix and one failed fold, so share one per conversation. Both memories are keyed by a hash of the prefix they cover, so a Transform shared between conversations never sends one the other's fold or skips a fold because the other's failed; it only folds more often than one per conversation would, as each forgets the other's.
func New ¶
New builds a Transform that folds through c's compaction endpoint. The result is the endpoint's output, a compaction item the same server expands on the next call.
func NewLocal ¶ added in v0.0.4
func NewLocal(model openresponses.Streamer, opts ...Option) *Transform
NewLocal builds a Transform that folds by asking model for a summary with an ordinary call: the items to fold, then the summary prompt as a user message, with no tools. The reasoning items among the items to fold are left out of the request: their content is encrypted for the model that produced them, which a summariser cannot read, and a provider refuses one another model produced. The result is one message carrying the summary (see WithSummaryItem). It works against any server, including those that answer 404 to the compaction endpoint. The model is named by WithModel; leave it empty to let the server pick its default. WithRequest edits the rest of the request.
The summary request's MaxOutputTokens is half the budget, at most DefaultSummaryMaxOutputTokens, unless WithRequest sets it otherwise. A summary response with no text, such as one that ends in a function call or in reasoning alone, a response the server ended incomplete, such as one cut at MaxOutputTokens, and a summary whose text is estimated at no fewer tokens than the items it folds are model errors rather than server ones, so the summary is asked once more. A second such answer is never applied, since there is nothing to apply, a summary cut short has lost part of what it folds and one too large would grow the request the fold exists to shrink: the fold is reported failed with ErrSummaryNoText, with ErrSummaryIncomplete and the server's reason, or with ErrSummaryTooLarge, the transcript is sent unfolded, and the transform backs off from that prefix as Transform.Transform describes. The summary's size is that of its text, the estimate of its item less that of an item with no text, so the wrapper WithSummaryItem puts around every summary does not count against it. The fold, failed or not, reports the last call's response ID, output types and request, the usage of every call summed, and the number of attempts.
func (*Transform) Last ¶
func (t *Transform) Last() openresponses.Item
Last returns the item that most recently replaced a folded prefix, or nil: the compaction item from New, the summary message from NewLocal. WithOnFold reports the same item with the split that produced it.
func (*Transform) Transform ¶
func (t *Transform) Transform(ctx context.Context, items agentturn.Transcript) (agentturn.Transcript, error)
Transform returns the transcript to send for this call: unchanged when it fits the budget or when its older part is smaller than WithMinFold, otherwise the fold of its older part followed by the recent items.
A fold that failed with ErrSummaryTooLarge, ErrSummaryIncomplete or ErrSummaryNoText sends the transcript unfolded, and the transform remembers it: the length and hash of the prefix it would have folded, and the estimate that triggered it. While a later transcript still begins with that prefix, it is not folded again until the part to fold has grown by at least WithKeepLast items (at least one) or the estimate by at least a quarter of the budget: the same prefix would most likely fail the same way, and asking on every turn would cost two summary calls and a failed fold each time. Such a call goes over budget, folds nothing and reports nothing to WithOnFold. A transcript that does not begin with the failed prefix, such as another conversation's or one rewound before it, folds as usual. WithBackOff turns the back-off off.