bonnie

module
v0.2.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 13, 2026 License: MIT

README

BONNIE

BONNIE

Durable agent runs for Go.
Survive a crash. Wait days for a human. Answer over HTTP.

Go Reference MIT


An agent turn normally lives and dies with your process. Kill it mid-tool-call and the work is gone. Ask the user a question and you have to hold the process open until they answer.

BONNIE fixes that. It wraps the Kit agent SDK so a run becomes durable:

run, _ := runner.Start(ctx, "deploy-42", runtime.Input{Text: "Deploy the app."})

if run.State == runtime.RunWaiting {
    fmt.Println(run.Suspend.Prompt) // "Which region?"
    os.Exit(0)                      // ← the process can end here
}

Come back tomorrow, in a different process, and finish it:

run, _ := runner.Resume(ctx, "deploy-42",
    []runtime.InputResponse{{Text: "eu-west-1"}})

fmt.Println(run.Response) // "Deployed to eu-west-1."

The agent remembers the whole conversation, including which tools it already called — so it does not repeat a side effect it has already performed.

Contents

Install

As a library:

go get github.com/mark3labs/bonnie

As a CLI:

go install github.com/mark3labs/bonnie/cmd/bonnie@latest

With Nix. This gives you the CLI with the microsandbox CLI (msb) already on its PATH:

nix profile install github:mark3labs/bonnie   # or: nix run github:mark3labs/bonnie

Set a provider key. BONNIE uses whatever Kit is configured for:

export ANTHROPIC_API_KEY=sk-ant-...   # or OPENAI_API_KEY, or GEMINI_API_KEY

Requires Go 1.27+. Sandboxing is optional and needs Docker or msb.

Development shell

The flake also gives you a shell with Go 1.27, golangci-lint, goreleaser, and the microsandbox CLI:

nix develop
go test -race ./...

The repository ships an .envrc, so direnv allow enters the same shell on cd.

Other flake outputs:

Output What it is
packages.default, packages.bonnie the BONNIE CLI
packages.microsandbox the msb CLI plus its libkrunfw
apps.msb nix run github:mark3labs/bonnie#msb
overlays.default both packages, for your own nixpkgs

Quickstart: no Go required

Scaffold an agent, edit one file, serve it. No Go toolchain, no build.

bonnie init my-agent --model anthropic/claude-sonnet-4-5
cd my-agent
# edit instructions.md — that file is the agent's system prompt
bonnie serve --agent .

The tree is four things: agent.yaml (the manifest: model, sandbox, channels), instructions.md (the system prompt, read fresh at every start), and skills/ and workspace/ (seed directories — files under workspace/ are mirrored into every run's sandbox, and a file the model already wrote is never overwritten).

# talk to it over HTTP
curl -s localhost:8080/runs -d '{"text":"What are you?"}'

Every serve setting comes from the manifest or a flag, and the startup banner says which source won. Flags win: --model, --addr, --sandbox, and the rest override the manifest when both are given.

The manifest is strict: an unknown key, an unknown apiVersion, or two manifests in one directory is an error that names what is wrong — never a silent default. A tree that carries Go tools cannot be served this way; serve refuses it and names bonnie build (the codegen path, coming in v0.2).

Quickstart

A durable run in 20 lines. The journal on disk is what makes it durable.

package main

import (
	"context"
	"fmt"
	"log"

	"github.com/mark3labs/bonnie/runtime"
	kit "github.com/mark3labs/kit/pkg/kit"
)

func main() {
	// Every message is journalled here before it is kept.
	journal, err := runtime.OpenFileJournal(".bonnie")
	if err != nil {
		log.Fatal(err)
	}
	defer journal.Close()

	runner := runtime.NewRunner(journal, runtime.KitAgent(
		kit.WithModel("anthropic/claude-sonnet-4-5"),
	))

	run, err := runner.Start(context.Background(), "run-1",
		runtime.Input{Text: "In one sentence, what is a durable agent run?"})
	if err != nil {
		log.Fatal(err)
	}
	fmt.Println(run.Response)
}

Run it again with a different message and the same run ID. BONNIE replays the conversation first, so the agent remembers:

run, _ := runner.Start(ctx, "run-1", runtime.Input{Text: "What did I just ask?"})
// `You asked me "In one sentence, what is a durable agent run?"`

Inspect what happened, without a server:

bonnie runs list --journal .bonnie
bonnie runs show --journal .bonnie run-1
RUN    STATE      STEPS  LAST
run-1  completed  2      You asked me "In one sentence, what is a durable agent run?"

Park and resume

This is the headline feature. An agent asks a question, your process exits, and a completely new process finishes the job.

BONNIE ships two tools for this. Register them and the model can call them:

Tool Parks the run to...
ask_human ask the operator a question
request_approval get approval before a risky action

runtime.KitAgent registers both automatically.

runner := runtime.NewRunner(journal, runtime.KitAgent(
	kit.WithModel("anthropic/claude-sonnet-4-5"),
	kit.WithSystemPrompt("Before you deploy anything, use ask_human to ask "+
		"which region to deploy to."),
))

run, err := runner.Start(ctx, "deploy-42", runtime.Input{Text: "Deploy the app."})
if err != nil {
	log.Fatal(err)
}

if run.State == runtime.RunWaiting {
	fmt.Println("agent asks:", run.Suspend.Prompt)
	return // nothing is holding compute — the process may exit
}

Later, anywhere, as long as it can read the same journal:

run, err := runner.Resume(ctx, "deploy-42",
	[]runtime.InputResponse{{Text: "eu-west-1"}})

A parked run holds no compute. It costs nothing to wait a week.

See examples/hitl-restart for a runnable version that genuinely calls os.Exit between the two phases:

go run ./examples/hitl-restart -phase ask
# ...process exits, run is parked on disk...
go run ./examples/hitl-restart -phase answer -answer "eu-west-1"

Your own tools

A BONNIE tool is a Kit tool. Pass it through and it joins the sandboxed and human-in-the-loop sets:

type chargeInput struct {
	Amount int    `json:"amount" description:"Amount in cents."`
	UserID string `json:"user_id" description:"Who to charge."`
}

chargeCard := kit.NewTool("charge_card", "Charge a customer's card.",
	func(ctx context.Context, in chargeInput) (kit.ToolOutput, error) {
		// Runs in YOUR process, with your secrets. The model sees only
		// the result you return.
		if err := stripe.Charge(in.UserID, in.Amount); err != nil {
			return kit.ErrorResult(err.Error()), nil
		}
		return kit.TextResult("Charged."), nil
	})

runner := runtime.NewRunner(journal, runtime.KitAgent(
	kit.WithModel("anthropic/claude-sonnet-4-5"),
	kit.WithExtraTools(chargeCard),
))

You can also write a tool that parks the run — that is all ask_human is:

return kit.ToolOutput{
	Content: "Waiting for the finance team.",
	Halt:    true,
	FinalValue: runtime.SuspendRequest{
		Kind:   "approval",
		Prompt: "Approve a $4,000 refund?",
	},
}, nil

The run stops, Start returns with State == RunWaiting, and run.Suspend.Prompt carries your question.

Prefer scaffolding over hand-wiring? bonnie init --tools creates a tree with one sample tool and a main.go you own; tools there live in tools/<name>/tool.go as func Tool() kit.Tool, and the directory name is the tool's name. Codegen that regenerates the wiring (and bonnie dev/ bonnie build) lands in v0.2; the scaffold builds today.

Sandboxing

By default, tool calls run as your process — your files, your network, your credentials. For anything untrusted, put them in a sandbox:

provider := sandbox.Docker(sandbox.WithDockerImage("python:3.12-slim"))

runner := runtime.NewRunner(journal, sandbox.Agent(provider,
	kit.WithModel("anthropic/claude-sonnet-4-5"),
))

The model now gets bash, read_file, write_file, and list_files that run inside a container rooted at /workspace.

Backend Isolation You install Extra Go deps
sandbox.Local() none — dev only 0
sandbox.Docker() container namespaces Docker 0
sandbox.Microsandbox() microVM, guest kernel msb 0

All three drive a CLI, so BONNIE stays a single static binary.

The microsandbox adapter is verified on Linux with KVM (msb 0.6.18, all 18 conformance cases, network policies enforced with real egress). It has not been run on macOS with Apple Silicon, and its network policy is fixed at create time: reattaching under a different policy fails with ErrPolicyMismatch rather than silently using the old rules.

Lock down the network:

provider := sandbox.Docker()
provider.SetNetworkPolicy(sandbox.NetworkPolicy{Mode: sandbox.NetworkDenyAll})

Pick the best backend available, without silently falling back to no isolation:

provider, err := sandbox.Select(ctx, sandbox.Microsandbox(), sandbox.Docker())

The sandbox opens on the first tool call that needs it, so a parked run holds no container. Read docs/SANDBOX.md before deploying.

Serve over HTTP

bonnie serve --journal .bonnie --model anthropic/claude-sonnet-4-5 --sandbox docker

Or serve a discovered agent tree — manifest, instructions, sandbox and channels from files, no build. See Quickstart: no Go required:

bonnie serve --agent my-agent

Or mount it in your own server:

runner := runtime.NewRunner(journal, runtime.KitAgent(opts...))
http.ListenAndServe(":8080", bonniehttp.New(runner).Handler())
Route Does
POST /runs start a run, or route to the one serving an address
GET /runs/{id} report a run's durable state
POST /runs/{id} send a message to an existing run
POST /runs/{id}/respond answer a parked run
POST /runs/{id}/cancel stop the turn in flight
GET /runs/{id}/stream NDJSON event stream, resumable via ?cursor=
# Start a run. It parks on a question.
curl -s localhost:8080/runs -d '{"text":"Deploy the app. Ask me the region first."}'
# {"run_id":"run-e8b3fa...","state":"waiting",
#  "suspend":{"kind":"question","prompt":"Which region?"}}

# Answer it.
curl -s localhost:8080/runs/run-e8b3fa.../respond \
  -d '{"responses":[{"text":"eu-west-1"}]}'
# {"run_id":"run-e8b3fa...","state":"completed","response":"Deployed to eu-west-1."}

Addresses

Chat platforms have threads, not run IDs. Pass an address and BONNIE keeps the mapping in the journal, so a restart does not orphan a conversation:

curl -s localhost:8080/runs -d '{"address":"slack:C123/T456","text":"hi"}'

The same address always resolves to the same run. POST /runs/{id} is the opposite: it targets one exact run and returns 404 rather than creating one.

Chat channels

Slack, Discord, and Telegram put the same durable runs into a conversation. Enable one in the manifest, put its credentials in the environment, and serve:

# agent.yaml
channels:
  slack: {}
  discord: {}
  telegram:
    username: mybot
export SLACK_BOT_TOKEN=xoxb-... SLACK_SIGNING_SECRET=...
export DISCORD_BOT_TOKEN=... DISCORD_PUBLIC_KEY=...
export TELEGRAM_BOT_TOKEN=... TELEGRAM_WEBHOOK_SECRET=...
bonnie serve --agent .

Each channel mounts one webhook (/slack/events, /discord/interactions, /telegram), verifies its platform's signature — Slack's v0 HMAC, Discord's Ed25519, Telegram's shared secret — and answers within the platform's ACK deadline while the turn runs on. The reply posts back to the thread; a parked run posts its question, and the next message on the thread is the answer.

The details are in docs/CHANNELS.md: the per-platform setup, the dispatch and steering rules, and what is deliberately not implemented (streaming edits, button HITL, attachments, gateway transports).

CLI

bonnie init     Scaffold an agent tree (manifest, instructions, seeds)
bonnie serve    Mount the HTTP channel and serve durable runs
bonnie runs     List and inspect durable runs
bonnie sandbox  Reclaim the sandboxes of terminal runs (prune)
bonnie version  Print the version
bonnie init my-agent --model anthropic/claude-sonnet-4-5
bonnie init .                      adopt this directory; never overwrites
bonnie init my-agent --format toml # or json; yaml is the default
bonnie init my-agent --tools       add a Go module with a sample tool

bonnie serve --agent my-agent      serve a discovered tree, no build
bonnie serve --agent . --config agent.toml
bonnie serve --addr :8080 --journal .bonnie \
             --model anthropic/claude-sonnet-4-5 \
             --sandbox docker --sandbox-deny-network

bonnie runs list --journal .bonnie --state waiting
bonnie runs show --journal .bonnie run-1
bonnie runs show --journal .bonnie run-1 --json | jq '.[] | select(.kind=="message")'

runs reads the journal directly, so it works while the server is stopped — which is exactly when you need it.

Storage

The journal is the durability seam. Two ship in the box:

runtime.NewMemoryJournal()          // tests and ephemeral runs
runtime.OpenFileJournal(".bonnie")  // JSONL, one file per run

FileJournal writes <root>/runs/<run-id>.jsonl, one JSON object per line, fsynced by default. It is plain text, so jq works on it:

jq -c 'select(.kind=="message") | {seq, role}' .bonnie/runs/run-1.jsonl

Bring your own by implementing seven methods:

type Journal interface {
	Append(ctx context.Context, rec Record) (seq int, err error)
	Replay(ctx context.Context, runID string) ([]Record, error)
	Checkpoint(ctx context.Context, runID string, state RunState) error
	State(ctx context.Context, runID string) (RunState, error)
	Runs(ctx context.Context, state RunState) ([]string, error)
	Persisted() bool
	Close() error
}

A table-driven conformance suite in runtime/journal_conformance_test.go runs against every implementation. Add yours to it and it inherits the whole suite.

Streaming

Subscribe to a run's events in-process:

events, unsubscribe := runner.Events().Subscribe("run-1", 0)
defer unsubscribe()

for ev := range events {
	fmt.Println(ev.Seq, ev.Type, ev.Text)
}

Or over HTTP, one JSON object per line:

curl -sN localhost:8080/runs/run-1/stream

Every event carries a monotonic seq. If a client drops, reconnect with the last one it saw and lose nothing:

curl -sN "localhost:8080/runs/run-1/stream?cursor=12"

Steer and cancel

runner.Steer("run-1", "actually, use eu-west-1")  // joins the running turn
runner.Cancel("run-1")                            // stops it

Cancelling keeps every completed step, so the run restores to a valid conversation and can be continued with Start.

Run states

pending → running → completed
                  → waiting    (parked for a human; resume with Resume)
                  → cancelled  (stopped by an operator; continue with Start)
                  → failed

How it works

BONNIE is built on four public Kit extension points. No fork, no patched SDK:

Need Kit API
Journal every message Options.SessionManager
Checkpoint each step Kit.OnStepFinish
Inject replayed context Kit.OnContextPrepare
Park for a human kit.ToolOutput{Halt, FinalValue}

Three properties are worth knowing, because they are the difference between a demo and something you can deploy:

  • Replay is lossless. Journalled messages keep their typed parts, so a resumed run knows which tools it called and what came back. It will not repeat a side effect it already performed.
  • A step commits atomically. A tool call and its result reach the journal as one write and one fsync (kit.StepAppender, adopted from Kit v0.106.0), so a crash cannot leave an unanswered tool call. If a torn step still reaches disk — from an older journal, a non-batching journal, or a short write — restore drops that incomplete step and records the repair.
  • Cancelling keeps finished work. Steps are persisted before the context is checked, so a cancelled turn loses only the step in flight.

Full detail, with the Kit citations, in docs/SPEC.md.

Limits

Stated plainly, because the failure modes are not obvious:

  • Sandboxing is opt-in. Without it, tool calls run as your process.
  • Docker is namespaces, not a kernel. Use microsandbox for hostile code.
  • microsandbox is verified on Linux/KVM only — not on macOS with Apple Silicon. Every network policy mode is enforced, but the policy is fixed at create time; reattaching under a different policy fails with ErrPolicyMismatch.
  • Sandbox egress is open unless you set a policy.
  • No auth verification on the HTTP channel. It carries a Principal; it does not check one. Authenticate in front of it. The chat channels are different: each verifies its platform's signature, and a channel without its credentials refuses to serve. That verifies the platform, not the person — a user ID inside a verified Slack event is Slack's word.
  • Run ownership is per host. The file journal locks each run with flock, so a second writer to the same run is refused rather than allowed to corrupt it. That lock does not work on a network filesystem, and it does not make two writers coordinate — one process still owns each run.
  • Events are journal-anchored. The stream replays the journal past the in-memory backlog, so a reconnect — even after a restart — has no gap. Live-only deltas are the exception, marked as such.
  • Sandbox lifecycle is journalled, and reclaiming is manual. bonnie sandbox prune deletes the sandboxes of terminal runs; serve does not sweep them on its own yet.
  • The mark3labs modules are publicly fetchable. A scaffolded module (bonnie init --tools) runs go mod tidy and resolves bonnie and kit from the proxy; no go.work, no GOPRIVATE. The zero-Go path (serve --agent) needs neither Go nor a module.

Examples

Example Shows
examples/minimal one durable run, start to finish
examples/hitl-restart park, exit the process, resume
go run ./examples/minimal -text "What is a durable agent run?"

See examples/README.md for copy-pasteable commands.

Documentation

Document Purpose
docs/HANDOVER.md Picking up the project: state, pitfalls, what to do next
docs/SANDBOX.md Sandbox backends and the contracts an adapter must honour
docs/SPEC.md Specification: scope, verified Kit facts, known risks, invariants
docs/L2.md Draft spec for v0.2: the agent tree, manifest, init/dev/build
docs/TASKS.md Open work, and an archive of what shipped
docs/UPSTREAM.md The home for BONNIE's future asks of Kit; answered ones live in docs/archive/
CONTRIBUTING.md Boundary rule, workspace setup, commands
SECURITY.md Disclosure, and what v0.1.0 does not protect you from

Contributing

go build ./...
go test -race ./...
golangci-lint run

Or run the same loop with task check, and CI parity with task ci.

The live-model tests are behind a build tag and need a provider key:

go test -race -tags integration ./runtime ./sandbox

They skip, never fail, when no key is present. One rule matters above the rest: BONNIE uses the public Kit SDK only. See CONTRIBUTING.md.

License

MIT — see LICENSE.

Directories

Path Synopsis
Package agent implements BONNIE's L2 discovery: the authored agent tree.
Package agent implements BONNIE's L2 discovery: the authored agent tree.
Package channel defines BONNIE's inbound transport abstraction (L3).
Package channel defines BONNIE's inbound transport abstraction (L3).
chat
Package chat holds the plumbing every chat-platform channel shares: the journalled address map, per-run turn locks, the channel.SessionRef implementation, and the dispatch rule that sends a message to the right runner entry point.
Package chat holds the plumbing every chat-platform channel shares: the journalled address map, per-run turn locks, the channel.SessionRef implementation, and the dispatch rule that sends a message to the right runner entry point.
discord
Package discord is BONNIE's Discord inbound transport (L3).
Package discord is BONNIE's Discord inbound transport (L3).
http
Package http is BONNIE's HTTP inbound transport (L3).
Package http is BONNIE's HTTP inbound transport (L3).
slack
Package slack is BONNIE's Slack inbound transport (L3).
Package slack is BONNIE's Slack inbound transport (L3).
telegram
Package telegram is BONNIE's Telegram inbound transport (L3).
Package telegram is BONNIE's Telegram inbound transport (L3).
Package channeltest is the conformance suite for channel adapters, in the same spirit as the journal and sandbox suites.
Package channeltest is the conformance suite for channel adapters, in the same spirit as the journal and sandbox suites.
cmd
bonnie command
Command bonnie is the BONNIE developer CLI.
Command bonnie is the BONNIE developer CLI.
examples
hitl-restart command
Command hitl-restart is the headline demonstration: a run parks for a human, the process exits, and a completely new process finishes the run.
Command hitl-restart is the headline demonstration: a run parks for a human, the process exits, and a completely new process finishes the run.
minimal command
Command minimal starts one durable run and prints the answer.
Command minimal starts one durable run and prints the answer.
Package runtime is BONNIE's durable execution layer (L1).
Package runtime is BONNIE's durable execution layer (L1).
Package sandbox gives a BONNIE run an isolated place to run tool calls.
Package sandbox gives a BONNIE run an isolated place to run tool calls.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL