Documentation
¶
Overview ¶
Package danger classifies shell commands by risk level and provides a configurable approval system for dangerous operations.
Classification is token-based (not regex) — it respects quotes, pipes, redirects, compound commands (&&, ||, ;), and multi-line input. Each command is classified into one of 9 risk classes, and the user can configure which actions (allow/prompt/deny) apply to each class.
The gate fails CLOSED. A command whose program name is recognised but used benignly classifies as Safe (allow); a command whose verb is NOT recognised classifies as Unknown and is denied by default. The set of recognised-safe commands (safeCommands) is therefore an explicit read-only allowlist — extend it, or the per-profile allowlist, to permit a tool rather than relying on it slipping through unclassified.
Threat model ¶
The classifier is an adversarial filter, not a parser for well-behaved input. It assumes a prompt-injected agent is actively trying to make a dangerous command read as harmless so it slips past the approval gate. The design therefore errs toward the worse class when in doubt, and is built in layers that each close a category of evasion:
Normalisation (see normalize) rewrites the command so token-level analysis can see through shell tricks before classification runs: - $'…' ANSI-C escapes decodeANSIC ($'\x72\x6d' → rm) - $IFS word-splitting expandIFS (rm$IFS-rf$IFS/ → rm -rf /) - {a,b,c} brace expansion expandBraces ({rm,-rf,/} → rm -rf /) - $(…)/`…`/<(…)/>(…) subst. extractSubstitutions (bodies classified too) - command/exec/builtin stripCommandWrappers - \-escapes (r\m, \rm) collapseUnquotedBackslashes - absolute paths (/bin/rm) basenameFirstToken + commandName The tokenizer additionally treats quote boundaries as NON word boundaries, so empty/adjacent quotes like r""m and "rm" still resolve to the single word `rm`.
Structural decomposition. A command is split into segments (on ;, &&, ||), each segment into pipe stages (on |), and EVERY stage is classified — not just the head — so `true | dd of=/dev/sda` and `echo x | sudo rm -rf /home` are seen for what their later stages do. A stage that pipes INTO xargs running a destructive/system verb has its upstream literal payload (echo/printf args) composed onto the inner command (`echo "/" | xargs rm -rf` classifies like `rm -rf /`); when the payload is not statically determinable the pipeline fails closed (unknown → deny). The worst class across all parts wins (see rank).
Wrapper unwrapping (unwrapWrappers). Leading execution wrappers (env, xargs, nohup, nice, setsid, timeout, …) are stripped so the real command underneath is classified; privileged wrappers (sudo, doas, pkexec) additionally impose a system_write floor and then let the inner command escalate further (sudo rm -rf /var → destructive).
Verb-independent resource scanning (classifyResourceToken). Some resources are dangerous regardless of the command touching them: /dev/tcp and /dev/udp pseudo-devices (reverse-shell channels) and sensitive credential paths (~/.ssh, /etc/shadow, ~/.aws/credentials, /proc/self/environ, …). These are flagged wherever they appear.
Payload re-classification. Shell -c strings (bash -c '…') and the bodies of command/process substitutions are themselves classified by re-entering Classify, so nested commands cannot hide a level deeper.
Limitations ¶
This is a heuristic defence-in-depth layer, NOT a sandbox or a complete shell interpreter. It does not, and cannot, catch everything:
- Variable indirection: `X=rm; $X -rf /` — the value of $X is not tracked. Note the fail-closed default turns this from a silent bypass into a denial: the unrecognised `$X` verb classifies as Unknown.
- Fully dynamic construction from runtime data, command output, or environment the classifier cannot evaluate.
- Arbitrary value transformations beyond the enumerated encodings (e.g. a secret piped through gzip/openssl before exfiltration).
- Interpreter escape hatches we do not special-case. Common ones ARE covered: awk/ed/vi/emacs invocations that carry a script or file operand classify as code_execution (embeddedShellInterpreters), as do non-shell interpreters fed from a pipe (`curl … | python`). Language-specific eval paths or editor command-mode shells we have not enumerated may still read as a known command used benignly — the known verb is the gap, not an unknown one.
Because these gaps exist, the classifier is paired with other controls: non-interactive denial, output redaction (internal/redact), and — for strong isolation — the container sandbox. When tuning, remember that over-classification only costs an extra prompt, while under-classification can let a destructive or exfiltrating command through silently; prefer the former.
Index ¶
- Constants
- func ContainsInvisible(s string) bool
- func FoldHomoglyphs(text string) string
- func HasConfusableScript(s string) bool
- func HostIsImplicitlyInternal(host string) bool
- func IsBlockedIP(ip net.IP) bool
- func IsPersistencePath(path string) bool
- func IsSafe(content string) bool
- func NormalizeForScan(text string) string
- func Rank(cls RiskClass) int
- func RecordRead(path string)
- func ResetReadLedgerForTest()
- func ResetTTYFrictionStateForTest()
- func SetTTYPathForTest(path string)
- func TrustShortcutAllowed(cls RiskClass) bool
- func UnreadScriptTargets(cmd string) []string
- func WasRead(path string) bool
- func WasReadFresh(path string) bool
- type Action
- type Approver
- type DangerousConfig
- type InjectionPattern
- type RiskClass
- func Classify(cmd string) RiskClass
- func ClassifyPath(path string) RiskClass
- func ClassifyPathWrite(path string) RiskClass
- func ClassifyScriptGate(cmd string) (RiskClass, []string)
- func ClassifyURL(rawURL string) RiskClass
- func LifecycleContentClass(path, content string, base RiskClass) (RiskClass, bool)
- type ScanResult
- type TTYApprover
- type ToolOperation
Constants ¶
const ToolBatchClass = RiskClass("tool_batch")
ToolBatchClass is the synthetic risk class the agent loop uses when it prompts once for a batch of tool calls together. A batch aggregates multiple tools whose individual classes may differ, so it must never be session-trustable: one Trust grant would blanket-approve the contents of every later batch (and, via the loop's SetTrustAll, every per-tool prompt in the session).
Variables ¶
This section is empty.
Functions ¶
func ContainsInvisible ¶ added in v1.8.0
ContainsInvisible reports whether s contains any invisible character that NormalizeForScan would strip. It is used to flag stealth-character evasion even when the normalized text does not match a blacklist pattern.
func FoldHomoglyphs ¶ added in v1.8.0
FoldHomoglyphs returns text with common Unicode confusables (Cyrillic/Greek look-alikes) replaced by their ASCII equivalents. It is used as an extra scan surface to catch mixed-script homoglyph attacks.
func HasConfusableScript ¶ added in v1.8.0
HasConfusableScript reports whether s mixes Latin script with characters from scripts that contain visually confusable letters (Cyrillic/Greek) or CJK. This is a separate signal from pattern matching: it catches pure homoglyph attacks even when the normalized text does not match a blacklist pattern.
func HostIsImplicitlyInternal ¶ added in v1.7.0
HostIsImplicitlyInternal reports whether the literal host string already resolves to an internal target by inspection alone — i.e. ClassifyURL returns SystemWrite for it with no DNS lookup (a literal internal IP, in any browser encoding, or a known-internal hostname). The dial-time SSRF guard uses this to tell apart a target that was *already* surfaced to the policy gate as internal (and dialed under that decision) from one that presented as external and must be re-validated against its resolved IPs.
func IsBlockedIP ¶ added in v1.7.0
IsBlockedIP reports whether ip falls in a range that the agent's web tools must never reach: loopback (127/8, ::1), RFC1918 / RFC4193 private (incl. IPv6 ULA fc00::/7), link-local (169.254/16 — which covers the 169.254.169.254 cloud-metadata endpoint — and fe80::/10), RFC 6598 CGNAT (100.64/10), RFC 2544 benchmark testing (198.18/15), or the unspecified address (0.0.0.0, ::). It is the single source of truth shared by both ClassifyURL's literal-host gate and the dial-time SSRF guard, so the two cannot drift apart. A nil IP is treated as blocked (fail closed).
func IsPersistencePath ¶ added in v1.27.0
IsPersistencePath reports whether path names a deferred-execution target. It is direction-agnostic; callers gating reads should keep using ClassifyPath (reads of these files stay at their existing class) and reserve the persistence escalation for writes via ClassifyPathWrite.
func IsSafe ¶
IsSafe returns true if no injection threats are detected in content. This is the primary gate used before injecting untrusted content into the system prompt.
func NormalizeForScan ¶ added in v1.8.0
NormalizeForScan returns a lower-cased, whitespace-normalized form of text with invisible characters removed. It does NOT fold homoglyphs so that non-English patterns (e.g., Russian, French) still match.
func Rank ¶ added in v1.1.0
Rank returns the severity order for priority comparison. Exported so consumers that enforce risk caps (e.g. the sub-agent maxRisk clamp) share this single ordering instead of mirroring it — a mirror silently drifts when a class is added, as happened with Unknown.
func RecordRead ¶ added in v1.27.0
func RecordRead(path string)
RecordRead marks path as read this session. Paths are normalised to absolute/cleaned form. Successful writes through the file tools also record here — content the agent authored is content it has seen.
The entry is fingerprinted at record time; WasReadFresh re-verifies the on-disk state at gate time so a post-read mutation re-fires the H-6 gate.
func ResetReadLedgerForTest ¶ added in v1.27.0
func ResetReadLedgerForTest()
ResetReadLedgerForTest clears the session ledger.
func ResetTTYFrictionStateForTest ¶ added in v1.14.0
func ResetTTYFrictionStateForTest()
ResetTTYFrictionStateForTest clears the process-wide approval log used by friction mode. It is intended for tests that need a clean approval baseline.
func SetTTYPathForTest ¶ added in v1.29.1
func SetTTYPathForTest(path string)
SetTTYPathForTest overrides the device path TTYApprover opens for interactive approvals. Test-only: a test process running on an operator's dev machine has a real controlling terminal, so any tool call that lands in the Prompt fallback would render a live approval prompt and block on the operator's keyboard forever (observed 2026-08-30 with browser tests on macOS). Pointing the path at a nonexistent device makes every fallback open fail, which engages the documented NonInteractiveAction fallback — exactly the semantics the non-interactive tests were authored against. Tests that script their own TTY set TTYPath directly after construction and are unaffected.
func TrustShortcutAllowed ¶ added in v1.25.1
TrustShortcutAllowed reports whether cls may be session-trusted via the "trust" shortcut. Destructive, Blocked, and Unknown must never be (fail-closed catch-alls; blanket-trusting them is carte blanche), and neither may ToolBatchClass — see its doc comment. Persistence is also excluded: its writes execute later, outside the session where the trust was granted, so a one-time "trust" must not cover every future hook, profile, and CI-workflow write (H-5). UnreadExec is excluded because the entire point of the gate is per-script review — trusting it once would blanket-approve every unread script for the session (H-6).
func UnreadScriptTargets ¶ added in v1.27.0
UnreadScriptTargets returns the script-file operands of cmd that execute code and have not been read this session. Empty when nothing gates.
Verb-aware by design: only stages whose command IS an execution context (interpreter, source, direct script invocation) are scanned, so `grep pattern build.sh` — a read — never triggers the gate.
func WasRead ¶ added in v1.27.0
WasRead reports whether path was recorded as read this session, regardless of whether the bytes have changed since. Licensing checks must use WasReadFresh.
func WasReadFresh ¶ added in v1.27.1
WasReadFresh reports whether path was read this session AND the bytes on disk are still the state that was displayed (or authored) at record time: same size, same mtime, and — for files up to readFingerprintMaxBytes — the same sha256 digest. A read that is no longer fresh does not license execution; the H-6 gate re-fires until the mutated content is re-read (which renews the fingerprint, because now the model has seen THAT).
Types ¶
type Action ¶
type Action string
Action represents what to do when a command of a given risk class is detected.
const ( Allow Action = "allow" Prompt Action = "prompt" Deny Action = "deny" // ReadOnly is not a per-class action — it is a non_interactive mode // (H-7): without a TTY, read-only inspection proceeds while writes, // execution, and egress stay denied. Containment via inability is not // safe-and-useful; read_only keeps headless agents useful enough that // nobody reaches for "allow". ReadOnly Action = "read_only" )
func ParseNonInteractiveAction ¶ added in v1.14.0
ParseNonInteractiveAction parses the non_interactive config value. It accepts "allow", "deny", and "read_only"; "prompt" and any other value are rejected because prompting is impossible without a TTY.
type Approver ¶
type Approver interface {
// PromptCommand asks the user to approve or deny a shell command.
// cls is the risk class (system_write, network_egress, etc.).
// Returns nil on approve, error on deny or timeout.
PromptCommand(cls RiskClass, cmd, description string) error
// PromptOperation asks the user to approve or deny a native tool operation
// (read_file on /etc, browser to external URL, etc.).
PromptOperation(op ToolOperation) error
}
Approver is the interface for user approval of dangerous operations. Two implementations exist:
- TTYApprover — opens /dev/tty for interactive approval (CLI mode)
- WSApprover — sends approval requests via WebSocket (serve mode)
When nil (no approver configured), calls fall back to non-interactive behavior (NonInteractiveAction). Tools MUST inject an approver to get interactive approval in any mode.
type DangerousConfig ¶
type DangerousConfig struct {
// Classes maps risk classes to their configured action.
// Only overrides for non-default values need to be set.
Classes map[RiskClass]Action `json:"classes,omitempty"`
// Allowlist is a list of command strings that are always allowed,
// regardless of their risk classification. Exact match only.
// Takes priority over Denylist.
Allowlist []string `json:"allowlist,omitempty"`
// Denylist is a list of command strings that are always denied,
// regardless of their risk classification. Prefix match (after trimming).
Denylist []string `json:"denylist,omitempty"`
// DefaultAction is the global default action applied to ALL risk classes
// when set. Per-class overrides in Classes still win.
// "allow" → YOLO mode (everything runs without prompt)
// "deny" → lockdown (everything denied unless explicitly allowed)
// Not set → uses built-in defaults per class
DefaultAction *string `json:"action,omitempty"`
// NonInteractive specifies what to do when running without a TTY.
// "read_only" (default) — read-only inspection proceeds, writes/exec/
// egress are denied; "deny" — block all prompted ops; "allow" — run
// everything. The read_only default keeps headless/CI usage useful
// enough that flipping to "allow" is never the path of least
// resistance (H-7): under deny, an agent under a restrictive posture
// cannot even `ls`, and containment via inability just gets turned off.
NonInteractive *string `json:"non_interactive,omitempty"`
// Approver handles interactive approval prompts for dangerous operations.
// When set, all Prompt-class operations use this instead of /dev/tty.
// Tools can inject their own approver (e.g., WebSocket-based for odek serve).
// When nil, CheckOperation falls back to /dev/tty (CLI-compatible default).
Approver Approver `json:"-"`
}
DangerousConfig defines how dangerous operations are handled. Configurable via the standard 4-layer odek config chain.
Default behavior per class (no sandbox):
safe → allow, local_write → allow, system_write → prompt, destructive → deny, network_egress → prompt, code_execution → prompt, install → prompt, blocked → deny, unknown → deny
The classifier fails closed: a command whose program name is not recognised classifies as Unknown and is denied by default. Set "unknown": "prompt" (or add trusted tools to the allowlist) to soften this for a given profile.
func (*DangerousConfig) ActionFor ¶
func (c *DangerousConfig) ActionFor(cls RiskClass) Action
ActionFor returns the configured action for the given risk class. Per-class overrides in Classes win first, then the global default action (the "action" field), then built-in defaults, then Prompt.
func (*DangerousConfig) ActionForCommand ¶
func (c *DangerousConfig) ActionForCommand(cmd string) Action
ActionForCommand returns the action for a specific command string. Allowlist and denylist are checked first (exact match for allowlist, prefix match for denylist), then falls back to the risk-class-based action.
func (*DangerousConfig) CheckOperation ¶
func (c *DangerousConfig) CheckOperation(op ToolOperation, trustedClasses map[RiskClass]bool) error
CheckOperation checks whether a tool operation is allowed, denied, or needs approval. Returns nil on allow, error on deny, and prompts the user on prompt. Uses the configured Approver when set; falls back to /dev/tty (TTYApprover) when no approver is configured.
func (*DangerousConfig) NonInteractiveAction ¶
func (c *DangerousConfig) NonInteractiveAction() Action
NonInteractiveAction returns the action to use when no TTY is available.
Unset → ReadOnly (H-7): read-only inspection proceeds, every mutation fails closed — useful enough that flipping to "allow" is never the path of least resistance.
An explicitly set but INVALID value fails closed to Deny: a typo must never silently loosen the gate.
type InjectionPattern ¶
InjectionPattern groups a compiled regex with a human-readable label describing what threat it detects.
type RiskClass ¶
type RiskClass string
RiskClass represents the risk level of a shell command.
const ( Safe RiskClass = "safe" LocalWrite RiskClass = "local_write" SystemWrite RiskClass = "system_write" Persistence RiskClass = "persistence" Destructive RiskClass = "destructive" NetworkEgress RiskClass = "network_egress" CodeExecution RiskClass = "code_execution" Install RiskClass = "install" Blocked RiskClass = "blocked" // Unknown is the fall-through class for a command whose program name the // classifier does not recognise. It defaults to Deny (same as // Destructive): the gate fails CLOSED rather than open, so a novel or // obfuscated verb that dodged every known-dangerous check cannot run // unprompted. Recognised-but-benign usage classifies as Safe instead. Unknown RiskClass = "unknown" )
const UnreadExec RiskClass = "unread_exec"
UnreadExec is the class for executing a script file that has not been read (or written) in this session. Same rank tier as SystemWrite: always prompts by default, never eligible for session trust shortcuts.
func Classify ¶
Classify determines the risk class of a shell command using token-level heuristics. Returns the highest-severity class detected.
Priority (highest to lowest): blocked > destructive > system_write > code_execution > network_egress > install > local_write > safe
Pipeline (see the package doc for the full evasion model):
raw cmd ─▶ isRawBlocked ─▶ normalize ─┬─▶ classifyOne(main) ─┐
└─▶ Classify(sub) ⟳ ───┴─▶ worst wins
normalize neutralises shell evasion tricks (ANSI-C/$IFS/brace expansion, $(…)/`…`/<(…) substitutions, command/exec wrappers, backslash escapes, absolute-path basenames) and returns the rewritten command plus any substitution bodies. classifyOne then splits into segments and pipe stages and classifies each (see classifyPipeline/classifyStage). Every extracted sub-expression is re-classified through Classify so nested commands cannot hide one level deeper; the worst class across the whole tree is returned.
func ClassifyPath ¶
ClassifyPath returns a RiskClass for a filesystem path.
Classification rules (highest wins):
- /boot, /dev, /proc, /sys, /mnt, /media → destructive
- / (the filesystem root itself) → system_write
- /tmp, $TMPDIR → local_write
- /etc, /root, /var, /run, /lib, /usr, /bin, /sbin, /opt, /srv → system_write
- $HOME/.ssh, .config, .gnupg, .aws, .kube, .docker, .gitconfig, .env → system_write
- $HOME/.odek/config.json, secrets.env, IDENTITY.md, skills/, sessions/, audit/, plans/, schedules.json, schedule-state.json, mcp_approvals.json, mcp_tool_approvals.json, restart.json, telegram.lock, etc. → system_write (odek trust anchors; rewriting them can disable the sandbox, persist attacker control, or leak secrets)
- $HOME shell rc/profile files (.bashrc, .zshrc, .profile, .zshenv, etc.) → system_write
- everything else → local_write
macOS: /private/{etc,var,tmp} are transparently normalised before matching.
func ClassifyPathWrite ¶ added in v1.27.0
ClassifyPathWrite classifies a filesystem WRITE target. It wraps ClassifyPath and additionally escalates deferred-execution targets to Persistence (rank above SystemWrite, default action Prompt, never eligible for session trust shortcuts). Reads keep using ClassifyPath — reading a CI workflow file must stay as frictionless as before.
func ClassifyScriptGate ¶ added in v1.27.0
ClassifyScriptGate classifies cmd with the unread-script rule layered on top of the standard classifier: when an unread script executes, the class becomes UnreadExec for everything at or below the SystemWrite tier — so "code_execution": "allow" and trusted-class grants cannot bypass it. Stronger findings (persistence, unknown, destructive, blocked) keep their own class; they already gate harder and are never trust-shortcuttable.
func ClassifyURL ¶
ClassifyURL returns a RiskClass for a browser URL. Internal IPs → system_write; external → network_egress. Uses proper IP parsing (handles decimal, octal, hex, IPv6 compressed, short forms like 127.1, and all other representations that browsers accept via inet_aton-style parsing) instead of string prefix matching which was trivially bypassable.
func LifecycleContentClass ¶ added in v1.27.0
LifecycleContentClass inspects content written to path for lifecycle hooks (H-5). It returns (Persistence, true) when the content plants deferred execution into package.json / conftest.py, else (base, false). Only ever escalates — base is returned unchanged otherwise.
type ScanResult ¶
type ScanResult struct {
Label string // human-readable threat label
Pattern string // the regexp pattern that matched (for debugging)
}
ScanResult describes a single detected injection threat.
func ScanInjection ¶
func ScanInjection(content string) []ScanResult
ScanInjection checks content for prompt injection attempts. Returns nil if no threats detected, or a list of found threats. Each threat includes a label describing what was found.
type TTYApprover ¶
type TTYApprover struct {
DangerousConfig *DangerousConfig
TrustedClasses map[RiskClass]bool
TTYPath string // overridden in tests
// Approval-fatigue mitigation. After FrictionThreshold approvals of
// the same class within FrictionWindow, the next prompt requires
// the user to type the literal word "approve" (no single-letter
// shortcut) and prints a 1.5s pause before accepting input. This
// breaks reflexive click-through and gives the user a moment to
// notice they have approved an unusual number of dangerous calls.
FrictionThreshold int
FrictionWindow time.Duration
// contains filtered or unexported fields
}
TTYApprover implements Approver by reading from /dev/tty. This is the default approver used in CLI mode (odek run, odek repl). When /dev/tty is not available (piped stdin, CI), it falls back to the configured NonInteractiveAction.
func NewTTYApprover ¶
func NewTTYApprover(cfg *DangerousConfig) *TTYApprover
NewTTYApprover creates a TTYApprover with the given config.
func (*TTYApprover) PromptCommand ¶
func (a *TTYApprover) PromptCommand(cls RiskClass, cmd, description string) error
func (*TTYApprover) PromptOperation ¶
func (a *TTYApprover) PromptOperation(op ToolOperation) error
func (*TTYApprover) SetTrustAll ¶
func (a *TTYApprover) SetTrustAll(enabled bool)
SetTrustAll enables or disables blanket trust for all risk classes. When enabled, PromptCommand returns nil for every call (used by batch approval).
func (*TTYApprover) SetTrustedClasses ¶
func (a *TTYApprover) SetTrustedClasses(m map[RiskClass]bool)
SetTrustedClasses atomically sets the trusted classes map. Takes ownership of the provided map — caller must not write to it after calling.
type ToolOperation ¶
ToolOperation describes a native tool call for approval checking.