Documentation
¶
Overview ¶
Package onboard ingests historical logs from external sources and streams them to the codag server's /v1/onboard/warm endpoint to pre-populate the org's template_cache. The result: subsequent `codag wrap` / `codag compact` calls hit the regex template cache instead of paying the LLM templater + classifier round-trip on every novel line shape.
Architecture: each Source is a fetcher that streams log lines (already fetched from the source's CLI or REST API) into a Go channel. The Runner buffers those into batches and POSTs them to the server. Credentials live entirely client-side in ~/.config/codag/config.json (no server-side credential store).
Index ¶
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
func EstimateCost ¶
EstimateCost is a rough single-knob translation. The real cost depends on which templater backend the server is running (MLX local = $0, hosted = ~$0.001/call). We assume hosted to be conservative.
func FormatPlan ¶
FormatPlan renders the table shown to the user before warm.
Types ¶
type FetchOpts ¶
type FetchOpts struct {
// Since is the lookback window. Most sources translate this to a
// `--since=...` flag or an API query parameter. If a source can't
// honor an arbitrary duration (e.g. only supports 1h/1d/7d
// presets), it should round UP to the next supported value.
Since time.Duration
// MaxLines is a hard cap on lines emitted. Sources MUST stop fetching
// once they've emitted this many; the runner enforces a second cap
// downstream but doing so wastes API calls / bandwidth.
MaxLines int
// ScrubPII, when true, asks sources to mask emails / IPv4 / 32+-char
// hex secrets in line bodies before emitting. The runner re-applies
// the same scrub as a defense-in-depth step.
ScrubPII bool
}
FetchOpts is what every Source receives. Sources should respect both Since and MaxLines as best-effort bounds — exceeding either burns user money and triggers the server's per-org daily quota.
func DefaultFetchOpts ¶
func DefaultFetchOpts() FetchOpts
DefaultFetchOpts mirrors the plan: 7d window, 250k lines/source. The retained setup warm path overrides MaxLines down to 50k.
type Plan ¶
type Plan struct {
Source string
Available bool
Hint string // shown when Available=false
EstLines int // best-effort estimate (often == MaxLines if unknown)
EstLLMCalls int // estimated cold templater calls (~5% of lines)
EstUSD float64 // very rough; surfaced as "~$X" to set expectations
}
Plan is a per-source pre-flight: what we'd ingest if the user said yes. Used by warm flows and `codag onboard --dry-run` to print a summary table before charging the user's LLM credits.
type ReplayStats ¶
ReplayStats holds the post-warm /v1/parse hit rate for a source's captured sample. Empty if replay was skipped or failed.
type Runner ¶
type Runner struct {
Client *api.Client
BatchSize int // how many lines per POST (default 1000)
Workers int // parallel sources (default 3)
OnBatch func(source string, batch int, resp *api.WarmResponse) // optional progress callback
// SampleSize controls how many lines per source we capture for the
// post-warm replay that produces the "realized next-incident hit
// rate". 0 disables the replay (default 100). Replay runs through
// /v1/parse (read-only) so it doesn't add to the daily quota.
SampleSize int
}
Runner orchestrates the bulk warm. It fetches from each Source in parallel (capped concurrency) and POSTs batches of lines to /v1/onboard/warm. Per-source results land in Stats; the caller decides how to render them.
func (*Runner) Run ¶
func (r *Runner) Run(ctx context.Context, cfg *config.Config, sources []Source, opts FetchOpts) (*Stats, error)
Run executes onboarding for the given sources. Returns an aggregate Stats and the first error encountered (subsequent errors are recorded per-source in Stats but don't abort the whole run — partial successes still warm what they got).
type Source ¶
type Source interface {
// Name returns the canonical key (e.g. "vercel", "datadog"). Used
// in user-facing tables and as the `source` tag on warm requests.
Name() string
// Detect reports whether this source is reachable from the current
// machine — typically by checking that the underlying CLI binary is
// on PATH or that credentials exist in cfg.Providers. The hint is
// shown to the user when Detect returns false (e.g. "install the
// vercel CLI" or "run `codag setup` to configure Datadog keys").
Detect(cfg *config.Config) (available bool, hint string)
// Fetch begins fetching and returns a channel of raw log lines.
// Lines are pushed as soon as they're available; the channel closes
// when the source is exhausted, MaxLines is reached, or ctx is
// cancelled. An error is returned ONLY for setup failures (missing
// binary, bad credentials) — runtime errors mid-fetch should close
// the channel with what was collected so far.
Fetch(ctx context.Context, cfg *config.Config, opts FetchOpts) (<-chan string, error)
}
Source is one log origin (vercel, aws, datadog, ...). Implementations live in `cli/internal/onboard/sources/`.
func AvailableSources ¶
AvailableSources filters to those that Detect()-ed successfully.
type SourceStats ¶
type Stats ¶
type Stats struct {
// contains filtered or unexported fields
}
Stats accumulates per-source warm results. Concurrent-safe.
func (*Stats) Add ¶
func (s *Stats) Add(source string, ss SourceStats)
func (*Stats) AllSamples ¶
AllSamples returns a snapshot for the post-warm replay loop.