onboard

package
v0.1.10 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: May 24, 2026 License: MIT Imports: 10 Imported by: 0

Documentation

Overview

Package onboard ingests historical logs from external sources and streams them to the codag server's /v1/onboard/warm endpoint to pre-populate the org's template_cache. The result: subsequent `codag wrap` / `codag compact` calls hit the regex template cache instead of paying the LLM templater + classifier round-trip on every novel line shape.

Architecture: each Source is a fetcher that streams log lines (already fetched from the source's CLI or REST API) into a Go channel. The Runner buffers those into batches and POSTs them to the server. Credentials live entirely client-side in ~/.config/codag/config.json (no server-side credential store).

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

func EstimateCost

func EstimateCost(lines int) (llmCalls int, usd float64)

EstimateCost is a rough single-knob translation. The real cost depends on which templater backend the server is running (MLX local = $0, hosted = ~$0.001/call). We assume hosted to be conservative.

func FormatPlan

func FormatPlan(plans []Plan) string

FormatPlan renders the table shown to the user before warm.

func Register

func Register(s Source)

Register installs a source under its Name(). Last registration wins (useful for tests that want to substitute a mock).

func Scrub

func Scrub(line string) string

Scrub returns line with emails/IPs/long-hex/bearer-style tokens masked. Idempotent.

Types

type FetchOpts

type FetchOpts struct {
	// Since is the lookback window. Most sources translate this to a
	// `--since=...` flag or an API query parameter. If a source can't
	// honor an arbitrary duration (e.g. only supports 1h/1d/7d
	// presets), it should round UP to the next supported value.
	Since time.Duration

	// MaxLines is a hard cap on lines emitted. Sources MUST stop fetching
	// once they've emitted this many; the runner enforces a second cap
	// downstream but doing so wastes API calls / bandwidth.
	MaxLines int

	// ScrubPII, when true, asks sources to mask emails / IPv4 / 32+-char
	// hex secrets in line bodies before emitting. The runner re-applies
	// the same scrub as a defense-in-depth step.
	ScrubPII bool
}

FetchOpts is what every Source receives. Sources should respect both Since and MaxLines as best-effort bounds — exceeding either burns user money and triggers the server's per-org daily quota.

func DefaultFetchOpts

func DefaultFetchOpts() FetchOpts

DefaultFetchOpts mirrors the plan: 7d window, 250k lines/source. The retained setup warm path overrides MaxLines down to 50k.

type Plan

type Plan struct {
	Source      string
	Available   bool
	Hint        string  // shown when Available=false
	EstLines    int     // best-effort estimate (often == MaxLines if unknown)
	EstLLMCalls int     // estimated cold templater calls (~5% of lines)
	EstUSD      float64 // very rough; surfaced as "~$X" to set expectations
}

Plan is a per-source pre-flight: what we'd ingest if the user said yes. Used by warm flows and `codag onboard --dry-run` to print a summary table before charging the user's LLM credits.

func PlanFor

func PlanFor(sources []Source, cfg *config.Config, opts FetchOpts) []Plan

Plan computes per-source pre-flight info. Cheap (no actual fetching).

type ReplayStats

type ReplayStats struct {
	CacheHits  int
	SampleSize int
}

ReplayStats holds the post-warm /v1/parse hit rate for a source's captured sample. Empty if replay was skipped or failed.

type Runner

type Runner struct {
	Client    *api.Client
	BatchSize int                                                    // how many lines per POST (default 1000)
	Workers   int                                                    // parallel sources (default 3)
	OnBatch   func(source string, batch int, resp *api.WarmResponse) // optional progress callback
	// SampleSize controls how many lines per source we capture for the
	// post-warm replay that produces the "realized next-incident hit
	// rate". 0 disables the replay (default 100). Replay runs through
	// /v1/parse (read-only) so it doesn't add to the daily quota.
	SampleSize int
}

Runner orchestrates the bulk warm. It fetches from each Source in parallel (capped concurrency) and POSTs batches of lines to /v1/onboard/warm. Per-source results land in Stats; the caller decides how to render them.

func NewRunner

func NewRunner(client *api.Client) *Runner

NewRunner returns a Runner with the plan's recommended defaults.

func (*Runner) Run

func (r *Runner) Run(ctx context.Context, cfg *config.Config, sources []Source, opts FetchOpts) (*Stats, error)

Run executes onboarding for the given sources. Returns an aggregate Stats and the first error encountered (subsequent errors are recorded per-source in Stats but don't abort the whole run — partial successes still warm what they got).

type Source

type Source interface {
	// Name returns the canonical key (e.g. "vercel", "datadog"). Used
	// in user-facing tables and as the `source` tag on warm requests.
	Name() string

	// Detect reports whether this source is reachable from the current
	// machine — typically by checking that the underlying CLI binary is
	// on PATH or that credentials exist in cfg.Providers. The hint is
	// shown to the user when Detect returns false (e.g. "install the
	// vercel CLI" or "run `codag setup` to configure Datadog keys").
	Detect(cfg *config.Config) (available bool, hint string)

	// Fetch begins fetching and returns a channel of raw log lines.
	// Lines are pushed as soon as they're available; the channel closes
	// when the source is exhausted, MaxLines is reached, or ctx is
	// cancelled. An error is returned ONLY for setup failures (missing
	// binary, bad credentials) — runtime errors mid-fetch should close
	// the channel with what was collected so far.
	Fetch(ctx context.Context, cfg *config.Config, opts FetchOpts) (<-chan string, error)
}

Source is one log origin (vercel, aws, datadog, ...). Implementations live in `cli/internal/onboard/sources/`.

func All

func All() []Source

All returns every registered source, sorted by name.

func AvailableSources

func AvailableSources(cfg *config.Config) []Source

AvailableSources filters to those that Detect()-ed successfully.

func Get

func Get(name string) Source

Get returns one source by name, or nil if not registered.

type SourceStats

type SourceStats struct {
	LinesProcessed int
	TemplatesAdded int
	LLMCalls       int
	CacheHits      int
	Batches        int
	Duration       time.Duration
	Err            error
}

type Stats

type Stats struct {
	// contains filtered or unexported fields
}

Stats accumulates per-source warm results. Concurrent-safe.

func NewStats

func NewStats() *Stats

func (*Stats) Add

func (s *Stats) Add(source string, ss SourceStats)

func (*Stats) AllSamples

func (s *Stats) AllSamples() map[string][]string

AllSamples returns a snapshot for the post-warm replay loop.

func (*Stats) Report

func (s *Stats) Report() string

Report renders a one-screen summary of the run for the user.

func (*Stats) SetReplay

func (s *Stats) SetReplay(source string, cacheHits, sampleSize int)

SetReplay records the realized post-warm hit rate for a source.

func (*Stats) SetSample

func (s *Stats) SetSample(source string, sample []string)

SetSample records the captured sample for a source. Caller is the runner; called once per source after its warm completes.

Directories

Path Synopsis
Package sources contains the per-origin log fetchers used by `codag onboard`.
Package sources contains the per-origin log fetchers used by `codag onboard`.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL