kpi-collection-tool

module
v0.1.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 1, 2026 License: Apache-2.0

README

KPI Collection Tool

CLI tool to automate KPI collection on OpenShift clusters. Define what and how to collect in a single tasks.yaml file, point the tool at a cluster, and it takes care of the rest — querying Prometheus/Thanos, running node-level diagnostics, measuring latency, and capturing application recovery data.

What it does

kpi-collector runs tasks against an OpenShift cluster and produces KPI artifacts. You configure which tasks to run and how in a tasks.yaml file, then execute them with a single command:

kpi-collector run --tasks tasks.yaml \
  --cluster-name my-cluster --cluster-type ran \
  --kubeconfig ~/.kube/config

Four task types are available:

Task What it collects Output
prometheus PromQL metrics from Thanos SQLite/PostgreSQL database + Grafana
per-node-data top, /proc/meminfo, /proc/cmdline, pod & node describes Artifact files per node
oslat OS latency test results from a user-supplied pod oslat_logs.out
app-recovery-time Workload pod status after node reboot pod_status.out

See Validation Tasks Overview for a detailed introduction to each task, authentication requirements, and when to use which.

Tasks run sequentially by default. Orchestration (order, failure policy, parallel mode) is configurable in the tasks.yaml file.

All artifacts are stored under ./kpi-collector-artifacts/ by default (override with --artifacts-dir).

[!IMPORTANT] The per-node-data, oslat, and app-recovery-time tasks require --kubeconfig (they use the Kubernetes API to create pods, list nodes, etc.). The prometheus task can use either --kubeconfig (auto-discovers Thanos URL and creates a token) or --token + --thanos-url (manual). The kubeconfig must belong to a user with admin privileges — the non-prometheus tasks create privileged pods and list cluster nodes, and the prometheus task creates tokens for the prometheus-k8s service account in openshift-monitoring.

Who is this for

  • Telco partners validating KPIs on RAN, Core, or Hub OpenShift clusters
  • SREs and platform engineers who need to capture cluster health metrics over time for analysis or compliance
  • Anyone running OpenShift who wants a config-driven way to collect metrics and system-level data without scripting

Architecture

Architecture

Key features

  • Multi-task orchestration — run Prometheus collection, node diagnostics, latency tests, and recovery checks from one file
  • No cluster setup required — works with any OpenShift cluster out of the box
  • Kubeconfig auto-discovery — automatically finds Thanos and creates a service account token
  • Built-in KPI profiles — generate ready-to-use Prometheus KPI files for RAN, Core, or Hub clusters
  • Flexible storage — SQLite (default, zero config) or PostgreSQL for Prometheus metrics
  • Grafana dashboard — one command to launch a pre-configured dashboard for Prometheus data
  • Single binary — no runtime dependencies, runs on Linux and macOS
  • Ready-to-use task profiles — quickstart and full validation examples under task-profiles/

Quick start: multi-task run

Use a task profile to run all four tasks:

kpi-collector run \
  --tasks task-profiles/tasks-quickstart.yaml \
  --cluster-name my-cluster --cluster-type ran \
  --kubeconfig ~/.kube/config --once

⚠️ The app-recovery-time task reboots cluster nodes. Edit nodeNames in the recovery fragment before running, or remove it from the orchestration order.

Example: Prometheus-only collection

You can also run just the Prometheus task directly without a tasks.yaml:

kpi-collector run \
  --cluster-name my-cluster --cluster-type ran \
  --kubeconfig ~/.kube/config \
  --prom-kpis-config prom-kpi-profiles/kpis-quickstart.yaml --once

Query the collected data:

$ kpi-collector db show kpis --name node-cpu-usage --limit 3

ID  KPI_NAME        CLUSTER      VALUE   TIMESTAMP            LABELS
--- ---             ---          ---     ---                  ---
1   node-cpu-usage  my-cluster   0.0342  2026-04-12 14:30:00  {"instance":"worker-0"}
2   node-cpu-usage  my-cluster   0.0128  2026-04-12 14:30:00  {"instance":"worker-1"}
3   node-cpu-usage  my-cluster   0.0694  2026-04-12 14:30:00  {"instance":"master-0"}

Total results: 3
Grafana dashboard

Grafana Dashboard

The database, db show/db remove, and Grafana apply to the prometheus task only. Other tasks write artifact files under --artifacts-dir.

Command map
  • kpi-collector run --tasks <file>: run configured tasks (prometheus, per-node-data, oslat, app-recovery-time)
  • kpi-collector run --prom-kpis-config <file>: collect Prometheus metrics only (no tasks file needed)
  • kpi-collector kpis generate --profile <profile>: generate a Prometheus KPI file for a cluster profile (ran, core, hub)
  • kpi-collector db show: query collected data
  • kpi-collector db remove: remove stored data
  • kpi-collector grafana start|stop: manage local Grafana dashboard

Documentation

Task-specific configuration guides:

Prometheus task:

License

Apache License 2.0 — see LICENSE for details.

AI Skill for Cursor / Claude Code

An AI agent skill is included that teaches your coding assistant how to use kpi-collector and generate Telco-specific PromQL queries. See docs/ai-skill/ for installation and usage instructions.

Directories

Path Synopsis
cmd
kpi-collector command
internal
collector
Package collector orchestrates KPI metric collection from Prometheus/Thanos.
Package collector orchestrates KPI metric collection from Prometheus/Thanos.
config
Package config provides configuration types, validation, and loading for the KPI collector.
Package config provides configuration types, validation, and loading for the KPI collector.
database
Package database provides a unified interface for storing and retrieving KPI metrics.
Package database provides a unified interface for storing and retrieving KPI metrics.
database/schema
Package schema defines the database DDL (table creation, indexes, pragmas) for all supported backends.
Package schema defines the database DDL (table creation, indexes, pragmas) for all supported backends.
kubernetes
Package kubernetes provides integration with Kubernetes/OpenShift clusters.
Package kubernetes provides integration with Kubernetes/OpenShift clusters.
logger
Package logger provides file-based logging initialization for the KPI collector.
Package logger provides file-based logging initialization for the KPI collector.
output
Package output provides multi-format output rendering for CLI commands.
Package output provides multi-format output rendering for CLI commands.
prometheus
Package prometheus provides client functionality for querying Prometheus/Thanos metrics endpoints.
Package prometheus provides client functionality for querying Prometheus/Thanos metrics endpoints.
task
Package task defines runnable units executed by `kpi-collector run`.
Package task defines runnable units executed by `kpi-collector run`.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL