kubectl-gpugo

command module
v0.1.4 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: May 30, 2026 License: MIT Imports: 12 Imported by: 0

README

kubectl-gpugo

A top-like TUI for per-pod GPU usage on Kubernetes — zero cluster footprint.

kubectl-gpugo auto-discovers any GPU metrics exporter you already have running and scrapes it through the kube-apiserver pod-proxy subresource. No DaemonSets installed, no port-forwards on your machine, no firewall changes. Works with MIG.

Out of the box it understands two kinds of exporters:

  • dcgm-exporter — NVIDIA's standard GPU metrics exporter (from the GPU Operator or the standalone chart).
  • Per-process GPU exporters — anything that emits gpu_process_memory_bytes with pod/namespace/container labels. Used when workloads bypass the device plugin via NVIDIA_VISIBLE_DEVICES=all and dcgm-exporter can't attribute by itself.

What it shows

Column Meaning
NAMESPACE Workload pod's namespace
POD Workload pod consuming the GPU
NODE Node hosting the exporter that reported the metric
GPU GPU index(es) used — 2 (single), 0,1 (two cards), 0:8 (MIG slice 8 of GPU 0)
GPU% Activity across the pod's GPUs (colored by intensity)
VRAM USED Used / Total framebuffer, summed across the pod's GPUs / slices
POWER Power draw in Watts (proportional share when GPUs are shared across pods)
image

Rows are sorted by physical GPU, then by MIG slice ID. A blank line separates each physical card so pile-ups are obvious at a glance.

Requirements

On the cluster
  • An NVIDIA GPU exporter the tool recognises:
    • dcgm-exporter — image name contains dcgm-exporter, or its /metrics emits DCGM_FI_DEV_* family names.
    • A per-process GPU exporter — anything whose /metrics emits gpu_process_memory_bytes with pod/namespace/container labels. Tools that produce this format work without code changes.
    • Use --exporter to point at any pod by namespace/name:port if image-name and metric-content auto-detection don't catch your setup.
  • For per-pod attribution from dcgm-exporter alone: dcgm-exporter must run with --kubernetes enabled, so metrics carry namespace/pod labels. The NVIDIA GPU Operator does this by default. Without it, you fall back to per-(node, GPU) rows.
  • For workloads using NVIDIA_VISIBLE_DEVICES=all to bypass the device plugin, dcgm-exporter can't attribute (kubelet's pod-resources API knows nothing). For that case you need a per-process exporter deployed alongside.
  • MIG: dcgm-exporter on MIG-configured cards emits DCGM_FI_PROF_GR_ENGINE_ACTIVE per slice instead of DCGM_FI_DEV_GPU_UTIL. We handle both.
Your kubeconfig user / service account needs
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: kubectl-gpugo-reader
rules:
- apiGroups: [""]
  resources: ["pods"]
  verbs: ["list", "get"]
- apiGroups: [""]
  resources: ["pods/proxy"]
  verbs: ["get"]

The list pods on all namespaces is for auto-discovery. The pods/proxy is what actually fetches /metrics. If your account can't be granted cluster-wide list pods, use the --exporter flag (below) to scope the tool to a single pod you know about — that only needs pods/proxy in the exporter's namespace.

Won't work with
  • Lens / freelens kubeconfigs — their local proxy doesn't pass pod-proxy subresource requests through. Use a direct kubeconfig.
  • Restricted RBAC tokens that can't get pods/proxy. There's no way around this: the entire scrape path goes through that subresource.

Install

If you already have krew installed, this is a one-liner:

kubectl krew install gpugo
kubectl gpugo

That's it. Krew downloads the right binary for your OS/arch, sha256-verifies it, drops it on $PATH, and kubectl krew upgrade keeps it current along with every other plugin.

Don't have krew yet?

macOS (Homebrew):

brew install krew
echo 'export PATH="${KREW_ROOT:-$HOME/.krew}/bin:$PATH"' >> ~/.zshrc
source ~/.zshrc

macOS / Linux (no Homebrew):

(
  set -x; cd "$(mktemp -d)" &&
  OS="$(uname | tr '[:upper:]' '[:lower:]')" &&
  ARCH="$(uname -m | sed -e 's/x86_64/amd64/' -e 's/\(arm\)\(64\)\?.*/\1\2/' -e 's/aarch64$/arm64/')" &&
  KREW="krew-${OS}_${ARCH}" &&
  curl -fsSLO "https://github.com/kubernetes-sigs/krew/releases/latest/download/${KREW}.tar.gz" &&
  tar zxvf "${KREW}.tar.gz" &&
  ./"${KREW}" install krew
)
echo 'export PATH="${KREW_ROOT:-$HOME/.krew}/bin:$PATH"' >> ~/.zshrc   # or ~/.bashrc
source ~/.zshrc

Windows has a separate installer.

Then run kubectl krew install gpugo from the recommended section above.

Alternative: from source
go install github.com/Tal-Naeh/kubectl-gpugo@latest
# put the resulting binary on your PATH, renamed to `kubectl-gpugo`
mv "$(go env GOPATH)/bin/kubectl-gpugo" /usr/local/bin/kubectl-gpugo
kubectl gpugo

Useful if you want to track main instead of tagged releases, or your environment can't reach github.com/kubernetes-sigs/krew.

Alternative: pre-built archive

Download the archive matching your OS/arch from the releases page, extract kubectl-gpugo, and drop it anywhere on $PATH.

How kubectl finds the plugin

kubectl discovers plugins by looking for any binary named kubectl-<something> on $PATH. The standard kubectl flags (--context, --kubeconfig, -n, etc.) are routed automatically via genericclioptions — they behave exactly like in vanilla kubectl.

Flags

Flag Purpose
--kubeconfig, --context Standard kubectl flags
--exporter ns/pod:port Skip auto-discovery and scrape a specific pod. Comma-separate or repeat the flag for multiple pods.
--dump Print raw /metrics from every auto-discovered exporter and exit. Useful for debugging label conventions.
--dump-pod ns/pod:port Print raw /metrics from one specific pod and exit.

Keys inside the TUI

Key Action
q / Ctrl-C / Esc Quit
r Force-refresh now
/ k Scroll up one row
/ j Scroll down
PgUp / b Page up
PgDn / Space / f Page down
Home / g Jump to top
End / G Jump to bottom

Local dev

go build ./...
./kubectl-gpugo

The tick interval is 20s; force a refresh with r.

Limitations / known issues

  • The tool scrapes through the apiserver pod-proxy subresource. Slow control planes (e.g. RKE2 on a busy DGX with many MIG slices) can take 5–10s per scrape. If the apiserver gets wedged, restart the dcgm-exporter pod.
  • Image-name classification is a short hardcoded list. Forked images with unusual names still get caught by the /metrics probe fallback, but if both the name AND the metric families look unusual the pod is skipped. Use --exporter as an override.
  • Power on shared GPUs (multiple pods on one card via MIG or NVIDIA_VISIBLE_DEVICES=all) is attributed proportionally to each pod's VRAM share — the most honest split DCGM data supports.

Documentation

Overview

kubectl-gpugo is a top-like TUI for per-pod GPU utilisation on a Kubernetes cluster. It discovers existing dcgm-exporter pods (no DaemonSet install, zero cluster footprint) and scrapes them through the kube-apiserver pod proxy subresource — no local port-forward, no firewall surprises.

Directories

Path Synopsis
internal
k8s
Package k8s discovers GPU-metrics-bearing pods in the cluster.
Package k8s discovers GPU-metrics-bearing pods in the cluster.
scraper
Package scraper turns a list of dcgm-exporter pods into a GPU snapshot.
Package scraper turns a list of dcgm-exporter pods into a GPU snapshot.
tui
Package tui is the Bubbletea TUI for kubectl-gpugo.
Package tui is the Bubbletea TUI for kubectl-gpugo.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL