cli

package
v0.1.4 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 31, 2026 License: Apache-2.0 Imports: 29 Imported by: 0

Documentation

Overview

Package cli holds the cobra commands.

Index

Constants

View Source
const HeartbeatGrace = 30 * time.Second

HeartbeatGrace is how long a device's worker may go without heartbeating before it is treated as out of contact. It gates both the reaper's Sweep threshold below and rc devices' "no contact" annotation (internal/cli/ps.go) so the two thresholds cannot silently drift apart.

Variables

This section is empty.

Functions

func NewAttachCmd

func NewAttachCmd() *cobra.Command

NewAttachCmd re-streams a job's output from the beginning. It is read-only: unlike `rc run`, exiting attach (Ctrl-C or otherwise) never affects the job — there is no lease to protect and nothing to detach from, since attach never held anything in the first place. Ctrl-C is therefore the normal, expected way to stop watching: it exits quietly with the conventional SIGINT status instead of surfacing the raw "context canceled" plumbing error that StreamLogs returns once its request is aborted.

func NewClearCmd

func NewClearCmd() *cobra.Command

NewClearCmd returns a quarantined device to the pool.

The endpoint has existed since quarantine did, but only the dashboard ever called it, so clearing a device from a terminal meant hand-writing a curl with an admin token — which is exactly what happened on 2026-08-17 after a worker restart quarantined dgx:gpu0 mid-lease. Documentation told operators a human had to clear it and then gave them no command to do it with.

Deliberately NOT reachable with a client token: a device is quarantined because something went wrong, and deciding the hardware is healthy again is an operator's call, not a queued job's.

func NewCpCmd

func NewCpCmd() *cobra.Command

NewCpCmd is `rc cp`: move a file you name, onto a box you hold, once.

It is not a deployment tool. It does not sync, watch, reconcile or install, and anything recurring or large belongs on the NAS mount the worker already sees — every byte of a copy crosses the controller twice, and the controller is a scheduler with a one-connection database, not a file server.

func NewDescribeCmd

func NewDescribeCmd() *cobra.Command

func NewDevicesCmd

func NewDevicesCmd() *cobra.Command

func NewHoldCmd

func NewHoldCmd() *cobra.Command

NewHoldCmd claims a device for a human, not a job: "I need a shell on this box", not "run this command". Under the hood it is a job with kind "hold" whose command the worker itself chooses (a sleeper) — see internal/worker's execute — so it goes through the exact same allocation, queue, and expiry machinery as rc run, and shows up in rc ps and rc devices with no new rendering. See the task 8 design note for why this is deliberately not a second lease mechanism.

It blocks until granted, prints the device and when it expires, then stays attached: unlike rc run, Ctrl-C on a GRANTED hold releases it (see the WaitTerminal branch below) rather than merely detaching, because a hold's whole purpose is that a human is present, and leaving means they're done with it.

func NewJobsCmd

func NewJobsCmd() *cobra.Command

func NewKillCmd

func NewKillCmd() *cobra.Command

NewKillCmd cancels a queued job outright, or asks the worker to terminate a running one. The controller enforces ownership: only the submitter (or an admin token) may kill a given job.

func NewLogsCmd

func NewLogsCmd() *cobra.Command

NewLogsCmd reads output that the controller stored for a job. By default it returns the current snapshot; --follow replays that snapshot and then waits for new output until the job finishes.

func NewPsCmd

func NewPsCmd() *cobra.Command

func NewReleaseCmd

func NewReleaseCmd() *cobra.Command

NewReleaseCmd ends a hold — or, for that matter, any job — early. It is a thin alias over rc kill's ownership-checked path (see (*client.Client).Release), named separately because "release" is the verb that matches what rc hold prints and what a human reaches for once they're done with a device, not "kill".

func NewRetireCmd

func NewRetireCmd() *cobra.Command

NewRetireCmd removes a device from the fleet: the card was pulled, the box was decommissioned. It is the counterpart to a worker registering one, and it exists because the only alternative was editing the controller's database by hand.

func NewRunCmd

func NewRunCmd() *cobra.Command

func NewServeCmd

func NewServeCmd() *cobra.Command

func NewWorkerCmd

func NewWorkerCmd() *cobra.Command

func RenderDescribe

func RenderDescribe(w io.Writer, out *server.DescribeResponse) error

RenderDescribe prints everything an agent needs to trust (or distrust) a device before writing commands for it: the device and its state and holder; every label grouped by key with its source and age, a conflicting value shown rather than hidden; the usage sheet and its age; and recent job history with outcome and duration.

Every write goes through one tabwriter so a broken pipe or full disk on the final Flush is reported rather than silently dropped, matching RenderDevices and RenderJobs.

func RenderDevices

func RenderDevices(w io.Writer, views []server.DeviceView) error

RenderDevices prints the fleet as a table. Heartbeat age is shown for any device whose worker has gone quiet for longer than HeartbeatGrace, regardless of the device's own state — a device can be marked unhealthy, have the network partition heal, and start heartbeating again well before anyone runs `rc devices clear`; annotating it "no contact 0s" forever in that window would contradict the freshest information this command has.

It returns tabwriter's Flush error rather than discarding it: tabwriter buffers every row until Flush, so on a broken pipe (`rc devices | head`) or a full disk the table can be silently dropped while this function (and so `rc devices`) would otherwise still report success.

func RenderJobHistory

func RenderJobHistory(w io.Writer, jobs []model.Job, now time.Time) error

func RenderJobs

func RenderJobs(w io.Writer, state *server.StateResponse) error

RenderJobs prints the active jobs followed by the queue. Queued jobs are listed because `rc ps` is the only way to find a job's ID, and `rc kill` (which the README tells operators to reach for) needs one: a queued job nobody can name is a job nobody can cancel except by waiting for it to start first.

Position is computed here rather than fetched per job: state.Queued already arrives in scheduling order (priority, then FIFO), so counting along it per device gives the same 1-based, per-device position the controller reports for a single job — without one HTTP request per queued job.

It returns tabwriter's Flush error for the same reason RenderDevices does: the whole table is buffered until Flush, so a discarded error means a silently empty `rc ps` on a broken pipe.

Types

This section is empty.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL