Documentation
¶
Overview ¶
Package serverlifecycle restarts and upgrades the server as durable, idempotent operations: JSON records on disk that outlive the process they act on, so a caller can hand off an operation id and ask about it later.
Index ¶
- Constants
- Variables
- func ExecutorBinary() (string, error)
- func HealthURL(mode, address string) string
- func NewID() string
- func SameBuild(versionA, commitA, versionB, commitB string) bool
- func UnitActive(ctx context.Context, unit string) bool
- func UnitName(opID string) string
- func ValidateID(id string) error
- type Action
- type ContainerLauncher
- type ContainerRestarter
- type DataBackup
- type DataRestore
- type Executor
- func (e *Executor) Run(ctx context.Context, id string) (*Operation, error)
- func (e *Executor) WithDataBackup(b DataBackup) *Executor
- func (e *Executor) WithDownloader(d release.Downloader) *Executor
- func (e *Executor) WithInstaller(i release.Installer) *Executor
- func (e *Executor) WithProber(p Prober) *Executor
- func (e *Executor) WithRestarter(r Restarter) *Executor
- type HealthProber
- type InstanceProber
- type LaunchChecker
- type Launcher
- type NodeStep
- type Operation
- func Abandon(store *Store, id string) (*Operation, error)
- func AwaitAdoption(ctx context.Context, store *Store, id string, timeout, interval time.Duration) (*Operation, error)
- func NewOperation(action Action, requestedBy string) *Operation
- func Start(ctx context.Context, store *Store, launcher Launcher, op *Operation) (*Operation, bool, error)
- type Options
- type Phase
- type Prober
- type Restarter
- type RestoreResult
- type Snapshot
- type Store
- func (s *Store) Active() (*Operation, error)
- func (s *Store) Create(op *Operation) error
- func (s *Store) CreateOrGet(op *Operation) (existing *Operation, created bool, err error)
- func (s *Store) Dir() string
- func (s *Store) Get(id string) (*Operation, error)
- func (s *Store) List() ([]*Operation, error)
- func (s *Store) LockOperation(id string) (func(), error)
- func (s *Store) PendingRestore() (*Operation, error)
- func (s *Store) ReadRestoreResult(id string) (*RestoreResult, error)
- func (s *Store) RecordRestoreAttempt(result *RestoreResult) (*RestoreResult, error)
- func (s *Store) Remove(id string) error
- func (s *Store) Update(op *Operation) error
- func (s *Store) WriteRestoreResult(result *RestoreResult) error
- type SystemdLauncher
- type SystemdRestarter
- type UnsupervisedLauncher
Constants ¶
const ( DefaultDir = "/var/lib/miren/server/lifecycle" RunnerDir = "/var/lib/miren/runner/lifecycle" )
DefaultDir is the server's ledger; RunnerDir is the runner's. They are separate ledgers with separate busy slots: a host that runs both daemons (the coordinator in standalone mode does not, but nothing forbids it) can have one operation on each.
const AdoptionTimeout = 2 * time.Minute
AdoptionTimeout is how long a handed-off operation may wait for a server to take it over before the hand-off is treated as failed.
const DefaultExitAfter = 8 * time.Minute
DefaultExitAfter is past the server's own shutdown timeout, so it only ever fires on a stop that is stuck rather than slow.
const DefaultHealthURL = "https://127.0.0.1:443/.well-known/miren/health"
Variables ¶
var ( ServerExecutorCommand = []string{"server", "operations", "run"} RunnerExecutorCommand = []string{"runner", "operations", "run"} )
The executor entrypoints, one per ledger. Each opens its own daemon's ledger and probes its own daemon, so an operation id alone does not say which one it belongs to; the launcher has to.
var ErrBusy = errors.New("another lifecycle operation is in progress")
var ErrInvalidID = errors.New("lifecycle operation id must be a ULID")
ErrInvalidID is returned for an operation id that is not a ULID. The id is a file name and the ledger's sort key, so the format is not negotiable.
var ErrLauncherStopping = errors.New("server is shutting down; try again once it is back")
ErrLauncherStopping is returned by Launch once the server is going down.
var ErrLocked = errors.New("lifecycle operation is being executed by another process")
ErrLocked is returned when another executor already holds an operation.
var ErrNotFound = errors.New("lifecycle operation not found")
var ErrNothingToAbandon = errors.New("operation is finished and left nothing pending")
ErrNothingToAbandon is returned by Abandon for an operation that is finished and left nothing pending.
var ErrUnsupervised = errors.New("this server is not supervised in a way that can restart it: install it as a systemd service, or run the container image through `miren server container install` so its restart policy brings it back")
ErrUnsupervised is what a restart or upgrade gets on a server nothing can bring back: not a systemd unit, and not a container that booted through the image's entrypoint.
Functions ¶
func ExecutorBinary ¶
ExecutorBinary is the binary a launcher should run the executor from: this process's own, resolved through any symlink so the transient unit captures the real file. Running our own build rather than the installed server's matters when they differ: an older server may predate the executor entirely.
func HealthURL ¶
HealthURL is where a server with this ingress mode and address answers /.well-known/miren/health from its own host.
func SameBuild ¶
SameBuild reports whether two builds are the same: commits decide when both are known, otherwise version strings.
func UnitActive ¶
UnitActive reports whether a unit is running or starting.
func ValidateID ¶
ValidateID checks that id is a ULID as NewID would mint one.
Types ¶
type ContainerLauncher ¶
type ContainerLauncher struct {
// NewExecutor builds an executor over the store; it is called once per
// launch or resume so each run has its own downloaded-artifact state.
NewExecutor func() (*Executor, error)
Log *slog.Logger
// contains filtered or unexported fields
}
ContainerLauncher runs the executor inside the server process. Nothing outlives the server in a container, so the executor dies with it at the restart it asked for. That is fine: the operation record in the volume is the checkpoint, and the next instance's Resume picks it up where it was, verifying itself as the new instance.
func NewContainerLauncher ¶
func NewContainerLauncher(ctx context.Context, log *slog.Logger, newExecutor func() (*Executor, error)) *ContainerLauncher
NewContainerLauncher binds runs to ctx. Wait returns once every run launched under it has returned.
func (*ContainerLauncher) Launch ¶
func (l *ContainerLauncher) Launch(_ context.Context, opID string) error
func (*ContainerLauncher) Resume ¶
func (l *ContainerLauncher) Resume(ctx context.Context, store *Store) error
Resume relaunches the operation the previous instance left unfinished, if there is one. It answers the question a systemd install never has to ask: the transient unit there outlives the restart, but here the executor was the process that just exited.
func (*ContainerLauncher) Wait ¶
func (l *ContainerLauncher) Wait()
Wait refuses further launches and blocks until every launched run has returned. Runs return promptly once the launcher's context is cancelled.
type ContainerRestarter ¶
type ContainerRestarter struct {
// Shutdown asks this process to stop. The default sends SIGTERM to
// itself, which the server handles the same way as one from outside: the
// boot graph's stop path takes the nested stack down cleanly.
Shutdown func() error
// ExitAfter bounds the graceful stop: a restart that never finishes
// stopping would leave the container running the build it was meant to
// replace, with no supervisor to notice. 0 means DefaultExitAfter.
ExitAfter time.Duration
}
ContainerRestarter restarts a container install by ending the process. There is no supervisor inside the container to ask; the one outside it (the container runtime's restart policy) starts a new container when this one exits, and container-boot in the image execs whatever the volume's release directory holds by then.
type DataBackup ¶
type DataBackup interface {
// Backup takes a snapshot for the operation and returns its reference.
Backup(ctx context.Context, opID string) (string, error)
}
DataBackup snapshots the server's data before an upgrade. The executor does not know what the data is: it stores the reference it gets back on the operation, and on rollback hands it to the restarted server, which does know how to put it back before serving.
type DataRestore ¶
type DataRestore struct {
BackupRef string `json:"backup_ref"`
// ForVersion and ForCommit name the build the request is for: the one
// being rolled back to. The request is persisted before the binary is
// swapped, and the build being rolled back from may still be crash
// looping under systemd at that point. If it honored the request it
// would restore, migrate the data again, and fail, leaving a settled
// request and migrated data for the build that actually needed it.
ForVersion string `json:"for_version,omitempty"`
ForCommit string `json:"for_commit,omitempty"`
RestoredAt *time.Time `json:"restored_at,omitempty"`
Error string `json:"error,omitempty"`
}
DataRestore is the restore request a rollback records on the operation, and its outcome.
func (*DataRestore) MeantFor ¶
func (r *DataRestore) MeantFor(version, commit string) bool
MeantFor reports whether the build identified by version and commit is the one this request is for. A request that could not name a build (the server was unreachable when the operation began) is for whoever boots.
type Executor ¶
type Executor struct {
// contains filtered or unexported fields
}
Executor drives an operation through its phases, persisting each transition. Everything needed to resume is in the Operation record.
func (*Executor) Run ¶
Run drives the operation to a terminal phase, to the hand-off to the server, or until ctx ends. A finished operation is returned unchanged; an interrupted one resumes where it was.
func (*Executor) WithDataBackup ¶
func (e *Executor) WithDataBackup(b DataBackup) *Executor
func (*Executor) WithDownloader ¶
func (e *Executor) WithDownloader(d release.Downloader) *Executor
func (*Executor) WithProber ¶
func (*Executor) WithRestarter ¶
type HealthProber ¶
type HealthProber struct {
URL string
// ServerName is the SNI to send. Empty works with the default autocert
// ingress, which answers SNI-less handshakes with its self-signed cert.
ServerName string
// contains filtered or unexported fields
}
HealthProber reads the "server" block of /.well-known/miren/health.
func NewHealthProber ¶
func NewHealthProber(url, serverName string) *HealthProber
NewHealthProber accepts any certificate: the point is to reach the process on this host, not to authenticate it.
type InstanceProber ¶
type InstanceProber struct {
Source *serverinfo.Source
}
InstanceProber reads this process's own serverinfo instead of asking the health endpoint. Inside the container the executor and the server are the same process, so there is nothing to reach over the network, and the answer is available before ingress is listening.
type LaunchChecker ¶
LaunchChecker is an optional Launcher capability: after Launch reported an error, it says whether the executor is running anyway. systemd-run can fail after the unit was submitted, and an executor that is running owns the record, so the caller must not mark it failed on top of it.
type NodeStep ¶
type NodeStep struct {
Name string `json:"name"`
RunnerID string `json:"runner_id"`
// OperationID is the runner-side operation, minted before the request
// so a repeated request finds it rather than starting another.
OperationID string `json:"operation_id,omitempty"`
// Phase is the runner operation's phase as last observed, PhasePending
// before the runner was asked, or StepSkipped.
Phase Phase `json:"phase"`
Error string `json:"error,omitempty"`
Progress string `json:"progress,omitempty"`
PreviousVersion string `json:"previous_version,omitempty"`
NewVersion string `json:"new_version,omitempty"`
// Cordoned records that the walk cordoned the node for this step, so a
// resumed walk knows to uncordon it. An operator's cordon is left alone.
Cordoned bool `json:"cordoned,omitempty"`
StartedAt *time.Time `json:"started_at,omitempty"`
FinishedAt *time.Time `json:"finished_at,omitempty"`
}
NodeStep is one runner's part of an upgrade. The runner runs its own operation in its own ledger; the step mirrors what the server last saw of it, so the cluster record stands on its own.
type Operation ¶
type Operation struct {
ID string `json:"id"`
Action Action `json:"action"`
RequestedBy string `json:"requested_by,omitempty"`
// TargetVersion is as requested ("latest", "main", a tag); ResolvedVersion
// and ResolvedCommit are what it became.
TargetVersion string `json:"target_version,omitempty"`
ResolvedVersion string `json:"resolved_version,omitempty"`
ResolvedCommit string `json:"resolved_commit,omitempty"`
// ArtifactType ("base" or "release"), NoRollback, and ReadyTimeoutSeconds
// override the executor defaults for one upgrade.
ArtifactType string `json:"artifact_type,omitempty"`
NoRollback bool `json:"no_rollback,omitempty"`
ReadyTimeoutSeconds int `json:"ready_timeout_seconds,omitempty"`
Phase Phase `json:"phase"`
Error string `json:"error,omitempty"`
// BackupRef names the data snapshot taken before an upgrade; empty when
// none was taken. DataRestore is set on rollback and asks the restarted
// server to put BackupRef back before it serves data; the server's answer
// lands in it once the executor has read the RestoreResult.
BackupRef string `json:"backup_ref,omitempty"`
DataRestore *DataRestore `json:"data_restore,omitempty"`
// Progress is a short note for the current phase, e.g. a download percentage.
Progress string `json:"progress,omitempty"`
// DrivenBy is the server instance that took the operation over for the
// runner phase, and Nodes is one step per runner it found. Both are set
// together, so DrivenBy on a record with no Nodes means a cluster with
// nothing to walk.
DrivenBy string `json:"driven_by,omitempty"`
Nodes []*NodeStep `json:"nodes,omitempty"`
// A successful operation ends with NewInstanceID != PreviousInstanceID;
// that change is how we know the restart actually happened.
PreviousInstanceID string `json:"previous_instance_id,omitempty"`
PreviousVersion string `json:"previous_version,omitempty"`
PreviousCommit string `json:"previous_commit,omitempty"`
NewInstanceID string `json:"new_instance_id,omitempty"`
NewVersion string `json:"new_version,omitempty"`
// Components are the runtime versions (containerd, runc, ...) the server
// reported once it was up. A base upgrade replaces those binaries next to
// miren, and this is the record that the restarted server is on them.
Components map[string]string `json:"components,omitempty"`
// RollbackFrom names the instance that asked for the rollback's restart,
// so a resumed rollback can tell a restart that already took (the
// instance answering now is a different one) from one still to do. The
// systemd executor outlives the restart and never resumes; the container
// executor is that instance and dies at the restart it asks for.
// container-boot, rolling back a build that never answered, writes its
// own name.
RollbackFrom string `json:"rollback_from,omitempty"`
// BootAttempts counts container boots of the new build since the restart
// phase, kept by container-boot; it is what catches a build that crashes
// before the executor inside it can run.
BootAttempts int `json:"boot_attempts,omitempty"`
CreatedAt time.Time `json:"created_at"`
UpdatedAt time.Time `json:"updated_at"`
FinishedAt *time.Time `json:"finished_at,omitempty"`
}
Operation is the durable record of one restart or upgrade.
func Abandon ¶
Abandon is the operator's way out of an operation nobody will finish: an executor that died mid-operation, or a rollback whose data restore keeps failing and keeps the server from booting. It marks an unfinished operation failed and settles a pending restore request as abandoned, so the next boot starts on the data as it is. It refuses while an executor still holds the operation.
func AwaitAdoption ¶
func AwaitAdoption(ctx context.Context, store *Store, id string, timeout, interval time.Duration) (*Operation, error)
AwaitAdoption waits for the server to take over a handed-off operation. The executor cannot know whether the build it just verified drives runner upgrades: a target older than that support would leave the record in upgrading_runners with nobody to finish it. If nothing has claimed it by the deadline, the operation is failed with that diagnosis. The wait polls rather than holds the operation lock, since the server needs the lock to adopt it.
func NewOperation ¶
func Start ¶
func Start(ctx context.Context, store *Store, launcher Launcher, op *Operation) (*Operation, bool, error)
Start records op and launches its executor. It is closed over the id: a second Start with an id already on disk returns that record untouched, so a caller that lost the first reply (the reply may well be lost, since the operation restarts the server answering it) can retry without starting a second operation. Two such calls arriving together are settled under the store's lock, so exactly one creates and the other gets the record. A record that exists with a different action or target is a caller bug and is refused rather than silently returned. The bool reports whether this call is the one that started it.
The launch runs on a context cut loose from the caller's. The caller is an RPC or an uplink session about to be torn down by the very restart it asked for, and a launch cancelled midway is the worst outcome: the unit may be running while the caller believes it is not.
type Options ¶
type Options struct {
ServiceName string
// Daemon names what is being restarted in progress notes and errors:
// "server" or "runner".
Daemon string
// StateDir is the daemon's state directory, passed through to the
// resource-limit refresh that precedes a restart.
StateDir string
InstallPath string
TempDir string
ArtifactType release.ArtifactType
// ReadyTimeout bounds how long a restarted server gets to report ready
// before the operation fails (and, for upgrades, rolls back).
ReadyTimeout time.Duration
ProbeInterval time.Duration
AutoRollback bool
// PathSymlink, when set, is kept pointing at InstallPath after a
// successful upgrade so the CLI on $PATH tracks the server.
PathSymlink string
// UpgradeRunners hands a verified upgrade to the daemon for the
// upgrading_runners phase instead of finishing it. On for the server,
// which has runners; off for a runner, which is one.
UpgradeRunners bool
}
Options tunes an Executor. Start from DefaultOptions; a zero Options is not a working configuration.
func DefaultOptions ¶
func DefaultOptions() Options
func RunnerOptions ¶
func RunnerOptions() Options
RunnerOptions is DefaultOptions for the runner daemon: its unit, its state directory, and no data backup, since a runner keeps no data of its own that a rollback would need to put back. The prober is the caller's to set; a runner has no health URL.
type Phase ¶
type Phase string
Phase is a checkpoint: an executor resumes from the recorded phase, and each phase's work is safe to repeat. The list is meant to grow: the pre-upgrade backup today is an etcd snapshot, and the full RFD-75 bundle slots into the same phase and BackupRef.
const ( PhasePending Phase = "pending" PhaseDownloading Phase = "downloading" // PhaseBackingUp sits after the download so the snapshot is as fresh as // possible when the restart happens, and so a failed download or an // already-installed target never costs a snapshot. PhaseBackingUp Phase = "backing_up" PhaseInstalling Phase = "installing" PhaseRestarting Phase = "restarting" PhaseVerifying Phase = "verifying" // PhaseUpgradingRunners follows a verified server upgrade on a cluster // with runners. The executor hands the operation to the server here: // the server has the node inventory and the RPC to each runner, and it // is the new build, which is the one that knows the runner protocol. PhaseUpgradingRunners Phase = "upgrading_runners" PhaseRollingBack Phase = "rolling_back" PhaseSucceeded Phase = "succeeded" PhaseFailed Phase = "failed" PhaseRolledBack Phase = "rolled_back" // StepSkipped is terminal and step-only: the runner was not asked to // upgrade, and the step's Error says why (not ready, predates managed // upgrades). An operation never has it. StepSkipped Phase = "skipped" )
type Prober ¶
Prober observes the running server: which process is there before acting, and when the new one is up.
type RestoreResult ¶
type RestoreResult struct {
OperationID string `json:"operation_id"`
BackupRef string `json:"backup_ref"`
RestoredAt time.Time `json:"restored_at"`
Error string `json:"error,omitempty"`
// Abandoned records that an operator gave up on the restore and told the
// server to start on the data as it is.
Abandoned bool `json:"abandoned,omitempty"`
}
RestoreResult is what the server writes after acting on a DataRestore request during boot. It is a file beside the operation record rather than a field in it because the executor keeps its own copy of the record and rewrites it while the server boots; a second writer would be overwritten.
A result with an Error does not settle the request: the server refuses to start on data it was asked to replace, and tries again on its next boot. Only a successful restore or an operator's Abandon ends that.
func (*RestoreResult) Settled ¶
func (r *RestoreResult) Settled() bool
Settled reports whether the request this result answers is over, one way or the other.
type Store ¶
type Store struct {
// contains filtered or unexported fields
}
Store keeps one JSON file per operation, written atomically so a process dying mid-write leaves the previous record intact.
func (*Store) Create ¶
Create refuses while another operation is running: two executors racing on one binary and one systemd unit cannot both be right. The check and the write happen under a directory lock so two concurrent callers cannot both pass it.
func (*Store) CreateOrGet ¶
CreateOrGet records op unless a record with its id already exists, in which case that record is returned instead and created is false. The lookup and the write share the directory lock, so two callers carrying one id cannot both miss and cannot both create; exactly one of them creates. Like Create, it refuses to create while another operation is running.
func (*Store) List ¶
List returns every operation, oldest first. A record that cannot be read is left out rather than hiding the rest; callers that must not miss one use listStrict.
func (*Store) LockOperation ¶
LockOperation claims id for one executor. It does not wait: a second executor on the same operation gets ErrLocked immediately.
func (*Store) PendingRestore ¶
PendingRestore returns the operation whose rollback asked the server to restore data and has not been answered, or nil. The server calls this early in boot, before the data it would restore is in use.
The request outlives its operation on purpose. An executor that waited on a server refusing to boot marks the operation failed and exits, and if that ended the request the very next boot would start on the data the rollback was meant to replace. Only a restore or Abandon settles it.
Records are read strictly: a request in a record that cannot be decoded is still a request, and the caller fails closed on the error.
func (*Store) ReadRestoreResult ¶
func (s *Store) ReadRestoreResult(id string) (*RestoreResult, error)
func (*Store) RecordRestoreAttempt ¶
func (s *Store) RecordRestoreAttempt(result *RestoreResult) (*RestoreResult, error)
RecordRestoreAttempt writes the outcome of a restore attempt, unless an operator abandoned the request while the attempt ran: a failed attempt must not reopen a request that was just settled, or the server would go back to refusing to boot right after being told not to. A successful attempt is always recorded. It returns the result now on disk.
func (*Store) Remove ¶
Remove deletes a record. It exists for one case: a record that was created but whose executor never started and whose failure could not be written, which would otherwise hold the busy slot forever.
func (*Store) WriteRestoreResult ¶
func (s *Store) WriteRestoreResult(result *RestoreResult) error
WriteRestoreResult records the server's answer to a DataRestore request.
type SystemdLauncher ¶
type SystemdLauncher struct {
Binary string
// Command is the subcommand that runs the executor, without the binary
// and without the operation flag. Nil means ServerExecutorCommand.
Command []string
}
SystemdLauncher runs the executor as a transient unit, independent of the daemon's own unit, so restarting the daemon does not take it down and journald keeps its output. The binary is captured by inode at exec, so the executor keeps running the build it started with after replacing the file on disk.
type SystemdRestarter ¶
type UnsupervisedLauncher ¶
type UnsupervisedLauncher struct{}
UnsupervisedLauncher refuses every launch with ErrUnsupervised. It stands in where neither systemd nor container-boot is present, so the refusal says why instead of a failed systemd-run.