Documentation
¶
Overview ¶
Package journal is a host's write-ahead log of flushed blocks, kept on a network disk of its own (plans/fsync-journal-2026-10-06.md). A flush is answered once the blocks it covers are on that disk. If the host is lost, another host opens the disk, reads the journal back, and serves its entries to the hosts that recover the lost host's VMs.
The disk holds two header slots and a ring. Entries go onto the ring in batches, one batch in flight at a time, each written and synced before its commits are answered. A position is an entry's logical byte position on the ring; it maps to the ring's start plus the position modulo the ring's length, and it only grows. An entry never wraps: a pad fills the ring's end. A batch ends on a 4 KiB boundary, with a pad before it where needed.
Reading back starts at the header's tail hint and stops at the first entry whose magic, version, generation, position or checksum is wrong. That position is the head. So the writer never lets the ring overwrite a position at or after the tail hint on the disk, and a batch whose write or sync failed is padded over by the next one: otherwise reading back would stop at the failed range and miss what follows.
Nothing here knows about pagers or control records. A commit's capture makes its entries, and the holder trims the journal by the covered positions the control records name.
Index ¶
- Constants
- Variables
- func EntryBytes(vm, volume string, blocks int) int64
- func Probes() []string
- func Sites() []string
- type Capture
- type Config
- type Entry
- type Held
- type Hooks
- type Journal
- func (j *Journal) CheckLease(ctx context.Context) error
- func (j *Journal) Close(ctx context.Context) (bool, error)
- func (j *Journal) Commit(ctx context.Context, room int64, capture Capture) ([]uint64, error)
- func (j *Journal) CommitHooked(ctx context.Context, room int64, capture Capture, hooks Hooks) ([]uint64, error)
- func (j *Journal) Formatted() bool
- func (j *Journal) Generation() uint64
- func (j *Journal) Held() []Held
- func (j *Journal) Identity() rank.Identity
- func (j *Journal) Oldest(bytes int64) []string
- func (j *Journal) Read(ctx context.Context, request ReadRequest, yield func(Entry) error) error
- func (j *Journal) Shortfall(room int64) int64
- func (j *Journal) Trim(vm string, covered map[uint64]uint64)
- func (j *Journal) Usage() Usage
- type Lease
- type ReadRequest
- type Usage
Constants ¶
const ( // BlockBytes is the size of a block: 4 KiB of a page. BlockBytes = 4096 // MaxEntryBytes bounds an entry of blocks. A pad before it at the ring's // end is shorter than it, so nothing on the ring is longer than // maxStoredBytes. MaxEntryBytes = maxStoredBytes - BlockBytes // MaxBatchBytes is the room a batch takes commits into. A batch holds at // least one commit, however large. MaxBatchBytes = 8 << 20 )
const ( // ProbeFormatted is a disk with no header, formatted as it opened. ProbeFormatted = "journal/formatted" // ProbeLeaseRefused is a disk refused as it opened, because its lease is // newer than the one it was opened under. ProbeLeaseRefused = "journal/lease-refused" // ProbeLeaseLost is an open journal that found its lease taken. ProbeLeaseLost = "journal/lease-lost" // ProbeTornEntry is a read back that ended at an entry whose head holds // together and whose checksum does not: a torn batch. ProbeTornEntry = "journal/torn-entry" // ProbeFailedRangePadded is a batch that padded over a failed one. ProbeFailedRangePadded = "journal/failed-range-padded" // ProbeRingEndPadded is a pad that filled the ring's end. ProbeRingEndPadded = "journal/ring-end-padded" // ProbeFull is a commit that waited for room on the ring. ProbeFull = "journal/full" // ProbeGrouped is a batch that held more than one commit. ProbeGrouped = "journal/grouped" // ProbeFenced is a commit refused because a read fenced its VM. ProbeFenced = "journal/fenced" )
The probes the journal marks.
const ( // BuggifyWriteSlow holds a batch's write for up to 50 ms, so commits pile // up behind it. BuggifyWriteSlow = "journal/write-slow" // BuggifyWriteFails fails a batch's write before it reaches the disk. BuggifyWriteFails = "journal/write-fails" // BuggifyWriteTorn writes the first half of a batch, then fails. BuggifyWriteTorn = "journal/write-torn" // BuggifySyncFails fails a batch's sync. BuggifySyncFails = "journal/sync-fails" // BuggifyHeaderWriteFails fails a write of the header. BuggifyHeaderWriteFails = "journal/header-write-fails" )
The fault-injection sites of the journal's writes.
Variables ¶
var ( // ErrLeased reports a journal disk leased under a newer assignment than // the one it was opened under, or to another member. ErrLeased = errors.New("journal: the disk is leased under a newer assignment") // ErrIdentity reports a disk whose header names another journal disk. ErrIdentity = errors.New("journal: the disk is another journal's") // ErrClosed reports a journal that was closed. ErrClosed = errors.New("journal: the journal is closed") // ErrTooLarge reports a commit or an entry larger than the journal holds. ErrTooLarge = errors.New("journal: too large for the journal") )
var ( // ErrFenced reports an entry of a VM that a read fenced at a newer // epoch. ErrFenced = errors.New("journal: the VM is fenced at a newer epoch") // ErrGeneration reports a read of a generation the journal is not: its // disk was formatted again, and the entries asked for are gone. ErrGeneration = errors.New("journal: the journal is of another generation") // ErrTrimmed reports an entry trimmed and written over while it was read. ErrTrimmed = errors.New("journal: an entry was trimmed while it was read") )
var ErrVersion = errors.New("journal: the disk is of another format version")
ErrVersion reports a journal disk written under another format version. The error names the version.
Functions ¶
func EntryBytes ¶
EntryBytes is how many bytes an entry of blocks blocks takes in the ring.
Types ¶
type Capture ¶
Capture makes the entries of one commit. The writer calls it as it forms the batch the commit joins, after the batch before has completed, so a capture takes everything stored up to then. Its entries must fit in the room the commit asked for, as EntryBytes counts them.
type Config ¶
type Config struct {
// Identity is the disk's. A disk whose header names another is refused.
Identity rank.Identity
// Lease is the assignment the disk is opened under.
Lease Lease
// Clock times the tail hint's writes. Nil is the wall clock.
Clock platform.Clock
// Entropy draws the generation of a disk that is formatted. Nil is the
// operating system's.
Entropy platform.Entropy
}
Config is what a journal is opened with.
type Entry ¶
type Entry struct {
// VM is the identity of the VM whose disk the blocks are of.
VM string
// Volume is the name of the volume the blocks are of.
Volume string
// Epoch is the VM's epoch when its blocks were taken.
Epoch uint64
// Blocks are the block numbers: each block's byte offset in the volume
// divided by BlockBytes.
Blocks []uint64
// Data holds the blocks, BlockBytes each, in the order Blocks names them.
Data []byte
// Position is where the entry is in its journal. Read sets it; a commit
// ignores it.
Position uint64
}
Entry is the changed blocks of one memory region, taken by one capture.
type Held ¶
type Held struct {
VM string
Epoch uint64
// Entries and Bytes count the live entries.
Entries int
Bytes int64
// First and Last are the positions of the first and last of them.
First, Last uint64
}
Held is what the journal holds of one VM and epoch.
type Hooks ¶
type Hooks struct {
// Placed runs once the commit's entries have positions, before they are
// written: the start of each, none where the capture failed or a read had
// fenced the commit's VM.
Placed func(positions []uint64)
// Failed runs when a commit whose capture made entries does not land:
// its batch's write or sync failed, or a read had fenced its VM. Its
// entries may be on the disk all the same, so whatever the capture took
// has to be taken again by the next.
Failed func()
// Full runs once the commit is the next to place and the ring has no
// room for it: whatever trims the ring has to be asked now. It runs once
// a commit, on the writer, with nothing held.
Full func()
}
Hooks are what a commit's caller learns on the writer, in batch order and before the writer forms its next batch, which no answer can promise: an answer reaches its caller whenever that caller runs.
type Journal ¶
type Journal struct {
// contains filtered or unexported fields
}
Journal is one journal disk, opened by its holder. It has one writer, which runs until Close.
func Open ¶
Open opens the journal on file, a journal disk attached to this machine, under config's lease. A disk with no header in either slot is formatted under a new generation. A disk whose header is of another format version, names another disk, or carries a newer lease is refused. Otherwise Open takes the lease, writing it into the header, and reads the journal back. The journal's writer runs until Close. The file stays the caller's to close.
func (*Journal) CheckLease ¶
CheckLease reads the header and reports ErrLeased if another assignment has taken the disk. From then on the journal refuses every commit and read. The holder calls it on every pass.
func (*Journal) Close ¶
Close stops the writer, after the batch in flight, and fails every commit still waiting. It then writes the tail into the header, with the empty flag when no live entry is left, and reports whether it was. The file stays the caller's.
func (*Journal) Commit ¶
Commit puts the entries capture makes onto the ring and returns their positions once the batch that holds them is written and synced. Room is the most bytes the entries may take. A commit waits while a batch is in flight, and while the ring has no room for it; it fails when its capture fails, when its batch's write or sync fails, when a read has fenced the VM of one of its entries at a newer epoch, and when the journal is closed or its lease taken. Batches complete in position order, so commits are answered in that order too. A read may still find the entries of a commit that failed, as a power loss may keep stores no flush covered.
func (*Journal) CommitHooked ¶
func (j *Journal) CommitHooked(ctx context.Context, room int64, capture Capture, hooks Hooks) ([]uint64, error)
CommitHooked is Commit with hooks the writer runs as the commit is placed and if it fails.
func (*Journal) Generation ¶
Generation is the journal's generation, drawn when its disk was formatted.
func (*Journal) Oldest ¶
Oldest is the VMs holding a live entry within bytes of the tail, the one holding the oldest first. Only the tail moving frees room on the ring, so freeing that much takes every one of those entries trimmed.
func (*Journal) Read ¶
Read is the server side of JOURNAL_READ. It fences the VM at the reader's epoch, so no entry of an older epoch is placed from then on, and every commit of one fails. It waits for the batches placed before the fence, so every entry any commit of the VM was answered for is among what it reads. Then it reads the VM's live entries of the epoch after the covered position from the disk, checks each, and hands them to yield in position order. Read stops at the first error yield returns.
func (*Journal) Shortfall ¶
Shortfall is how many bytes of live entries trimming has to free before a commit of room fits on the ring by itself: zero once it fits. A commit needs room for its entries and for the pad that keeps them from wrapping, which may be as large.
func (*Journal) Trim ¶
Trim says which of a VM's entries are still live. An entry is live while the VM's control record names this journal and the entry's epoch, and the entry's position is after the covered position the record names for it. covered maps each epoch the record names to its covered position; every entry of an epoch it does not name is dead, and so is every entry at or before its covered position. The tail moves on to the oldest live entry, and the ring before it is free once the tail hint reaches the header.
type Lease ¶
Lease is the assignment a journal disk is opened under: the generation of the membership that assigned it, and the member it was assigned to.
type ReadRequest ¶
type ReadRequest struct {
VM string
// Epoch is the epoch whose entries are asked for.
Epoch uint64
// After is the covered position: only entries after it are returned.
After uint64
// Generation is the journal's generation the control record names.
Generation uint64
// Reader is the reader's own epoch, newer than Epoch. The VM is fenced
// at it.
Reader uint64
}
ReadRequest asks for one VM's entries of one epoch: JOURNAL_READ.
type Usage ¶
type Usage struct {
// Ring is the ring's length.
Ring int64
// Used is the bytes from the tail to the next position: what trimming
// has not yet freed.
Used int64
// Next is the next position. It starts at the ring's length when the
// journal is formatted and grows by every byte the journal places, pads
// included.
Next int64
}
Usage is how full the ring is.