failed

package
v0.37.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 9, 2026 License: MIT Imports: 14 Imported by: 0

Documentation

Overview

Package failed is where a job goes when it gives up.

It is a FailedJobProvider interface and the implementations of it, so a dead letter list can live somewhere other than the queue that produced it.

The default is that it does not move

A github.com/arandu-io/hesape/queue.DatabaseQueue marks the job failed in the jobs table it was already in, and queue.Queue.Failed lists them and queue.Queue.Retry puts one back. Every driver implements those two, so an application that never wires a provider still has a dead letter list, one store and one answer to "is this job still queued".

What this package is for is the case that arrangement cannot serve: a queue whose store is not durable enough to keep failures, or one that is flushed by something other than the application. A RESP queue is both. The provider is then a deliberate second place, wired on purpose, and the cost -- two stores, two retentions -- is paid knowingly.

The commands read a provider and nothing else

`aru queue:failed`, `queue:retry`, `queue:forget` and `queue:flush` are built over a FailedJobProvider. An application that registers them and gives the worker no provider -- see queue.Worker.SetFailedJobs -- has four commands over a list nothing writes, and finds that out during the incident they were meant for. Wire the same provider into both.

Every read takes a Grant

A failed job carries a customer's payload, so listing them is a read like any other: every method takes an auth.Grant and filters by auth.Tenant(g). A provider that answered across tenants would leak the arguments of every job every customer ever queued.

DatabaseUUIDFailedJobProvider is an alias of DatabaseFailedJobProvider, because the id is the uuid and the two would run the same query.

Index

Constants

View Source
const Action auth.Action = "queue:failed"

Action is the permission a failed job list is reached under.

One spelling, because a Grant is issued for one action and refused on any other (auth.Grant.Check): the worker that records a parked job and the five commands that read it back have to name the same permission, and they are in different packages. A second spelling is a console that cannot see what the worker wrote.

View Source
const DefaultTable = "failed_jobs"

DefaultTable is where failures are logged when no table name is given.

Variables

View Source
var ErrNoTenant = errors.New("queue/failed: the Grant carries no tenant, and a failed job list cannot be scoped without one")

ErrNoTenant is returned when a Grant carries no tenant.

It mirrors jobs.ErrNoTenant and exists for the same reason: a query with no tenant to filter by is a query that reads every customer's failures.

View Source
var ErrNotFound = errors.New("queue/failed: no such failed job")

ErrNotFound is what Find wraps when there is no such failed job.

It is an error rather than a zero value, because "no such job" and "a job with no payload" are different answers and a command that retries the second one silently does nothing.

Functions

This section is empty.

Types

type AddActionToFailedJobsTable added in v0.22.0

type AddActionToFailedJobsTable struct {
	migrations.BaseMigration

	// Table is the table to alter. Empty means DefaultTable.
	Table string
}

AddActionToFailedJobsTable adds the column that carries the permission a job was pushed under, which the record used to drop.

A retry rebuilds the job from this row, and the Grant it runs under is built from an action. Without the column the only action a retry could name was the one the dead letter list is read with, so the work came back as an administrator rather than as itself.

It is its own migration rather than an edit to CreateFailedJobsTable: a published migration is not changed, because a database that already applied it would not apply it again and the two would be one name over two schemas.

func (AddActionToFailedJobsTable) Down added in v0.22.0

Down drops the column.

func (AddActionToFailedJobsTable) GetName added in v0.22.0

GetName returns the migration's name.

func (AddActionToFailedJobsTable) Up added in v0.22.0

Up adds the column.

Nullable, so the release running while this is applied keeps inserting without it: a NOT NULL column with no default added to a table that has rows fails on every row already there, and the previous binary's insert names nine columns rather than ten.

type CountableFailedJobProvider

type CountableFailedJobProvider interface {
	// Count is how many failed jobs there are. An empty connection or queue
	// means every one.
	Count(ctx context.Context, g auth.Grant, connection, queue string) (int, error)
}

CountableFailedJobProvider is a provider that can say how many.

It is a second interface rather than a sixth method on FailedJobProvider: counting is what a monitor does every minute, and a provider that would have to load every row to answer should not pretend it can.

type CreateFailedJobsTable added in v0.5.0

type CreateFailedJobsTable struct {
	migrations.BaseMigration

	// Table is the table to create. Empty means DefaultTable.
	Table string
}

CreateFailedJobsTable creates the table a DatabaseFailedJobProvider logs to.

The table name is a field because the provider's is: an application that keeps its failures somewhere other than failed_jobs migrates the name it wired.

func (CreateFailedJobsTable) Down added in v0.5.0

Down drops the failed jobs table, and the index with it.

func (CreateFailedJobsTable) GetName added in v0.5.0

func (CreateFailedJobsTable) GetName() string

GetName returns the migration's name.

func (CreateFailedJobsTable) Up added in v0.5.0

Up creates the failed jobs table and the index every read of it uses.

Portable types only: TEXT, INTEGER and TIMESTAMP mean the same thing on SQLite, Postgres and MySQL.

type DatabaseFailedJobProvider

type DatabaseFailedJobProvider struct {
	// contains filtered or unexported fields
}

DatabaseFailedJobProvider keeps the failed jobs in a table.

It is what an application wires when it wants the dead letter list to survive the queue -- a job that failed on a Redis queue that was then flushed is still a job somebody has to answer for.

The table is its own, and that is the point of the type: a queue whose store is not a database has nowhere else to keep them.

func NewDatabaseFailedJobProvider

func NewDatabaseFailedJobProvider(db *database.DB, table string) *DatabaseFailedJobProvider

NewDatabaseFailedJobProvider returns the provider over db.

An empty table means DefaultTable.

func (*DatabaseFailedJobProvider) All

All is this tenant's failed jobs, newest first.

func (*DatabaseFailedJobProvider) Count

func (p *DatabaseFailedJobProvider) Count(ctx context.Context, g auth.Grant, connectionName, queue string) (int, error)

Count is how many of this tenant's jobs have failed.

func (*DatabaseFailedJobProvider) Find

Find is one of this tenant's failed jobs.

func (*DatabaseFailedJobProvider) Flush

Flush removes this tenant's failed jobs older than age, or all of them when age is zero.

func (*DatabaseFailedJobProvider) Forget

Forget removes one of this tenant's failed jobs.

func (*DatabaseFailedJobProvider) GetTable

func (p *DatabaseFailedJobProvider) GetTable() string

GetTable is the name of the table this provider reads.

It returns the name rather than a query over it: a query handed out is a caller writing its own SQL against a table it does not own, and the tenant filter is not optional.

func (*DatabaseFailedJobProvider) IDs

IDs is the identifiers of this tenant's failed jobs, newest first.

func (*DatabaseFailedJobProvider) Log

Log records a job that gave up, once.

The id is the job's own uuid and it is the primary key, so a second record of one failure is refused by the table rather than written. That refusal is the answer the caller wanted -- the failure is listed -- so it is read back as such, and only an insert that failed for any other reason is returned as an error.

func (*DatabaseFailedJobProvider) Migrations

Migrations returns the failed jobs table.

The schema is on the provider rather than on the queue module because it belongs to whoever wired this provider: an application that keeps its failures in the jobs table declares nothing here.

func (*DatabaseFailedJobProvider) Prune

func (p *DatabaseFailedJobProvider) Prune(ctx context.Context, g auth.Grant, before time.Time) (int, error)

Prune removes this tenant's failed jobs that failed before an instant, and returns how many went.

type DatabaseUUIDFailedJobProvider

type DatabaseUUIDFailedJobProvider = DatabaseFailedJobProvider

DatabaseUUIDFailedJobProvider is DatabaseFailedJobProvider under the name for the provider keyed by the job's uuid.

It is an alias because there is nothing left to distinguish: the id is the uuid (see database.NewID), so a provider that looked up by uuid would run the same query.

type FailedJob

type FailedJob struct {
	// ID identifies this failure. It is the job's own UUID, because the id is
	// minted by the application and a second one would only be a second thing
	// to quote at somebody.
	ID string
	// UUID is the job's identifier, which is the same string as ID. It is kept
	// because it is what a provider indexes and what `aru queue:retry` takes.
	UUID string
	// TenantID is who the work belonged to, and it is not optional: a failed
	// job list that crossed customers would be one customer reading another's
	// payloads.
	TenantID string
	// Connection is the queue connection it was on.
	Connection string
	// Queue is the queue it was on.
	Queue string
	// Name is what routes the job to a handler.
	Name string
	// Action is the permission the job was pushed under, and it is the half of
	// the envelope that cannot be rebuilt from anything else here.
	//
	// The worker reissues a job's Grant from its action -- jobs.GrantFor is
	// auth.SystemGrant(action, tenant) -- so a record that lost it can only put
	// the job back under the action the dead letter list itself is read with.
	// That is an administrative permission: every Policy that checks the job's
	// own action refuses the work, and every Policy that does not lets it do
	// more than the push ever authorized.
	//
	// Empty on a record written before the column existed. A retry that finds
	// none is refused rather than guessed at: an action is exactly the thing
	// nothing else in the record implies.
	Action string
	// Payload is the job's arguments, as they were stored.
	//
	// As they were, and not masked. Masking them here would cost the retry and
	// protect nothing. A driver that parks in place still holds the same bytes
	// -- the database queue leaves the row and only marks failed_at, the redis
	// queue leaves the job hash and only moves the id -- so a masked copy would
	// sit one table over from the original, written by the same park. And when
	// the store no longer holds the job, this is the last copy of the arguments
	// there is: masked, a retry has nothing to put back.
	//
	// What is enforced is the boundary the bytes can actually cross. Reading
	// them takes a Grant and they are scoped to one tenant, and they never go
	// into a log line -- a log is shipped, retained and read without a Grant,
	// which a table is not. Keeping them unreadable at rest as well is
	// encryption, not redaction: it is one decision for both stores, with a key
	// to manage and rotate, and doing it to this record alone would leave the
	// same payload in the clear beside it.
	Payload []byte
	// Exception is why it gave up.
	Exception string
	// FailedAt is when.
	FailedAt time.Time
}

FailedJob is one job that gave up.

It is the record a dead letter list holds: what the job was, whose it was, why it gave up and when.

type FailedJobProvider

type FailedJobProvider interface {
	// Log records a job that gave up, and returns the id it was recorded under.
	//
	// Recording the same job twice records it once. The id is the job's own
	// uuid, so the second call is the same failure arriving again -- a worker
	// that retried the write, a replay -- and a dead letter list that answered
	// with two rows would have an operator retry the work twice.
	Log(ctx context.Context, g auth.Grant, job FailedJob) (string, error)

	// IDs is the identifiers of the failed jobs, newest first. An empty queue
	// means every queue.
	IDs(ctx context.Context, g auth.Grant, queue string) ([]string, error)

	// All is the failed jobs, newest first.
	All(ctx context.Context, g auth.Grant) ([]FailedJob, error)

	// Find is one failed job, or an error wrapping [ErrNotFound].
	Find(ctx context.Context, g auth.Grant, id string) (FailedJob, error)

	// Forget removes one failed job, and reports whether there was one.
	Forget(ctx context.Context, g auth.Grant, id string) (bool, error)

	// Flush removes the failed jobs older than age. A zero age removes all of
	// them, which is what `queue:flush` with no --hours does.
	Flush(ctx context.Context, g auth.Grant, age time.Duration) error
}

FailedJobProvider is where a job goes when it gives up.

Every method takes a context and an auth.Grant.

The Grant is not decoration. A failed job carries a customer's payload, and "list the failed jobs" is a read like any other -- so it takes a Grant and every implementation filters by auth.Tenant(g). A provider that answered across tenants would be the one query in the collection that leaks, and it would leak the arguments of every job every customer ever queued.

type FileFailedJobProvider

type FileFailedJobProvider struct {
	// contains filtered or unexported fields
}

FileFailedJobProvider keeps the failed jobs in a JSON file.

It exists for the deployment that has a queue and no database: a single process on one machine, draining a RESP queue, that still wants a dead letter list to survive a restart.

It holds the newest Limit failures and drops the rest, which is the right behaviour for a file: a dead letter list that grows without bound on a disk nobody is watching is an outage waiting to be filed as a disk-full alert.

It is not for more than one process. The file is rewritten whole under a mutex this process holds, and two of them would lose each other's writes -- which is exactly the case DatabaseFailedJobProvider is for.

func NewFileFailedJobProvider

func NewFileFailedJobProvider(path string, limit int) *FileFailedJobProvider

NewFileFailedJobProvider returns the provider over path.

A limit of zero or less means a hundred.

func (*FileFailedJobProvider) All

All is this tenant's failed jobs, newest first. It answers all().

func (*FileFailedJobProvider) Count

func (p *FileFailedJobProvider) Count(ctx context.Context, g auth.Grant, connectionName, queue string) (int, error)

Count is how many of this tenant's jobs have failed. It answers count().

func (*FileFailedJobProvider) Find

Find is one of this tenant's failed jobs. It answers find().

func (*FileFailedJobProvider) Flush

Flush removes this tenant's failed jobs older than age, or all of them when age is zero. It answers flush().

func (*FileFailedJobProvider) Forget

Forget removes one of this tenant's failed jobs. It answers forget().

func (*FileFailedJobProvider) IDs

func (p *FileFailedJobProvider) IDs(ctx context.Context, g auth.Grant, queue string) ([]string, error)

IDs is the identifiers of this tenant's failed jobs, newest first. It answers ids().

func (*FileFailedJobProvider) Log

Log records a job that gave up, once. It answers log().

The file has no key to refuse a duplicate with, so the check is here: a record already under this id is this failure, already listed.

func (*FileFailedJobProvider) Prune

func (p *FileFailedJobProvider) Prune(_ context.Context, g auth.Grant, before time.Time) (int, error)

Prune removes this tenant's failed jobs that failed before an instant, and returns how many went. It answers prune().

type NullFailedJobProvider

type NullFailedJobProvider struct{}

NullFailedJobProvider accepts every failure and keeps none of them.

It is what an application wires when the dead letter list lives somewhere else entirely -- an error tracker, a log pipeline -- and keeping a second copy in the database would be a table nobody reads.

It is a value and not a pointer, for the reason NullQueue is: a nil pointer that silently swallows every failure is a worse mistake than the one this type exists to make cheap.

func (NullFailedJobProvider) All

All is empty.

func (NullFailedJobProvider) Count

Count is zero.

func (NullFailedJobProvider) Find

Find never finds anything.

func (NullFailedJobProvider) Flush

Flush has nothing to flush.

func (NullFailedJobProvider) Forget

Forget has nothing to forget, and reports true: the caller asked for the job to be gone, and it is.

func (NullFailedJobProvider) IDs

IDs is empty.

func (NullFailedJobProvider) Log

Log discards the failure and returns no id.

type PrunableFailedJobProvider

type PrunableFailedJobProvider interface {
	// Prune removes the entries that failed before this instant, and returns
	// how many went.
	Prune(ctx context.Context, g auth.Grant, before time.Time) (int, error)
}

PrunableFailedJobProvider is a provider that can drop old entries.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL