Documentation
¶
Overview ¶
Package llm registers `k6/x/llm`, a k6 extension for LLM-aware load testing. See the project README for the metric set and per-request semantics.
Index ¶
Constants ¶
const ( MetaTTFTMs = "llm.ttft_ms" MetaITLMeanMs = "llm.itl_mean_ms" MetaITLP50Ms = "llm.itl_p50_ms" MetaITLMaxMs = "llm.itl_max_ms" MetaITLSamples = "llm.itl_samples" MetaTPOTMs = "llm.tpot_ms" MetaChunks = "llm.chunks" MetaThinkingTokens = "llm.thinking_tokens" MetaCachedTokens = "llm.cached_tokens" MetaThinkingChunks = "llm.thinking_chunks" MetaTTFTextMs = "llm.ttf_text_ms" MetaResponseHeadersMs = "llm.response_headers_ms" MetaCostUSD = "llm.cost_usd" MetaEnergyJ = "llm.energy_j" MetaGoodput = "llm.goodput" MetaAborted = "llm.aborted" )
Metadata keys for measurements the Generation schema has no field for.
Agent Observability records gen_ai.client.time_to_first_token and the total operation duration, but nothing about the shape of the stream in between: a response that streams smoothly at 40 tok/s and one that stalls for two seconds mid-generation are indistinguishable in the schema today. Until ITL and TPOT have real fields, they travel as metadata.
const SyntheticProducerTagKey = "agento11y.synthetic.producer"
SyntheticProducerTagKey identifies which generator produced the record, so a consumer can tell a k6 canary apart from other synthetic sources.
const SyntheticTagKey = "agento11y.synthetic"
SyntheticTagKey marks a generation as produced by a load generator rather than by real traffic.
A consumer that treats synthetic records as real traffic will overstate usage and cost. Until Agent Observability reserves a tag for this, the convention has to live somewhere public — hence a constant rather than a string literal in an example script.
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Agento11yConfig ¶ added in v0.2.0
type Agento11yConfig struct {
// Endpoint is the Agent Observability generation-export base URL. The SDK
// appends the export path. Empty disables export.
Endpoint string
// Protocol is "http" (default), "grpc", or "none".
Protocol string
// AuthMode is "none" (default), "tenant", "bearer", or "basic".
AuthMode string
TenantID string
BearerToken string
BasicUser string
BasicPassword string
// Insecure permits cleartext export. Defaults to true for a loopback
// endpoint and false otherwise, so a local demo works without ceremony
// while a remote endpoint is not silently downgraded.
Insecure *bool
AgentName string
AgentVersion string
// Synthetic marks every exported generation as load-generated. Default
// true: traffic from a load generator is synthetic by construction, and a
// consumer that cannot distinguish it will corrupt its own cost and usage
// reporting.
Synthetic bool
// CaptureContent sends prompt and completion text. Default false —
// benchmark corpora are usually uninteresting and sometimes proprietary.
CaptureContent bool
// Tags are merged into every exported generation.
Tags map[string]string
// FlushInterval bounds how long a record waits before export.
FlushInterval time.Duration
}
Agento11yConfig configures export of one Grafana Agent Observability generation record per chat call.
The extension owns its own k6 metrics and OTel-free timing, so the SDK client is built with no-op tracer and meter: only the generation-export path runs.
type Client ¶
type Client struct {
// contains filtered or unexported fields
}
Client is the JS-facing OpenAI-compatible chat client.
func (*Client) Chat ¶
Chat sends a streaming chat completion request. Returns a Promise resolving to a result object (see chatResult.toJSObject for the full shape) or rejecting with the categorized error.
JS:
client.chat({
messages: [...],
max_tokens: 256,
// control fields (peeled off before the upstream POST):
slo: { ttft_ms: 500, tpot_ms: 50, e2el_ms: 5000 },
cache_state: "cold",
tags: { region: "us-east", shape: "short" },
})
func (*Client) Embed ¶ added in v0.2.0
Embed dispatches a request to /v1/embeddings. The `input` field accepts a string or an array of strings; the response always exposes `embeddings` as number[][] in input order.
JS:
client.embed({ input: "hello", model?: "nomic-embed-text", tags?: {...} })
client.embed({ input: ["hello", "world"] })
type CostModel ¶
CostModel parameterises a per-request USD cost estimate, computed from server-reported token counts:
usd = prompt_tokens * usd_per_million_input_tokens / 1e6
+ completion_tokens * usd_per_million_output_tokens / 1e6
Use hosted-API published rates directly; for self-hosted inference, compute your effective $/M-token rate offline (idle GPU $/hr divided by sustained throughput, plus marginal electricity) and plug it in here.
type Dataset ¶
type Dataset struct {
// contains filtered or unexported fields
}
Dataset is a deterministic, replayable corpus of chat requests. Loaded once per process from a JSONL file and shared across VUs via dsCache; per-Dataset instances each carry their own cursor and shuffle permutation so two VUs reading from the same file do not see the same order unless they share a (path, seed) pair.
func (*Dataset) At ¶
At returns the i-th request (modulo dataset size, negative-safe) without advancing the cursor. Use this when the caller wants to derive the index from __VU and __ITER for fully reproducible workloads.
func (*Dataset) Next ¶
Next advances the internal cursor and returns the next request, wrapping at the end. Concurrency-safe across VUs (within a single process) when the same Dataset instance is reused; in k6 each VU constructs its own instance, so "wrap" semantics apply per VU.
type EnergyModel ¶
EnergyModel parameterises a per-request energy estimate. The math, per request:
dynamic_j = prompt_tokens * j_per_input_token + completion_tokens * j_per_output_token static_j = idle_w * (duration_s) total_j = dynamic_j + static_j
Coefficients must be measured for your (GPU, model, batch regime) tuple. Under concurrent load the static term over-attributes idle power; divide idle_w by your expected per-VU concurrency for wall-plug accuracy. This is a budgeting metric, not a measurement.
func (*EnergyModel) Empty ¶
func (e *EnergyModel) Empty() bool
Empty reports whether the model would produce zero for any request.
type Options ¶
type Options struct {
BaseURL string
APIKey string
Model string
Timeout time.Duration
IgnoreEOS bool
// Wire selects the request encoding: "openai" (default) or "anthropic".
// Set it to drive an Anthropic Messages endpoint, including a gateway that
// fronts one.
Wire Wire
// Headers are sent on every request. Use for custom auth schemes, gateway
// routing keys (e.g. OpenRouter "HTTP-Referer"), or observability headers.
Headers map[string]string
// DefaultSLO applies to every chat() call that doesn't supply its own.
DefaultSLO *SLOPredicate
// Energy, when set, enables per-request energy estimation. See EnergyModel.
Energy *EnergyModel
// Cost, when set, enables per-request USD estimation. See CostModel.
Cost *CostModel
// Agento11y, when set, exports one Grafana Agent Observability generation
// record per chat call. See Agento11yConfig.
Agento11y *Agento11yConfig
}
Options configures an llm.Client.
type SLOPredicate ¶
SLOPredicate is the per-request SLO used to compute goodput and per-SLO attainment Rates. A zero field disables that SLO (it always passes).
Semantics match vLLM's `--goodput ttft:X tpot:Y e2el:Z` flag (PR #9338, shipped v0.6.4).
func (*SLOPredicate) Empty ¶
func (s *SLOPredicate) Empty() bool
Empty reports whether the predicate is functionally a no-op.
type Session ¶ added in v0.2.0
type Session struct {
// contains filtered or unexported fields
}
Session wraps a Client and carries conversation state for a single multi-turn dialogue. Each Send appends the user message, calls chat, appends the assistant reply, and auto-tags `session_id`, `turn`, and `cache_state` so dashboards can roll up per-session.
func (*Session) Id ¶ added in v0.2.0
Id is named `Id` not `ID` because Sobek's default field-name mapper lowercases only the first letter; `ID` would surface in JS as `iD()`.
func (*Session) Messages ¶ added in v0.2.0
Messages returns a deep copy of the history; callers may mutate freely.
func (*Session) Reset ¶ added in v0.2.0
func (s *Session) Reset()
Reset clears conversation history (keeping the system prompt) and rewinds the turn counter. Token totals are preserved across resets; use a new Session for fresh accounting.
func (*Session) Send ¶ added in v0.2.0
Send appends a user message and dispatches a chat call. The argument is either a string (treated as the user content) or an object with a `content` field plus any chat() option (max_tokens, temperature, slo, tags, etc.). On resolution the assistant reply is appended to history.
type ToolCall ¶ added in v0.2.0
ToolCall is one assembled tool invocation from a streamed assistant turn. Arguments is the raw JSON string the model produced; callers parse it.
type Wire ¶ added in v0.2.0
type Wire string
Wire selects the request/response encoding a Client speaks.
const WireResponses Wire = "responses"
WireResponses is the OpenAI Responses encoding (POST /v1/responses), which Codex and newer OpenAI clients use.
It is a different API from chat completions, not a variant: messages become `input`, `max_tokens` becomes `max_output_tokens`, a system message becomes `instructions`, tool definitions are flat rather than wrapped in a `function` object, and the stream is a taxonomy of named `response.*` events rather than choice deltas.

