Documentation
¶
Overview ¶
Package jsonflat flattens JSON documents and reshapes them with a JSON config: it normalises and corrects keys, renames, drops, merges and defaults them, splits batches into records and routes records to named outputs, with zero heap allocations on the hot path.
A config is compiled once into a Transformer, which is immutable and safe for concurrent use:
t := jsonflat.MustCompile([]byte(`{
"flatten": {"separator": "_"},
"keys": {"normalize": "snake"},
"rules": [{"op": "drop", "path": "debug"}]
}`))
An empty config, {}, flattens the whole document with "." between key segments. Every section of the config is optional and they combine freely; see Config for all of them.
One document, one row ¶
Append and Transform turn one document into one flat object. Append writes into a buffer that the caller reuses, which is what keeps the call free of allocations:
out, err = t.Append(out[:0], src)
One document, many rows ¶
Each is for configs that split a batch into records (input.explode) or route a record to several outputs (outputs). It calls a function for every row and hands it the name of the output, the JSON, and the values the config asked to expose, for example a partition key:
err := t.Each(body, jsonflat.Options{}, func(r *jsonflat.Record) error {
if r.Err != nil {
return deadLetter(r.Index, r.Fields, r.Err)
}
return put(r.Name, r.Fields[0], r.JSON)
})
A record that fails is reported through Record.Err and never stops the batch; none of its rows are delivered, so a record never arrives in part. The Record and its slices are valid only until the callback returns.
Two kinds of path ¶
Source paths address the input (rule path and from, section from, input.explode, and every Source). They always use "." between segments, whatever flatten.separator is, and use the keys as the input spells them. Output keys are what gets written (rule to, the path of a default rule, column to, alias to, keys.keep, keys.drop) and are used exactly as given.
The full config reference, recipes and the preset documentation are at https://nayan9229.github.io/jsonflat/ and in the docs/ directory of the repository.
Guarantees ¶
The output is always valid JSON and never aliases the input. Number text is copied through unchanged, so no precision is lost, and is checked against RFC 8259 because the parser is more lenient than that. On error, Append returns dst unchanged.
Vendor layouts are configs, not code: see the presets package.
Example ¶
package main
import (
"fmt"
"github.com/nayan9229/jsonflat"
)
func main() {
t := jsonflat.MustCompile([]byte(`{
"rules": [
{"op": "rename", "from": "user.id", "to": "uid"},
{"op": "merge", "from": ["user.first", "user.last"], "to": "user.name", "sep": " "},
{"op": "drop", "path": "debug"},
{"op": "default", "path": "env", "value": "prod"}
]
}`))
src := []byte(`{"user":{"id":7,"first":"Ada","last":"Lovelace"},"tags":["a","b"],"debug":{"x":1}}`)
var out []byte // reuse this buffer between calls
out, err := t.Append(out[:0], src)
if err != nil {
panic(err)
}
fmt.Println(string(out))
}
Output: {"uid":7,"tags.0":"a","tags.1":"b","user.name":"Ada Lovelace","env":"prod"}
Example (LazyCompile) ¶
package main
import (
"fmt"
"sync"
"github.com/nayan9229/jsonflat"
)
// lazy compiles the config on first use and hands every later caller the
// same Transformer and error. This is the whole of what a program needs for
// a config that is read from disk or a flag at run time; a config that is
// part of the program is simpler still as a package-level
// jsonflat.MustCompile, which runs at init.
var lazy = sync.OnceValues(func() (*jsonflat.Transformer, error) {
config := []byte(`{"keys": {"normalize": "snake"}}`) // os.ReadFile in a real program
return jsonflat.Compile(config)
})
func main() {
t, err := lazy()
if err != nil {
fmt.Println("config:", err)
return
}
out, err := t.Transform([]byte(`{"userId": 7, "orderTotal": 9.5}`))
if err != nil {
fmt.Println(err)
return
}
fmt.Println(string(out))
}
Output: {"user_id":7,"order_total":9.5}
Example (RudderStackPreset) ¶
A preset is plain config. Unmarshal it, add your own key corrections, and build the Transformer.
package main
import (
"encoding/json"
"fmt"
"time"
"github.com/nayan9229/jsonflat"
"github.com/nayan9229/jsonflat/presets"
)
func main() {
var cfg jsonflat.Config
if err := json.Unmarshal(presets.RudderStack, &cfg); err != nil {
panic(err)
}
cfg.Keys.Aliases = []jsonflat.Alias{{To: "product_id", From: []string{"prodcutId", "pid"}}}
cfg.Keys.Drop = []string{"context_ip"}
t, err := jsonflat.New(cfg)
if err != nil {
panic(err)
}
body := []byte(`{"batch":[{
"type": "track", "event": "Product Added", "messageId": "m-1", "anonymousId": "anon-1",
"context": {"ip": "1.2.3.4", "app": {"name": "Shop"}},
"properties": {"prodcutId": "P1", "Total Price": 499.75, "user_id": "ignored"}
}]}`)
opt := jsonflat.Options{Now: time.Date(2026, 9, 21, 10, 0, 0, 0, time.UTC)}
err = t.Each(body, opt, func(r *jsonflat.Record) error {
if r.Err != nil {
fmt.Println("bad record", r.Index, r.Err)
return nil
}
// r.Fields is messageId, anonymousId, userId. Copy what you keep.
fmt.Printf("%s key=%s -> %s\n", r.Name, r.Fields[1], r.JSON)
return nil
})
if err != nil {
panic(err)
}
}
Output: tracks key=anon-1 -> {"id":"m-1","anonymous_id":"anon-1","received_at":"2026-09-21T10:00:00.000Z","timestamp":"2026-09-21T10:00:00.000Z","event":"product_added","event_text":"Product Added","context_app_name":"Shop"} product_added key=anon-1 -> {"id":"m-1","anonymous_id":"anon-1","received_at":"2026-09-21T10:00:00.000Z","timestamp":"2026-09-21T10:00:00.000Z","event":"product_added","event_text":"Product Added","context_app_name":"Shop","product_id":"P1","total_price":499.75}
Index ¶
Examples ¶
Constants ¶
const ( OpRename = "rename" OpDrop = "drop" OpMerge = "merge" OpDefault = "default" MergeConcat = "concat" MergeArray = "array" MergeFirst = "first" ArraysIndex = "index" // flatten arrays as key.0, key.1 ArraysRaw = "raw" // keep arrays as JSON values ArraysString = "string" // write arrays as a JSON string, for typed columns NormalizeNone = "none" NormalizeSnake = "snake" CollisionKeep = "keep" // write duplicates as they come (default) CollisionFirst = "first" // the first value wins, later ones are counted CollisionError = "error" // the record fails with ErrCollision AsRaw = "raw" // column value is written as it is AsString = "string" // strings and numbers are written as strings AsNumber = "number" // the first text that is a JSON number, as a number AsBool = "bool" // the first text that is true or false, as a boolean FnClockSkew = "clock_skew" FnMap = "map" FnLower = "lower" )
Rule operations, merge modes and the other config keywords.
Variables ¶
var ( // ErrRootNotContainer means the record is a scalar, not an object or array. ErrRootNotContainer = errors.New("jsonflat: record is not an object or an array") // ErrInvalidNumber means the record holds a number that is not valid JSON, // such as 00 or NaN. The parser accepts these; the output must not. ErrInvalidNumber = errors.New("jsonflat: invalid number") // ErrCollision means two keys got the same name under on_collision "error". ErrCollision = errors.New("jsonflat: key collision") // ErrRequired means a required derived value came out empty. ErrRequired = errors.New("jsonflat: required value is empty") // ErrNoOutput means no configured output matched the record. ErrNoOutput = errors.New("jsonflat: no output matches the record") )
Errors that concern one record. Each reports them in Record.Err and carries on with the next record; Append returns them.
var ( // ErrExplode means the input.explode path exists but is not an array. ErrExplode = errors.New("jsonflat: input.explode path is not an array") // ErrNotSimple is returned by Append and Transform for a config that can // produce several rows: one with input.explode or with outputs. Use Each. ErrNotSimple = errors.New("jsonflat: the config can produce several rows; use Each") )
Errors that concern the whole call.
Functions ¶
This section is empty.
Types ¶
type Alias ¶
type Alias struct {
To string `json:"to"`
From []string `json:"from"`
When *Cond `json:"when,omitempty"`
}
Alias maps wrong output keys to the right one. With snake normalisation the From entries are normalised when the config is compiled, so they can be written the way the producer writes them.
type Column ¶
type Column struct {
To string `json:"to"`
// From lists sources in order of preference; the first with a value wins.
From SourceList `json:"from"`
// As is AsRaw (default) or AsString.
As string `json:"as,omitempty"`
When *Cond `json:"when,omitempty"`
// Quiet stops keys dropped in favour of this column from being counted.
Quiet bool `json:"quiet,omitempty"`
}
Column is an explicit output key written before any section. Column names are reserved: a flattened key with the same name is dropped, whether or not the column had a value.
type Cond ¶
type Cond struct {
Path SourceList `json:"path"`
Equals *string `json:"equals,omitempty"`
In []string `json:"in,omitempty"`
Exists *bool `json:"exists,omitempty"`
}
Cond is a test on one value of the record. Exactly one of Equals, In and Exists must be set.
type Config ¶
type Config struct {
Input InputConfig `json:"input,omitempty"`
Flatten FlattenConfig `json:"flatten,omitempty"`
Keys KeysConfig `json:"keys,omitempty"`
Derive map[string]Derive `json:"derive,omitempty"`
Columns []Column `json:"columns,omitempty"`
Sections []Section `json:"sections,omitempty"`
Rules []Rule `json:"rules,omitempty"`
Outputs []Output `json:"outputs,omitempty"`
Expose SourceList `json:"expose,omitempty"`
// Newline appends "\n" to every record.
Newline bool `json:"newline,omitempty"`
// MaxPooledInput is the largest input, in bytes, whose per-call state is
// kept for reuse after the call. That state holds about eight times the
// input size, once per active P. 0 means 4 MiB.
MaxPooledInput int `json:"max_pooled_input,omitempty"`
}
Config describes a transformation. Every section is optional; an empty Config flattens the whole document with "." as the separator.
type Derive ¶
type Derive struct {
From SourceList `json:"from"`
Normalize string `json:"normalize,omitempty"`
DigitPrefix string `json:"digit_prefix,omitempty"`
Reserved []string `json:"reserved,omitempty"`
ReservedPrefix string `json:"reserved_prefix,omitempty"`
// Required fails the record with ErrRequired when the value is empty.
Required bool `json:"required,omitempty"`
When *Cond `json:"when,omitempty"`
}
Derive defines a named value computed once per record and usable as "$name" in sources and output names.
type FlattenConfig ¶
type FlattenConfig struct {
// Separator joins output key segments. Default ".".
Separator string `json:"separator,omitempty"`
// Arrays is ArraysIndex (default), ArraysRaw or ArraysString.
Arrays string `json:"arrays,omitempty"`
// MaxDepth limits how many container levels are flattened. 0 = no limit.
MaxDepth int `json:"max_depth,omitempty"`
// DropNulls leaves out keys whose value is null.
DropNulls bool `json:"drop_nulls,omitempty"`
// DropEmptyObjects leaves out keys whose value is {}.
DropEmptyObjects bool `json:"drop_empty_objects,omitempty"`
}
FlattenConfig controls how nested values become flat keys.
type InputConfig ¶
type InputConfig struct {
// Explode is the path of an array whose elements are the records. A
// document without that path is treated as a single record.
Explode string `json:"explode,omitempty"`
// Inherit lists top-level keys of the enclosing document that a record
// falls back to when it lacks them, for example a batch-level "sentAt".
Inherit []string `json:"inherit,omitempty"`
}
InputConfig says how one document becomes one or more records.
type KeysConfig ¶
type KeysConfig struct {
// Normalize is NormalizeNone (default) or NormalizeSnake.
Normalize string `json:"normalize,omitempty"`
// DigitPrefix is put in front of output keys that start with a digit.
DigitPrefix string `json:"digit_prefix,omitempty"`
// OnCollision is CollisionKeep (default), CollisionFirst or CollisionError.
OnCollision string `json:"on_collision,omitempty"`
// SegmentAliases replace one key segment wherever it appears.
SegmentAliases map[string]string `json:"segment_aliases,omitempty"`
// Aliases replace a full output key.
Aliases []Alias `json:"aliases,omitempty"`
// Drop lists output keys to leave out. A trailing * matches a prefix.
Drop []string `json:"drop,omitempty"`
// Keep lists the output keys to write; every other flattened key is left
// out. Empty keeps all. A trailing * matches a prefix. A key must match
// Keep and not match Drop.
Keep []string `json:"keep,omitempty"`
// Rest names a column that collects the keys Keep left out, as a JSON
// string of one flat object, so that nothing is lost. Written at the end
// of the row, only when something was left out. Needs Keep.
Rest string `json:"rest,omitempty"`
}
KeysConfig corrects and polices output keys.
type Options ¶
type Options struct {
// Now is the value of "$now". The zero value means time.Now().
Now time.Time
// Vars holds the values of "$name" sources that are not derived values.
// The map is only read, so one map can serve many concurrent calls.
Vars map[string]string
}
Options are the per-call inputs of Each.
type Output ¶
type Output struct {
// Name is a literal, or "$name" for a derived value.
Name string `json:"name"`
When *Cond `json:"when,omitempty"`
// Sections limits the output to the sections with these IDs.
Sections []string `json:"sections,omitempty"`
}
Output routes a record to a named destination. A record can match several outputs and then produces one row for each.
type Record ¶
type Record struct {
// Index is the position of the source record in the exploded array.
Index int
// Err is set when the source record failed. JSON and Name are then nil,
// and no rows of that source record are delivered.
Err error
// Name is the name of the matching output. It is nil when the config has
// no outputs.
Name []byte
// JSON is the flat JSON object.
JSON []byte
// Fields holds the values of the "expose" sources, in config order. An
// entry is nil when the source has no value.
Fields [][]byte
// Dropped counts the keys left out because their output key was taken.
Dropped int
}
Record is one output row, or one failed source record. It belongs to the Transformer: the Record and all of its slices are valid only until the callback returns. Copy what you keep.
type Rule ¶
type Rule struct {
Op string `json:"op"`
Path string `json:"path,omitempty"`
From StringList `json:"from,omitempty"`
To string `json:"to,omitempty"`
Sep string `json:"sep,omitempty"`
Mode string `json:"mode,omitempty"`
KeepSources bool `json:"keep_sources,omitempty"`
Value json.RawMessage `json:"value,omitempty"`
}
Rule is one source-side step. Paths in Path and From are source paths and always use "." between segments. To (and Path of a default rule) are output keys and are written exactly as given.
type Section ¶
type Section struct {
ID string `json:"id,omitempty"`
From string `json:"from"`
Prefix string `json:"prefix,omitempty"`
When *Cond `json:"when,omitempty"`
// Quiet stops collisions inside this section from being counted or
// turned into errors.
Quiet bool `json:"quiet,omitempty"`
}
Section flattens one subtree of the record under a key prefix. Without any sections the whole record is flattened with no prefix.
type Source ¶
type Source struct {
Path string `json:"-"`
Fn string `json:"fn,omitempty"`
Sent string `json:"sent,omitempty"`
Original string `json:"original,omitempty"`
// From is the path or "$" variable a map or lower reads.
From string `json:"from,omitempty"`
// Values maps the text of From to any JSON value (map).
Values map[string]json.RawMessage `json:"values,omitempty"`
// Else is the value of a map whose text is not in Values; absent, the
// source is missing on a miss.
Else json.RawMessage `json:"else,omitempty"`
}
Source is where a value comes from: a path in the record ("user.id"), a variable ("$now", "$name" of a derived value, or a key of Options.Vars), or a function ({"fn":"clock_skew","sent":"sentAt","original":"originalTimestamp"}, {"fn":"map","from":"dt","values":{"1":"mobile"},"else":"other"}, {"fn":"lower","from":"$os"}). Paths always use "." between segments.
func (Source) MarshalJSON ¶
MarshalJSON writes a path source as a string.
func (*Source) UnmarshalJSON ¶
UnmarshalJSON accepts a string or a function object.
type SourceList ¶
type SourceList []Source
SourceList unmarshals from one source or an array of sources.
func (*SourceList) UnmarshalJSON ¶
func (l *SourceList) UnmarshalJSON(b []byte) error
UnmarshalJSON implements json.Unmarshaler.
type StringList ¶
type StringList []string
StringList unmarshals from a JSON string or an array of strings.
func (*StringList) UnmarshalJSON ¶
func (l *StringList) UnmarshalJSON(b []byte) error
UnmarshalJSON implements json.Unmarshaler.
type Transformer ¶
type Transformer struct {
// contains filtered or unexported fields
}
Transformer applies one compiled config. It is immutable and safe for concurrent use by any number of goroutines.
func Compile ¶
func Compile(config []byte) (*Transformer, error)
Compile parses a JSON config and builds a Transformer from it. Unknown fields and data after the config object are errors, so a misspelt section name is reported instead of being ignored.
func MustCompile ¶
func MustCompile(config []byte) *Transformer
MustCompile is like Compile but panics on error. It is meant for configs that are part of the program, such as package-level variables.
func New ¶
func New(cfg Config) (*Transformer, error)
New validates cfg and builds a Transformer from it. The Transformer keeps no reference to cfg, so cfg can be changed and used again afterwards.
Example ¶
A Config can be built in Go instead of parsed from JSON. New validates it the same way Compile does. A single source is a SourceList with one entry.
package main
import (
"fmt"
"github.com/nayan9229/jsonflat"
)
func main() {
cfg := jsonflat.Config{
Flatten: jsonflat.FlattenConfig{Separator: "_", DropNulls: true},
Keys: jsonflat.KeysConfig{
Normalize: jsonflat.NormalizeSnake,
OnCollision: jsonflat.CollisionFirst,
Aliases: []jsonflat.Alias{{To: "product_id", From: []string{"prodcutId"}}},
Keep: []string{"id", "product_id", "price_*"},
},
Columns: []jsonflat.Column{
{To: "id", From: jsonflat.SourceList{{Path: "messageId"}}, As: jsonflat.AsString},
},
}
t, err := jsonflat.New(cfg)
if err != nil {
panic(err)
}
src := []byte(`{"messageId": 42, "prodcutId": "P1", "price": {"amount": 9.5, "currency": "EUR"}, "debug": null, "note": "dropped by keep"}`)
out, err := t.Transform(src)
if err != nil {
panic(err)
}
fmt.Println(string(out))
}
Output: {"id":"42","product_id":"P1","price_amount":9.5,"price_currency":"EUR"}
func (*Transformer) Append ¶
func (t *Transformer) Append(dst, src []byte) ([]byte, error)
Append flattens the document src and appends the result to dst. It is for configs that produce exactly one row per document: those without input.explode and without outputs. For any other config it returns ErrNotSimple.
On error, dst is returned unchanged. The result does not alias src.
func (*Transformer) Each ¶
Each turns the document src into rows and calls fn for every row. With input.explode, every element of that array is a source record; a source record that matches several outputs gives one row for each.
A source record that fails is reported as one Record with Err set, and none of its rows are delivered; the remaining records are still processed. Each itself returns an error only when src is not valid JSON, for ErrExplode, or when fn returns an error, which stops the iteration and is returned as is.
The Record passed to fn, and every slice in it, is valid only until fn returns.
Example ¶
Each splits a document into records and routes every record to the outputs whose condition it meets. The callback sees one Record per row, with the output name, the exposed fields and the flat JSON. A record that fails is reported through Record.Err and does not stop the batch.
package main
import (
"fmt"
"github.com/nayan9229/jsonflat"
)
func main() {
t := jsonflat.MustCompile([]byte(`{
"input": {"explode": "items", "inherit": ["orderId"]},
"flatten": {"separator": "_"},
"keys": {"normalize": "snake"},
"columns": [{"to": "order_id", "from": "orderId", "as": "string"}],
"outputs": [
{"name": "items"},
{"name": "gifts", "when": {"path": "gift", "equals": "true"}}
],
"expose": ["sku"]
}`))
body := []byte(`{"orderId": 981, "items": [
{"sku": "A1", "unitPrice": 499.75},
{"sku": "B7", "unitPrice": 5, "gift": "true"},
"not a record"
]}`)
err := t.Each(body, jsonflat.Options{}, func(r *jsonflat.Record) error {
if r.Err != nil {
fmt.Printf("record %d failed: %v\n", r.Index, r.Err)
return nil // one bad record does not fail the batch
}
// r.Name, r.Fields and r.JSON are valid until this function returns.
fmt.Printf("%s sku=%s %s\n", r.Name, r.Fields[0], r.JSON)
return nil
})
if err != nil {
panic(err)
}
}
Output: items sku=A1 {"order_id":"981","sku":"A1","unit_price":499.75} items sku=B7 {"order_id":"981","sku":"B7","unit_price":5,"gift":"true"} gifts sku=B7 {"order_id":"981","sku":"B7","unit_price":5,"gift":"true"} record 2 failed: jsonflat: record is not an object or an array
Directories
¶
| Path | Synopsis |
|---|---|
|
Command example flattens RudderStack-style track events with jsonflat: one flat JSON row per event, with the standard fields as columns and duplicate keys and duplicate fields merged by the config in config.json.
|
Command example flattens RudderStack-style track events with jsonflat: one flat JSON row per event, with the standard fields as columns and duplicate keys and duplicate fields merged by the config in config.json. |
|
internal
|
|
|
doccheck
command
Command doccheck reports exported identifiers without a doc comment, doc comments that do not start with the identifier's name, and packages without a package comment.
|
Command doccheck reports exported identifiers without a doc comment, doc comments that do not start with the identifier's name, and packages without a package comment. |
|
jsonenc
Package jsonenc holds the JSON string escaper and the RFC 8259 number check used by jsonflat.
|
Package jsonenc holds the JSON string escaper and the RFC 8259 number check used by jsonflat. |
|
Package presets holds ready-made jsonflat configs.
|
Package presets holds ready-made jsonflat configs. |