jsonflat

package module
v0.3.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 7, 2026 License: MIT Imports: 13 Imported by: 0

README

jsonflat: zero-allocation JSON flattening and transformation for Go

jsonflat turns nested JSON into flat rows, driven entirely by a JSON config, without allocating on the hot path. Flatten, normalise keys to snake_case, fix misspelt keys, rename, drop, merge and default, whitelist the columns you want, split a batch into records, and route each record to one or more named outputs. It is built for high-throughput pipelines that feed columnar storage. Vendor layouts are configs, not code: the bundled RudderStack preset produces rows that follow the RudderStack warehouse schema, ready for Kinesis, Firehose, Parquet and Athena.

Go Reference Release Go Report Card Go version

Documentation: nayan9229.github.io/jsonflat — config reference, recipes, presets, performance, FAQ. The same pages are in docs/. API reference on pkg.go.dev.

Install

go get github.com/nayan9229/jsonflat

Go 1.24 or newer. The only dependency is github.com/valyala/fastjson.

Quick start

A config is a JSON object; compile it once and share the Transformer.

{
  "rules": [
    {"op": "rename",  "from": "user.id", "to": "uid"},
    {"op": "merge",   "from": ["user.first", "user.last"], "to": "user.name", "sep": " "},
    {"op": "drop",    "path": "debug"},
    {"op": "default", "path": "env", "value": "prod"}
  ]
}
in:  {"user":{"id":7,"first":"Ada","last":"Lovelace"},"tags":["a","b"],"debug":{"x":1}}
out: {"uid":7,"tags.0":"a","tags.1":"b","user.name":"Ada Lovelace","env":"prod"}
t := jsonflat.MustCompile(config)

var out []byte // reuse this buffer between calls
out, err := t.Append(out[:0], src)
if err != nil {
	return err
}
fmt.Println(string(out))

Two ways to run a Transformer:

  • Append(dst, src) / Transform(src): one document in, one flat object out.
  • Each(src, opt, fn): one document in, any number of rows out, one callback per row with the output's name. Needed for configs with input.explode or outputs.

Recipes

Inputs, configs and outputs below are taken from the tests, and docs_test.go replays them, so they are known to be right. More in the recipes page.

Clean keys for a warehouse
{
  "flatten": {"separator": "_", "arrays": "string", "drop_nulls": true, "drop_empty_objects": true},
  "keys": {
    "normalize": "snake",
    "digit_prefix": "_",
    "on_collision": "first",
    "segment_aliases": {"shiping": "shipping"},
    "aliases": [{"to": "product_id", "from": ["prodcutId", "pid"]}],
    "keep": ["product_id", "shipping_*", "_2fa", "products"]
  }
}
in:  {"prodcutId":"P1","shiping":{"zipCode":"380001"},"2fa":true,"coupon":null,"products":[{"sku":"A1"}],"internal":{"trace":"t"}}
out: {"product_id":"P1","shipping_zip_code":"380001","_2fa":true,"products":"[{\"sku\":\"A1\"}]"}

Alias from entries are written the way the producer writes them; keep and drop are written as final column names. Add "rest": "extra" to collect whatever keep left out into one JSON-string column instead of losing it. There is no fuzzy matching, on purpose: a wrong guess would corrupt a column silently.

One row per line item, with inherited order fields
{
  "input": {"explode": "items", "inherit": ["orderId", "currency"]},
  "flatten": {"separator": "_"},
  "keys": {"normalize": "snake"},
  "columns": [
    {"to": "order_id", "from": "orderId", "as": "string"},
    {"to": "currency", "from": ["currency", "$default_currency"]}
  ],
  "expose": ["sku", "orderId"]
}
var default_currency=INR
in: {"orderId": 981, "items": [
      {"sku": "A1", "unitPrice": 499.75, "dims": {"W": 2, "H": 3}},
      {"sku": "B7", "unitPrice": 500.00, "currency": "USD"},
      7
    ]}
row: {"order_id":"981","currency":"INR","sku":"A1","unit_price":499.75,"dims_w":2,"dims_h":3}
  fields: A1,981
row: {"order_id":"981","currency":"USD","sku":"B7","unit_price":500.00}
  dropped: 1
err: ErrRootNotContainer

The item's own currency is a column name, so the flattened copy is dropped and counted. 7 is not a record; it is reported and the batch carries on.

Route logs to several outputs
{
  "columns": [{"to": "ts", "from": ["time", "$now"]}, {"to": "level", "from": "level"}],
  "sections": [
    {"id": "msg",  "from": "payload"},
    {"id": "http", "from": "http", "prefix": "http", "when": {"path": "http", "exists": true}},
    {"id": "err",  "from": "error", "prefix": "error"}
  ],
  "outputs": [
    {"name": "errors", "when": {"path": "level", "in": ["error", "fatal"]}},
    {"name": "access", "sections": ["http"], "when": {"path": "http.status", "exists": true}}
  ]
}
in: {"level":"error","time":"T1","payload":{"msg":"boom"},"http":{"status":500,"path":"/x"},"error":{"kind":"io"},"ignored":1}
row errors: {"ts":"T1","level":"error","msg":"boom","http.status":500,"http.path":"/x","error.kind":"io"}
row access: {"ts":"T1","level":"error","http.status":500,"http.path":"/x"}
RudderStack, via the preset
var cfg jsonflat.Config
if err := json.Unmarshal(presets.RudderStack, &cfg); err != nil {
	panic(err)
}
cfg.Keys.Aliases = []jsonflat.Alias{{To: "product_id", From: []string{"prodcutId", "pid"}}}
cfg.Keys.Drop = []string{"context_ip"}
t, err := jsonflat.New(cfg)
if err != nil {
	panic(err)
}

err = t.Each(body, jsonflat.Options{}, func(r *jsonflat.Record) error {
	if r.Err != nil {
		return deadLetter(r.Index, r.Fields, r.Err) // Fields: messageId, anonymousId, userId
	}
	return putRecord(r.Name, r.Fields[1], r.JSON) // valid until this function returns
})

A track event gives a tracks row and a row in a table named after the event; see presets for the full table, the choices the preset makes, and a warning about Firehose and Glue.

A complete program

example/ turns RudderStack SDK track events into one flat row each: a fixed set of columns with fallback sources, a drop list for copies and junk, and a rest column that catches whatever no column maps. It ships with seven sample events, so it runs with no arguments, and it reads gzip or plain NDJSON files or stdin.

go run ./example -pretty
go run ./example -in events.ndjson.gz -out flat.ndjson -verify

Config reference

Every section and field, defaults, the order things happen in, validation errors and a troubleshooting checklist: config reference (docs/config-reference.md).

The one rule to know before reading it: source paths and output keys are different things. Source paths address the input (rules[].from, sections[].from, every source) and always use . between segments, with the keys spelt as the producer spells them. Output keys are what gets written (rules[].to, columns[].to, aliases[].to, keys.keep, keys.drop) and are used exactly as given, after normalisation.

Guarantees

  • Zero heap allocations in Append and Each in the steady state, which means from the third call on a pooled state, once its buffers have grown to the size of your traffic. TestZeroAllocs, TestRudderZeroAllocs and TestKeepZeroAllocs enforce it, and the release workflow refuses a build whose benchmarks allocate.
  • The output is always valid JSON. The package has its own RFC 8259 string escaper and number check because the parser's are more lenient; a number such as 00 fails the record with ErrInvalidNumber. FuzzAppend and FuzzEach assert json.Valid on every output.
  • A Transformer is immutable and safe for concurrent use. TestConcurrent, TestSharedTransformerMixed; see Concurrency.
  • The output never aliases the input, and Append returns dst unchanged on error. TestAppendKeepsPrefixAndInput, TestErrors.
  • All-or-nothing rows per record. One bad record never fails the batch, and a record never arrives in part: all of its rows are built first, and if any fails, one Record with Err set is delivered instead. TestRudderBatch.
  • A Record is valid only until the callback returns. Copy what you keep. TestRecordIsACopy.

Details, including memory behaviour and the 64-bit hash used for collision tracking, are on the performance page.

Concurrency

Compile once, share everywhere. A Transformer is read-only after Compile or New; all per-call state lives in a sync.Pool it owns, so any number of goroutines can call Append and Each on the same one with no lock in the hot path. TestSharedTransformerMixed runs 16 goroutines mixing both methods and different Options under the race detector, and BenchmarkAppendParallel scales 3.9× on 4 cores and 6.1× on 10 (4 performance + 6 efficiency cores; tables on the performance page).

  • A package-level var t = jsonflat.MustCompile(config) runs at init, which is once-only and goroutine-safe. No sync.Once is needed.
  • To compile lazily, for example from a file read on first use, wrap it in your code: sync.OnceValues(func() (*jsonflat.Transformer, error) { ... }) (Example_lazyCompile).
  • Options.Vars is only read. Build one map per route and share it; a map per call costs two allocations.
  • A Record and its slices are valid until the callback returns. After that the pooled buffers are reused by the next call, possibly on another goroutine (TestRetainedRecordIsInvalid shows the bytes change). Copy what you keep.
  • Zero allocations is a steady-state property: the first two calls on a freshly built pooled state allocate (17–42 KB in total), and a GC cycle drops pooled states. A pooled state holds about 8× the largest input it has seen; max_pooled_input and a shrink rule bound that (performance page).

When not to use this

  • Documents above a few MB: the whole document is parsed into memory, and a pooled state keeps about 8× that size.
  • Streaming from an io.Reader: the API takes a complete []byte (read lines and call Each per line, as example/ does).
  • You need to keep the nesting: this package only flattens.
  • You need type casting today: values keep the type they came with.

Versioning

Semantic versioning, tags vMAJOR.MINOR.PATCH, one changelog entry per version.

  • v0.x: a minor version may change the config language or the Go API, and the changelog says so; a patch version only fixes. Pin a minor (go get github.com/nayan9229/jsonflat@v0.1) if you want no surprises.
  • v1.0.0 freezes the public API and the meaning of every config field. Additions come in minors.
  • A breaking change after v1 ships as a new module path, github.com/nayan9229/jsonflat/v2, and the v1 line keeps receiving fixes.
  • Presets follow the same rules: a change to what a preset produces is a minor before v1 and a major after.

Reporting bugs and security issues

  • Bugs and feature requests: open an issue. The templates ask for the config, the input and the expected output, which is exactly what becomes a test case.
  • Vulnerabilities: privately, via GitHub security advisories. See SECURITY.md for what counts and the known limits.

Contributing

See CONTRIBUTING.md: make all runs the same checks as CI, and there is a section on how to propose a preset. Maintainers release with RELEASING.md.

License

MIT

Documentation

Overview

Package jsonflat flattens JSON documents and reshapes them with a JSON config: it normalises and corrects keys, renames, drops, merges and defaults them, splits batches into records and routes records to named outputs, with zero heap allocations on the hot path.

A config is compiled once into a Transformer, which is immutable and safe for concurrent use:

t := jsonflat.MustCompile([]byte(`{
  "flatten": {"separator": "_"},
  "keys":    {"normalize": "snake"},
  "rules":   [{"op": "drop", "path": "debug"}]
}`))

An empty config, {}, flattens the whole document with "." between key segments. Every section of the config is optional and they combine freely; see Config for all of them.

One document, one row

Append and Transform turn one document into one flat object. Append writes into a buffer that the caller reuses, which is what keeps the call free of allocations:

out, err = t.Append(out[:0], src)

One document, many rows

Each is for configs that split a batch into records (input.explode) or route a record to several outputs (outputs). It calls a function for every row and hands it the name of the output, the JSON, and the values the config asked to expose, for example a partition key:

err := t.Each(body, jsonflat.Options{}, func(r *jsonflat.Record) error {
	if r.Err != nil {
		return deadLetter(r.Index, r.Fields, r.Err)
	}
	return put(r.Name, r.Fields[0], r.JSON)
})

A record that fails is reported through Record.Err and never stops the batch; none of its rows are delivered, so a record never arrives in part. The Record and its slices are valid only until the callback returns.

Two kinds of path

Source paths address the input (rule path and from, section from, input.explode, and every Source). They always use "." between segments, whatever flatten.separator is, and use the keys as the input spells them. Output keys are what gets written (rule to, the path of a default rule, column to, alias to, keys.keep, keys.drop) and are used exactly as given.

The full config reference, recipes and the preset documentation are at https://nayan9229.github.io/jsonflat/ and in the docs/ directory of the repository.

Guarantees

The output is always valid JSON and never aliases the input. Number text is copied through unchanged, so no precision is lost, and is checked against RFC 8259 because the parser is more lenient than that. On error, Append returns dst unchanged.

Vendor layouts are configs, not code: see the presets package.

Example
package main

import (
	"fmt"

	"github.com/nayan9229/jsonflat"
)

func main() {
	t := jsonflat.MustCompile([]byte(`{
	  "rules": [
	    {"op": "rename",  "from": "user.id", "to": "uid"},
	    {"op": "merge",   "from": ["user.first", "user.last"], "to": "user.name", "sep": " "},
	    {"op": "drop",    "path": "debug"},
	    {"op": "default", "path": "env", "value": "prod"}
	  ]
	}`))

	src := []byte(`{"user":{"id":7,"first":"Ada","last":"Lovelace"},"tags":["a","b"],"debug":{"x":1}}`)

	var out []byte // reuse this buffer between calls
	out, err := t.Append(out[:0], src)
	if err != nil {
		panic(err)
	}
	fmt.Println(string(out))
}
Output:
{"uid":7,"tags.0":"a","tags.1":"b","user.name":"Ada Lovelace","env":"prod"}
Example (LazyCompile)
package main

import (
	"fmt"
	"sync"

	"github.com/nayan9229/jsonflat"
)

// lazy compiles the config on first use and hands every later caller the
// same Transformer and error. This is the whole of what a program needs for
// a config that is read from disk or a flag at run time; a config that is
// part of the program is simpler still as a package-level
// jsonflat.MustCompile, which runs at init.
var lazy = sync.OnceValues(func() (*jsonflat.Transformer, error) {
	config := []byte(`{"keys": {"normalize": "snake"}}`) // os.ReadFile in a real program
	return jsonflat.Compile(config)
})

func main() {
	t, err := lazy()
	if err != nil {
		fmt.Println("config:", err)
		return
	}
	out, err := t.Transform([]byte(`{"userId": 7, "orderTotal": 9.5}`))
	if err != nil {
		fmt.Println(err)
		return
	}
	fmt.Println(string(out))
}
Output:
{"user_id":7,"order_total":9.5}
Example (RudderStackPreset)

A preset is plain config. Unmarshal it, add your own key corrections, and build the Transformer.

package main

import (
	"encoding/json"
	"fmt"
	"time"

	"github.com/nayan9229/jsonflat"
	"github.com/nayan9229/jsonflat/presets"
)

func main() {
	var cfg jsonflat.Config
	if err := json.Unmarshal(presets.RudderStack, &cfg); err != nil {
		panic(err)
	}
	cfg.Keys.Aliases = []jsonflat.Alias{{To: "product_id", From: []string{"prodcutId", "pid"}}}
	cfg.Keys.Drop = []string{"context_ip"}
	t, err := jsonflat.New(cfg)
	if err != nil {
		panic(err)
	}

	body := []byte(`{"batch":[{
	  "type": "track", "event": "Product Added", "messageId": "m-1", "anonymousId": "anon-1",
	  "context": {"ip": "1.2.3.4", "app": {"name": "Shop"}},
	  "properties": {"prodcutId": "P1", "Total Price": 499.75, "user_id": "ignored"}
	}]}`)

	opt := jsonflat.Options{Now: time.Date(2026, 9, 21, 10, 0, 0, 0, time.UTC)}
	err = t.Each(body, opt, func(r *jsonflat.Record) error {
		if r.Err != nil {
			fmt.Println("bad record", r.Index, r.Err)
			return nil
		}
		// r.Fields is messageId, anonymousId, userId. Copy what you keep.
		fmt.Printf("%s key=%s -> %s\n", r.Name, r.Fields[1], r.JSON)
		return nil
	})
	if err != nil {
		panic(err)
	}
}
Output:
tracks key=anon-1 -> {"id":"m-1","anonymous_id":"anon-1","received_at":"2026-09-21T10:00:00.000Z","timestamp":"2026-09-21T10:00:00.000Z","event":"product_added","event_text":"Product Added","context_app_name":"Shop"}
product_added key=anon-1 -> {"id":"m-1","anonymous_id":"anon-1","received_at":"2026-09-21T10:00:00.000Z","timestamp":"2026-09-21T10:00:00.000Z","event":"product_added","event_text":"Product Added","context_app_name":"Shop","product_id":"P1","total_price":499.75}

Index

Examples

Constants

View Source
const (
	OpRename  = "rename"
	OpDrop    = "drop"
	OpMerge   = "merge"
	OpDefault = "default"

	MergeConcat = "concat"
	MergeArray  = "array"
	MergeFirst  = "first"

	ArraysIndex  = "index"  // flatten arrays as key.0, key.1
	ArraysRaw    = "raw"    // keep arrays as JSON values
	ArraysString = "string" // write arrays as a JSON string, for typed columns

	NormalizeNone  = "none"
	NormalizeSnake = "snake"

	CollisionKeep  = "keep"  // write duplicates as they come (default)
	CollisionFirst = "first" // the first value wins, later ones are counted
	CollisionError = "error" // the record fails with ErrCollision

	AsRaw    = "raw"    // column value is written as it is
	AsString = "string" // strings and numbers are written as strings
	AsNumber = "number" // the first text that is a JSON number, as a number
	AsBool   = "bool"   // the first text that is true or false, as a boolean

	FnClockSkew = "clock_skew"
	FnMap       = "map"
	FnLower     = "lower"
)

Rule operations, merge modes and the other config keywords.

Variables

View Source
var (
	// ErrRootNotContainer means the record is a scalar, not an object or array.
	ErrRootNotContainer = errors.New("jsonflat: record is not an object or an array")
	// ErrInvalidNumber means the record holds a number that is not valid JSON,
	// such as 00 or NaN. The parser accepts these; the output must not.
	ErrInvalidNumber = errors.New("jsonflat: invalid number")
	// ErrCollision means two keys got the same name under on_collision "error".
	ErrCollision = errors.New("jsonflat: key collision")
	// ErrRequired means a required derived value came out empty.
	ErrRequired = errors.New("jsonflat: required value is empty")
	// ErrNoOutput means no configured output matched the record.
	ErrNoOutput = errors.New("jsonflat: no output matches the record")
)

Errors that concern one record. Each reports them in Record.Err and carries on with the next record; Append returns them.

View Source
var (
	// ErrExplode means the input.explode path exists but is not an array.
	ErrExplode = errors.New("jsonflat: input.explode path is not an array")
	// ErrNotSimple is returned by Append and Transform for a config that can
	// produce several rows: one with input.explode or with outputs. Use Each.
	ErrNotSimple = errors.New("jsonflat: the config can produce several rows; use Each")
)

Errors that concern the whole call.

Functions

This section is empty.

Types

type Alias

type Alias struct {
	To   string   `json:"to"`
	From []string `json:"from"`
	When *Cond    `json:"when,omitempty"`
}

Alias maps wrong output keys to the right one. With snake normalisation the From entries are normalised when the config is compiled, so they can be written the way the producer writes them.

type Column

type Column struct {
	To string `json:"to"`
	// From lists sources in order of preference; the first with a value wins.
	From SourceList `json:"from"`
	// As is AsRaw (default) or AsString.
	As   string `json:"as,omitempty"`
	When *Cond  `json:"when,omitempty"`
	// Quiet stops keys dropped in favour of this column from being counted.
	Quiet bool `json:"quiet,omitempty"`
}

Column is an explicit output key written before any section. Column names are reserved: a flattened key with the same name is dropped, whether or not the column had a value.

type Cond

type Cond struct {
	Path   SourceList `json:"path"`
	Equals *string    `json:"equals,omitempty"`
	In     []string   `json:"in,omitempty"`
	Exists *bool      `json:"exists,omitempty"`
}

Cond is a test on one value of the record. Exactly one of Equals, In and Exists must be set.

type Config

type Config struct {
	Input    InputConfig       `json:"input,omitempty"`
	Flatten  FlattenConfig     `json:"flatten,omitempty"`
	Keys     KeysConfig        `json:"keys,omitempty"`
	Derive   map[string]Derive `json:"derive,omitempty"`
	Columns  []Column          `json:"columns,omitempty"`
	Sections []Section         `json:"sections,omitempty"`
	Rules    []Rule            `json:"rules,omitempty"`
	Outputs  []Output          `json:"outputs,omitempty"`
	Expose   SourceList        `json:"expose,omitempty"`
	// Newline appends "\n" to every record.
	Newline bool `json:"newline,omitempty"`
	// MaxPooledInput is the largest input, in bytes, whose per-call state is
	// kept for reuse after the call. That state holds about eight times the
	// input size, once per active P. 0 means 4 MiB.
	MaxPooledInput int `json:"max_pooled_input,omitempty"`
}

Config describes a transformation. Every section is optional; an empty Config flattens the whole document with "." as the separator.

type Derive

type Derive struct {
	From           SourceList `json:"from"`
	Normalize      string     `json:"normalize,omitempty"`
	DigitPrefix    string     `json:"digit_prefix,omitempty"`
	Reserved       []string   `json:"reserved,omitempty"`
	ReservedPrefix string     `json:"reserved_prefix,omitempty"`
	// Required fails the record with ErrRequired when the value is empty.
	Required bool  `json:"required,omitempty"`
	When     *Cond `json:"when,omitempty"`
}

Derive defines a named value computed once per record and usable as "$name" in sources and output names.

type FlattenConfig

type FlattenConfig struct {
	// Separator joins output key segments. Default ".".
	Separator string `json:"separator,omitempty"`
	// Arrays is ArraysIndex (default), ArraysRaw or ArraysString.
	Arrays string `json:"arrays,omitempty"`
	// MaxDepth limits how many container levels are flattened. 0 = no limit.
	MaxDepth int `json:"max_depth,omitempty"`
	// DropNulls leaves out keys whose value is null.
	DropNulls bool `json:"drop_nulls,omitempty"`
	// DropEmptyObjects leaves out keys whose value is {}.
	DropEmptyObjects bool `json:"drop_empty_objects,omitempty"`
}

FlattenConfig controls how nested values become flat keys.

type InputConfig

type InputConfig struct {
	// Explode is the path of an array whose elements are the records. A
	// document without that path is treated as a single record.
	Explode string `json:"explode,omitempty"`
	// Inherit lists top-level keys of the enclosing document that a record
	// falls back to when it lacks them, for example a batch-level "sentAt".
	Inherit []string `json:"inherit,omitempty"`
}

InputConfig says how one document becomes one or more records.

type KeysConfig

type KeysConfig struct {
	// Normalize is NormalizeNone (default) or NormalizeSnake.
	Normalize string `json:"normalize,omitempty"`
	// DigitPrefix is put in front of output keys that start with a digit.
	DigitPrefix string `json:"digit_prefix,omitempty"`
	// OnCollision is CollisionKeep (default), CollisionFirst or CollisionError.
	OnCollision string `json:"on_collision,omitempty"`
	// SegmentAliases replace one key segment wherever it appears.
	SegmentAliases map[string]string `json:"segment_aliases,omitempty"`
	// Aliases replace a full output key.
	Aliases []Alias `json:"aliases,omitempty"`
	// Drop lists output keys to leave out. A trailing * matches a prefix.
	Drop []string `json:"drop,omitempty"`
	// Keep lists the output keys to write; every other flattened key is left
	// out. Empty keeps all. A trailing * matches a prefix. A key must match
	// Keep and not match Drop.
	Keep []string `json:"keep,omitempty"`
	// Rest names a column that collects the keys Keep left out, as a JSON
	// string of one flat object, so that nothing is lost. Written at the end
	// of the row, only when something was left out. Needs Keep.
	Rest string `json:"rest,omitempty"`
}

KeysConfig corrects and polices output keys.

type Options

type Options struct {
	// Now is the value of "$now". The zero value means time.Now().
	Now time.Time
	// Vars holds the values of "$name" sources that are not derived values.
	// The map is only read, so one map can serve many concurrent calls.
	Vars map[string]string
}

Options are the per-call inputs of Each.

type Output

type Output struct {
	// Name is a literal, or "$name" for a derived value.
	Name string `json:"name"`
	When *Cond  `json:"when,omitempty"`
	// Sections limits the output to the sections with these IDs.
	Sections []string `json:"sections,omitempty"`
}

Output routes a record to a named destination. A record can match several outputs and then produces one row for each.

type Record

type Record struct {
	// Index is the position of the source record in the exploded array.
	Index int
	// Err is set when the source record failed. JSON and Name are then nil,
	// and no rows of that source record are delivered.
	Err error
	// Name is the name of the matching output. It is nil when the config has
	// no outputs.
	Name []byte
	// JSON is the flat JSON object.
	JSON []byte
	// Fields holds the values of the "expose" sources, in config order. An
	// entry is nil when the source has no value.
	Fields [][]byte
	// Dropped counts the keys left out because their output key was taken.
	Dropped int
}

Record is one output row, or one failed source record. It belongs to the Transformer: the Record and all of its slices are valid only until the callback returns. Copy what you keep.

type Rule

type Rule struct {
	Op          string          `json:"op"`
	Path        string          `json:"path,omitempty"`
	From        StringList      `json:"from,omitempty"`
	To          string          `json:"to,omitempty"`
	Sep         string          `json:"sep,omitempty"`
	Mode        string          `json:"mode,omitempty"`
	KeepSources bool            `json:"keep_sources,omitempty"`
	Value       json.RawMessage `json:"value,omitempty"`
}

Rule is one source-side step. Paths in Path and From are source paths and always use "." between segments. To (and Path of a default rule) are output keys and are written exactly as given.

type Section

type Section struct {
	ID     string `json:"id,omitempty"`
	From   string `json:"from"`
	Prefix string `json:"prefix,omitempty"`
	When   *Cond  `json:"when,omitempty"`
	// Quiet stops collisions inside this section from being counted or
	// turned into errors.
	Quiet bool `json:"quiet,omitempty"`
}

Section flattens one subtree of the record under a key prefix. Without any sections the whole record is flattened with no prefix.

type Source

type Source struct {
	Path     string `json:"-"`
	Fn       string `json:"fn,omitempty"`
	Sent     string `json:"sent,omitempty"`
	Original string `json:"original,omitempty"`
	// From is the path or "$" variable a map or lower reads.
	From string `json:"from,omitempty"`
	// Values maps the text of From to any JSON value (map).
	Values map[string]json.RawMessage `json:"values,omitempty"`
	// Else is the value of a map whose text is not in Values; absent, the
	// source is missing on a miss.
	Else json.RawMessage `json:"else,omitempty"`
}

Source is where a value comes from: a path in the record ("user.id"), a variable ("$now", "$name" of a derived value, or a key of Options.Vars), or a function ({"fn":"clock_skew","sent":"sentAt","original":"originalTimestamp"}, {"fn":"map","from":"dt","values":{"1":"mobile"},"else":"other"}, {"fn":"lower","from":"$os"}). Paths always use "." between segments.

func (Source) MarshalJSON

func (s Source) MarshalJSON() ([]byte, error)

MarshalJSON writes a path source as a string.

func (*Source) UnmarshalJSON

func (s *Source) UnmarshalJSON(b []byte) error

UnmarshalJSON accepts a string or a function object.

type SourceList

type SourceList []Source

SourceList unmarshals from one source or an array of sources.

func (*SourceList) UnmarshalJSON

func (l *SourceList) UnmarshalJSON(b []byte) error

UnmarshalJSON implements json.Unmarshaler.

type StringList

type StringList []string

StringList unmarshals from a JSON string or an array of strings.

func (*StringList) UnmarshalJSON

func (l *StringList) UnmarshalJSON(b []byte) error

UnmarshalJSON implements json.Unmarshaler.

type Transformer

type Transformer struct {
	// contains filtered or unexported fields
}

Transformer applies one compiled config. It is immutable and safe for concurrent use by any number of goroutines.

func Compile

func Compile(config []byte) (*Transformer, error)

Compile parses a JSON config and builds a Transformer from it. Unknown fields and data after the config object are errors, so a misspelt section name is reported instead of being ignored.

func MustCompile

func MustCompile(config []byte) *Transformer

MustCompile is like Compile but panics on error. It is meant for configs that are part of the program, such as package-level variables.

func New

func New(cfg Config) (*Transformer, error)

New validates cfg and builds a Transformer from it. The Transformer keeps no reference to cfg, so cfg can be changed and used again afterwards.

Example

A Config can be built in Go instead of parsed from JSON. New validates it the same way Compile does. A single source is a SourceList with one entry.

package main

import (
	"fmt"

	"github.com/nayan9229/jsonflat"
)

func main() {
	cfg := jsonflat.Config{
		Flatten: jsonflat.FlattenConfig{Separator: "_", DropNulls: true},
		Keys: jsonflat.KeysConfig{
			Normalize:   jsonflat.NormalizeSnake,
			OnCollision: jsonflat.CollisionFirst,
			Aliases:     []jsonflat.Alias{{To: "product_id", From: []string{"prodcutId"}}},
			Keep:        []string{"id", "product_id", "price_*"},
		},
		Columns: []jsonflat.Column{
			{To: "id", From: jsonflat.SourceList{{Path: "messageId"}}, As: jsonflat.AsString},
		},
	}
	t, err := jsonflat.New(cfg)
	if err != nil {
		panic(err)
	}

	src := []byte(`{"messageId": 42, "prodcutId": "P1", "price": {"amount": 9.5, "currency": "EUR"}, "debug": null, "note": "dropped by keep"}`)
	out, err := t.Transform(src)
	if err != nil {
		panic(err)
	}
	fmt.Println(string(out))
}
Output:
{"id":"42","product_id":"P1","price_amount":9.5,"price_currency":"EUR"}

func (*Transformer) Append

func (t *Transformer) Append(dst, src []byte) ([]byte, error)

Append flattens the document src and appends the result to dst. It is for configs that produce exactly one row per document: those without input.explode and without outputs. For any other config it returns ErrNotSimple.

On error, dst is returned unchanged. The result does not alias src.

func (*Transformer) Each

func (t *Transformer) Each(src []byte, opt Options, fn func(*Record) error) error

Each turns the document src into rows and calls fn for every row. With input.explode, every element of that array is a source record; a source record that matches several outputs gives one row for each.

A source record that fails is reported as one Record with Err set, and none of its rows are delivered; the remaining records are still processed. Each itself returns an error only when src is not valid JSON, for ErrExplode, or when fn returns an error, which stops the iteration and is returned as is.

The Record passed to fn, and every slice in it, is valid only until fn returns.

Example

Each splits a document into records and routes every record to the outputs whose condition it meets. The callback sees one Record per row, with the output name, the exposed fields and the flat JSON. A record that fails is reported through Record.Err and does not stop the batch.

package main

import (
	"fmt"

	"github.com/nayan9229/jsonflat"
)

func main() {
	t := jsonflat.MustCompile([]byte(`{
	  "input":   {"explode": "items", "inherit": ["orderId"]},
	  "flatten": {"separator": "_"},
	  "keys":    {"normalize": "snake"},
	  "columns": [{"to": "order_id", "from": "orderId", "as": "string"}],
	  "outputs": [
	    {"name": "items"},
	    {"name": "gifts", "when": {"path": "gift", "equals": "true"}}
	  ],
	  "expose": ["sku"]
	}`))

	body := []byte(`{"orderId": 981, "items": [
	  {"sku": "A1", "unitPrice": 499.75},
	  {"sku": "B7", "unitPrice": 5, "gift": "true"},
	  "not a record"
	]}`)

	err := t.Each(body, jsonflat.Options{}, func(r *jsonflat.Record) error {
		if r.Err != nil {
			fmt.Printf("record %d failed: %v\n", r.Index, r.Err)
			return nil // one bad record does not fail the batch
		}
		// r.Name, r.Fields and r.JSON are valid until this function returns.
		fmt.Printf("%s sku=%s %s\n", r.Name, r.Fields[0], r.JSON)
		return nil
	})
	if err != nil {
		panic(err)
	}
}
Output:
items sku=A1 {"order_id":"981","sku":"A1","unit_price":499.75}
items sku=B7 {"order_id":"981","sku":"B7","unit_price":5,"gift":"true"}
gifts sku=B7 {"order_id":"981","sku":"B7","unit_price":5,"gift":"true"}
record 2 failed: jsonflat: record is not an object or an array

func (*Transformer) Transform

func (t *Transformer) Transform(src []byte) ([]byte, error)

Transform is Append with a nil dst.

Directories

Path Synopsis
Command example flattens RudderStack-style track events with jsonflat: one flat JSON row per event, with the standard fields as columns and duplicate keys and duplicate fields merged by the config in config.json.
Command example flattens RudderStack-style track events with jsonflat: one flat JSON row per event, with the standard fields as columns and duplicate keys and duplicate fields merged by the config in config.json.
internal
doccheck command
Command doccheck reports exported identifiers without a doc comment, doc comments that do not start with the identifier's name, and packages without a package comment.
Command doccheck reports exported identifiers without a doc comment, doc comments that do not start with the identifier's name, and packages without a package comment.
jsonenc
Package jsonenc holds the JSON string escaper and the RFC 8259 number check used by jsonflat.
Package jsonenc holds the JSON string escaper and the RFC 8259 number check used by jsonflat.
Package presets holds ready-made jsonflat configs.
Package presets holds ready-made jsonflat configs.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL