i8

package
v1.4.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Jul 18, 2026 License: MIT Imports: 2 Imported by: 1

Documentation

Overview

Package i8 provides SIMD-accelerated operations on int8 slices.

int8 is the 8-bit signed integer workhorse of quantized numeric pipelines. Its narrow range (-128..127) makes element-wise arithmetic overflow almost immediately, so this package does not mirror the wrapping arithmetic of the i16/i32 packages one-to-one. Instead it ships the operations that are genuinely high-impact and well-defined at 8-bit width:

  • Saturating arithmetic (AddSaturate, SubSaturate, and the scalar-broadcast AddScalarSaturate, SubScalarSaturate): single hardware instructions (PADDSB/PSUBSB, SQADD/SQSUB) that clamp to [-128, 127] instead of wrapping, which is what 8-bit arithmetic almost always wants.
  • int32-accumulated reductions (Sum, DotProduct, SumAbs, SAD): widen to int32 so the running total has headroom. DotProduct is the inner loop of quantized matmul/conv; it uses ARM64 SDOT (FEAT_DotProd) where available and AVX2 VPMADDWD otherwise. SumAbs is the L1 norm and SAD the sum of absolute differences (block matching), both via PSADBW on AVX2.
  • Signed min/max (MinMax reduction; element-wise two-slice Min/Max).
  • Element-wise Clamp (activation clipping) and saturating Abs/Neg, where -128 maps to 127 (SQABS/SQNEG on NEON; saturating constructions on AVX2).
  • Saturating AbsDiff (|a-b| clamped to [0,127]) and MaxAbs (the per-tensor abs-max for dynamic quantization, returned as int because |-128| = 128).
  • Sign-extending widening (ToInt16, ToInt32) to hand off to the wider integer or float packages.

Sum and DotProduct accumulate in int32 with two's-complement wraparound, exactly like their pure-Go references. int32 wrapping addition is associative and commutative modulo 2^32, so the lane-parallel SIMD reductions are bit-identical to the scalar reference regardless of summation order. The intermediate products never overflow their SIMD lane (|int8 * int8| <= 16384), so only the final running total can wrap, and it wraps identically.

All functions automatically select the optimal implementation based on runtime CPU feature detection and fall back to a pure-Go implementation on unsupported architectures.

Thread Safety: All functions are safe for concurrent use. Memory: All functions are zero-allocation (no heap allocations).

Index

Examples

Constants

This section is empty.

Variables

This section is empty.

Functions

func Abs

func Abs(dst, a []int8)

Abs writes the saturating absolute value dst[i] = |a[i]| for i in [0, n), n = min(len(dst), len(a)). abs(-128) saturates to 127 (SQABS on NEON; on AVX2 max(a, saturating(0-a))). Any trailing capacity in dst is left untouched.

func AbsDiff

func AbsDiff(dst, a, b []int8)

AbsDiff writes the saturating absolute difference dst[i] = |a[i] - b[i]|, clamped to [0, 127], for i in [0, n), n = min(len(dst), len(a), len(b)). |127 - (-128)| = 255 saturates to 127, consistent with Abs. It uses max(saturating(a-b), saturating(b-a)) on AVX2 and SABD then an unsigned min with 127 on NEON. Any trailing capacity in dst is left untouched.

func AddSaturate

func AddSaturate(dst, a, b []int8)

AddSaturate writes dst[i] = clamp(int(a[i]) + int(b[i]), -128, 127) for i in [0, n), n = min(len(dst), len(a), len(b)). The add saturates to the int8 range instead of wrapping, so 100 + 100 = 127 and -100 + -100 = -128. Any trailing capacity in dst is left untouched.

Example
package main

import (
	"fmt"

	"github.com/tphakala/simd/i8"
)

func main() {
	dst := make([]int8, 3)
	i8.AddSaturate(dst, []int8{100, -100, 1}, []int8{100, -100, 2})
	// Saturates to the int8 range instead of wrapping.
	fmt.Println(dst)
}
Output:
[127 -128 3]

func AddScalarSaturate

func AddScalarSaturate(dst, a []int8, s int8)

AddScalarSaturate writes dst[i] = clamp(int(a[i]) + int(s), -128, 127) for i in [0, n), n = min(len(dst), len(a)). It broadcasts the scalar s and adds with signed saturation (VPADDSB on AVX2, SQADD on NEON). Any trailing capacity in dst is left untouched.

func Clamp

func Clamp(dst, src []int8, lo, hi int8)

Clamp writes dst[i] = min(max(src[i], lo), hi) (signed) for i in [0, n), n = min(len(dst), len(src)). It is the activation-clipping primitive. If lo > hi every element maps to hi (max-then-min ordering). Any trailing capacity in dst is left untouched.

func DotProduct

func DotProduct(a, b []int8) int32

DotProduct returns sum_i int32(a[i]) * int32(b[i]) over i in [0, n), n = min(len(a), len(b)), accumulated in int32 with two's-complement wraparound. An empty operand returns 0. a and b are read-only; the call allocates nothing.

This is the inner loop of quantized matmul/convolution. On ARM64 with FEAT_DotProd it uses SDOT (16 int8 multiply-accumulates per instruction); on AVX2 it widens with VPMOVSXBW and reduces with VPMADDWD.

Example
package main

import (
	"fmt"

	"github.com/tphakala/simd/i8"
)

func main() {
	a := []int8{1, 2, 3, 4, -5}
	b := []int8{10, 20, 30, 40, 50}
	// int32 accumulation: 10 + 40 + 90 + 160 - 250 = 50.
	fmt.Println(i8.DotProduct(a, b))
}
Output:
50

func Max

func Max(dst, a, b []int8)

Max writes dst[i] = max(a[i], b[i]) (signed) for i in [0, n), n = min(len(dst), len(a), len(b)). This is the element-wise two-slice maximum (PMAXSB/SMAX), distinct from the MinMax reduction. Any trailing capacity in dst is left untouched.

func MaxAbs

func MaxAbs(a []int8) int

MaxAbs returns max_i |a[i]| accumulated as int (range [0, 128], since |-128| = 128 does not fit int8). It is the per-tensor scale for dynamic quantization (PABSB+PMAXUB on AVX2; ABS+UMAXV on NEON). An empty a returns 0. a is read-only; the call allocates nothing.

func Min

func Min(dst, a, b []int8)

Min writes dst[i] = min(a[i], b[i]) (signed) for i in [0, n), n = min(len(dst), len(a), len(b)). This is the element-wise two-slice minimum (PMINSB/SMIN), distinct from the MinMax reduction. Any trailing capacity in dst is left untouched.

func MinMax

func MinMax(a []int8) (minVal, maxVal int8)

MinMax returns the smallest and largest int8 in a:

minVal = min_i a[i],  maxVal = max_i a[i]

Both are signed comparisons. An empty a returns (0, 0). a is read-only; the call allocates nothing.

Example
package main

import (
	"fmt"

	"github.com/tphakala/simd/i8"
)

func main() {
	lo, hi := i8.MinMax([]int8{0, -128, 127, 3, -1})
	fmt.Println(lo, hi)
}
Output:
-128 127

func Neg

func Neg(dst, a []int8)

Neg writes the saturating negation dst[i] = -a[i] for i in [0, n), n = min(len(dst), len(a)). neg(-128) saturates to 127 (SQNEG on NEON; saturating(0-a) via VPSUBSB on AVX2). Any trailing capacity in dst is left untouched.

func SAD

func SAD(a, b []int8) int32

SAD returns sum_i |a[i] - b[i]| (the sum of absolute differences) over i in [0, n), n = min(len(a), len(b)), accumulated in int32 with two's-complement wraparound. The per-element difference is the true |a-b| in [0, 255] (not saturated), so |127 - (-128)| contributes 255. An empty operand returns 0. a and b are read-only; the call allocates nothing.

SAD is the block-matching / feature-distance reduction (the scalar companion to AbsDiff). On AVX2 it offsets both operands by 128 and uses PSADBW; on NEON, SABD then UADDLP/UADALP widen-accumulate.

func SubSaturate

func SubSaturate(dst, a, b []int8)

SubSaturate writes dst[i] = clamp(int(a[i]) - int(b[i]), -128, 127) for i in [0, n), n = min(len(dst), len(a), len(b)). The subtract saturates to the int8 range instead of wrapping. Any trailing capacity in dst is left untouched.

func SubScalarSaturate

func SubScalarSaturate(dst, a []int8, s int8)

SubScalarSaturate writes dst[i] = clamp(int(a[i]) - int(s), -128, 127) for i in [0, n), n = min(len(dst), len(a)). It broadcasts the scalar s and subtracts with signed saturation (VPSUBSB on AVX2, SQSUB on NEON). Any trailing capacity in dst is left untouched.

func Sum

func Sum(a []int8) int32

Sum returns the sum of all elements of a, accumulated in int32 with two's-complement wraparound. An empty a returns 0. a is read-only; the call allocates nothing.

func SumAbs

func SumAbs(a []int8) int32

SumAbs returns sum_i |a[i]| (the L1 norm), accumulated in int32 with two's-complement wraparound. |-128| = 128 contributes the full 128. An empty a returns 0. a is read-only; the call allocates nothing.

On AVX2 it uses PABSB then PSADBW (sum of absolute differences against zero); on NEON, ABS then UADDLP/UADALP widen-accumulate.

func ToInt16

func ToInt16(dst []int16, src []int8)

ToInt16 sign-extends src into dst: dst[i] = int16(src[i]) for i in [0, n), n = min(len(dst), len(src)). It is exact (int8 fits in int16). Any trailing capacity in dst is left untouched.

func ToInt32

func ToInt32(dst []int32, src []int8)

ToInt32 sign-extends src into dst: dst[i] = int32(src[i]) for i in [0, n), n = min(len(dst), len(src)). It is exact (int8 fits in int32). Any trailing capacity in dst is left untouched.

Types

This section is empty.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL