f32

package
v1.0.19 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Dec 10, 2025 License: MIT Imports: 2 Imported by: 18

Documentation

Overview

Package f32 provides SIMD-accelerated operations on float32 slices.

All functions automatically select the optimal implementation based on runtime CPU feature detection. Functions gracefully fall back to pure Go implementations on unsupported architectures.

Thread Safety: All functions are safe for concurrent use. Memory: All functions are zero-allocation (no heap allocations).

Index

Examples

Constants

This section is empty.

Variables

This section is empty.

Functions

func Abs

func Abs(dst, a []float32)

Abs computes element-wise absolute value: dst[i] = |a[i]|.

func AccumulateAdd

func AccumulateAdd(dst, src []float32, offset int)

AccumulateAdd adds src to dst starting at offset: dst[offset:offset+len(src)] += src. This is a key primitive for overlap-add in FFT-based convolution.

Panics if offset+len(src) > len(dst) or if offset < 0.

func Add

func Add(dst, a, b []float32)

Add computes element-wise addition: dst[i] = a[i] + b[i].

Example
package main

import (
	"fmt"

	"github.com/tphakala/simd/f32"
)

func main() {
	a := []float32{1, 2, 3, 4}
	b := []float32{5, 6, 7, 8}
	dst := make([]float32, len(a))

	f32.Add(dst, a, b)
	fmt.Println(dst)
}
Output:
[6 8 10 12]

func AddScalar

func AddScalar(dst, a []float32, s float32)

AddScalar adds a scalar to each element: dst[i] = a[i] + s.

func AddScaled added in v1.0.9

func AddScaled(dst []float32, alpha float32, s []float32)

AddScaled adds scaled values to dst: dst[i] += alpha * s[i]. This is the AXPY operation from BLAS Level 1. Processes min(len(dst), len(s)) elements.

func Clamp

func Clamp(dst, a []float32, minVal, maxVal float32)

Clamp clamps each element to [min, max].

Example
package main

import (
	"fmt"

	"github.com/tphakala/simd/f32"
)

func main() {
	a := []float32{-5, 0, 5, 10, 15}
	dst := make([]float32, len(a))

	f32.Clamp(dst, a, 0, 10)
	fmt.Println(dst)
}
Output:
[0 0 5 10 10]

func ClampScale added in v1.0.15

func ClampScale(dst, src []float32, minVal, maxVal, scale float32)

ClampScale performs fused clamp and scale: dst[i] = (clamp(src[i], min, max) - min) * scale. This is useful for normalizing data to a specific range. Processes min(len(dst), len(src)) elements.

Uses AVX on AMD64 (8x float32), NEON on ARM64 (4x float32).

func ConvolveValid

func ConvolveValid(dst, signal, kernel []float32)

ConvolveValid computes valid convolution of signal with kernel. dst[i] = sum(signal[i+j] * kernel[j]) for j in 0..len(kernel)-1. Output length is len(signal) - len(kernel) + 1.

This is equivalent to applying a FIR filter without zero-padding.

func ConvolveValidMulti added in v1.0.8

func ConvolveValidMulti(dsts [][]float32, signal []float32, kernels [][]float32)

ConvolveValidMulti applies multiple kernels to the same signal. dsts[k][i] = sum(signal[i+j] * kernels[k][j]) for each kernel k. All kernels must have the same length.

This is a convenience wrapper that calls ConvolveValid for each kernel. For polyphase resampling with multiple filter phases, this provides a clean API without additional overhead.

Panics if kernels have different lengths or if dsts/kernels lengths don't match.

func CubicInterpDot added in v1.0.14

func CubicInterpDot(hist, a, b, c, d []float32, x float32) float32

CubicInterpDot computes the fused cubic interpolation dot product:

Σ hist[i] * (a[i] + x*(b[i] + x*(c[i] + x*d[i])))

This is the hot inner loop for polyphase resampling with cubic coefficient interpolation. The polynomial a + x*(b + x*(c + x*d)) is evaluated using Horner's method for numerical stability, then multiplied by hist and summed.

Parameters:

  • hist: history buffer (signal samples)
  • a, b, c, d: cubic polynomial coefficient arrays
  • x: fractional phase, typically in [0, 1)

All slices must have equal length. Returns 0 for empty slices.

This fused operation is more efficient than 4 separate DotProduct calls because it reads the hist array only once (37% less memory bandwidth).

Uses AVX+FMA on AMD64, NEON on ARM64, with pure Go fallback.

func CubicInterpDotUnsafe added in v1.0.14

func CubicInterpDotUnsafe(hist, a, b, c, d []float32, x float32) float32

CubicInterpDotUnsafe computes the fused cubic interpolation dot product without length validation.

PRECONDITIONS (caller must ensure):

  • len(hist) == len(a) == len(b) == len(c) == len(d)
  • len(hist) > 0

Violating these preconditions results in undefined behavior. Use CubicInterpDot for safe operation with automatic length handling.

func CumulativeSum added in v1.0.9

func CumulativeSum(dst, a []float32)

CumulativeSum computes the cumulative sum: dst[i] = sum(a[0:i+1]). Processes min(len(dst), len(a)) elements.

func Deinterleave2 added in v1.0.8

func Deinterleave2(a, b, src []float32)

Deinterleave2 deinterleaves a slice: a[0]=src[0], b[0]=src[1], a[1]=src[2], b[1]=src[3], ... Processes min(len(a), len(b), len(src)/2) pairs. This is the inverse of Interleave2, useful for splitting stereo audio to channels.

func Div

func Div(dst, a, b []float32)

Div computes element-wise division: dst[i] = a[i] / b[i].

func DotProduct

func DotProduct(a, b []float32) float32

DotProduct computes the dot product of two float32 slices. Returns sum(a[i] * b[i]) for i in 0..min(len(a), len(b)).

Uses AVX+FMA on AMD64 (8x float32), NEON on ARM64 (4x float32).

Example
package main

import (
	"fmt"

	"github.com/tphakala/simd/f32"
)

func main() {
	a := []float32{1, 2, 3, 4}
	b := []float32{5, 6, 7, 8}

	result := f32.DotProduct(a, b)
	fmt.Printf("%.0f\n", result)
}
Output:
70

func DotProductBatch

func DotProductBatch(results []float32, rows [][]float32, vec []float32)

DotProductBatch computes multiple dot products against the same vector. results[i] = DotProduct(rows[i], vec) for each row. This is more cache-efficient than calling DotProduct in a loop because vec stays hot in L1 cache across all dot products.

func DotProductUnsafe added in v1.0.12

func DotProductUnsafe(a, b []float32) float32

DotProductUnsafe computes the dot product without length validation. This is a low-overhead variant for performance-critical code paths.

PRECONDITIONS (caller must ensure):

  • len(a) == len(b)
  • len(a) > 0

Violating these preconditions results in undefined behavior (panic or incorrect results). Use DotProduct for safe operation with automatic length handling.

func EuclideanDistance added in v1.0.9

func EuclideanDistance(a, b []float32) float32

EuclideanDistance computes the Euclidean distance between two vectors. Returns sqrt(sum((a[i] - b[i])^2)) for i in 0..min(len(a), len(b)).

func Exp added in v1.0.15

func Exp(dst, src []float32)

Exp computes the exponential function: dst[i] = e^src[i]. Uses polynomial approximation for reasonable accuracy and performance. Processes min(len(dst), len(src)) elements.

Uses AVX+FMA on AMD64 (8x float32), NEON on ARM64 (4x float32).

func ExpInPlace added in v1.0.15

func ExpInPlace(a []float32)

ExpInPlace computes exp in-place: a[i] = e^a[i].

func FMA

func FMA(dst, a, b, c []float32)

FMA computes fused multiply-add: dst[i] = a[i] * b[i] + c[i].

Example
package main

import (
	"fmt"

	"github.com/tphakala/simd/f32"
)

func main() {
	a := []float32{1, 2, 3}
	b := []float32{2, 2, 2}
	c := []float32{1, 1, 1}
	dst := make([]float32, len(a))

	// dst[i] = a[i] * b[i] + c[i]
	f32.FMA(dst, a, b, c)
	fmt.Println(dst)
}
Output:
[3 5 7]

func Int32ToFloat32Scale added in v1.0.18

func Int32ToFloat32Scale(dst []float32, src []int32, scale float32)

Int32ToFloat32Scale converts int32 samples to float32 and scales in one pass. dst[i] = float32(src[i]) * scale

This is optimized for audio processing where PCM samples need to be converted to normalized floating-point. For example, 16-bit audio uses scale = 1.0/32768.0 and 32-bit audio uses scale = 1.0/2147483648.0.

Processes min(len(dst), len(src)) elements.

Uses AVX on AMD64 (8x int32), NEON on ARM64 (4x int32).

func Int32ToFloat32ScaleUnsafe added in v1.0.18

func Int32ToFloat32ScaleUnsafe(dst []float32, src []int32, scale float32)

Int32ToFloat32ScaleUnsafe converts int32 samples to float32 and scales without length validation. This is a low-overhead variant for performance-critical code paths.

PRECONDITIONS (caller must ensure):

  • len(dst) >= len(src)
  • len(src) > 0

Violating these preconditions results in undefined behavior. Use Int32ToFloat32Scale for safe operation with automatic length handling.

func Interleave2 added in v1.0.8

func Interleave2(dst, a, b []float32)

Interleave2 interleaves two slices: dst[0]=a[0], dst[1]=b[0], dst[2]=a[1], dst[3]=b[1], ... Processes min(len(a), len(b), len(dst)/2) pairs. This is useful for converting separate channels to interleaved stereo audio.

func Max

func Max(a []float32) float32

Max returns the maximum value.

func MaxIdx added in v1.0.9

func MaxIdx(a []float32) int

MaxIdx returns the index of the maximum value in the slice. Returns -1 for empty slices.

func Mean added in v1.0.9

func Mean(a []float32) float32

Mean computes the arithmetic mean of a slice. Returns 0 for empty slices.

func Min

func Min(a []float32) float32

Min returns the minimum value.

func MinIdx added in v1.0.9

func MinIdx(a []float32) int

MinIdx returns the index of the minimum value in the slice. Returns -1 for empty slices.

func Mul

func Mul(dst, a, b []float32)

Mul computes element-wise multiplication: dst[i] = a[i] * b[i].

Example
package main

import (
	"fmt"

	"github.com/tphakala/simd/f32"
)

func main() {
	a := []float32{1, 2, 3, 4}
	b := []float32{2, 2, 2, 2}
	dst := make([]float32, len(a))

	f32.Mul(dst, a, b)
	fmt.Println(dst)
}
Output:
[2 4 6 8]

func Neg

func Neg(dst, a []float32)

Neg computes element-wise negation: dst[i] = -a[i].

func Normalize added in v1.0.9

func Normalize(dst, a []float32)

Normalize normalizes a vector to unit length: dst = a / ||a||. If the magnitude is zero or very small (< 1e-7), copies the input unchanged. Processes min(len(dst), len(a)) elements.

func ReLU added in v1.0.15

func ReLU(dst, src []float32)

ReLU computes the Rectified Linear Unit: dst[i] = max(0, src[i]). This is commonly used as an activation function in neural networks. Processes min(len(dst), len(src)) elements.

Uses AVX on AMD64 (8x float32), NEON on ARM64 (4x float32).

func ReLUInPlace added in v1.0.15

func ReLUInPlace(a []float32)

ReLUInPlace computes ReLU in-place: a[i] = max(0, a[i]).

func Reciprocal added in v1.0.9

func Reciprocal(dst, a []float32)

Reciprocal computes element-wise reciprocal: dst[i] = 1/a[i]. Processes min(len(dst), len(a)) elements.

func Scale

func Scale(dst, a []float32, s float32)

Scale multiplies each element by a scalar: dst[i] = a[i] * s.

Example
package main

import (
	"fmt"

	"github.com/tphakala/simd/f32"
)

func main() {
	a := []float32{1, 2, 3, 4}
	dst := make([]float32, len(a))

	f32.Scale(dst, a, 3.0)
	fmt.Println(dst)
}
Output:
[3 6 9 12]

func Sigmoid added in v1.0.15

func Sigmoid(dst, src []float32)

Sigmoid computes the sigmoid activation function: dst[i] = 1 / (1 + e^(-src[i])). This is commonly used as an activation function in neural networks. Processes min(len(dst), len(src)) elements.

Uses AVX+FMA on AMD64 (8x float32), NEON on ARM64 (4x float32).

func SigmoidInPlace added in v1.0.15

func SigmoidInPlace(a []float32)

SigmoidInPlace computes the sigmoid activation function in-place: a[i] = 1 / (1 + e^(-a[i])). This is commonly used as an activation function in neural networks.

Uses AVX+FMA on AMD64 (8x float32), NEON on ARM64 (4x float32).

func Sqrt added in v1.0.9

func Sqrt(dst, a []float32)

Sqrt computes element-wise square root: dst[i] = sqrt(a[i]). Processes min(len(dst), len(a)) elements.

func StdDev added in v1.0.9

func StdDev(a []float32) float32

StdDev computes the population standard deviation of a slice. Returns 0 for empty slices.

func Sub

func Sub(dst, a, b []float32)

Sub computes element-wise subtraction: dst[i] = a[i] - b[i].

func Sum

func Sum(a []float32) float32

Sum returns the sum of all elements.

Example
package main

import (
	"fmt"

	"github.com/tphakala/simd/f32"
)

func main() {
	a := []float32{1, 2, 3, 4, 5}

	result := f32.Sum(a)
	fmt.Printf("%.0f\n", result)
}
Output:
15

func Tanh added in v1.0.15

func Tanh(dst, src []float32)

Tanh computes the hyperbolic tangent: dst[i] = tanh(src[i]). Uses fast approximation: tanh(x) ≈ x / (1 + |x|) for |x| < 1, sign(x) for |x| >= 2.5, polynomial otherwise. Processes min(len(dst), len(src)) elements.

Uses AVX on AMD64 (8x float32), NEON on ARM64 (4x float32).

func TanhInPlace added in v1.0.15

func TanhInPlace(a []float32)

TanhInPlace computes tanh in-place: a[i] = tanh(a[i]).

func Variance added in v1.0.9

func Variance(a []float32) float32

Variance computes the population variance of a slice. Returns 0 for empty slices.

Types

This section is empty.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL