Documentation
¶
Overview ¶
Package f32 provides SIMD-accelerated operations on float32 slices.
All functions automatically select the optimal implementation based on runtime CPU feature detection. Functions gracefully fall back to pure Go implementations on unsupported architectures.
Thread Safety: All functions are safe for concurrent use. Memory: All functions are zero-allocation (no heap allocations).
Index ¶
- func Abs(dst, a []float32)
- func AccumulateAdd(dst, src []float32, offset int)
- func Add(dst, a, b []float32)
- func AddScalar(dst, a []float32, s float32)
- func AddScaled(dst []float32, alpha float32, s []float32)
- func Clamp(dst, a []float32, minVal, maxVal float32)
- func ClampScale(dst, src []float32, minVal, maxVal, scale float32)
- func ConvolveValid(dst, signal, kernel []float32)
- func ConvolveValidMulti(dsts [][]float32, signal []float32, kernels [][]float32)
- func CubicInterpDot(hist, a, b, c, d []float32, x float32) float32
- func CubicInterpDotUnsafe(hist, a, b, c, d []float32, x float32) float32
- func CumulativeSum(dst, a []float32)
- func Deinterleave2(a, b, src []float32)
- func Div(dst, a, b []float32)
- func DotProduct(a, b []float32) float32
- func DotProductBatch(results []float32, rows [][]float32, vec []float32)
- func DotProductUnsafe(a, b []float32) float32
- func EuclideanDistance(a, b []float32) float32
- func Exp(dst, src []float32)
- func ExpInPlace(a []float32)
- func FMA(dst, a, b, c []float32)
- func Int32ToFloat32Scale(dst []float32, src []int32, scale float32)
- func Int32ToFloat32ScaleUnsafe(dst []float32, src []int32, scale float32)
- func Interleave2(dst, a, b []float32)
- func Max(a []float32) float32
- func MaxIdx(a []float32) int
- func Mean(a []float32) float32
- func Min(a []float32) float32
- func MinIdx(a []float32) int
- func Mul(dst, a, b []float32)
- func Neg(dst, a []float32)
- func Normalize(dst, a []float32)
- func ReLU(dst, src []float32)
- func ReLUInPlace(a []float32)
- func Reciprocal(dst, a []float32)
- func Scale(dst, a []float32, s float32)
- func Sigmoid(dst, src []float32)
- func SigmoidInPlace(a []float32)
- func Sqrt(dst, a []float32)
- func StdDev(a []float32) float32
- func Sub(dst, a, b []float32)
- func Sum(a []float32) float32
- func Tanh(dst, src []float32)
- func TanhInPlace(a []float32)
- func Variance(a []float32) float32
Examples ¶
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
func AccumulateAdd ¶
AccumulateAdd adds src to dst starting at offset: dst[offset:offset+len(src)] += src. This is a key primitive for overlap-add in FFT-based convolution.
Panics if offset+len(src) > len(dst) or if offset < 0.
func Add ¶
func Add(dst, a, b []float32)
Add computes element-wise addition: dst[i] = a[i] + b[i].
Example ¶
package main
import (
"fmt"
"github.com/tphakala/simd/f32"
)
func main() {
a := []float32{1, 2, 3, 4}
b := []float32{5, 6, 7, 8}
dst := make([]float32, len(a))
f32.Add(dst, a, b)
fmt.Println(dst)
}
Output: [6 8 10 12]
func AddScaled ¶ added in v1.0.9
AddScaled adds scaled values to dst: dst[i] += alpha * s[i]. This is the AXPY operation from BLAS Level 1. Processes min(len(dst), len(s)) elements.
func Clamp ¶
Clamp clamps each element to [min, max].
Example ¶
package main
import (
"fmt"
"github.com/tphakala/simd/f32"
)
func main() {
a := []float32{-5, 0, 5, 10, 15}
dst := make([]float32, len(a))
f32.Clamp(dst, a, 0, 10)
fmt.Println(dst)
}
Output: [0 0 5 10 10]
func ClampScale ¶ added in v1.0.15
ClampScale performs fused clamp and scale: dst[i] = (clamp(src[i], min, max) - min) * scale. This is useful for normalizing data to a specific range. Processes min(len(dst), len(src)) elements.
Uses AVX on AMD64 (8x float32), NEON on ARM64 (4x float32).
func ConvolveValid ¶
func ConvolveValid(dst, signal, kernel []float32)
ConvolveValid computes valid convolution of signal with kernel. dst[i] = sum(signal[i+j] * kernel[j]) for j in 0..len(kernel)-1. Output length is len(signal) - len(kernel) + 1.
This is equivalent to applying a FIR filter without zero-padding.
func ConvolveValidMulti ¶ added in v1.0.8
ConvolveValidMulti applies multiple kernels to the same signal. dsts[k][i] = sum(signal[i+j] * kernels[k][j]) for each kernel k. All kernels must have the same length.
This is a convenience wrapper that calls ConvolveValid for each kernel. For polyphase resampling with multiple filter phases, this provides a clean API without additional overhead.
Panics if kernels have different lengths or if dsts/kernels lengths don't match.
func CubicInterpDot ¶ added in v1.0.14
CubicInterpDot computes the fused cubic interpolation dot product:
Σ hist[i] * (a[i] + x*(b[i] + x*(c[i] + x*d[i])))
This is the hot inner loop for polyphase resampling with cubic coefficient interpolation. The polynomial a + x*(b + x*(c + x*d)) is evaluated using Horner's method for numerical stability, then multiplied by hist and summed.
Parameters:
- hist: history buffer (signal samples)
- a, b, c, d: cubic polynomial coefficient arrays
- x: fractional phase, typically in [0, 1)
All slices must have equal length. Returns 0 for empty slices.
This fused operation is more efficient than 4 separate DotProduct calls because it reads the hist array only once (37% less memory bandwidth).
Uses AVX+FMA on AMD64, NEON on ARM64, with pure Go fallback.
func CubicInterpDotUnsafe ¶ added in v1.0.14
CubicInterpDotUnsafe computes the fused cubic interpolation dot product without length validation.
PRECONDITIONS (caller must ensure):
- len(hist) == len(a) == len(b) == len(c) == len(d)
- len(hist) > 0
Violating these preconditions results in undefined behavior. Use CubicInterpDot for safe operation with automatic length handling.
func CumulativeSum ¶ added in v1.0.9
func CumulativeSum(dst, a []float32)
CumulativeSum computes the cumulative sum: dst[i] = sum(a[0:i+1]). Processes min(len(dst), len(a)) elements.
func Deinterleave2 ¶ added in v1.0.8
func Deinterleave2(a, b, src []float32)
Deinterleave2 deinterleaves a slice: a[0]=src[0], b[0]=src[1], a[1]=src[2], b[1]=src[3], ... Processes min(len(a), len(b), len(src)/2) pairs. This is the inverse of Interleave2, useful for splitting stereo audio to channels.
func DotProduct ¶
DotProduct computes the dot product of two float32 slices. Returns sum(a[i] * b[i]) for i in 0..min(len(a), len(b)).
Uses AVX+FMA on AMD64 (8x float32), NEON on ARM64 (4x float32).
Example ¶
package main
import (
"fmt"
"github.com/tphakala/simd/f32"
)
func main() {
a := []float32{1, 2, 3, 4}
b := []float32{5, 6, 7, 8}
result := f32.DotProduct(a, b)
fmt.Printf("%.0f\n", result)
}
Output: 70
func DotProductBatch ¶
DotProductBatch computes multiple dot products against the same vector. results[i] = DotProduct(rows[i], vec) for each row. This is more cache-efficient than calling DotProduct in a loop because vec stays hot in L1 cache across all dot products.
func DotProductUnsafe ¶ added in v1.0.12
DotProductUnsafe computes the dot product without length validation. This is a low-overhead variant for performance-critical code paths.
PRECONDITIONS (caller must ensure):
- len(a) == len(b)
- len(a) > 0
Violating these preconditions results in undefined behavior (panic or incorrect results). Use DotProduct for safe operation with automatic length handling.
func EuclideanDistance ¶ added in v1.0.9
EuclideanDistance computes the Euclidean distance between two vectors. Returns sqrt(sum((a[i] - b[i])^2)) for i in 0..min(len(a), len(b)).
func Exp ¶ added in v1.0.15
func Exp(dst, src []float32)
Exp computes the exponential function: dst[i] = e^src[i]. Uses polynomial approximation for reasonable accuracy and performance. Processes min(len(dst), len(src)) elements.
Uses AVX+FMA on AMD64 (8x float32), NEON on ARM64 (4x float32).
func ExpInPlace ¶ added in v1.0.15
func ExpInPlace(a []float32)
ExpInPlace computes exp in-place: a[i] = e^a[i].
func FMA ¶
func FMA(dst, a, b, c []float32)
FMA computes fused multiply-add: dst[i] = a[i] * b[i] + c[i].
Example ¶
package main
import (
"fmt"
"github.com/tphakala/simd/f32"
)
func main() {
a := []float32{1, 2, 3}
b := []float32{2, 2, 2}
c := []float32{1, 1, 1}
dst := make([]float32, len(a))
// dst[i] = a[i] * b[i] + c[i]
f32.FMA(dst, a, b, c)
fmt.Println(dst)
}
Output: [3 5 7]
func Int32ToFloat32Scale ¶ added in v1.0.18
Int32ToFloat32Scale converts int32 samples to float32 and scales in one pass. dst[i] = float32(src[i]) * scale
This is optimized for audio processing where PCM samples need to be converted to normalized floating-point. For example, 16-bit audio uses scale = 1.0/32768.0 and 32-bit audio uses scale = 1.0/2147483648.0.
Processes min(len(dst), len(src)) elements.
Uses AVX on AMD64 (8x int32), NEON on ARM64 (4x int32).
func Int32ToFloat32ScaleUnsafe ¶ added in v1.0.18
Int32ToFloat32ScaleUnsafe converts int32 samples to float32 and scales without length validation. This is a low-overhead variant for performance-critical code paths.
PRECONDITIONS (caller must ensure):
- len(dst) >= len(src)
- len(src) > 0
Violating these preconditions results in undefined behavior. Use Int32ToFloat32Scale for safe operation with automatic length handling.
func Interleave2 ¶ added in v1.0.8
func Interleave2(dst, a, b []float32)
Interleave2 interleaves two slices: dst[0]=a[0], dst[1]=b[0], dst[2]=a[1], dst[3]=b[1], ... Processes min(len(a), len(b), len(dst)/2) pairs. This is useful for converting separate channels to interleaved stereo audio.
func MaxIdx ¶ added in v1.0.9
MaxIdx returns the index of the maximum value in the slice. Returns -1 for empty slices.
func Mean ¶ added in v1.0.9
Mean computes the arithmetic mean of a slice. Returns 0 for empty slices.
func MinIdx ¶ added in v1.0.9
MinIdx returns the index of the minimum value in the slice. Returns -1 for empty slices.
func Mul ¶
func Mul(dst, a, b []float32)
Mul computes element-wise multiplication: dst[i] = a[i] * b[i].
Example ¶
package main
import (
"fmt"
"github.com/tphakala/simd/f32"
)
func main() {
a := []float32{1, 2, 3, 4}
b := []float32{2, 2, 2, 2}
dst := make([]float32, len(a))
f32.Mul(dst, a, b)
fmt.Println(dst)
}
Output: [2 4 6 8]
func Normalize ¶ added in v1.0.9
func Normalize(dst, a []float32)
Normalize normalizes a vector to unit length: dst = a / ||a||. If the magnitude is zero or very small (< 1e-7), copies the input unchanged. Processes min(len(dst), len(a)) elements.
func ReLU ¶ added in v1.0.15
func ReLU(dst, src []float32)
ReLU computes the Rectified Linear Unit: dst[i] = max(0, src[i]). This is commonly used as an activation function in neural networks. Processes min(len(dst), len(src)) elements.
Uses AVX on AMD64 (8x float32), NEON on ARM64 (4x float32).
func ReLUInPlace ¶ added in v1.0.15
func ReLUInPlace(a []float32)
ReLUInPlace computes ReLU in-place: a[i] = max(0, a[i]).
func Reciprocal ¶ added in v1.0.9
func Reciprocal(dst, a []float32)
Reciprocal computes element-wise reciprocal: dst[i] = 1/a[i]. Processes min(len(dst), len(a)) elements.
func Scale ¶
Scale multiplies each element by a scalar: dst[i] = a[i] * s.
Example ¶
package main
import (
"fmt"
"github.com/tphakala/simd/f32"
)
func main() {
a := []float32{1, 2, 3, 4}
dst := make([]float32, len(a))
f32.Scale(dst, a, 3.0)
fmt.Println(dst)
}
Output: [3 6 9 12]
func Sigmoid ¶ added in v1.0.15
func Sigmoid(dst, src []float32)
Sigmoid computes the sigmoid activation function: dst[i] = 1 / (1 + e^(-src[i])). This is commonly used as an activation function in neural networks. Processes min(len(dst), len(src)) elements.
Uses AVX+FMA on AMD64 (8x float32), NEON on ARM64 (4x float32).
func SigmoidInPlace ¶ added in v1.0.15
func SigmoidInPlace(a []float32)
SigmoidInPlace computes the sigmoid activation function in-place: a[i] = 1 / (1 + e^(-a[i])). This is commonly used as an activation function in neural networks.
Uses AVX+FMA on AMD64 (8x float32), NEON on ARM64 (4x float32).
func Sqrt ¶ added in v1.0.9
func Sqrt(dst, a []float32)
Sqrt computes element-wise square root: dst[i] = sqrt(a[i]). Processes min(len(dst), len(a)) elements.
func StdDev ¶ added in v1.0.9
StdDev computes the population standard deviation of a slice. Returns 0 for empty slices.
func Sub ¶
func Sub(dst, a, b []float32)
Sub computes element-wise subtraction: dst[i] = a[i] - b[i].
func Sum ¶
Sum returns the sum of all elements.
Example ¶
package main
import (
"fmt"
"github.com/tphakala/simd/f32"
)
func main() {
a := []float32{1, 2, 3, 4, 5}
result := f32.Sum(a)
fmt.Printf("%.0f\n", result)
}
Output: 15
func Tanh ¶ added in v1.0.15
func Tanh(dst, src []float32)
Tanh computes the hyperbolic tangent: dst[i] = tanh(src[i]). Uses fast approximation: tanh(x) ≈ x / (1 + |x|) for |x| < 1, sign(x) for |x| >= 2.5, polynomial otherwise. Processes min(len(dst), len(src)) elements.
Uses AVX on AMD64 (8x float32), NEON on ARM64 (4x float32).
func TanhInPlace ¶ added in v1.0.15
func TanhInPlace(a []float32)
TanhInPlace computes tanh in-place: a[i] = tanh(a[i]).
Types ¶
This section is empty.