simd

A high-performance SIMD (Single Instruction, Multiple Data) library for Go providing vectorized operations on float64 and float32 slices.
Features
- Pure Go assembly - No CGO required, simple cross-compilation
- Runtime CPU detection - Automatically selects optimal implementation (AVX+FMA, NEON, or pure Go)
- Zero allocations - All operations work on pre-allocated slices
- 23 operations - Arithmetic, reduction, statistical, vector, and signal processing operations
- Multi-architecture - AMD64 (AVX+FMA) and ARM64 (NEON) with pure Go fallback
- Thread-safe - All functions are safe for concurrent use
Installation
go get github.com/tphakala/simd
Requires Go 1.25+
Quick Start
package main
import (
"fmt"
"github.com/tphakala/simd/pkg/simd/cpu"
"github.com/tphakala/simd/pkg/simd/f64"
)
func main() {
fmt.Println("SIMD:", cpu.Info())
// Vector operations
a := []float64{1, 2, 3, 4, 5, 6, 7, 8}
b := []float64{8, 7, 6, 5, 4, 3, 2, 1}
// Dot product
dot := f64.DotProduct(a, b)
fmt.Println("Dot product:", dot) // 120
// Element-wise operations
dst := make([]float64, len(a))
f64.Add(dst, a, b)
fmt.Println("Sum:", dst) // [9, 9, 9, 9, 9, 9, 9, 9]
// Statistical operations
mean := f64.Mean(a)
stddev := f64.StdDev(a)
fmt.Printf("Mean: %.2f, StdDev: %.2f\n", mean, stddev)
// Vector operations
f64.Normalize(dst, a)
fmt.Println("Normalized:", dst)
// Distance calculation
dist := f64.EuclideanDistance(a, b)
fmt.Println("Distance:", dist)
}
Packages
cpu - CPU Feature Detection
import "github.com/tphakala/simd/pkg/simd/cpu"
fmt.Println(cpu.Info()) // "AMD64 AVX2+FMA" or "ARM64 NEON"
fmt.Println(cpu.HasAVX()) // true/false
fmt.Println(cpu.HasNEON()) // true/false
f64 - float64 Operations
| Category |
Function |
Description |
SIMD Width |
| Arithmetic |
Add(dst, a, b) |
Element-wise addition |
4x (AVX) / 2x (NEON) |
|
Sub(dst, a, b) |
Element-wise subtraction |
4x / 2x |
|
Mul(dst, a, b) |
Element-wise multiplication |
4x / 2x |
|
Div(dst, a, b) |
Element-wise division |
4x / 2x |
|
Scale(dst, a, s) |
Multiply by scalar |
4x / 2x |
|
AddScalar(dst, a, s) |
Add scalar |
4x / 2x |
|
FMA(dst, a, b, c) |
Fused multiply-add: a*b+c |
4x / 2x |
| Unary |
Abs(dst, a) |
Absolute value |
4x / 2x |
|
Neg(dst, a) |
Negation |
4x / 2x |
|
Sqrt(dst, a) |
Square root |
4x / 2x |
|
Reciprocal(dst, a) |
Reciprocal (1/x) |
4x / 2x |
| Reduction |
DotProduct(a, b) |
Dot product |
4x / 2x |
|
Sum(a) |
Sum of elements |
4x / 2x |
|
Min(a) |
Minimum value |
4x / 2x |
|
Max(a) |
Maximum value |
4x / 2x |
| Statistical |
Mean(a) |
Arithmetic mean |
4x / 2x |
|
Variance(a) |
Population variance |
4x / 2x |
|
StdDev(a) |
Standard deviation |
4x / 2x |
| Vector |
EuclideanDistance(a, b) |
L2 distance |
4x / 2x |
|
Normalize(dst, a) |
Unit vector normalization |
4x / 2x |
|
CumulativeSum(dst, a) |
Running sum |
Sequential |
| Range |
Clamp(dst, a, min, max) |
Clamp to range |
4x / 2x |
| Batch |
DotProductBatch(r, rows, v) |
Multiple dot products |
4x / 2x |
| Signal |
ConvolveValid(dst, sig, k) |
FIR filter / convolution |
4x / 2x |
f32 - float32 Operations
Same API as f64 but for float32 with wider SIMD:
| Architecture |
SIMD Width |
| AMD64 (AVX) |
8x float32 |
| ARM64 (NEON) |
4x float32 |
Benchmarks on AMD64 with AVX+FMA (Intel/AMD processor):
float64 Operations
| Operation |
Size |
Time |
Throughput |
| DotProduct |
277 |
34.9 ns |
127 GB/s |
| DotProduct |
1000 |
159 ns |
100 GB/s |
| Add |
1000 |
106 ns |
226 GB/s |
| Mul |
1000 |
109 ns |
220 GB/s |
| FMA |
1000 |
121 ns |
263 GB/s |
| Sum |
1000 |
90.6 ns |
88 GB/s |
| Mean |
1000 |
89.6 ns |
89 GB/s |
| Variance |
1000 |
540 ns |
15 GB/s |
| EuclideanDistance |
100 |
90 ns |
18 GB/s |
| Normalize |
100 |
42.4 ns |
38 GB/s |
| Sqrt |
100 |
130 ns |
12 GB/s |
| Reciprocal |
100 |
86.6 ns |
18 GB/s |
float32 Operations
| Operation |
Size |
Time |
Throughput |
| DotProduct |
100 |
7.2 ns |
111 GB/s |
| DotProduct |
1000 |
70 ns |
114 GB/s |
| Add |
1000 |
47.3 ns |
253 GB/s |
| Mul |
1000 |
47.5 ns |
252 GB/s |
| FMA |
1000 |
63.5 ns |
252 GB/s |
Comparison vs Pure Go
| Operation |
SIMD |
Pure Go |
Speedup |
| DotProduct (1000 f64) |
159 ns |
890 ns |
5.6x |
| Add (1000 f64) |
106 ns |
420 ns |
4.0x |
| FMA (1000 f64) |
121 ns |
1100 ns |
9.1x |
| DotProduct (1000 f32) |
70 ns |
450 ns |
6.4x |
Architecture Support
| Architecture |
Instruction Set |
Status |
| AMD64 |
AVX + FMA |
Full SIMD support |
| AMD64 |
SSE2 only |
Pure Go fallback |
| ARM64 |
NEON/ASIMD |
Full SIMD support |
| ARM64 |
SVE/SVE2 |
Planned |
| Other |
- |
Pure Go fallback |
Design Principles
- No CGO - Pure Go assembly for maximum portability and easy cross-compilation
- Runtime dispatch - CPU features detected once at init time, zero runtime overhead
- Zero allocations - No heap allocations in hot paths
- Safe defaults - Gracefully falls back to pure Go on unsupported CPUs
- Boundary safe - Handles any slice length, not just SIMD-aligned sizes
Testing
The library includes comprehensive tests validated against a C reference implementation:
# Run all tests
go test ./...
# Run benchmarks
go test ./pkg/simd/f64 -bench=. -benchmem
# Generate test expectations from C reference
cd testdata && gcc -O2 -march=native -o generate_expectations generate_expectations.c -lm
./generate_expectations
Contributing
Contributions are welcome! Please see CONTRIBUTING.md for guidelines.
License
This project is licensed under the MIT License - see the LICENSE file for details.