Documentation
¶
Overview ¶
Package cpu provides CPU feature detection for SIMD operations.
Example ¶
package main
import (
"fmt"
"github.com/tphakala/simd/cpu"
)
func main() {
// Check available SIMD features
fmt.Println("CPU Info:", cpu.Info())
fmt.Println("AVX:", cpu.HasAVX())
fmt.Println("AVX2:", cpu.HasAVX2())
fmt.Println("FMA:", cpu.HasFMA())
fmt.Println("NEON:", cpu.HasNEON())
}
Output:
Index ¶
Examples ¶
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
func HasAVX ¶
func HasAVX() bool
HasAVX returns true if AVX is available.
Example ¶
package main
import (
"fmt"
"github.com/tphakala/simd/cpu"
)
func main() {
if cpu.HasAVX() {
fmt.Println("AVX is available")
} else {
fmt.Println("AVX is not available")
}
}
Output:
func HasAVXVNNI ¶ added in v1.4.0
func HasAVXVNNI() bool
HasAVXVNNI returns true if the VEX-encoded AVX-VNNI instructions (VPDPWSSD / VPDPBUSD and their saturating variants) are available. They fuse the widen-multiply-accumulate that int16/int8 dot products otherwise spell as a VPMADDWD/VPMADDUBSW plus a separate VPADDD. Shipped on Intel Alder Lake and later and on AMD Zen 4 and later, in VEX form so both hybrid core types carry it. Detection gates on AVX2 (the VEX form runs on YMM state).
func HasDOTPROD ¶ added in v1.4.0
func HasDOTPROD() bool
HasDOTPROD returns true if the ARM64 int8 dot-product instructions (SDOT/UDOT, FEAT_DotProd) are available. They accelerate int8 dot products and quantized matmul inner loops.
func HasF16C ¶ added in v1.2.0
func HasF16C() bool
HasF16C returns true if the x86 F16C half-precision conversion instructions (VCVTPH2PS / VCVTPS2PH) are available. They accelerate Float16 <-> float32 slice conversion. F16C provides conversion only, not half-precision arithmetic.
func HasFP16 ¶ added in v1.0.19
func HasFP16() bool
HasFP16 returns true if ARM FP16 (half-precision) is available.
func HasNEON ¶
func HasNEON() bool
HasNEON returns true if ARM NEON is available.
Example ¶
package main
import (
"fmt"
"github.com/tphakala/simd/cpu"
)
func main() {
if cpu.HasNEON() {
fmt.Println("NEON is available")
} else {
fmt.Println("NEON is not available")
}
}
Output:
func HasPCLMULQDQ ¶ added in v1.2.0
func HasPCLMULQDQ() bool
HasPCLMULQDQ returns true if the x86 carry-less multiply instruction (PCLMULQDQ) is available. It is used to accelerate CRC folding.
func HasPMULL ¶ added in v1.2.0
func HasPMULL() bool
HasPMULL returns true if the ARM64 polynomial multiply instruction (PMULL) is available. It is used to accelerate CRC folding.
func Info ¶
func Info() string
Info returns a string describing the available SIMD features.
Example ¶
package main
import (
"fmt"
"github.com/tphakala/simd/cpu"
)
func main() {
// Returns a string like "AMD64 AVX+FMA" or "ARM64 NEON"
info := cpu.Info()
fmt.Printf("CPU supports: %s\n", info)
}
Output:
Types ¶
type Features ¶
type Features struct {
// x86/AMD64 features
SSE bool
SSE2 bool
SSE3 bool
SSSE3 bool
SSE41 bool
SSE42 bool
AVX bool
AVX2 bool
AVXVNNI bool // AVX-VNNI (VEX-encoded VPDPWSSD/VPDPBUSD); Alder Lake+ and Zen 4+
AVX512F bool
AVX512VL bool
FMA bool
BMI1 bool
BMI2 bool
POPCNT bool
PCLMULQDQ bool // carry-less multiply (CLMUL) - used for CRC folding
F16C bool // half<->single float conversion (VCVTPH2PS/VCVTPS2PH)
// ARM64 features
NEON bool
FP16 bool // ARM64 half-precision floating point (FEAT_FP16)
SVE bool
SVE2 bool
PMULL bool // polynomial multiply (FEAT_PMULL) - used for CRC folding
DOTPROD bool // int8 dot product (FEAT_DotProd) - SDOT/UDOT
}
Features contains detected CPU SIMD capabilities.
var ARM64 Features
ARM64 contains ARM64 CPU features (populated on arm64).
var X86 Features
X86 contains x86/AMD64 CPU features (populated on amd64).