README
¶
gopherllm-wasm
Compiles GopherLLM's inference engine to WebAssembly so it can run entirely
client-side in a browser tab — no server. It runs on the existing pure-Go
scalar kernel tier (the same one that already backs every non-amd64/
non-arm64 build) everywhere, and additionally offloads Q4_K/Q6_K matvecs to
a hand-written WebGPU compute backend (internal/webgpu/) whenever
navigator.gpu is available — one binary, runtime-detected, no build flag.
No third-party JS/WASM code is used anywhere in this path — just the Go
compiler's own js/wasm target, Go's own wasm_exec.js bootstrap, and the
browser's native WebGPU API.
Building
make wasm-build
This produces bin/gopherllm.wasm and stages bin/wasm_exec.js (copied
from the Go toolchain's own lib/wasm/ directory — not a project file).
On this dev machine specifically, the go resolved from PATH is a stale
go1.16 install; the working toolchain is at
C:\Users\waldherr\go-sdk\go\bin\go.exe. Pass it explicitly if make picks
up the wrong one, e.g. make GO=/c/Users/waldherr/go-sdk/go/bin/go.exe wasm-build.
This is a local environment quirk, not a repo requirement.
The demo
demo/ is the actual user-facing page: pick a text-model GGUF (and,
optionally, a matching Pixtral-style vision projector GGUF) via native file
inputs, then chat — with images, if a vision projector was loaded — entirely
in the tab. Serve it via the devserver (below) at /demo/. It calls
gopherllm_loadModelWithVision/gopherllm_hasVision/gopherllm_generate
(images as base64 in each message's images field) — the same bridge
functions the vision-in-browser correctness harness below exercises via
fetch, since a native OS file-picker dialog can't be driven by browser
automation tooling (verified indirectly that way instead — see
webgpu-vision-test.html).
Model files are read entirely client-side via FileReader; nothing is
uploaded anywhere. For a large (multi-GB) model, keep in mind a browser tab
has to hold the whole file as one contiguous buffer at least twice, briefly
(the read buffer, then GopherLLM's own copy) — see the 2GB-model finding
below.
Trying the correctness harnesses in a browser
testdata/harness/ holds correctness-verification pages, not the demo
(above) — each fetches its own tiny synthetic fixture and checks the result
against a known-correct value, rather than taking user input.
make wasm-build
GOPHERLLM_GEN_WASM_FIXTURE=1 go test -run TestGenerateWasmFixture -v .
go build -o bin/devserver.exe ./cmd/gopherllm-wasm/devserver
./bin/devserver.exe
Then open http://127.0.0.1:8090/. The page fetches
testdata/harness/tiny-model.gguf (regenerated by the go test command
above — not committed, it's a tiny synthetic fixture, see
wasm_fixture_test.go) and /bin/gopherllm.wasm//bin/wasm_exec.js, then
runs a deterministic (temperature-0) generation and checks it against the
known-correct output from a native run of the same bytes
(TestWasmFixtureGoldenOutputPrint-style comparison, hardcoded in
testdata/harness/app.js as GOLDEN_TEXT).
What's here
main.go/bridge.go/promise.go— the wasm entry point: registersgopherllm_loadModel/gopherllm_loadModelWithVision/gopherllm_hasVision/gopherllm_generate(its per-message wire format has an optionalimages: string[]of base64-encoded PNG/JPEG bytes) /gopherllm_stopGeneration/gopherllm_isGenerating/gopherllm_isModelLoadedon the JS global object viasyscall/js, bridging Go's goroutine model to JS Promises.gopherllm_generaterejects immediately if a generation is already in progress rather than silently queuing behind it;gopherllm_loadModel/gopherllm_loadModelWithVisionsignal any in-flight generation to stop before waiting for the outgoing model to close, so switching models mid-generation stays responsive.demo/— the polished, user-facing page (see "The demo" above).devserver/— a trivialnet/httpstatic file server for manual testing; not part of the wasm build itself. Mounts/(harness),/bin/(wasm build output),/demo/, and optionally/models/(point-modelsat a directory holding real, large GGUFs — e.g. the repo root — to test against them without copying multi-GB files into this directory).testdata/harness/— correctness-verification pages:index.html/app.js— the CPU-only golden-output check above.webgpu-smoke.html— validates theinternal/webgpuplumbing itself (device acquisition, buffer upload, shader compile, dispatch, readback) with a trivial shader, independent of any model.webgpu-kernel-test.html— compares the hand-written Q4_K/Q6_K WGSL matvec kernels against this project's existing portable Go reference, using the real quantizer to generate test data (no external oracle to check against, so verification is against this project's own, already-tested CPU implementation).webgpu-forward-pass-test.html— the real end-to-end check: loads a tiny quantized (Q4_K/Q6_K) GGUF (GOPHERLLM_GEN_WASM_GPU_FIXTURE=1 go test -run TestGenerateWasmGPUFixture -v .regenerates it) and confirms a full multi-layer, multi-token generation through the realWeight.GPU/MatvecIntowiring matches the native CPU-path output exactly.webgpu-vision-test.html— the same kind of check for the vision path: loads a tiny text+vision GGUF pair (GOPHERLLM_GEN_WASM_VISION_FIXTURE=1 go test -run TestGenerateWasmDemoVisionFixture -v .regenerates both) viagopherllm_loadModelWithVision, sends an image withgopherllm_generate, and checks the result against the native golden output — this is what actually exercisesgopherllm_loadModelWithVisionend to end, since the demo's own file-picker inputs can't be driven by browser automation tooling.real-model-bench.html— loads a real (large) model via/models/and times generation; see the 2GB-model finding below for what happened the one time this was tried against the actual Ministral 3B file.
webgpu_smoke.go/webgpu_kernel_check.go— the Go-side logic those test-only pages call into (kept out of_test.gofiles since those are excluded from a normalgo build, only compiled undergo test).
WebGPU backend
internal/webgpu/ is a from-scratch wrapper over the browser's WebGPU API
(device/buffer/shader/dispatch, mirroring how internal/metal wraps Metal),
plus hand-written WGSL reimplementations of this project's own Q4_K/Q6_K
dequantizing matvec kernels (shader_q4k.go, shader_q6k.go) — including a
from-scratch IEEE-754 half-float decoder, since WGSL has no dependable
shader-f16 extension to lean on. Verified against the portable Go
reference kernels at ~1e-6 relative error (float32 summation-order noise,
not a bug).
At the root package, webgpu_js.go/webgpu_stub.go mirror
metal_darwin.go/metal_stub.go's discipline exactly: Weight.GPU is
populated automatically for Q4_K/Q6_K tensors whenever a WebGPU device is
available (no LoadOptions flag needed — WebGPU only ever exists under
GOOS=js, where Metal never does, so the two can't conflict), and
Weight.MatvecInto tries it right alongside the existing Metal check.
Known limitation, not yet solved: a tensor larger than the device's
maxStorageBufferBindingSize (commonly 128MiB) — e.g. Ministral 3B's tied
~300MB+ output/embedding projection — cannot be uploaded as one GPU buffer
yet. prepareWebGPUWeight catches this and falls back to the CPU path for
that specific tensor rather than failing the load; splitting an oversized
tensor across multiple buffers is designed (see the project plan) but not
implemented.
Also not yet done: per-layer dispatch batching. The current integration
calls the GPU kernel once per matvec (one CPU↔GPU round trip each), which is
correct — verified via webgpu-forward-pass-test.html producing output
bit-identical to the CPU path, and via webgpu-vision-test.html for the
vision path too — but not yet the "~3 sync points per layer" batched design
the project plan calls for, which needs restructuring ForwardBodyInto's
per-layer loop itself, not just this leaf-level hook.
A real bug worth knowing about if you touch internal/webgpu/buffer.go:
GPUQueue.writeBuffer (and buffer/map/range sizes generally) require byte
counts to be multiples of 4. GGUF's Q6_K block is 210 bytes, so plenty of
real Q6_K tensors (e.g. 37 rows × 210B = 7770B) have a total size that
isn't 4-aligned — this isn't hypothetical, it broke an actual browser run
during development (Failed to execute 'writeBuffer' ... must be a multiple of 4). CreateStorageBuffer/CreateUniformBuffer/WriteBuffer/ReadBuffer
all round up via roundUpTo4 now; Buffer.Size() still reports the
caller's logical (unrounded) size.
2GB-model finding: the one attempt at loading the real
Ministral-3-3B-Instruct-2512-Q4_K_M.gguf (2,147,023,008 bytes) in a
sandboxed test browser failed at a plain JS new Uint8Array(...) allocation
— before Go/wasm was ever involved — with "Array buffer allocation
failed". Binary-searching in isolation found that sandbox's single-buffer
ceiling sits at ~2,145.3–2,145.9 MB, just under the file's size by only
~1–1.7 MB. That reads as a limit of that specific sandboxed container, not a
demonstrated wasm32/browser-architecture limit — real desktop Chrome/Firefox
on ordinary hardware routinely allocates single buffers well past 2GiB. The
"does the real file fit in a real browser" question is genuinely still open;
try real-model-bench.html yourself on real desktop hardware to find out.
If it doesn't, the right fix is a streaming GGUF loader that never
materializes the whole file as one JS/Go buffer (parse the header, then
stream each tensor's bytes into its own allocation as it arrives) — not
implemented yet.
Documentation
¶
Overview ¶
Command gopherllm-wasm compiles GopherLLM's inference engine to WebAssembly for in-browser use. It registers a small set of JS-callable functions (see bridge.go) and otherwise does nothing on its own — all control comes from the host page via wasm_exec.js.
Source Files
¶
Directories
¶
| Path | Synopsis |
|---|---|
|
Command devserver is a tiny static file server for manually exercising the wasm browser harness (cmd/gopherllm-wasm/testdata/harness) during development.
|
Command devserver is a tiny static file server for manually exercising the wasm browser harness (cmd/gopherllm-wasm/testdata/harness) during development. |