bm25

package module
v1.3.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Jun 27, 2026 License: MIT Imports: 28 Imported by: 0

README

bm25 — FastEmbed Qdrant/bm25 sparse encoder for Go

BM25 sparse vectors in Go for Qdrant hybrid search, matching Python FastEmbed byte-for-byte.

CI Go Reference Go Report Card

A Go port of FastEmbed's Qdrant/bm25 sparse text encoder (aka fastembed-bm25-go). It builds the BM25 sparse vectors a Go service needs for Qdrant hybrid search, and its output matches the Python encoder exactly — so the token ids line up with a corpus that was indexed in Python with FastEmbed. It's the sparse counterpart to the dense fastembed-go port.

Why a Go BM25 encoder for Qdrant

Qdrant hybrid search combines a dense vector with a sparse BM25 vector and fuses the two result sets. The sparse vectors are usually produced in Python by FastEmbed's SparseTextEmbedding("Qdrant/bm25"). Each token id in a sparse vector is a hash of a stemmed token, so a query only matches the stored corpus if it is tokenized, stemmed, and hashed in exactly the same way.

There was no Go encoder that did this. The dense FastEmbed port, anush008/fastembed-go, does not cover sparse models. This package fills that gap and is checked against output captured from FastEmbed itself.

Why this library vs. Qdrant server-side BM25

Since Qdrant v1.15.2, Qdrant can convert text to Qdrant/bm25 sparse vectors server-side (via the Inference API / Qdrant Cloud Inference), and the official Go client can pass Document{Model: "qdrant/bm25", Text: ...}. So why compute BM25 vectors in Go? Use this library when you want:

  • Client-side control — tokenization and vectorization happen entirely in your Go process: no Inference API, no extra network round-trip, no dependency on a Qdrant version or feature flag.
  • Self-hosted or older Qdrant — works against any Qdrant, or none.
  • Byte-for-byte FastEmbed parity — produces the same sparse vectors as Python FastEmbed's Qdrant/bm25, so a Go service can query a corpus indexed by a Python/FastEmbed pipeline. Verified with golden tests.
  • Non-Qdrant / offline use — plain BM25 sparse vectors for any sparse retrieval backend, or for testing and evaluation without a server.

Install

go get github.com/harsh04/bm25

Building BM25 sparse query vectors in Go

package main

import (
	"fmt"

	"github.com/harsh04/bm25"
)

func main() {
	enc := bm25.New()
	indices, values := enc.Encode("What is the cancellation policy?")
	fmt.Println(indices, values)
}

indices are uint32 token ids and values are float32 BM25 term-frequency weights. The two slices are aligned and ordered by first occurrence of each unique stemmed token.

Languages

New() is English. For other languages use NewWithLanguage:

enc, err := bm25.NewWithLanguage("german")
indices, values := enc.Encode("Wann ist die Stornierungsfrist?")

SupportedLanguages() returns the 16 languages whose output is verified byte-for-byte against FastEmbed: danish, dutch, english, finnish, french, german, hungarian, italian, norwegian, portuguese, romanian, russian, spanish, swedish, tamil, turkish. (Arabic and Greek are not yet supported — Arabic lacks a Snowball validation vocabulary and Greek lacks a pure-Go stemmer.)

WithK, WithB, and WithAvgLen override the BM25 parameters (changing them breaks parity with a FastEmbed-indexed corpus — leave them alone unless you're building your own index). WithoutStemmer() disables stemming and, matching FastEmbed's disable_stemmer, also disables stopword removal.

Qdrant hybrid search (dense + sparse vectors)

Use the sparse vector from this package as the sparse half of a hybrid query with the Qdrant Go client:

indices, values := enc.Encode(query)

sparse := &qdrant.Vector{
	Indices: &qdrant.SparseIndices{Data: indices},
	Data:    values,
}

(Field names vary slightly between client versions; adjust as needed.) Pair this with a dense embedding and Qdrant's RRF fusion to assemble full hybrid search. See _examples/qdrant_hybrid for a complete runnable program (collection setup, upsert, and dense+sparse RRF query).

For a higher-level helper, the separate qhybrid module (github.com/harsh04/bm25/qhybrid) wraps collection setup, upsert, and dense+sparse RRF search behind a small API — you supply a dense embedder, it handles the sparse side. It's a separate module so the core stays dep-light.

Scope

  • Term frequencies only. IDF is applied by Qdrant on the server through the sparse vector's IDF modifier, the way FastEmbed expects. Do not apply IDF on the client.
  • Stateless. k = 1.2, b = 0.75, and an assumed average length of 256 are fixed, matching the model. No corpus statistics are needed.
  • Empty, whitespace-only, and stopword-only input return empty slices.
  • Query vs document. Encode is the document path (BM25 term-frequency weights). EncodeQuery is FastEmbed's query_embed path — every unique token weighted 1.0 — for query vectors.
  • Hot paths. Encode is allocation-light (~1.7µs/query) and concurrency-safe to share. For high-QPS loops, EncodeInto(text, idx, val) reuses caller buffers for identical output with fewer allocations.

Concurrency & large corpora

A single Encoder is safe to share across goroutines. For indexing large corpora, Go (no GIL) encodes across all cores in-process:

// Parallel batch — same output as EncodeBatch, work spread over a worker pool.
indices, values := enc.EncodeBatchParallel(docs, 0) // 0 = runtime.NumCPU()

// Streaming pipeline — results as they finish, with context cancellation.
for r := range enc.EncodeStream(ctx, textCh, 0) {
    upsert(r.Text, r.Indices, r.Values)
}

EncodeBatchParallel is ~Nx faster on multicore (e.g. ~2.9x on 8 cores) and produces byte-identical output. EncodeStream emits StreamResults out of order (correlate via Text) and stops when the input closes or ctx is cancelled.

Performance vs Python FastEmbed

Same machine, same model, same corpus, byte-identical output (8-core Apple Silicon; full methodology in docs/BENCHMARKS.md):

Metric Python FastEmbed Go (this lib)
Throughput, 1 core 14,785 docs/sec 51,303 docs/sec (~3.5x)
Throughput, 8 cores (parallel) 170,314 docs/sec (~11.5x)
Cold start → first result 1,435 ms 2.9 ms (~500x)

A single ~3 MB static binary — no model download, no ONNX Runtime, no import cost. It also compiles to WebAssembly (js/wasm for the browser, wasip1/wasm for edge runtimes), so you can encode queries client-side or at the edge with no server round-trip — see _examples/wasm. The Python encoder can't.

FastEmbed Go compatibility

This is the core of the contract: output is byte-for-byte identical to FastEmbed's Qdrant/bm25 model — verified across fastembed 0.4.0 through 0.8.0 (the model output is unchanged across those releases; the golden vectors are pinned to 0.7.1). Token ids are produced by hashing stemmed tokens, so they line up with any corpus indexed by a compatible FastEmbed. A CI matrix re-runs the parity check against pinned versions plus the unpinned latest, so a future release that changes the model is caught. The pipeline is:

lowercase
  → split on non-word characters
  → drop stopwords, single-character punctuation, and tokens longer than 40 chars
  → Snowball (English) stemming
  → BM25 term-frequency weighting
  → abs(murmurhash3_x86_32(token, seed 0)) token ids

How parity is tested

Parity is verified against data captured from FastEmbed rather than assumed:

  • testdata/golden_vectors.json — 50 queries (punctuation, repeated tokens, mixed scripts, emoji, long tokens, stopword-only, underscores) with the exact indices and values FastEmbed produces.
  • testdata/intermediate_stages.json — per-stage output (tokens, stems, term-frequency map) for debugging individual cases.
  • testdata/stemmer_golden.json — 12,469 English word/stem pairs from py_rust_stemmers; the Go stemmer (blevesearch/snowballstem) matches all of them, and all 16 languages match across 1.1M+ words of the Snowball test vocab.
  • testdata/multilang_goldens.json — per-language encode goldens for all 16 supported languages.
  • testdata/query_goldens.jsonEncodeQuery parity against FastEmbed's query_embed.
go test ./...

Regenerate the fixtures (requires Python with fastembed==0.7.1):

python tools/gen_goldens.py            # English vectors + stemmer corpus
python tools/gen_multilang_goldens.py  # per-language vectors

Notes

  • Python's regex \w is Unicode-aware and treats combining marks as separators (for example, Devanagari and Tamil vowel signs split a word). Go's regexp \w is ASCII only, so tokenization is done by hand to match.
  • mmh3.hash returns a signed 32-bit value; token ids are its absolute value. abs(-2^31) is 2^31, which does not fit in int32 but does fit in uint32.

Dependencies

License

MIT

Documentation

Overview

Package bm25 is a byte-for-byte Go port of Python FastEmbed's "Qdrant/bm25" sparse-text encoder (fastembed==0.7.1).

It turns text into a Qdrant-compatible sparse vector — parallel arrays of token-id indices and BM25 term-frequency values — so a Go service can build the sparse half of a Qdrant hybrid search against a corpus that was indexed in Python with FastEmbed. The token ids match FastEmbed's exactly, so queries line up with the stored vectors.

Usage

Construct an Encoder with New, then call Encoder.Encode:

enc := bm25.New()
indices, values := enc.Encode("What is the cancellation policy?")

indices are uint32 token ids and values are float32 BM25 term-frequency weights; the two slices are aligned and ordered by first occurrence of each unique stemmed token. Use Encoder.EncodeBatch for many texts at once.

Languages

New is English. For other languages use NewWithLanguage, which returns an error for unsupported languages:

enc, err := bm25.NewWithLanguage("german")

SupportedLanguages lists the 16 languages whose output is verified byte-for-byte against FastEmbed (danish, dutch, english, finnish, french, german, hungarian, italian, norwegian, portuguese, romanian, russian, spanish, swedish, tamil, turkish). WithK, WithB, WithAvgLen, and WithoutStemmer override defaults.

Query vs document encoding

Encoder.Encode is the document path (BM25 term-frequency weights), used for indexing and matching a FastEmbed-indexed corpus. Encoder.EncodeQuery is FastEmbed's query_embed path (every token weighted 1.0) for query vectors.

Concurrency

A single Encoder is safe to share across goroutines. Encoder.EncodeBatchParallel spreads batch encoding over a worker pool (Go has no GIL, so this scales across cores), and Encoder.EncodeStream is a context-cancellable channel pipeline for streaming ETL — both for indexing large corpora.

Performance

On an 8-core machine, encoding is ~3.5x faster per core than Python FastEmbed and ~11.5x with EncodeBatchParallel, with ~500x faster cold start (a single ~3 MB static binary, no model download or ONNX Runtime). The encoder is pure Go and also compiles to WebAssembly (js/wasm and wasip1/wasm). See docs/BENCHMARKS.md and _examples/wasm in the repository.

Term frequencies only (no IDF)

The encoder emits term frequencies only — never IDF. IDF is applied server-side by Qdrant via the sparse vector's IDF modifier, exactly as FastEmbed expects. Do not apply IDF on the client.

Pipeline

Encoding mirrors fastembed.sparse.bm25.Bm25's document path:

lowercase
  -> split on non-word characters (Unicode-aware)
  -> drop stopwords, single-character punctuation, tokens longer than 40 runes
  -> Snowball (English) stemming
  -> BM25 term-frequency weighting (k=1.2, b=0.75, avg_len=256)
  -> abs(murmurhash3_x86_32(token, seed 0)) token ids

Empty, whitespace-only, and stopword-only input return empty slices.

Compatibility

Output is verified byte-for-byte against FastEmbed's "Qdrant/bm25" model across fastembed 0.4.0–0.8.0 (the model output is unchanged across those releases; goldens are pinned to 0.7.1). See the project README for a Qdrant Go client example and the full parity-testing setup: https://github.com/harsh04/bm25

Example

Encode a single query into a Qdrant-compatible sparse vector.

package main

import (
	"fmt"

	"github.com/harsh04/bm25"
)

func main() {
	enc := bm25.New()
	indices, values := enc.Encode("What is the cancellation policy?")

	// Stopwords ("what", "is", "the") are dropped; "cancellation" and "policy"
	// are stemmed, then hashed to token ids. Values are term frequencies (no IDF).
	fmt.Println("terms:", len(indices))
	for i := range indices {
		fmt.Printf("%d %.4f\n", indices[i], values[i])
	}
}
Output:
terms: 2
1818586924 1.6832
203139330 1.6832

Index

Examples

Constants

View Source
const (
	DefaultK        = 1.2   // term-frequency saturation
	DefaultB        = 0.75  // document-length normalization
	DefaultAvgLen   = 256.0 // assumed average document length
	DefaultLanguage = "english"
)

BM25 hyperparameters for the "Qdrant/bm25" model. These are fixed (stateless); the encoder does not see a corpus, so avgLen is a constant, not a measured mean.

Variables

This section is empty.

Functions

func SupportedLanguages added in v0.2.0

func SupportedLanguages() []string

SupportedLanguages returns the languages this package can encode, sorted. Each one matches FastEmbed's "Qdrant/bm25" output for that language.

Types

type Encoder

type Encoder struct {
	// contains filtered or unexported fields
}

Encoder produces FastEmbed-compatible BM25 sparse vectors. Construct one with New (English) or NewWithLanguage. An Encoder is safe for concurrent use; all fields are read-only after construction, so a single Encoder can be shared across goroutines.

func New

func New() *Encoder

New returns an English encoder with FastEmbed's "Qdrant/bm25" defaults (k=1.2, b=0.75, avgLen=256, English stopwords + Snowball stemmer).

func NewWithLanguage added in v0.2.0

func NewWithLanguage(language string, opts ...Option) (*Encoder, error)

NewWithLanguage returns an encoder for the given language (see SupportedLanguages), with optional overrides. It returns an error if the language is not supported.

Example

NewWithLanguage builds an encoder for a non-English language.

package main

import (
	"fmt"

	"github.com/harsh04/bm25"
)

func main() {
	enc, err := bm25.NewWithLanguage("german")
	if err != nil {
		panic(err)
	}
	indices, _ := enc.Encode("Wann ist der Check-in?")
	fmt.Println(len(indices), "terms")
}
Output:
2 terms

func (*Encoder) Encode

func (e *Encoder) Encode(text string) (indices []uint32, values []float32)

Encode turns a single text into a sparse vector: indices are token ids (abs(murmurhash3_x86_32(stemmed_token, seed=0))) and values are the BM25 term-frequency weights. Empty, whitespace-only, and stopword-only text yields empty slices.

The arrays are aligned and ordered by first occurrence of each unique stemmed token, matching FastEmbed's output ordering exactly. Values are term frequencies only; Qdrant applies IDF server-side. To encode many texts, use Encoder.EncodeBatch.

The returned slices are freshly allocated and owned by the caller (safe to retain or mutate), and are non-nil even when empty. For a buffer-reusing variant on hot paths, see Encoder.EncodeInto.

Example

Stemming collapses inflected forms to one term, and repeats are counted together, so "running"/"runs" both fold into a single entry.

package main

import (
	"fmt"

	"github.com/harsh04/bm25"
)

func main() {
	enc := bm25.New()
	indices, _ := enc.Encode("running runs ran runner")

	// run (from running+runs), ran, runner -> 3 unique terms.
	fmt.Println(len(indices), "unique terms")
}
Output:
3 unique terms
Example (StopwordsOnly)

Empty, whitespace-only, and stopword-only text produce an empty sparse vector.

package main

import (
	"fmt"

	"github.com/harsh04/bm25"
)

func main() {
	enc := bm25.New()
	indices, values := enc.Encode("the and of is")
	fmt.Println(len(indices), len(values))
}
Output:
0 0

func (*Encoder) EncodeBatch

func (e *Encoder) EncodeBatch(texts []string) (indices [][]uint32, values [][]float32)

EncodeBatch encodes each text independently with Encoder.Encode. The returned slices are aligned with texts: indices[i]/values[i] are the sparse vector for texts[i]. An empty input yields empty (non-nil) results.

Example

EncodeBatch encodes several texts in one call; results are aligned by index.

package main

import (
	"fmt"

	"github.com/harsh04/bm25"
)

func main() {
	enc := bm25.New()
	indices, _ := enc.EncodeBatch([]string{"wifi password", "free breakfast"})

	for i, idx := range indices {
		fmt.Printf("doc %d: %d terms\n", i, len(idx))
	}
}
Output:
doc 0: 2 terms
doc 1: 2 terms

func (*Encoder) EncodeBatchParallel added in v1.3.0

func (e *Encoder) EncodeBatchParallel(texts []string, workers int) (indices [][]uint32, values [][]float32)

EncodeBatchParallel is Encoder.EncodeBatch across a worker pool — Go has no GIL, so this scales encoding across all cores in-process, which matters when indexing large corpora. Output is identical to EncodeBatch (results stay aligned with texts by index); only the work is parallelized.

workers <= 0 uses runtime.NumCPU(). A single Encoder is safe to share here.

func (*Encoder) EncodeInto added in v1.2.0

func (e *Encoder) EncodeInto(text string, indices []uint32, values []float32) ([]uint32, []float32)

EncodeInto is Encoder.Encode that appends into caller-provided buffers, avoiding allocations in hot loops. Pass the slices returned by the previous call (or nil); they are truncated and refilled, and the grown slices returned:

var idx []uint32
var val []float32
for _, q := range queries {
	idx, val = enc.EncodeInto(q, idx, val)
	// use idx/val before the next call — they are overwritten
}

Output is identical to Encoder.Encode; the returned slices are non-nil even when empty.

func (*Encoder) EncodeQuery added in v1.1.0

func (e *Encoder) EncodeQuery(text string) (indices []uint32, values []float32)

EncodeQuery encodes a query the way FastEmbed's query_embed does: each unique stemmed token gets a weight of 1.0, with no term-frequency weighting. Use this for QUERY vectors when your corpus was indexed with FastEmbed's standard document/query split; use Encoder.Encode (term frequencies) when both your index and queries go through the document path.

Indices are deduplicated token ids in first-occurrence order and values are all 1.0. (FastEmbed emits them in unspecified set order; Qdrant treats a sparse vector as a map, so order does not affect retrieval.) The returned slices are non-nil even when empty.

Example

EncodeQuery weights every unique token 1.0 (FastEmbed's query_embed path), rather than by term frequency.

package main

import (
	"fmt"

	"github.com/harsh04/bm25"
)

func main() {
	enc := bm25.New()
	_, values := enc.EncodeQuery("breakfast breakfast pool")
	fmt.Println(values)
}
Output:
[1 1]

func (*Encoder) EncodeStream added in v1.3.0

func (e *Encoder) EncodeStream(ctx context.Context, texts <-chan string, workers int) <-chan StreamResult

EncodeStream encodes a stream of texts concurrently, emitting results as they finish (order not preserved — use Text to correlate). It's a building block for pipelined ETL: feed documents in, upsert sparse vectors out, with backpressure and cancellation via ctx.

The output channel is closed when texts is drained or ctx is cancelled. workers <= 0 uses runtime.NumCPU().

func (*Encoder) Language added in v1.0.0

func (e *Encoder) Language() string

Language returns the encoder's configured language.

type Option added in v0.2.0

type Option func(*config)

Option configures an Encoder built by NewWithLanguage. The set of options is closed by design; use WithK, WithB, WithAvgLen, or WithoutStemmer.

func WithAvgLen added in v1.0.0

func WithAvgLen(avgLen float64) Option

WithAvgLen overrides the assumed average document length (default DefaultAvgLen). See WithK for the parity caveat.

func WithB added in v1.0.0

func WithB(b float64) Option

WithB overrides the BM25 document-length normalization parameter (default DefaultB). See WithK for the parity caveat.

func WithK added in v1.0.0

func WithK(k float64) Option

WithK overrides the BM25 term-frequency saturation parameter (default DefaultK). Changing k/b/avgLen breaks parity with a FastEmbed-indexed corpus — leave them alone unless you are building an independent index.

func WithoutStemmer added in v0.2.0

func WithoutStemmer() Option

WithoutStemmer disables stemming. This mirrors FastEmbed's disable_stemmer=True, which also drops stopword removal — tokens are only lowercased, filtered by length/punctuation, then hashed.

type StreamResult added in v1.3.0

type StreamResult struct {
	Text    string
	Indices []uint32
	Values  []float32
}

StreamResult is one encoded document from Encoder.EncodeStream. Text echoes the input so callers can correlate results, which may arrive out of order.

Directories

Path Synopsis
qhybrid module

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL