Documentation
¶
Overview ¶
Package bm25 is a byte-for-byte Go port of Python FastEmbed's "Qdrant/bm25" sparse-text encoder (fastembed==0.7.1).
It turns text into a Qdrant-compatible sparse vector — parallel arrays of token-id indices and BM25 term-frequency values — so a Go service can build the sparse half of a Qdrant hybrid search against a corpus that was indexed in Python with FastEmbed. The token ids match FastEmbed's exactly, so queries line up with the stored vectors.
Usage ¶
Construct an Encoder with New, then call Encoder.Encode:
enc := bm25.New()
indices, values := enc.Encode("What is the cancellation policy?")
indices are uint32 token ids and values are float32 BM25 term-frequency weights; the two slices are aligned and ordered by first occurrence of each unique stemmed token. Use Encoder.EncodeBatch for many texts at once.
Languages ¶
New is English. For other languages use NewWithLanguage, which returns an error for unsupported languages:
enc, err := bm25.NewWithLanguage("german")
SupportedLanguages lists the 16 languages whose output is verified byte-for-byte against FastEmbed (danish, dutch, english, finnish, french, german, hungarian, italian, norwegian, portuguese, romanian, russian, spanish, swedish, tamil, turkish). WithK, WithB, WithAvgLen, and WithoutStemmer override defaults.
Query vs document encoding ¶
Encoder.Encode is the document path (BM25 term-frequency weights), used for indexing and matching a FastEmbed-indexed corpus. Encoder.EncodeQuery is FastEmbed's query_embed path (every token weighted 1.0) for query vectors.
Term frequencies only (no IDF) ¶
The encoder emits term frequencies only — never IDF. IDF is applied server-side by Qdrant via the sparse vector's IDF modifier, exactly as FastEmbed expects. Do not apply IDF on the client.
Pipeline ¶
Encoding mirrors fastembed.sparse.bm25.Bm25's document path:
lowercase -> split on non-word characters (Unicode-aware) -> drop stopwords, single-character punctuation, tokens longer than 40 runes -> Snowball (English) stemming -> BM25 term-frequency weighting (k=1.2, b=0.75, avg_len=256) -> abs(murmurhash3_x86_32(token, seed 0)) token ids
Empty, whitespace-only, and stopword-only input return empty slices.
Compatibility ¶
Output is verified byte-for-byte against fastembed==0.7.1 (model "Qdrant/bm25") with golden tests. See the project README for a Qdrant Go client example and the full parity-testing setup: https://github.com/harsh04/bm25
Example ¶
Encode a single query into a Qdrant-compatible sparse vector.
package main
import (
"fmt"
"github.com/harsh04/bm25"
)
func main() {
enc := bm25.New()
indices, values := enc.Encode("What is the cancellation policy?")
// Stopwords ("what", "is", "the") are dropped; "cancellation" and "policy"
// are stemmed, then hashed to token ids. Values are term frequencies (no IDF).
fmt.Println("terms:", len(indices))
for i := range indices {
fmt.Printf("%d %.4f\n", indices[i], values[i])
}
}
Output: terms: 2 1818586924 1.6832 203139330 1.6832
Index ¶
Examples ¶
Constants ¶
const ( DefaultK = 1.2 // term-frequency saturation DefaultB = 0.75 // document-length normalization DefaultAvgLen = 256.0 // assumed average document length DefaultLanguage = "english" )
BM25 hyperparameters for the "Qdrant/bm25" model. These are fixed (stateless); the encoder does not see a corpus, so avgLen is a constant, not a measured mean.
Variables ¶
This section is empty.
Functions ¶
func SupportedLanguages ¶ added in v0.2.0
func SupportedLanguages() []string
SupportedLanguages returns the languages this package can encode, sorted. Each one matches FastEmbed's "Qdrant/bm25" output for that language.
Types ¶
type Encoder ¶
type Encoder struct {
// contains filtered or unexported fields
}
Encoder produces FastEmbed-compatible BM25 sparse vectors. Construct one with New (English) or NewWithLanguage. An Encoder is safe for concurrent use; all fields are read-only after construction, so a single Encoder can be shared across goroutines.
func New ¶
func New() *Encoder
New returns an English encoder with FastEmbed's "Qdrant/bm25" defaults (k=1.2, b=0.75, avgLen=256, English stopwords + Snowball stemmer).
func NewWithLanguage ¶ added in v0.2.0
NewWithLanguage returns an encoder for the given language (see SupportedLanguages), with optional overrides. It returns an error if the language is not supported.
Example ¶
NewWithLanguage builds an encoder for a non-English language.
package main
import (
"fmt"
"github.com/harsh04/bm25"
)
func main() {
enc, err := bm25.NewWithLanguage("german")
if err != nil {
panic(err)
}
indices, _ := enc.Encode("Wann ist der Check-in?")
fmt.Println(len(indices), "terms")
}
Output: 2 terms
func (*Encoder) Encode ¶
Encode turns a single text into a sparse vector: indices are token ids (abs(murmurhash3_x86_32(stemmed_token, seed=0))) and values are the BM25 term-frequency weights. Empty, whitespace-only, and stopword-only text yields empty slices.
The arrays are aligned and ordered by first occurrence of each unique stemmed token, matching FastEmbed's output ordering exactly. Values are term frequencies only; Qdrant applies IDF server-side. To encode many texts, use Encoder.EncodeBatch.
The returned slices are freshly allocated and owned by the caller (safe to retain or mutate), and are non-nil even when empty.
Example ¶
Stemming collapses inflected forms to one term, and repeats are counted together, so "running"/"runs" both fold into a single entry.
package main
import (
"fmt"
"github.com/harsh04/bm25"
)
func main() {
enc := bm25.New()
indices, _ := enc.Encode("running runs ran runner")
// run (from running+runs), ran, runner -> 3 unique terms.
fmt.Println(len(indices), "unique terms")
}
Output: 3 unique terms
Example (StopwordsOnly) ¶
Empty, whitespace-only, and stopword-only text produce an empty sparse vector.
package main
import (
"fmt"
"github.com/harsh04/bm25"
)
func main() {
enc := bm25.New()
indices, values := enc.Encode("the and of is")
fmt.Println(len(indices), len(values))
}
Output: 0 0
func (*Encoder) EncodeBatch ¶
EncodeBatch encodes each text independently with Encoder.Encode. The returned slices are aligned with texts: indices[i]/values[i] are the sparse vector for texts[i]. An empty input yields empty (non-nil) results.
Example ¶
EncodeBatch encodes several texts in one call; results are aligned by index.
package main
import (
"fmt"
"github.com/harsh04/bm25"
)
func main() {
enc := bm25.New()
indices, _ := enc.EncodeBatch([]string{"wifi password", "free breakfast"})
for i, idx := range indices {
fmt.Printf("doc %d: %d terms\n", i, len(idx))
}
}
Output: doc 0: 2 terms doc 1: 2 terms
func (*Encoder) EncodeQuery ¶ added in v1.1.0
EncodeQuery encodes a query the way FastEmbed's query_embed does: each unique stemmed token gets a weight of 1.0, with no term-frequency weighting. Use this for QUERY vectors when your corpus was indexed with FastEmbed's standard document/query split; use Encoder.Encode (term frequencies) when both your index and queries go through the document path.
Indices are deduplicated token ids in first-occurrence order and values are all 1.0. (FastEmbed emits them in unspecified set order; Qdrant treats a sparse vector as a map, so order does not affect retrieval.) The returned slices are non-nil even when empty.
Example ¶
EncodeQuery weights every unique token 1.0 (FastEmbed's query_embed path), rather than by term frequency.
package main
import (
"fmt"
"github.com/harsh04/bm25"
)
func main() {
enc := bm25.New()
_, values := enc.EncodeQuery("breakfast breakfast pool")
fmt.Println(values)
}
Output: [1 1]
type Option ¶ added in v0.2.0
type Option func(*config)
Option configures an Encoder built by NewWithLanguage. The set of options is closed by design; use WithK, WithB, WithAvgLen, or WithoutStemmer.
func WithAvgLen ¶ added in v1.0.0
WithAvgLen overrides the assumed average document length (default DefaultAvgLen). See WithK for the parity caveat.
func WithB ¶ added in v1.0.0
WithB overrides the BM25 document-length normalization parameter (default DefaultB). See WithK for the parity caveat.
func WithK ¶ added in v1.0.0
WithK overrides the BM25 term-frequency saturation parameter (default DefaultK). Changing k/b/avgLen breaks parity with a FastEmbed-indexed corpus — leave them alone unless you are building an independent index.
func WithoutStemmer ¶ added in v0.2.0
func WithoutStemmer() Option
WithoutStemmer disables stemming. This mirrors FastEmbed's disable_stemmer=True, which also drops stopword removal — tokens are only lowercased, filtered by length/punctuation, then hashed.