bm25md

package module
v0.0.0-...-0bf9e79 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Jul 24, 2025 License: BSD-3-Clause Imports: 12 Imported by: 0

README

BM25md - Field-Aware BM25 for Markdown

Go Version Go Reference Go Report Card Tests

BM25md is a Go implementation of the BM25 text ranking algorithm with field-aware scoring for markdown documents. It augments traditional BM25 by assigning different weights to content based on its semantic importance in markdown structure (headers, bold text, code blocks, etc.).

Installation

go get github.com/chriscorrea/bm25md

Quick Start

package main

import (
    "fmt"
    "github.com/chriscorrea/bm25md"
)

func main() {
    // create a corpus with default config
    corpus := bm25md.NewCorpus()
    
    // create a parser for markdown documents
    parser := bm25md.NewMarkdownFieldParser()
    
    // add documents
    docs := []string{
        "# Introduction\nThis is a **key concept** in information retrieval.",
        "## Details\nThe algorithm uses `code examples` to demonstrate usage.",
        "### Summary\nBM25 is effective for ranking documents.",
    }
    
    for _, content := range docs {
        fields := parser.ParseDocument(content)
        corpus.AddDocument(bm25md.Document{
            Fields:   fields,
            Original: content,
        })
    }
    
    // search
    results := corpus.Search("key concept", 10)
    
    for _, result := range results {
        fmt.Printf("Score: %.3f, Document: %s\n", result.Score, result.Document.Original)
    }
}

Custom Configuration

The functional options API provides clean, extensible configuration:

// custom field weights
weights := map[bm25md.Field]float64{
    bm25md.FieldH1:   10.0,  // very heavily weight H1 headers
    bm25md.FieldCode: 2.0,   // increase code importance
    bm25md.FieldBody: 1.0,   
}

// custom BM25 parameters
params := bm25md.BM25Parameters{
    K1: 1.5,  // term frequency saturation
    B:  0.5,  // doc length normalization
}

// Create corpus with custom configuration
corpus := bm25md.NewCorpus(
    bm25md.WithFieldWeights(weights),
    bm25md.WithBM25Params(params),
)

Note that, for advanced use cases, you can specify different BM25 parameters for each field:

fieldParams := map[bm25md.Field]bm25md.BM25Parameters{
    bm25md.FieldH1:   {K1: 0.8, B: 0.9}, // long headers are penalized
    bm25md.FieldBody: {K1: 1.5, B: 0.75}, // higher saturation for body text
    bm25md.FieldCode: {K1: 1.2, B: 0.5},  // lenient on code block length
}

corpus := bm25md.NewCorpus(bm25md.WithFieldParams(fieldParams))
Custom Tokenizers

The default tokenizer is very simple. You can implement custom tokenization to apply stemming, normalization, or domain-specific processing.

Using a function-based tokenizer provides a lightweight approach:

customStemmingTokenizer := func(text string) []string {
    tokens := strings.Fields(strings.ToLower(text))
    // apply your custom stemming/normalization logic here
    return applyStemming(tokens)
}

corpus := bm25md.NewCorpus(bm25md.WithTokenizer(bm25md.TokenizerFunc(customStemmingTokenizer)))

Alternatively, you can implement the Tokenizer interface:

type MyTokenizer struct{}
func (t MyTokenizer) Tokenize(text string) []string {
    // apply your custom stemming/normalization logic here
    return customTokenize(text)
}

corpus := bm25md.NewCorpus(bm25md.WithTokenizer(MyTokenizer{}))

Dependencies

BM25md leverages goldmark for AST-based markdown parsing.

Contributing

Contributions and issues are welcome – please see the issues page.

License

This project is licensed under the BSD-3 License.

Documentation

Overview

Package bm25md provides BM25-based search functionality with Markdown field awareness.

This implements a field-weighted variant of the BM25 (Best Matching 25) ranking algorithm that gives different weights to content based on its semantic importance in Markdown documents (headers, bold text, body content, etc.).

The algorithm handles multiple fields, where each field can have its own weight, allowing for more nuanced ranking based on where terms appear in the document structure.

Index

Constants

This section is empty.

Variables

View Source
var DefaultFieldWeights = map[Field]float64{
	FieldH1:     5.0,
	FieldH2:     3.0,
	FieldH3:     2.0,
	FieldH4:     2.0,
	FieldH5:     2.0,
	FieldH6:     2.0,
	FieldBold:   1.5,
	FieldItalic: 1.2,
	FieldCode:   0.8,
	FieldBody:   1.0,
}

DefaultFieldWeights provides sensible default weights for markdown fields

Functions

func DefaultFieldBM25Parameters

func DefaultFieldBM25Parameters() map[Field]BM25Parameters

DefaultFieldBM25Parameters returns field-specific BM25 parameters optimized for each field type

Types

type BM25Parameters

type BM25Parameters struct {
	K1 float64 // controls term frequency saturation
	B  float64 // controls length normalization
}

BM25Parameters holds the tuning parameters for BM25 algorithm

func DefaultBM25Parameters

func DefaultBM25Parameters() BM25Parameters

DefaultBM25Parameters returns recommended BM25 parameters

type Corpus

type Corpus struct {
	// contains filtered or unexported fields
}

Corpus manages the BM25md search index for a corpus

func NewCorpus

func NewCorpus(opts ...CorpusOption) *Corpus

NewCorpus creates a new BM25md corpus with optional configuration

func (*Corpus) AddDocument

func (c *Corpus) AddDocument(doc Document)

AddDocument adds a document to the corpus

func (*Corpus) Score

func (c *Corpus) Score(query string, docIndex int) float64

Score calculates the BM25md score for a query against a specific document

func (*Corpus) Search

func (c *Corpus) Search(query string, limit int) []SearchResult

Search performs a BM25md search and returns ranked results

type CorpusOption

type CorpusOption func(*Corpus)

CorpusOption defines a function that configures a corpus

func WithBM25Params

func WithBM25Params(params BM25Parameters) CorpusOption

WithBM25Params sets custom BM25 parameters for the corpus

func WithFieldParams

func WithFieldParams(fieldParams map[Field]BM25Parameters) CorpusOption

WithFieldParams sets per-field BM25 parameters (BM25F mode)

func WithFieldWeights

func WithFieldWeights(fieldWeights map[Field]float64) CorpusOption

WithFieldWeights sets custom field weights for the corpus

func WithTokenizer

func WithTokenizer(tokenizer Tokenizer) CorpusOption

WithTokenizer sets a custom tokenizer for the corpus

type DefaultTokenizer

type DefaultTokenizer struct{}

DefaultTokenizer implements a basic default tokenizer

func (DefaultTokenizer) Tokenize

func (t DefaultTokenizer) Tokenize(text string) []string

Tokenize implements the Tokenizer interface

type Document

type Document struct {
	ID       int              // document identifier
	Fields   map[Field]string // content separated by field type
	Original string           // original document text
}

Document represents a parsed document with field-separated content

type Field

type Field string

Field represents a specific markdown field type with its weight

const (
	FieldH1     Field = "h1"
	FieldH2     Field = "h2"
	FieldH3     Field = "h3"
	FieldH4     Field = "h4"
	FieldH5     Field = "h5"
	FieldH6     Field = "h6"
	FieldBold   Field = "bold"
	FieldItalic Field = "italic"
	FieldCode   Field = "code"
	FieldBody   Field = "body"
)

type MarkdownFieldParser

type MarkdownFieldParser struct {
	// contains filtered or unexported fields
}

MarkdownFieldParser extracts content from markdown documents

func NewMarkdownFieldParser

func NewMarkdownFieldParser() *MarkdownFieldParser

NewMarkdownFieldParser creates new AST-based parser instance

func (*MarkdownFieldParser) ParseDocument

func (p *MarkdownFieldParser) ParseDocument(content string) map[Field]string

ParseDocument extracts field-specific content using AST traversal

func (*MarkdownFieldParser) ParseDocuments

func (p *MarkdownFieldParser) ParseDocuments(contents []string) []Document

ParseDocuments parses multiple markdown documents into BM25md Documents

type SearchResult

type SearchResult struct {
	Document Document
	Score    float64
	Index    int
}

SearchResult represents a document with its relevance score

type Tokenizer

type Tokenizer interface {
	Tokenize(text string) []string
}

Tokenizer defines the interface for text tokenization

type TokenizerFunc

type TokenizerFunc func(string) []string

TokenizerFunc is a func adapter that allows using functions as Tokenizers

func (TokenizerFunc) Tokenize

func (f TokenizerFunc) Tokenize(text string) []string

Tokenize implements the Tokenizer interface for function types

Directories

Path Synopsis
examples
basic command
custom command

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL