Documentation
¶
Overview ¶
Package veclite provides an embeddable vector database for Go applications.
VecLite stores vectors, text documents, and metadata in a single file, supports HNSW indexing for fast approximate nearest-neighbor search, and keeps the core storage and search APIs local-first. Optional integrations provide embedders, config, and MCP tooling.
Quick Start ¶
Open a database, create a collection, insert vectors, and search:
db, err := veclite.Open("data.veclite")
if err != nil {
log.Fatal(err)
}
defer db.Close()
coll := db.Collection("embeddings")
id, err := coll.Insert(vector, map[string]any{"text": "hello world"})
results, err := coll.Search(queryVector, veclite.TopK(10))
Collections ¶
Collections are namespaced containers for vectors. Create them with options:
coll, err := db.CreateCollection("docs",
veclite.WithDimension(384),
veclite.WithHNSW(16, 200),
veclite.WithDistanceType(veclite.DistanceCosine),
)
Search ¶
Vector search supports filtering, pagination, and hybrid vector+text search:
results, err := coll.Search(query,
veclite.TopK(20),
veclite.WithFilter(veclite.Equal("category", "science")),
veclite.Threshold(0.7),
)
Text-only documents can be stored for BM25-first workflows:
id, err := coll.InsertTextDocument("searchable text", map[string]any{"source": "timeline"})
Persistence ¶
Data is persisted to a single file using gob encoding with atomic writes. Use ":memory:" for an in-memory database:
db, err := veclite.Open(":memory:")
Thread Safety ¶
All operations on DB and Collection are safe for concurrent use. Multiple goroutines can read and write simultaneously.
Package veclite provides an embeddable vector database for Go. It stores vectors with metadata in a single file using gob encoding.
Basic usage:
db, err := veclite.Open("data.veclite")
if err != nil {
log.Fatal(err)
}
defer db.Close()
coll := db.Collection("embeddings")
id, err := coll.Insert(vector, map[string]any{"file": "main.go"})
results, err := coll.Search(queryVector, veclite.TopK(10))
Index ¶
- Constants
- Variables
- func ExpandEnvVars(s string) string
- func ExpandPath(path string) string
- func ParseDuration(s string, defaultDur time.Duration) time.Duration
- type Collection
- func (c *Collection) AddVectorSpace(config VectorSpaceConfig) error
- func (c *Collection) All() []*Record
- func (c *Collection) ArchiveRecord(id uint64) error
- func (c *Collection) CleanupExpired() (int, error)
- func (c *Collection) Clear() error
- func (c *Collection) Consolidate(config ConsolidationConfig) (*ConsolidationResult, error)
- func (c *Collection) Count() int
- func (c *Collection) CountExpired() int
- func (c *Collection) Delete(id uint64) error
- func (c *Collection) DeleteMetadataValue(key string) error
- func (c *Collection) DeleteWhere(filters ...Filter) (int, error)
- func (c *Collection) Dimension() int
- func (c *Collection) DistanceType() floats.DistanceType
- func (c *Collection) EmbeddingProfile() (EmbeddingProfile, bool)
- func (c *Collection) EnforceMemoryLimit(config MemoryConfig) int
- func (c *Collection) ExpandConsolidation(consolidationID uint64) ([]*Record, error)
- func (c *Collection) Find(filters ...Filter) ([]*Record, error)
- func (c *Collection) FindOne(filters ...Filter) (*Record, error)
- func (c *Collection) FindSimilarClusters(config ConsolidationConfig) ([]MemoryCluster, error)
- func (c *Collection) ForEach(fn func(*Record) bool)
- func (c *Collection) Get(id uint64) (*Record, error)
- func (c *Collection) GetArchived() ([]*Record, error)
- func (c *Collection) GetConsolidations() ([]*Record, error)
- func (c *Collection) GetSession(sessionID string) ([]*Record, error)
- func (c *Collection) GetSessionStats(sessionID string) (SessionStats, error)
- func (c *Collection) GetThread(chunkID uint64) ([]*Record, error)
- func (c *Collection) GetVector(id uint64) ([]float32, error)
- func (c *Collection) HasIndex() bool
- func (c *Collection) HasVectorSpace(name string) bool
- func (c *Collection) HybridSearch(query []float32, text string, opts ...SearchOption) ([]Result, error)
- func (c *Collection) HybridSearchSpace(space string, query []float32, text string, opts ...SearchOption) ([]Result, error)
- func (c *Collection) IndexStats() *HNSWStats
- func (c *Collection) IndexType() IndexType
- func (c *Collection) Insert(vector []float32, payload map[string]any) (uint64, error)
- func (c *Collection) InsertBatch(vectors [][]float32, payloads []map[string]any) ([]uint64, error)
- func (c *Collection) InsertDocument(vector []float32, content string, payload map[string]any) (uint64, error)
- func (c *Collection) InsertRecord(in RecordInput) (uint64, error)
- func (c *Collection) InsertText(text string, payload map[string]any) (uint64, error)
- func (c *Collection) InsertTextDocument(content string, payload map[string]any) (uint64, error)
- func (c *Collection) InsertTextDocumentWithOptions(content string, payload map[string]any, opts ...InsertOption) (uint64, error)
- func (c *Collection) InsertTurn(turn ConversationTurn) (uint64, error)
- func (c *Collection) InsertWithOptions(vector []float32, payload map[string]any, opts ...InsertOption) (uint64, error)
- func (c *Collection) Iterate(opts ...IterOption) *Iterator
- func (c *Collection) ListSessions() []string
- func (c *Collection) Metadata() map[string]any
- func (c *Collection) MultiSpaceSearch(queries map[string][]float32, opts ...SearchOption) ([]Result, error)
- func (c *Collection) Name() string
- func (c *Collection) Reset() error
- func (c *Collection) Search(query []float32, opts ...SearchOption) ([]Result, error)
- func (c *Collection) SearchExplain(query []float32, opts ...SearchOption) (*SearchExplanation, error)
- func (c *Collection) SearchInSession(sessionID string, query []float32, opts ...SearchOption) ([]Result, error)
- func (c *Collection) SearchSpace(space string, query []float32, opts ...SearchOption) ([]Result, error)
- func (c *Collection) SearchStream(query []float32, fn SearchFunc, opts ...SearchOption) error
- func (c *Collection) SearchText(text string, opts ...SearchOption) ([]Result, error)
- func (c *Collection) SetEmbeddingProfile(profile EmbeddingProfile) error
- func (c *Collection) SetMetadata(metadata map[string]any) error
- func (c *Collection) SetMetadataValue(key string, value any) error
- func (c *Collection) SetRecordVector(id uint64, space string, vector []float32) error
- func (c *Collection) StartMemoryLimiter(config MemoryConfig) func()
- func (c *Collection) Stats() CollectionStats
- func (c *Collection) Subscribe(query []float32, opts ...SubscriptionOption) (*Subscription, error)
- func (c *Collection) TextSearch(query string, opts ...SearchOption) ([]Result, error)
- func (c *Collection) UnarchiveRecord(id uint64) error
- func (c *Collection) Unsubscribe(subscriptionID string) error
- func (c *Collection) Update(id uint64, payload map[string]any) error
- func (c *Collection) UpdateDocument(id uint64, content string, payload map[string]any) error
- func (c *Collection) UpdateVector(id uint64, vector []float32) error
- func (c *Collection) Upsert(id uint64, vector []float32, payload map[string]any) (uint64, error)
- func (c *Collection) UpsertByKey(keyField string, keyValue any, vector []float32, payload map[string]any) (uint64, bool, error)
- func (c *Collection) UpsertRecordByKey(keyField string, keyValue any, in RecordInput) (uint64, bool, error)
- func (c *Collection) UpsertTextDocument(id uint64, content string, payload map[string]any) (uint64, error)
- func (c *Collection) UpsertTextDocumentByKey(keyField string, keyValue any, content string, payload map[string]any) (uint64, bool, error)
- func (c *Collection) VectorSpace(name string) (VectorSpaceInfo, error)
- func (c *Collection) VectorSpaces() []VectorSpaceInfo
- type CollectionOption
- func WithDimension(dim int) CollectionOption
- func WithDistanceType(t floats.DistanceType) CollectionOption
- func WithEmbedder(e Embedder) CollectionOption
- func WithEmbeddingProfile(profile EmbeddingProfile) CollectionOption
- func WithHNSW(m, efConstruction int) CollectionOption
- func WithHNSWConfig(config HNSWConfig) CollectionOption
- func WithMemoryLimits(config MemoryConfig) CollectionOption
- func WithTextIndex(fields ...string) CollectionOption
- func WithVectorSpace(config VectorSpaceConfig) CollectionOption
- type CollectionSnapshot
- type CollectionStats
- type Config
- type ConsolidationConfig
- type ConsolidationResult
- type ConversationTurn
- type DB
- func (db *DB) Close() error
- func (db *DB) Collection(name string) *Collection
- func (db *DB) Collections() []string
- func (db *DB) CreateCollection(name string, opts ...CollectionOption) (*Collection, error)
- func (db *DB) CreateEpisodeStore(memoriesCollectionName string) (*EpisodeStore, error)
- func (db *DB) CreateKnowledgeGraph(name string) (*KnowledgeGraph, error)
- func (db *DB) DeleteMetadataValue(key string) error
- func (db *DB) DropCollection(name string) error
- func (db *DB) GetCollection(name string) (*Collection, error)
- func (db *DB) GetEpisodeStore(name string) (*EpisodeStore, error)
- func (db *DB) GetKnowledgeGraph(name string) (*KnowledgeGraph, error)
- func (db *DB) HasCollection(name string) bool
- func (db *DB) IsClosed() bool
- func (db *DB) Metadata() map[string]any
- func (db *DB) Metrics() MetricsSnapshot
- func (db *DB) Path() string
- func (db *DB) Reload() error
- func (db *DB) SetMetadata(metadata map[string]any) error
- func (db *DB) SetMetadataValue(key string, value any) error
- func (db *DB) StartTTLCleaner(interval time.Duration, callback TTLCleanerCallback) func()
- func (db *DB) Stats() DatabaseStats
- func (db *DB) Sync() error
- type DatabaseSnapshot
- type DatabaseStats
- type DecayConfig
- type DecayType
- type DimensionError
- type DistanceType
- type Embedder
- type EmbedderConfig
- type EmbeddingProfile
- type Entity
- type EntitySnapshot
- type Episode
- type EpisodeConfig
- type EpisodeResult
- type EpisodeSnapshot
- type EpisodeStore
- func (es *EpisodeStore) CreateEpisode(recordIDs []uint64, title string) (*Episode, error)
- func (es *EpisodeStore) DeleteEpisode(episodeID string) error
- func (es *EpisodeStore) DetectEpisodes(config EpisodeConfig) ([]*Episode, error)
- func (es *EpisodeStore) ExpandEpisode(episodeID string) ([]*Record, error)
- func (es *EpisodeStore) FindRecordEpisode(recordID uint64) (*Episode, error)
- func (es *EpisodeStore) GetEpisode(episodeID string) (*Episode, error)
- func (es *EpisodeStore) ListEpisodes() []*Episode
- func (es *EpisodeStore) SearchEpisodes(query []float32, limit int) ([]*Episode, error)
- func (es *EpisodeStore) SearchWithEpisodeExpansion(query []float32, opts ...SearchOption) ([]EpisodeResult, error)
- type EpisodeStoreSnapshot
- type ExpandedSearchResult
- type Filter
- func AccessCountAbove(n uint64) Filter
- func AccessCountBelow(n uint64) Filter
- func AccessedAfter(t time.Time) Filter
- func AccessedBefore(t time.Time) Filter
- func AgeNewerThan(d time.Duration) Filter
- func AgeOlderThan(d time.Duration) Filter
- func And(filters ...Filter) Filter
- func Between(key string, min, max float64) Filter
- func Contains(key, substr string) Filter
- func CreatedAfter(t time.Time) Filter
- func CreatedBefore(t time.Time) Filter
- func Equal(key string, value any) Filter
- func Exists(key string) Filter
- func ExpiredBefore(t time.Time) Filter
- func GT(key string, value float64) Filter
- func GTE(key string, value float64) Filter
- func Glob(key, pattern string) Filter
- func GreaterThan(key string, value float64) Filter
- func GreaterThanOrEqual(key string, value float64) Filter
- func HasTTLFilter() Filter
- func ImportanceAbove(threshold float32) Filter
- func ImportanceBelow(threshold float32) Filter
- func ImportanceBetween(min, max float32) Filter
- func In(key string, values ...any) Filter
- func LT(key string, value float64) Filter
- func LTE(key string, value float64) Filter
- func LessThan(key string, value float64) Filter
- func LessThanOrEqual(key string, value float64) Filter
- func NeverAccessed() Filter
- func Not(filter Filter) Filter
- func NotEqual(key string, value any) Filter
- func NotExpired() Filter
- func NotIn(key string, values ...any) Filter
- func Or(filters ...Filter) Filter
- func Prefix(key, prefix string) Filter
- func Suffix(key, suffix string) Filter
- func UpdatedAfter(t time.Time) Filter
- func UpdatedBefore(t time.Time) Filter
- type FilterFunc
- type FuseOption
- type GraphSnapshot
- type HNSWConfig
- type HNSWStats
- type Index
- type IndexResult
- type IndexType
- type InsertOption
- type InvertedIndexSnapshot
- type IterOption
- type Iterator
- type KnowledgeGraph
- func (kg *KnowledgeGraph) AddEntity(entity Entity) error
- func (kg *KnowledgeGraph) AddRelationship(rel Relationship) error
- func (kg *KnowledgeGraph) DeleteEntity(entityID string) error
- func (kg *KnowledgeGraph) DeleteRelationship(relID string) error
- func (kg *KnowledgeGraph) GetEntity(entityID string) (*Entity, error)
- func (kg *KnowledgeGraph) GetRelationship(relID string) (*Relationship, error)
- func (kg *KnowledgeGraph) GetRelationships(entityID string, direction string) []*Relationship
- func (kg *KnowledgeGraph) ListEntities(entityType string) []*Entity
- func (kg *KnowledgeGraph) Name() string
- func (kg *KnowledgeGraph) SearchWithExpansion(query []float32, traversalConfig TraversalConfig, opts ...SearchOption) ([]ExpandedSearchResult, error)
- func (kg *KnowledgeGraph) Stats() KnowledgeGraphStats
- func (kg *KnowledgeGraph) Traverse(startIDs []string, config TraversalConfig) (*TraversalResult, error)
- func (kg *KnowledgeGraph) UpdateEntity(entity Entity) error
- type KnowledgeGraphStats
- type Logger
- type MatchEvent
- type MemoryCluster
- type MemoryConfig
- type MemoryLimiter
- type Metrics
- type MetricsSnapshot
- type NopLogger
- type NotFoundError
- type ONNXConfig
- type OllamaConfig
- type OpenAIConfig
- type Option
- type ProfiledEmbedder
- type Record
- type RecordInput
- type RecordSnapshot
- type Relationship
- type RelationshipSnapshot
- type Result
- type SearchExplanation
- type SearchFunc
- type SearchOption
- func Threshold(t float32) SearchOption
- func TopK(k int) SearchOption
- func WithAccessTracking(enabled bool) SearchOption
- func WithContent(include bool) SearchOption
- func WithDecay(decayType DecayType, halfLife time.Duration) SearchOption
- func WithEfSearch(ef int) SearchOption
- func WithFilter(f Filter) SearchOption
- func WithFilters(filters ...Filter) SearchOption
- func WithImportanceBoost(factor float32) SearchOption
- func WithLimit(n int) SearchOption
- func WithOffset(n int) SearchOption
- func WithTextWeight(w float64) SearchOption
- func WithVectorWeight(w float64) SearchOption
- type SessionStats
- type Storage
- type StorageError
- type Subscription
- type SubscriptionOption
- type TFEntry
- type TTLCleaner
- type TTLCleanerCallback
- type TimeRange
- type TimeRangeSnapshot
- type TraversalConfig
- type TraversalResult
- type VectorSpaceConfig
- type VectorSpaceInfo
Constants ¶
const ( // PayloadKeyArchived indicates if a record has been archived. PayloadKeyArchived = "_archived" // PayloadKeyConsolidationGroup is the ID of the consolidation group. PayloadKeyConsolidationGroup = "_consolidation_group" // PayloadKeyConsolidatedFrom contains IDs of records this was consolidated from. PayloadKeyConsolidatedFrom = "_consolidated_from" // PayloadKeyIsConsolidation indicates this record is a consolidation of others. PayloadKeyIsConsolidation = "_is_consolidation" )
Reserved payload keys for memory consolidation.
const ( // PayloadKeySessionID identifies which session/conversation a record belongs to. PayloadKeySessionID = "_session_id" // PayloadKeyTurnNumber is the sequential turn number within a session. PayloadKeyTurnNumber = "_turn_number" // PayloadKeyRole indicates the role (e.g., "user", "assistant", "system"). PayloadKeyRole = "_role" // PayloadKeyParentChunk links to the parent chunk ID for threaded conversations. PayloadKeyParentChunk = "_parent_chunk" // PayloadKeyChildChunks contains IDs of child chunks. PayloadKeyChildChunks = "_child_chunks" // PayloadKeyThreadRoot is the ID of the root chunk in a thread. PayloadKeyThreadRoot = "_thread_root" )
Reserved payload keys for conversation tracking.
const ( // DistanceCosine uses cosine similarity (higher = more similar). DistanceCosine = floats.DistanceCosine // DistanceDot uses dot product (higher = more similar). DistanceDot = floats.DistanceDot // DistanceEuclidean uses Euclidean distance (lower = more similar). DistanceEuclidean = floats.DistanceEuclidean // DistanceEuclideanSquared uses squared Euclidean distance (lower = more similar). // Faster than Euclidean since it avoids sqrt. DistanceEuclideanSquared = floats.DistanceEuclideanSquared )
const DefaultVectorSpace = "default"
DefaultVectorSpace is the reserved name of the implicit vector space backed by Record.Vector and the collection's primary dimension/distance/index. Every collection has this space, including those created before named vector spaces existed. It cannot be removed or redeclared with AddVectorSpace.
const DefaultWALCheckpointBytes int64 = 64 << 20 // 64 MiB
DefaultWALCheckpointBytes is the WAL size at which a WAL-enabled database automatically folds the log into a fresh snapshot (see WithWALCheckpoint).
const Version = "0.24.0"
Version is the library version.
Variables ¶
var ( // ErrNotFound is returned when a record or collection is not found. ErrNotFound = errors.New("veclite: not found") // ErrDimensionMismatch is returned when vector dimensions don't match. ErrDimensionMismatch = errors.New("veclite: dimension mismatch") // ErrEmptyVector is returned when an empty vector is provided. ErrEmptyVector = errors.New("veclite: empty vector") // ErrCollectionExists is returned when trying to create a collection that already exists. ErrCollectionExists = errors.New("veclite: collection already exists") // ErrDatabaseClosed is returned when operations are attempted on a closed database. ErrDatabaseClosed = errors.New("veclite: database closed") // ErrInvalidPath is returned when an invalid file path is provided. ErrInvalidPath = errors.New("veclite: invalid path") // ErrBatchSizeMismatch is returned when batch operation input sizes don't match. ErrBatchSizeMismatch = errors.New("veclite: batch size mismatch") // ErrReadOnly is returned when a write operation is attempted on a read-only database. ErrReadOnly = errors.New("veclite: database is read-only") // without WithReadOnly. A shared file lock is only safe for read-only access. ErrSharedReadRequiresReadOnly = errors.New("veclite: shared read requires read-only mode") // ErrVectorSpaceExists is returned when declaring a vector space whose name // is already in use (including the reserved "default" space). ErrVectorSpaceExists = errors.New("veclite: vector space already exists") // ErrVectorSpaceNotFound is returned when referencing an undeclared vector space. ErrVectorSpaceNotFound = errors.New("veclite: vector space not found") // ErrInvalidVectorSpace is returned when a vector-space declaration is invalid, // e.g. an empty name or an attempt to redeclare the reserved "default" space. ErrInvalidVectorSpace = errors.New("veclite: invalid vector space") // ErrProfileMismatch is returned when two embedding profiles are incompatible, // or when a vector does not match a declared embedding profile. ErrProfileMismatch = errors.New("veclite: embedding profile mismatch") )
Sentinel errors for common conditions.
var ( // ErrFileLocked is returned when the database file is locked by another process. ErrFileLocked = storage.ErrFileLocked // ErrChecksumMismatch is returned when the file checksum does not match. ErrChecksumMismatch = storage.ErrChecksumMismatch // ErrCorruptedFile is returned when the database file is corrupted. ErrCorruptedFile = storage.ErrCorruptedFile // ErrInvalidVersion is returned when the file version is not supported. ErrInvalidVersion = storage.ErrInvalidVersion )
Storage-level sentinel errors re-exported for consumer use.
var ErrNoEmbedder = errors.New("veclite: no embedder configured")
ErrNoEmbedder is returned when an embedding operation is attempted without an embedder configured on the collection.
Functions ¶
func ExpandEnvVars ¶ added in v0.18.0
ExpandEnvVars expands ${VAR} and ${VAR:-default} patterns in a string.
func ExpandPath ¶ added in v0.11.0
ExpandPath expands ~ and environment variables in a path.
Types ¶
type Collection ¶
type Collection struct {
// contains filtered or unexported fields
}
Collection represents a collection of vectors with the same dimension.
func (*Collection) AddVectorSpace ¶ added in v0.16.0
func (c *Collection) AddVectorSpace(config VectorSpaceConfig) error
AddVectorSpace declares an additional named vector space on the collection.
The space is independent of the default space and of other spaces: it has its own dimension, distance metric, and optional HNSW index. After declaring it, insert vectors into it with InsertRecord (or SetRecordVector) and query it with SearchSpace. The reserved name DefaultVectorSpace cannot be used.
func (*Collection) All ¶
func (c *Collection) All() []*Record
All returns all records in the collection.
func (*Collection) ArchiveRecord ¶ added in v0.8.0
func (c *Collection) ArchiveRecord(id uint64) error
ArchiveRecord marks a record as archived. Archived records are excluded from normal searches but can be retrieved with GetArchived.
func (*Collection) CleanupExpired ¶ added in v0.8.0
func (c *Collection) CleanupExpired() (int, error)
CleanupExpired removes all expired records from the collection. Returns the number of records removed.
func (*Collection) Clear ¶
func (c *Collection) Clear() error
Clear removes all records from the collection. It preserves the nextID counter to avoid ID reuse after reinsertion. Use Reset if you want to also reset the ID counter. Returns an error if the database is read-only.
func (*Collection) Consolidate ¶ added in v0.8.0
func (c *Collection) Consolidate(config ConsolidationConfig) (*ConsolidationResult, error)
Consolidate finds similar memory clusters and optionally creates consolidated records.
func (*Collection) Count ¶
func (c *Collection) Count() int
Count returns the number of records in the collection.
func (*Collection) CountExpired ¶ added in v0.8.0
func (c *Collection) CountExpired() int
CountExpired returns the number of expired records in the collection.
func (*Collection) Delete ¶
func (c *Collection) Delete(id uint64) error
Delete removes a record by ID.
func (*Collection) DeleteMetadataValue ¶ added in v0.15.0
func (c *Collection) DeleteMetadataValue(key string) error
DeleteMetadataValue removes one collection metadata value.
func (*Collection) DeleteWhere ¶
func (c *Collection) DeleteWhere(filters ...Filter) (int, error)
DeleteWhere removes all records matching the filters. Returns the number of deleted records.
func (*Collection) Dimension ¶
func (c *Collection) Dimension() int
Dimension returns the vector dimension. Returns 0 if no vectors have been inserted yet.
func (*Collection) DistanceType ¶
func (c *Collection) DistanceType() floats.DistanceType
DistanceType returns the distance metric type.
func (*Collection) EmbeddingProfile ¶ added in v0.16.0
func (c *Collection) EmbeddingProfile() (EmbeddingProfile, bool)
EmbeddingProfile returns the collection's default-space embedding profile and whether one is set.
func (*Collection) EnforceMemoryLimit ¶ added in v0.9.0
func (c *Collection) EnforceMemoryLimit(config MemoryConfig) int
EnforceMemoryLimit checks and enforces the memory limit for a collection. Call this after inserts if you want immediate enforcement without background monitoring.
func (*Collection) ExpandConsolidation ¶ added in v0.8.0
func (c *Collection) ExpandConsolidation(consolidationID uint64) ([]*Record, error)
ExpandConsolidation retrieves the original records that were consolidated into a consolidation record.
func (*Collection) Find ¶
func (c *Collection) Find(filters ...Filter) ([]*Record, error)
Find retrieves all records matching the filters.
func (*Collection) FindOne ¶
func (c *Collection) FindOne(filters ...Filter) (*Record, error)
FindOne retrieves the first record matching the filters.
func (*Collection) FindSimilarClusters ¶ added in v0.8.0
func (c *Collection) FindSimilarClusters(config ConsolidationConfig) ([]MemoryCluster, error)
FindSimilarClusters identifies clusters of similar memories using single-linkage clustering.
func (*Collection) ForEach ¶ added in v0.6.0
func (c *Collection) ForEach(fn func(*Record) bool)
ForEach iterates over all records in the collection, calling fn for each. If fn returns false, iteration stops early. Records are cloned before being passed to fn.
func (*Collection) Get ¶
func (c *Collection) Get(id uint64) (*Record, error)
Get retrieves a record by ID.
func (*Collection) GetArchived ¶ added in v0.8.0
func (c *Collection) GetArchived() ([]*Record, error)
GetArchived retrieves all archived records.
func (*Collection) GetConsolidations ¶ added in v0.8.0
func (c *Collection) GetConsolidations() ([]*Record, error)
GetConsolidations retrieves all consolidation records.
func (*Collection) GetSession ¶ added in v0.8.0
func (c *Collection) GetSession(sessionID string) ([]*Record, error)
GetSession retrieves all records belonging to a session. Returns records in turn number order.
func (*Collection) GetSessionStats ¶ added in v0.8.0
func (c *Collection) GetSessionStats(sessionID string) (SessionStats, error)
GetSessionStats returns statistics about a session.
func (*Collection) GetThread ¶ added in v0.8.0
func (c *Collection) GetThread(chunkID uint64) ([]*Record, error)
GetThread retrieves all records in a thread starting from the given chunk ID. Returns records in chronological order.
func (*Collection) GetVector ¶
func (c *Collection) GetVector(id uint64) ([]float32, error)
GetVector retrieves just the vector for a record.
func (*Collection) HasIndex ¶ added in v0.2.0
func (c *Collection) HasIndex() bool
HasIndex returns true if this collection has an index.
func (*Collection) HasVectorSpace ¶ added in v0.16.0
func (c *Collection) HasVectorSpace(name string) bool
HasVectorSpace reports whether the collection has the named space. The default space always exists.
func (*Collection) HybridSearch ¶ added in v0.6.0
func (c *Collection) HybridSearch(query []float32, text string, opts ...SearchOption) ([]Result, error)
HybridSearch performs both vector search and BM25 text search, then fuses results using Reciprocal Rank Fusion (RRF) with k=60. Requires text indexing to be enabled via WithTextIndex. Use WithVectorWeight and WithTextWeight to control the balance.
func (*Collection) HybridSearchSpace ¶ added in v0.17.0
func (c *Collection) HybridSearchSpace(space string, query []float32, text string, opts ...SearchOption) ([]Result, error)
HybridSearchSpace performs vector search over a named vector space and BM25 text search over the collection, then fuses the two result sets with Reciprocal Rank Fusion (k=60). It is the named-space analog of HybridSearch: use it when the query vector lives in a named space declared via AddVectorSpace rather than the default space.
Passing DefaultVectorSpace (or "") is equivalent to HybridSearch. Requires text indexing to be enabled via WithTextIndex. Use WithVectorWeight and WithTextWeight to control the fusion balance.
func (*Collection) IndexStats ¶ added in v0.2.0
func (c *Collection) IndexStats() *HNSWStats
IndexStats returns statistics about the collection's HNSW index. Returns nil if no HNSW index is configured.
func (*Collection) IndexType ¶ added in v0.2.0
func (c *Collection) IndexType() IndexType
IndexType returns the index type for this collection.
func (*Collection) Insert ¶
Insert adds a vector with optional payload to the collection. Returns the assigned record ID.
func (*Collection) InsertBatch ¶
InsertBatch adds multiple vectors with payloads to the collection. Returns the assigned record IDs. If payloads is nil or shorter than vectors, missing payloads are treated as nil.
func (*Collection) InsertDocument ¶ added in v0.6.0
func (c *Collection) InsertDocument(vector []float32, content string, payload map[string]any) (uint64, error)
InsertDocument inserts a vector with content text and payload. Content is automatically indexed for BM25 text search when text indexing is enabled.
func (*Collection) InsertRecord ¶ added in v0.16.0
func (c *Collection) InsertRecord(in RecordInput) (uint64, error)
InsertRecord inserts (or, when in.ID names an existing record, replaces) one logical record that may carry vectors in several named vector spaces at once, plus optional content and payload. Returns the record ID.
Each key of in.Vectors must be DefaultVectorSpace (or "") or a space declared via AddVectorSpace; unknown spaces return ErrVectorSpaceNotFound. Vectors are validated against each space's dimension and embedding profile before any state changes.
func (*Collection) InsertText ¶ added in v0.6.0
InsertText embeds the text using the configured embedder and inserts the result. Requires an embedder to be set via WithEmbedder.
func (*Collection) InsertTextDocument ¶ added in v0.15.0
InsertTextDocument inserts content and payload without a vector. Text-only records are indexed by BM25 when text indexing is enabled and are skipped by vector search, hybrid vector search, and vector subscriptions.
func (*Collection) InsertTextDocumentWithOptions ¶ added in v0.15.0
func (c *Collection) InsertTextDocumentWithOptions(content string, payload map[string]any, opts ...InsertOption) (uint64, error)
InsertTextDocumentWithOptions inserts content and payload without a vector, applying insert options such as TTL and importance.
func (*Collection) InsertTurn ¶ added in v0.8.0
func (c *Collection) InsertTurn(turn ConversationTurn) (uint64, error)
InsertTurn inserts a conversation turn with conversation metadata. Returns the record ID.
func (*Collection) InsertWithOptions ¶ added in v0.8.0
func (c *Collection) InsertWithOptions(vector []float32, payload map[string]any, opts ...InsertOption) (uint64, error)
InsertWithOptions adds a vector with optional payload and insert options. Use this method to set TTL, importance, and other options.
func (*Collection) Iterate ¶ added in v0.6.0
func (c *Collection) Iterate(opts ...IterOption) *Iterator
Iterate returns an iterator over collection records. Options can control offset and limit for pagination.
func (*Collection) ListSessions ¶ added in v0.8.0
func (c *Collection) ListSessions() []string
ListSessions returns all unique session IDs in the collection.
func (*Collection) Metadata ¶ added in v0.15.0
func (c *Collection) Metadata() map[string]any
Metadata returns a deep copy of the collection metadata.
func (*Collection) MultiSpaceSearch ¶ added in v0.16.0
func (c *Collection) MultiSpaceSearch(queries map[string][]float32, opts ...SearchOption) ([]Result, error)
MultiSpaceSearch runs one query per named vector space and fuses the result sets into a single ranking with Reciprocal Rank Fusion. Keys of queries are vector-space names (DefaultVectorSpace or "" targets the default space).
This is the multimodal entry point: e.g. fuse a "text" query and an "image" query for the same item. For weighted fusion or to also fold in BM25 text results, call SearchSpace / TextSearch and combine the sets with FuseRRF.
func (*Collection) Reset ¶ added in v0.14.0
func (c *Collection) Reset() error
Reset removes all records from the collection and resets the ID counter to 1. Unlike Clear, this allows ID reuse which may be desirable for testing or when the collection is being fully repopulated. Returns an error if the database is read-only.
func (*Collection) Search ¶
func (c *Collection) Search(query []float32, opts ...SearchOption) ([]Result, error)
Search finds the most similar vectors to the query vector.
func (*Collection) SearchExplain ¶ added in v0.2.0
func (c *Collection) SearchExplain(query []float32, opts ...SearchOption) (*SearchExplanation, error)
SearchExplain performs a search and returns detailed statistics.
func (*Collection) SearchInSession ¶ added in v0.8.0
func (c *Collection) SearchInSession(sessionID string, query []float32, opts ...SearchOption) ([]Result, error)
SearchInSession searches for similar vectors within a specific session.
func (*Collection) SearchSpace ¶ added in v0.16.0
func (c *Collection) SearchSpace(space string, query []float32, opts ...SearchOption) ([]Result, error)
SearchSpace searches a single named vector space. Passing DefaultVectorSpace (or "") is equivalent to Search. All standard SearchOptions apply (TopK, filters, threshold, pagination, efSearch).
func (*Collection) SearchStream ¶ added in v0.6.0
func (c *Collection) SearchStream(query []float32, fn SearchFunc, opts ...SearchOption) error
SearchStream performs a search and streams results to the callback function. For brute-force searches, results are yielded as they are found, enabling early termination without computing all scores. For HNSW searches, the full result set is fetched first since the index returns ordered results.
func (*Collection) SearchText ¶ added in v0.6.0
func (c *Collection) SearchText(text string, opts ...SearchOption) ([]Result, error)
SearchText embeds the text query using the configured embedder and searches. Requires an embedder to be set via WithEmbedder.
func (*Collection) SetEmbeddingProfile ¶ added in v0.16.0
func (c *Collection) SetEmbeddingProfile(profile EmbeddingProfile) error
SetEmbeddingProfile attaches (or replaces) the collection's first-class default-space embedding profile. Subsequent default-space inserts validate against the profile's dimension.
func (*Collection) SetMetadata ¶ added in v0.15.0
func (c *Collection) SetMetadata(metadata map[string]any) error
SetMetadata replaces the collection metadata.
func (*Collection) SetMetadataValue ¶ added in v0.15.0
func (c *Collection) SetMetadataValue(key string, value any) error
SetMetadataValue sets one collection metadata value.
func (*Collection) SetRecordVector ¶ added in v0.16.0
func (c *Collection) SetRecordVector(id uint64, space string, vector []float32) error
SetRecordVector sets or replaces a single space's vector on an existing record, leaving the record's other spaces, content, and payload untouched. Use DefaultVectorSpace (or "") to target the default space.
func (*Collection) StartMemoryLimiter ¶ added in v0.9.0
func (c *Collection) StartMemoryLimiter(config MemoryConfig) func()
StartMemoryLimiter starts background memory pressure monitoring for a collection. Returns a function to stop the limiter.
func (*Collection) Stats ¶
func (c *Collection) Stats() CollectionStats
Stats returns statistics about the collection.
func (*Collection) Subscribe ¶ added in v0.8.0
func (c *Collection) Subscribe(query []float32, opts ...SubscriptionOption) (*Subscription, error)
Subscribe creates a subscription that receives events when new records match the query. Returns the subscription which can be used to receive events and close the subscription.
func (*Collection) TextSearch ¶ added in v0.6.0
func (c *Collection) TextSearch(query string, opts ...SearchOption) ([]Result, error)
TextSearch performs BM25 full-text search over indexed fields. Requires text indexing to be enabled via WithTextIndex.
func (*Collection) UnarchiveRecord ¶ added in v0.8.0
func (c *Collection) UnarchiveRecord(id uint64) error
UnarchiveRecord removes the archived flag from a record.
func (*Collection) Unsubscribe ¶ added in v0.8.0
func (c *Collection) Unsubscribe(subscriptionID string) error
Unsubscribe removes a subscription by ID.
func (*Collection) Update ¶
func (c *Collection) Update(id uint64, payload map[string]any) error
Update updates the payload for a record.
func (*Collection) UpdateDocument ¶ added in v0.15.0
UpdateDocument updates the content and payload for a record without changing its vector.
func (*Collection) UpdateVector ¶ added in v0.4.0
func (c *Collection) UpdateVector(id uint64, vector []float32) error
UpdateVector updates the vector for a record.
func (*Collection) Upsert ¶ added in v0.4.0
Upsert inserts a new record or updates an existing one by ID. If the ID is 0, a new record is created with an auto-generated ID. If the ID exists, the vector and payload are updated. If the ID doesn't exist, a new record is created with that ID. Returns the record ID (either the provided one or newly generated).
func (*Collection) UpsertByKey ¶ added in v0.4.0
func (c *Collection) UpsertByKey(keyField string, keyValue any, vector []float32, payload map[string]any) (uint64, bool, error)
UpsertByKey inserts a new record or updates an existing one based on a key field. If a record with payload[keyField] == keyValue exists, it is updated. Otherwise, a new record is inserted. Returns the record ID and whether it was an insert (true) or update (false).
func (*Collection) UpsertRecordByKey ¶ added in v0.17.0
func (c *Collection) UpsertRecordByKey(keyField string, keyValue any, in RecordInput) (uint64, bool, error)
UpsertRecordByKey inserts or replaces the record identified by a payload key, carrying vectors across several named vector spaces atomically. It is the named-vector-space analog of UpsertTextDocumentByKey: it scans for an existing record whose payload[keyField] equals keyValue and, if found, replaces that record's content, payload, and vectors with in's; otherwise it inserts a new record.
On replace, the existing record's CreatedAt, ExpiresAt, Importance, AccessCount, and LastAccessedAt are preserved (matching UpsertTextDocumentByKey semantics). Vectors for every space the record previously participated in are removed from their indexes before the new vectors are inserted, so a space that the replacement no longer carries is correctly dropped.
Returns (id, inserted, error) where inserted is true when a new record was created and false when an existing record was replaced.
keyField must be a non-empty payload key. keyValue is compared with the same type-aware rules as payload filters (see Equal). Each key of in.Vectors must be DefaultVectorSpace (or "") or a space declared via AddVectorSpace.
func (*Collection) UpsertTextDocument ¶ added in v0.15.0
func (c *Collection) UpsertTextDocument(id uint64, content string, payload map[string]any) (uint64, error)
UpsertTextDocument inserts or updates a text-only document by ID. If id is 0, a new record is created with an auto-generated ID.
func (*Collection) UpsertTextDocumentByKey ¶ added in v0.15.0
func (c *Collection) UpsertTextDocumentByKey(keyField string, keyValue any, content string, payload map[string]any) (uint64, bool, error)
UpsertTextDocumentByKey inserts or updates a text-only document by payload key.
func (*Collection) VectorSpace ¶ added in v0.16.0
func (c *Collection) VectorSpace(name string) (VectorSpaceInfo, error)
VectorSpace returns information about a single vector space, or ErrVectorSpaceNotFound.
func (*Collection) VectorSpaces ¶ added in v0.16.0
func (c *Collection) VectorSpaces() []VectorSpaceInfo
VectorSpaces returns information about every vector space on the collection, starting with the default space, followed by named spaces in name order.
type CollectionOption ¶
type CollectionOption interface {
// contains filtered or unexported methods
}
CollectionOption configures a collection.
func WithDimension ¶
func WithDimension(dim int) CollectionOption
WithDimension sets the vector dimension for the collection. If set, all vectors must match this dimension. If not set (0), the dimension is determined by the first insert.
func WithDistanceType ¶
func WithDistanceType(t floats.DistanceType) CollectionOption
WithDistanceType sets the distance metric for the collection. Default is cosine similarity.
func WithEmbedder ¶ added in v0.6.0
func WithEmbedder(e Embedder) CollectionOption
WithEmbedder sets an auto-embedding plugin for the collection. When set, InsertText and SearchText methods become available.
func WithEmbeddingProfile ¶ added in v0.16.0
func WithEmbeddingProfile(profile EmbeddingProfile) CollectionOption
WithEmbeddingProfile attaches a first-class embedding profile to the collection's default vector space. Vectors inserted into the default space are then validated against the profile's declared dimension.
func WithHNSW ¶ added in v0.2.0
func WithHNSW(m, efConstruction int) CollectionOption
WithHNSW enables HNSW indexing for the collection. m is the maximum number of connections per node (default: 16, recommended: 12-48). efConstruction is the candidate list size during construction (default: 200).
func WithHNSWConfig ¶ added in v0.2.0
func WithHNSWConfig(config HNSWConfig) CollectionOption
WithHNSWConfig enables HNSW indexing with custom configuration.
func WithMemoryLimits ¶ added in v0.9.0
func WithMemoryLimits(config MemoryConfig) CollectionOption
WithMemoryLimits returns a collection option that enables memory pressure handling. When the collection exceeds MaxRecords, older/less important records are evicted.
func WithTextIndex ¶ added in v0.6.0
func WithTextIndex(fields ...string) CollectionOption
WithTextIndex enables BM25 full-text indexing on the specified payload fields. When enabled, string values in these fields are tokenized and indexed for text search. The Content field of records is always indexed when text indexing is enabled.
func WithVectorSpace ¶ added in v0.16.0
func WithVectorSpace(config VectorSpaceConfig) CollectionOption
WithVectorSpace declares an additional named vector space at collection creation time. May be passed multiple times to declare several spaces. The reserved DefaultVectorSpace name is rejected (the default space always exists). Equivalent to calling Collection.AddVectorSpace after creation.
type CollectionSnapshot ¶
type CollectionSnapshot = storage.CollectionSnapshot
CollectionSnapshot is the serializable state of a collection.
func NewCollectionSnapshot ¶
func NewCollectionSnapshot(name string, dimension int, distanceType floats.DistanceType) *CollectionSnapshot
NewCollectionSnapshot creates a new empty collection snapshot.
type CollectionStats ¶
type CollectionStats struct {
// Name is the collection name.
Name string
// Count is the number of records in the collection.
Count int
// VectorCount is the number of records with a vector.
VectorCount int
// TextOnlyCount is the number of records without a vector.
TextOnlyCount int
// Dimension is the vector dimension (0 if not yet set).
Dimension int
// DistanceType is the distance metric used.
DistanceType string
// IndexType is the index type (none, hnsw).
IndexType string
}
CollectionStats contains statistics about a collection.
type Config ¶ added in v0.11.0
type Config struct {
Embedder EmbedderConfig `yaml:"embedder"`
}
Config represents the full veclite configuration.
func DefaultConfig ¶ added in v0.11.0
func DefaultConfig() *Config
DefaultConfig returns a configuration with sensible defaults.
func LoadConfig ¶ added in v0.11.0
LoadConfig loads configuration from a YAML file. It searches for config files in the following order: 1. The explicit path provided (if not empty) 2. ./veclite.yaml 3. ~/.veclite/config.yaml If no config file is found, returns default configuration.
type ConsolidationConfig ¶ added in v0.8.0
type ConsolidationConfig struct {
// SimilarityThreshold is the minimum similarity for records to be grouped (0.0-1.0).
// Higher values mean stricter grouping.
SimilarityThreshold float32
// MinGroupSize is the minimum number of records to form a cluster.
MinGroupSize int
// MaxGroupSize is the maximum number of records in a cluster.
MaxGroupSize int
// SummaryGenerator is a function that creates a summary from a group of records.
// It returns the summary text, additional payload, and any error.
// If nil, no consolidation records are created (only clustering is performed).
SummaryGenerator func([]*Record) (string, map[string]any, error)
// Embedder is used to generate embeddings for consolidated summaries.
// Required if SummaryGenerator is provided.
Embedder Embedder
// ArchiveOriginals determines whether to archive original records after consolidation.
ArchiveOriginals bool
// Filters can be used to limit which records are considered for consolidation.
Filters []Filter
}
ConsolidationConfig configures memory consolidation.
type ConsolidationResult ¶ added in v0.8.0
type ConsolidationResult struct {
// ClustersFound is the number of clusters identified.
ClustersFound int
// RecordsConsolidated is the total number of records that were consolidated.
RecordsConsolidated int
// ConsolidatedRecordIDs are the IDs of newly created consolidation records.
ConsolidatedRecordIDs []uint64
// ArchivedRecordIDs are the IDs of records that were archived.
ArchivedRecordIDs []uint64
// Clusters contains details about each cluster found.
Clusters []MemoryCluster
}
ConsolidationResult contains the results of a consolidation operation.
type ConversationTurn ¶ added in v0.8.0
type ConversationTurn struct {
// SessionID identifies the conversation session.
SessionID string
// TurnNumber is the sequential turn number (1-indexed).
TurnNumber int
// Role is the speaker role (e.g., "user", "assistant").
Role string
// Content is the text content of the turn.
Content string
// Vector is the embedding vector. If nil, the collection's embedder is used.
Vector []float32
// ParentChunkID links to a parent chunk for threaded conversations.
ParentChunkID uint64
// Payload contains additional metadata.
Payload map[string]any
// Importance is the importance score (0.0-1.0).
Importance float32
// TTL is the time-to-live for this turn.
TTL time.Duration
}
ConversationTurn represents a single turn in a conversation.
type DB ¶
type DB struct {
// contains filtered or unexported fields
}
DB represents a VecLite database.
func Open ¶
Open opens or creates a VecLite database at the given path. Use ":memory:" for an in-memory database that won't be persisted.
func (*DB) Collection ¶
func (db *DB) Collection(name string) *Collection
Collection returns a collection by name, creating it if it doesn't exist. This is the preferred way to get collections for most use cases. In read-only mode, returns nil if the collection doesn't exist (cannot create).
func (*DB) Collections ¶
Collections returns the names of all collections.
func (*DB) CreateCollection ¶
func (db *DB) CreateCollection(name string, opts ...CollectionOption) (*Collection, error)
CreateCollection creates a new collection with the given options. Returns an error if the collection already exists.
func (*DB) CreateEpisodeStore ¶ added in v0.8.0
func (db *DB) CreateEpisodeStore(memoriesCollectionName string) (*EpisodeStore, error)
CreateEpisodeStore creates a new episode store for a collection.
func (*DB) CreateKnowledgeGraph ¶ added in v0.8.0
func (db *DB) CreateKnowledgeGraph(name string) (*KnowledgeGraph, error)
CreateKnowledgeGraph creates a new knowledge graph.
func (*DB) DeleteMetadataValue ¶ added in v0.15.0
DeleteMetadataValue removes one database metadata value.
func (*DB) DropCollection ¶
DropCollection removes a collection and all its data.
func (*DB) GetCollection ¶
func (db *DB) GetCollection(name string) (*Collection, error)
GetCollection returns an existing collection or ErrNotFound.
func (*DB) GetEpisodeStore ¶ added in v0.14.0
func (db *DB) GetEpisodeStore(name string) (*EpisodeStore, error)
func (*DB) GetKnowledgeGraph ¶ added in v0.14.0
func (db *DB) GetKnowledgeGraph(name string) (*KnowledgeGraph, error)
func (*DB) HasCollection ¶
HasCollection returns true if a collection exists.
func (*DB) Metrics ¶ added in v0.6.0
func (db *DB) Metrics() MetricsSnapshot
Metrics returns the current metrics snapshot.
func (*DB) Reload ¶ added in v0.20.0
Reload re-reads the database from storage, rebuilding all in-memory state (collections, HNSW indexes, BM25 inverted indexes, knowledge graphs, episode stores). It is intended for read-only databases opened with WithSharedRead so they can pick up writes performed by another process without closing and reopening.
Reload acquires the DB write lock for the duration of the rebuild. It is not safe to call concurrently with reads or writes on the same DB instance.
Background workers (TTL cleaner, memory limiter) are stopped before reload because they reference the old collection structs that are replaced during reload; callers that need those workers must restart them after Reload returns.
Reload returns an error if the database is closed, if the storage backend does not support reloading (e.g. in-memory), or if the reloaded snapshot is corrupt. On error, the in-memory state is left unchanged.
func (*DB) SetMetadata ¶ added in v0.15.0
SetMetadata replaces the database metadata.
func (*DB) SetMetadataValue ¶ added in v0.15.0
SetMetadataValue sets one database metadata value.
func (*DB) StartTTLCleaner ¶ added in v0.9.0
func (db *DB) StartTTLCleaner(interval time.Duration, callback TTLCleanerCallback) func()
StartTTLCleaner starts a background goroutine that periodically cleans up expired records from all collections in the database. Returns a function to stop the cleaner.
type DatabaseSnapshot ¶
type DatabaseSnapshot = storage.DatabaseSnapshot
DatabaseSnapshot is the serializable state of the database.
func NewDatabaseSnapshot ¶
func NewDatabaseSnapshot() *DatabaseSnapshot
NewDatabaseSnapshot creates a new empty database snapshot.
type DatabaseStats ¶
type DatabaseStats struct {
// Path is the database file path (":memory:" for in-memory).
Path string
// Collections is the number of collections.
Collections int
// TotalRecords is the total number of records across all collections.
TotalRecords int
// CollectionStats contains stats for each collection.
CollectionStats []CollectionStats
}
DatabaseStats contains statistics about the database.
type DecayConfig ¶ added in v0.8.0
type DecayConfig struct {
// Type is the decay function type.
Type DecayType
// HalfLife is the time after which the score is halved (for exponential decay).
// For linear decay, this is the maximum age after which the score is zero.
// For gaussian decay, this is the sigma parameter.
HalfLife time.Duration
}
DecayConfig holds the configuration for temporal decay.
type DecayType ¶ added in v0.8.0
type DecayType string
DecayType specifies the type of temporal decay function to apply.
const ( // DecayNone applies no temporal decay. DecayNone DecayType = "none" // DecayExponential applies exponential decay: score * 2^(-age/halfLife) DecayExponential DecayType = "exponential" // DecayLinear applies linear decay: score * max(0, 1 - age/maxAge) DecayLinear DecayType = "linear" // DecayGaussian applies Gaussian decay: score * exp(-0.5 * (age/sigma)^2) DecayGaussian DecayType = "gaussian" )
type DimensionError ¶
DimensionError provides details about dimension mismatches.
func (*DimensionError) Error ¶
func (e *DimensionError) Error() string
func (*DimensionError) Unwrap ¶
func (e *DimensionError) Unwrap() error
type DistanceType ¶ added in v0.2.0
type DistanceType = floats.DistanceType
Re-export distance types for external use.
type Embedder ¶ added in v0.6.0
type Embedder interface {
// Embed converts a single text into a vector embedding.
Embed(text string) ([]float32, error)
// EmbedBatch converts multiple texts into vector embeddings.
EmbedBatch(texts []string) ([][]float32, error)
// Dimension returns the output vector dimension.
Dimension() int
}
Embedder is the interface for auto-embedding text to vectors. Implementations live in separate modules to maintain zero-dependency core.
func NewEmbedderFromConfig ¶ added in v0.11.0
func NewEmbedderFromConfig(cfg EmbedderConfig) (Embedder, error)
NewEmbedderFromConfig creates an embedder based on the provided configuration. It returns an Embedder that can be used with veclite collections.
Supported providers:
- "openai": OpenAI embedding API (requires API key)
- "ollama": Ollama local embedding (requires Ollama running)
- "onnx": Local ONNX inference (requires onnx build tag)
Example:
cfg, _ := veclite.LoadConfig("veclite.yaml")
embedder, _ := veclite.NewEmbedderFromConfig(cfg.Embedder)
defer embedder.Close()
db, _ := veclite.Open("data.veclite")
coll, _ := db.CreateCollection("docs",
veclite.WithDimension(embedder.Dimension()),
veclite.WithEmbedder(embedder),
)
type EmbedderConfig ¶ added in v0.11.0
type EmbedderConfig struct {
Provider string `yaml:"provider"`
OpenAI OpenAIConfig `yaml:"openai"`
Ollama OllamaConfig `yaml:"ollama"`
ONNX ONNXConfig `yaml:"onnx"`
}
EmbedderConfig specifies which embedder provider to use and its settings.
type EmbeddingProfile ¶ added in v0.16.0
type EmbeddingProfile struct {
// Provider is the embedding provider, e.g. "openai" or "ollama".
Provider string
// Model is the embedding model identifier.
Model string
// Dimension is the expected vector dimension. 0 means "unspecified".
Dimension int
// Distance is the distance metric the embeddings are intended for.
Distance DistanceType
// Normalize records whether vectors are L2-normalized.
Normalize bool
// Version is an optional app-defined revision for the embedding pipeline.
Version string
}
EmbeddingProfile is a first-class description of how an embedding was produced.
It is the typed replacement for storing provider/model details loosely in metadata. Attaching a profile to a collection or vector space lets VecLite reject vectors that do not match the declared dimension (and, between two profiles, detect when a model/provider/distance change invalidates an index).
func EmbedderProfile ¶ added in v0.23.0
func EmbedderProfile(e Embedder) (EmbeddingProfile, bool)
EmbedderProfile extracts an EmbeddingProfile from an embedder when it self-describes. It supports both ProfiledEmbedder implementations and the built-in providers under embed/ (which return common.ProfileData). The second return value reports whether the embedder provided a profile.
func (EmbeddingProfile) Compatible ¶ added in v0.16.0
func (p EmbeddingProfile) Compatible(other EmbeddingProfile) error
Compatible reports whether two embedding profiles describe interchangeable vectors. It returns nil when compatible, or an error describing the first mismatch. Provider, Model, Dimension (when both set), Distance (when both set), and Normalize must agree. Version is advisory and never causes an error.
func (EmbeddingProfile) IsZero ¶ added in v0.16.0
func (p EmbeddingProfile) IsZero() bool
IsZero reports whether the profile carries no information.
type Entity ¶ added in v0.8.0
type Entity struct {
// ID is the unique identifier for this entity.
ID string
// Type categorizes the entity (e.g., "person", "company", "concept").
Type string
// Name is the human-readable name of the entity.
Name string
// Vector is the embedding that represents this entity.
Vector []float32
// Properties contains additional entity attributes.
Properties map[string]any
}
Entity represents a node in the knowledge graph.
type EntitySnapshot ¶ added in v0.14.0
type EntitySnapshot = storage.EntitySnapshot
EntitySnapshot is the serializable state of a knowledge graph entity.
type Episode ¶ added in v0.8.0
type Episode struct {
// ID is the unique identifier for this episode.
ID string
// Title is a human-readable summary of the episode.
Title string
// Vector is the embedding that represents the episode (centroid of contained records).
Vector []float32
// TimeRange is the span from the first to last record in the episode.
TimeRange TimeRange
// RecordIDs are the IDs of records that belong to this episode.
RecordIDs []uint64
// CreatedAt is when the episode was created.
CreatedAt time.Time
// Metadata contains additional episode information.
Metadata map[string]any
}
Episode represents a coherent group of related memories forming a discrete experience.
func (*Episode) RecordCount ¶ added in v0.8.0
RecordCount returns the number of records in the episode.
type EpisodeConfig ¶ added in v0.8.0
type EpisodeConfig struct {
// TimeGapThreshold is the maximum time gap between records in the same episode.
// Records separated by more than this are considered part of different episodes.
TimeGapThreshold time.Duration
// MinRecords is the minimum number of records to form an episode.
MinRecords int
// MaxRecords is the maximum number of records in an episode.
MaxRecords int
// SimilarityThreshold is the minimum similarity for records to be grouped.
// Used when temporal clustering alone isn't sufficient.
SimilarityThreshold float32
// Filters can be used to limit which records are considered for episodes.
Filters []Filter
}
EpisodeConfig configures episode detection.
type EpisodeResult ¶ added in v0.8.0
type EpisodeResult struct {
// Result is the underlying search result.
Result Result
// Episode is the episode containing this result (if any).
Episode *Episode
// EpisodeRecords are other records in the same episode.
EpisodeRecords []*Record
}
EpisodeResult represents a search result that includes episode context.
type EpisodeSnapshot ¶ added in v0.14.0
type EpisodeSnapshot = storage.EpisodeSnapshot
EpisodeSnapshot is the serializable state of a single episode.
type EpisodeStore ¶ added in v0.8.0
type EpisodeStore struct {
// contains filtered or unexported fields
}
EpisodeStore manages episodes for a collection.
func (*EpisodeStore) CreateEpisode ¶ added in v0.8.0
func (es *EpisodeStore) CreateEpisode(recordIDs []uint64, title string) (*Episode, error)
CreateEpisode manually creates an episode from a set of record IDs.
func (*EpisodeStore) DeleteEpisode ¶ added in v0.8.0
func (es *EpisodeStore) DeleteEpisode(episodeID string) error
DeleteEpisode removes an episode (does not delete the underlying records).
func (*EpisodeStore) DetectEpisodes ¶ added in v0.8.0
func (es *EpisodeStore) DetectEpisodes(config EpisodeConfig) ([]*Episode, error)
DetectEpisodes automatically detects episodes using temporal and similarity clustering.
func (*EpisodeStore) ExpandEpisode ¶ added in v0.8.0
func (es *EpisodeStore) ExpandEpisode(episodeID string) ([]*Record, error)
ExpandEpisode retrieves all records belonging to an episode.
func (*EpisodeStore) FindRecordEpisode ¶ added in v0.8.0
func (es *EpisodeStore) FindRecordEpisode(recordID uint64) (*Episode, error)
FindRecordEpisode finds the episode containing a specific record (if any).
func (*EpisodeStore) GetEpisode ¶ added in v0.8.0
func (es *EpisodeStore) GetEpisode(episodeID string) (*Episode, error)
GetEpisode retrieves an episode by ID.
func (*EpisodeStore) ListEpisodes ¶ added in v0.8.0
func (es *EpisodeStore) ListEpisodes() []*Episode
ListEpisodes returns all episodes.
func (*EpisodeStore) SearchEpisodes ¶ added in v0.8.0
func (es *EpisodeStore) SearchEpisodes(query []float32, limit int) ([]*Episode, error)
SearchEpisodes searches for episodes by their vector representation.
func (*EpisodeStore) SearchWithEpisodeExpansion ¶ added in v0.8.0
func (es *EpisodeStore) SearchWithEpisodeExpansion(query []float32, opts ...SearchOption) ([]EpisodeResult, error)
SearchWithEpisodeExpansion performs a search and includes episode context for results.
type EpisodeStoreSnapshot ¶ added in v0.14.0
type EpisodeStoreSnapshot = storage.EpisodeStoreSnapshot
EpisodeStoreSnapshot is the serializable state of an episode store.
type ExpandedSearchResult ¶ added in v0.8.0
type ExpandedSearchResult struct {
// Entity is the primary search result.
Entity *Entity
// Score is the similarity score.
Score float32
// RelatedEntities are entities connected to the primary result.
RelatedEntities []*Entity
// Relationships are the connections to related entities.
Relationships []*Relationship
}
ExpandedSearchResult contains search results with graph context.
type Filter ¶
type Filter interface {
// Match returns true if the record matches the filter criteria.
Match(r *Record) bool
}
Filter is an interface for filtering records based on payload values.
func AccessCountAbove ¶ added in v0.8.0
AccessCountAbove creates a filter that matches records accessed more than n times.
func AccessCountBelow ¶ added in v0.8.0
AccessCountBelow creates a filter that matches records accessed fewer than n times.
func AccessedAfter ¶ added in v0.8.0
AccessedAfter creates a filter that matches records last accessed after the given time.
func AccessedBefore ¶ added in v0.8.0
AccessedBefore creates a filter that matches records last accessed before the given time.
func AgeNewerThan ¶ added in v0.8.0
AgeNewerThan creates a filter that matches records created less than d ago.
func AgeOlderThan ¶ added in v0.8.0
AgeOlderThan creates a filter that matches records created more than d ago.
func Between ¶ added in v0.4.0
Between creates a filter that matches records where min <= payload[key] <= max.
func Contains ¶
Contains creates a filter that matches records where payload[key] contains the substring.
func CreatedAfter ¶ added in v0.8.0
CreatedAfter creates a filter that matches records created after the given time.
func CreatedBefore ¶ added in v0.8.0
CreatedBefore creates a filter that matches records created before the given time.
func ExpiredBefore ¶ added in v0.8.0
ExpiredBefore creates a filter that matches records that expired before the given time.
func GreaterThan ¶ added in v0.4.0
GreaterThan creates a filter that matches records where payload[key] > value.
func GreaterThanOrEqual ¶ added in v0.4.0
GreaterThanOrEqual creates a filter that matches records where payload[key] >= value.
func HasTTLFilter ¶ added in v0.8.0
func HasTTLFilter() Filter
HasTTLFilter creates a filter that matches records with a TTL set.
func ImportanceAbove ¶ added in v0.8.0
ImportanceAbove creates a filter that matches records with importance > threshold.
func ImportanceBelow ¶ added in v0.8.0
ImportanceBelow creates a filter that matches records with importance < threshold.
func ImportanceBetween ¶ added in v0.8.0
ImportanceBetween creates a filter that matches records with min <= importance <= max.
func LessThan ¶ added in v0.4.0
LessThan creates a filter that matches records where payload[key] < value.
func LessThanOrEqual ¶ added in v0.4.0
LessThanOrEqual creates a filter that matches records where payload[key] <= value.
func NeverAccessed ¶ added in v0.8.0
func NeverAccessed() Filter
NeverAccessed creates a filter that matches records that have never been accessed via search.
func NotEqual ¶
NotEqual creates a filter that matches records where payload[key] does not equal value.
func NotExpired ¶ added in v0.8.0
func NotExpired() Filter
NotExpired creates a filter that matches records that have not expired. This includes records with no TTL set.
func NotIn ¶
NotIn creates a filter that matches records where payload[key] is not in the given values.
func UpdatedAfter ¶ added in v0.8.0
UpdatedAfter creates a filter that matches records updated after the given time.
func UpdatedBefore ¶ added in v0.8.0
UpdatedBefore creates a filter that matches records updated before the given time.
type FilterFunc ¶
FilterFunc is a function adapter for the Filter interface.
func (FilterFunc) Match ¶
func (f FilterFunc) Match(r *Record) bool
Match implements Filter interface.
type FuseOption ¶ added in v0.16.0
type FuseOption interface {
// contains filtered or unexported methods
}
FuseOption configures result fusion (see FuseRRF).
func WithFusionTopK ¶ added in v0.16.0
func WithFusionTopK(n int) FuseOption
WithFusionTopK truncates the fused output to at most n results (0 = no limit).
func WithFusionWeights ¶ added in v0.16.0
func WithFusionWeights(weights ...float64) FuseOption
WithFusionWeights sets a per-result-set weight. The i-th weight scales the contribution of the i-th result set. Missing weights default to 1.0.
func WithRRFK ¶ added in v0.16.0
func WithRRFK(k int) FuseOption
WithRRFK sets the Reciprocal Rank Fusion constant k (default 60). Higher values flatten the influence of top ranks.
type GraphSnapshot ¶ added in v0.14.0
type GraphSnapshot = storage.GraphSnapshot
GraphSnapshot is the serializable state of a knowledge graph.
type HNSWConfig ¶ added in v0.2.0
type HNSWConfig = storage.HNSWConfig
HNSWConfig holds HNSW index configuration.
type HNSWStats ¶ added in v0.12.0
type HNSWStats struct {
// NodeCount is the number of vectors in the index.
NodeCount int `json:"node_count"`
// MaxLevel is the highest level in the HNSW graph.
MaxLevel int `json:"max_level"`
// EntryPointID is the ID of the entry point node.
EntryPointID uint64 `json:"entry_point_id"`
}
HNSWStats contains statistics about an HNSW index. This is the public wrapper for internal index statistics.
type Index ¶ added in v0.2.0
type Index interface {
// Insert adds a vector with the given ID to the index.
Insert(id uint64, vector []float32) error
// Delete removes a vector from the index.
Delete(id uint64) error
// Search finds the k nearest neighbors to the query vector.
// Returns IDs and distances/similarities.
Search(query []float32, k int) ([]IndexResult, error)
// SearchWithEf searches with a custom ef parameter (for HNSW).
// For indexes that don't support ef, this should behave like Search.
SearchWithEf(query []float32, k int, ef int) ([]IndexResult, error)
// Count returns the number of vectors in the index.
Count() int
// Type returns the index type name.
Type() string
// Clear removes all vectors from the index.
Clear()
}
Index is the interface for vector search indexes. Implementations can provide different algorithms (brute-force, HNSW, etc.).
type IndexResult ¶ added in v0.2.0
IndexResult represents a search result from an index.
type InsertOption ¶ added in v0.8.0
type InsertOption interface {
// contains filtered or unexported methods
}
InsertOption configures record insertion behavior.
func WithContentOption ¶ added in v0.8.0
func WithContentOption(content string) InsertOption
WithContentOption sets the content field for the record. This is an alternative to using InsertDocument.
func WithExpiresAt ¶ added in v0.8.0
func WithExpiresAt(t time.Time) InsertOption
WithExpiresAt sets an explicit expiration time for the record. This takes precedence over WithTTL if both are specified.
func WithImportance ¶ added in v0.8.0
func WithImportance(score float32) InsertOption
WithImportance sets the importance score for the record. Value should be between 0.0 and 1.0. Values outside this range are clamped.
func WithTTL ¶ added in v0.8.0
func WithTTL(d time.Duration) InsertOption
WithTTL sets a time-to-live duration for the record. The record will expire after this duration from the time of insertion.
type InvertedIndexSnapshot ¶ added in v0.6.0
type InvertedIndexSnapshot = storage.InvertedIndexSnapshot
InvertedIndexSnapshot is the serializable state of the BM25 inverted index.
type IterOption ¶ added in v0.6.0
type IterOption interface {
// contains filtered or unexported methods
}
IterOption configures the iterator.
func IterLimit ¶ added in v0.6.0
func IterLimit(n int) IterOption
IterLimit sets the maximum number of records to return.
func IterOffset ¶ added in v0.6.0
func IterOffset(n int) IterOption
IterOffset sets the number of records to skip.
type Iterator ¶ added in v0.6.0
type Iterator struct {
// contains filtered or unexported fields
}
Iterator allows iterating over collection records one at a time.
type KnowledgeGraph ¶ added in v0.8.0
type KnowledgeGraph struct {
// contains filtered or unexported fields
}
KnowledgeGraph provides a graph-based knowledge base with vector search.
func (*KnowledgeGraph) AddEntity ¶ added in v0.8.0
func (kg *KnowledgeGraph) AddEntity(entity Entity) error
AddEntity adds an entity to the knowledge graph.
func (*KnowledgeGraph) AddRelationship ¶ added in v0.8.0
func (kg *KnowledgeGraph) AddRelationship(rel Relationship) error
AddRelationship adds a relationship between two entities.
func (*KnowledgeGraph) DeleteEntity ¶ added in v0.8.0
func (kg *KnowledgeGraph) DeleteEntity(entityID string) error
DeleteEntity removes an entity and all its relationships.
func (*KnowledgeGraph) DeleteRelationship ¶ added in v0.8.0
func (kg *KnowledgeGraph) DeleteRelationship(relID string) error
DeleteRelationship removes a relationship.
func (*KnowledgeGraph) GetEntity ¶ added in v0.8.0
func (kg *KnowledgeGraph) GetEntity(entityID string) (*Entity, error)
GetEntity retrieves an entity by ID.
func (*KnowledgeGraph) GetRelationship ¶ added in v0.8.0
func (kg *KnowledgeGraph) GetRelationship(relID string) (*Relationship, error)
GetRelationship retrieves a relationship by ID.
func (*KnowledgeGraph) GetRelationships ¶ added in v0.8.0
func (kg *KnowledgeGraph) GetRelationships(entityID string, direction string) []*Relationship
GetRelationships returns all relationships for an entity.
func (*KnowledgeGraph) ListEntities ¶ added in v0.8.0
func (kg *KnowledgeGraph) ListEntities(entityType string) []*Entity
ListEntities returns all entities, optionally filtered by type.
func (*KnowledgeGraph) Name ¶ added in v0.8.0
func (kg *KnowledgeGraph) Name() string
Name returns the name of the knowledge graph.
func (*KnowledgeGraph) SearchWithExpansion ¶ added in v0.8.0
func (kg *KnowledgeGraph) SearchWithExpansion(query []float32, traversalConfig TraversalConfig, opts ...SearchOption) ([]ExpandedSearchResult, error)
SearchWithExpansion searches for similar entities and expands results with graph context.
func (*KnowledgeGraph) Stats ¶ added in v0.8.0
func (kg *KnowledgeGraph) Stats() KnowledgeGraphStats
Stats returns statistics about the knowledge graph.
func (*KnowledgeGraph) Traverse ¶ added in v0.8.0
func (kg *KnowledgeGraph) Traverse(startIDs []string, config TraversalConfig) (*TraversalResult, error)
Traverse performs a graph traversal starting from the given entity IDs.
func (*KnowledgeGraph) UpdateEntity ¶ added in v0.8.0
func (kg *KnowledgeGraph) UpdateEntity(entity Entity) error
UpdateEntity updates an existing entity.
type KnowledgeGraphStats ¶ added in v0.8.0
type KnowledgeGraphStats struct {
// EntityCount is the total number of entities.
EntityCount int
// RelationshipCount is the total number of relationships.
RelationshipCount int
// EntityTypes maps entity types to counts.
EntityTypes map[string]int
// RelationshipTypes maps relationship types to counts.
RelationshipTypes map[string]int
}
KnowledgeGraphStats contains statistics about a knowledge graph.
type Logger ¶ added in v0.6.0
type Logger interface {
// Debug logs a debug message with key-value pairs.
Debug(msg string, keysAndValues ...any)
// Info logs an informational message with key-value pairs.
Info(msg string, keysAndValues ...any)
// Error logs an error message with key-value pairs.
Error(msg string, keysAndValues ...any)
}
Logger is the interface for structured logging in VecLite. Implementations can bridge to any logging library (slog, zap, zerolog, etc.).
type MatchEvent ¶ added in v0.8.0
type MatchEvent struct {
// Record is the matching record.
Record *Record
// Score is the similarity score to the subscription query.
Score float32
// Timestamp is when the match was detected.
Timestamp time.Time
// SubscriptionID is the ID of the subscription that matched.
SubscriptionID string
}
MatchEvent represents an event when a new record matches a subscription.
type MemoryCluster ¶ added in v0.8.0
type MemoryCluster struct {
// ID is a unique identifier for this cluster.
ID string
// Records are the records in this cluster.
Records []*Record
// Centroid is the average vector of all records in the cluster.
Centroid []float32
// AverageImportance is the average importance score of records.
AverageImportance float32
// TimeRange is the span from oldest to newest record.
TimeRange TimeRange
}
MemoryCluster represents a group of similar memories that could be consolidated.
type MemoryConfig ¶ added in v0.9.0
type MemoryConfig struct {
// MaxRecords is the maximum number of records allowed in the collection.
// When exceeded, eviction is triggered.
MaxRecords int
// EvictionPolicy determines how records are selected for eviction.
// Supported values: "lru", "fifo", "importance"
EvictionPolicy string
// EvictionBatchSize is the number of records to evict at once when pressure is detected.
// Default is 10% of MaxRecords.
EvictionBatchSize int
// CleanupInterval is how often to check for memory pressure.
// If zero, cleanup only happens on insert.
CleanupInterval time.Duration
}
MemoryConfig configures memory pressure handling for a collection.
type MemoryLimiter ¶ added in v0.9.0
type MemoryLimiter struct {
// contains filtered or unexported fields
}
MemoryLimiter manages memory pressure for a collection.
func (*MemoryLimiter) Stop ¶ added in v0.9.0
func (ml *MemoryLimiter) Stop()
Stop stops the memory limiter and waits for it to finish.
type Metrics ¶ added in v0.6.0
type Metrics struct {
// contains filtered or unexported fields
}
Metrics provides observable counters for database operations. All operations are atomic and safe for concurrent access.
func (*Metrics) Snapshot ¶ added in v0.6.0
func (m *Metrics) Snapshot() MetricsSnapshot
Snapshot returns a point-in-time snapshot of the metrics.
type MetricsSnapshot ¶ added in v0.6.0
type MetricsSnapshot struct {
SearchCount int64 `json:"search_count"`
InsertCount int64 `json:"insert_count"`
DeleteCount int64 `json:"delete_count"`
AvgSearchTime time.Duration `json:"avg_search_time_ns"`
}
MetricsSnapshot is a point-in-time snapshot of database metrics.
type NopLogger ¶ added in v0.6.0
type NopLogger struct{}
NopLogger is a no-op logger that discards all messages. This is the default logger used when none is configured, ensuring zero overhead.
type NotFoundError ¶
type NotFoundError struct {
Type string // "record", "collection", etc.
ID string // Identifier that was not found
}
NotFoundError provides details about what was not found.
func (*NotFoundError) Error ¶
func (e *NotFoundError) Error() string
func (*NotFoundError) Unwrap ¶
func (e *NotFoundError) Unwrap() error
type ONNXConfig ¶ added in v0.11.0
ONNXConfig holds ONNX embedder configuration.
type OllamaConfig ¶ added in v0.11.0
type OllamaConfig struct {
BaseURL string `yaml:"base_url"`
Model string `yaml:"model"`
Timeout string `yaml:"timeout"`
}
OllamaConfig holds Ollama embedder configuration.
type OpenAIConfig ¶ added in v0.11.0
type OpenAIConfig struct {
APIKey string `yaml:"api_key"`
Model string `yaml:"model"`
BaseURL string `yaml:"base_url"`
Dimension int `yaml:"dimension"`
Timeout string `yaml:"timeout"`
}
OpenAIConfig holds OpenAI embedder configuration.
type Option ¶
type Option interface {
// contains filtered or unexported methods
}
Option configures the database.
func WithLogger ¶ added in v0.6.0
WithLogger sets a logger for the database. Pass nil or NopLogger{} to disable logging (default).
func WithReadOnly ¶
WithReadOnly opens the database in read-only mode. Write operations will return an error.
func WithSharedRead ¶ added in v0.20.0
WithSharedRead enables a shared file lock when opening a read-only database, allowing multiple processes to open the same database file simultaneously for read-only access. This is useful for scenarios where one process writes (e.g. an indexer) while other processes read (e.g. search tools).
SharedRead requires ReadOnly — opening a writable database with a shared lock would risk data loss from concurrent full-snapshot saves. If SharedRead is enabled without ReadOnly, Open returns an error.
Readers opened with SharedRead see a point-in-time snapshot taken at Open. Call Reload() to pick up writes from other processes.
func WithSyncOnWrite ¶
WithSyncOnWrite enables automatic sync after each write operation. This is slower but ensures durability.
func WithWAL ¶ added in v0.24.0
WithWAL enables the write-ahead log for a file-backed database. Every completed write appends the affected records to a sidecar log (path + ".wal") with a single fsync, so mutations survive a crash between full-snapshot saves without WithSyncOnWrite's cost of rewriting the whole database on every write. The log is replayed and folded into a fresh snapshot on the next open, and truncated by every successful Sync or Close.
The option is ignored for in-memory databases. Opens without WithWAL still recover a log left behind by a crashed WAL-enabled writer.
Durability caveat: a failed log append (disk full, I/O error) does not fail the write — it is logged, the affected records are retried on the next flush, and the data still persists on the next successful Sync or Close. Between a failed append and that point, those writes would not survive a crash.
func WithWALCheckpoint ¶ added in v0.24.0
WithWALCheckpoint sets the log size (in bytes) at which a WAL-enabled database automatically folds the log into a fresh snapshot and truncates it, bounding log growth for long-running writers that rarely call Sync. The default is DefaultWALCheckpointBytes (64 MiB). Pass 0 to disable automatic checkpointing (the log then grows until an explicit Sync or Close). Has no effect without WithWAL.
type ProfiledEmbedder ¶ added in v0.23.0
type ProfiledEmbedder interface {
Embedder
// Profile describes the embedder's provider, model, dimension,
// intended distance metric, and normalization behavior.
Profile() EmbeddingProfile
}
ProfiledEmbedder is an optional extension of Embedder for implementations that can describe how they produce vectors. It is not required: an Embedder may or may not support it, and callers discover support with a type assertion (or via EmbedderProfile, which also understands the built-in providers under embed/).
if pe, ok := e.(veclite.ProfiledEmbedder); ok {
profile := pe.Profile()
...
}
type Record ¶
type Record struct {
// ID is the unique identifier for this record.
ID uint64
// Vector is the embedding vector for the implicit "default" vector space.
Vector []float32
// Vectors holds embeddings for additional named vector spaces, keyed by
// space name (see Collection.AddVectorSpace). The "default" space is always
// represented by the Vector field above and never stored here. Nil when the
// record only uses the default space.
Vectors map[string][]float32
// Payload contains arbitrary metadata associated with the vector.
Payload map[string]any
// Content is the optional original text content associated with this record.
// Used for document-oriented storage and automatically indexed by BM25 when text indexing is enabled.
Content string
// CreatedAt is when the record was inserted.
CreatedAt time.Time
// UpdatedAt is when the record was last updated.
UpdatedAt time.Time
// ExpiresAt is when this record expires. Zero value means never expires.
ExpiresAt time.Time
// Importance is a score from 0.0 to 1.0 indicating how important this record is.
// Higher values make the record more likely to be returned in search results.
Importance float32
// AccessCount tracks how many times this record has been accessed via search.
AccessCount uint64
// LastAccessedAt is when this record was last accessed via search.
LastAccessedAt time.Time
}
Record represents a stored vector with its metadata.
func (*Record) HasTTL ¶ added in v0.8.0
HasTTL returns true if the record has an expiration time set.
func (*Record) IsExpired ¶ added in v0.8.0
IsExpired returns true if the record has a TTL set and has expired.
func (*Record) TTL ¶ added in v0.8.0
TTL returns the remaining time until expiration. Returns 0 if no TTL is set or if the record has already expired.
type RecordInput ¶ added in v0.16.0
type RecordInput struct {
// ID selects the record to upsert. 0 assigns a fresh auto-incremented ID.
ID uint64
// Content is optional text indexed by BM25 when text indexing is enabled.
Content string
// Payload is arbitrary metadata stored with the record.
Payload map[string]any
// Vectors maps vector-space name to embedding. See the type doc for the
// reserved "default" key.
Vectors map[string][]float32
}
RecordInput is one logical record that may carry vectors in several named vector spaces simultaneously, plus optional text content and payload.
Keys of Vectors are vector-space names. The reserved key DefaultVectorSpace (or the empty string) targets the default space (Record.Vector). Spaces other than "default" must already be declared via AddVectorSpace. A record may omit any space; it is then absent from that space's index. A record with no vectors at all behaves like a text-only document.
type RecordSnapshot ¶
type RecordSnapshot = storage.RecordSnapshot
RecordSnapshot is the serializable state of a record.
type Relationship ¶ added in v0.8.0
type Relationship struct {
// ID is the unique identifier for this relationship.
ID string
// SourceID is the ID of the source entity.
SourceID string
// TargetID is the ID of the target entity.
TargetID string
// Type describes the relationship (e.g., "works_at", "knows", "related_to").
Type string
// Weight is the strength of the relationship (0.0-1.0).
Weight float32
// Properties contains additional relationship attributes.
Properties map[string]any
// Bidirectional indicates if the relationship goes both ways.
Bidirectional bool
}
Relationship represents an edge between two entities in the knowledge graph.
func (*Relationship) Clone ¶ added in v0.8.0
func (r *Relationship) Clone() *Relationship
Clone creates a copy of the relationship.
type RelationshipSnapshot ¶ added in v0.14.0
type RelationshipSnapshot = storage.RelationshipSnapshot
RelationshipSnapshot is the serializable state of a knowledge graph relationship.
type Result ¶
type Result struct {
// Record is the matched record.
Record *Record
// Score is the similarity/distance score.
// For cosine/dot: higher is more similar.
// For euclidean: lower is more similar.
Score float32
}
Result represents a search result with its similarity score.
func FuseRRF ¶ added in v0.16.0
func FuseRRF(resultSets [][]Result, opts ...FuseOption) []Result
FuseRRF merges multiple ranked result sets into a single ranking using Reciprocal Rank Fusion. It is the public, modality-agnostic fusion primitive behind HybridSearch and Collection.MultiSpaceSearch: callers can fuse any mix of vector-space results, BM25 text results, or externally produced rankings.
Records are deduplicated by ID; a record appearing in several sets accumulates the (weighted) reciprocal-rank contribution of each. The returned results are sorted by fused score descending.
type SearchExplanation ¶ added in v0.2.0
type SearchExplanation struct {
// Results contains the search results.
Results []Result
// IndexType is the type of index used (none, hnsw).
IndexType string
// NodesVisited is the number of nodes visited during search.
// Only populated for HNSW searches.
NodesVisited int
// LayersVisited is the number of HNSW layers visited.
// Only populated for HNSW searches.
LayersVisited int
// Duration is how long the search took.
Duration time.Duration
// BruteForce indicates whether brute-force search was used.
BruteForce bool
}
SearchExplanation provides details about how a search was performed.
type SearchFunc ¶ added in v0.6.0
SearchFunc is a callback for streaming search results. Return false to stop receiving results.
type SearchOption ¶
type SearchOption interface {
// contains filtered or unexported methods
}
SearchOption configures search behavior.
func Threshold ¶
func Threshold(t float32) SearchOption
Threshold sets the minimum similarity score for results. For cosine/dot: results with score >= threshold are returned. For euclidean: results with score <= threshold are returned.
func TopK ¶
func TopK(k int) SearchOption
TopK sets the maximum number of results to return. Default is 10.
func WithAccessTracking ¶ added in v0.8.0
func WithAccessTracking(enabled bool) SearchOption
WithAccessTracking enables access tracking for search results. When enabled, accessing a record via search increments its access count and updates its last accessed timestamp.
func WithContent ¶ added in v0.6.0
func WithContent(include bool) SearchOption
WithContent controls whether the Content field is included in search results. By default, Content is included. Set to false to exclude it for smaller results.
func WithDecay ¶ added in v0.8.0
func WithDecay(decayType DecayType, halfLife time.Duration) SearchOption
WithDecay sets the temporal decay function for search results. The decay is applied based on record age (time since creation).
func WithEfSearch ¶ added in v0.2.0
func WithEfSearch(ef int) SearchOption
WithEfSearch sets the efSearch parameter for HNSW search. Higher values improve recall at the cost of speed. Has no effect on collections without HNSW index.
func WithFilter ¶
func WithFilter(f Filter) SearchOption
WithFilter adds a filter to the search. Multiple filters are combined with AND logic.
func WithFilters ¶
func WithFilters(filters ...Filter) SearchOption
WithFilters adds multiple filters to the search. All filters are combined with AND logic.
func WithImportanceBoost ¶ added in v0.8.0
func WithImportanceBoost(factor float32) SearchOption
WithImportanceBoost sets a boost factor for importance scores. The final score is: baseScore * (1 + importanceBoost * importance) A factor of 1.0 means importance can double the score.
func WithLimit ¶ added in v0.6.0
func WithLimit(n int) SearchOption
WithLimit sets the maximum number of results to return. This is an alias for TopK for use in pagination contexts.
func WithOffset ¶ added in v0.6.0
func WithOffset(n int) SearchOption
WithOffset sets the number of results to skip before returning. Use with TopK for pagination: WithOffset(20), TopK(10) returns results 21-30.
func WithTextWeight ¶ added in v0.6.0
func WithTextWeight(w float64) SearchOption
WithTextWeight sets the weight for the text search component in hybrid search. Default is 1.0.
func WithVectorWeight ¶ added in v0.6.0
func WithVectorWeight(w float64) SearchOption
WithVectorWeight sets the weight for the vector search component in hybrid search. Default is 1.0.
type SessionStats ¶ added in v0.8.0
type SessionStats struct {
// SessionID is the session identifier.
SessionID string
// TurnCount is the number of turns in the session.
TurnCount int
// FirstTurn is the timestamp of the first turn.
FirstTurn time.Time
// LastTurn is the timestamp of the last turn.
LastTurn time.Time
// Roles contains the count of turns by role.
Roles map[string]int
}
SessionStats returns statistics about a session.
type Storage ¶
Storage is the interface for database persistence. Implement this interface to provide custom storage backends.
type StorageError ¶
StorageError wraps storage-related errors with context.
type Subscription ¶ added in v0.8.0
type Subscription struct {
// ID is the unique identifier for this subscription.
ID string
// Query is the embedding vector to match against.
Query []float32
// Threshold is the minimum similarity score for a match.
Threshold float32
// Filters are additional filters to apply.
Filters []Filter
// contains filtered or unexported fields
}
Subscription represents an active subscription for matching records.
func (*Subscription) Close ¶ added in v0.8.0
func (s *Subscription) Close() error
Close closes the subscription and its event channel.
func (*Subscription) Events ¶ added in v0.8.0
func (s *Subscription) Events() <-chan MatchEvent
Events returns the channel for receiving match events.
func (*Subscription) IsClosed ¶ added in v0.8.0
func (s *Subscription) IsClosed() bool
IsClosed returns true if the subscription is closed.
type SubscriptionOption ¶ added in v0.8.0
type SubscriptionOption interface {
// contains filtered or unexported methods
}
SubscriptionOption configures a subscription.
func WithSubscriptionBufferSize ¶ added in v0.8.0
func WithSubscriptionBufferSize(size int) SubscriptionOption
WithSubscriptionBufferSize sets the event channel buffer size.
func WithSubscriptionFilter ¶ added in v0.8.0
func WithSubscriptionFilter(f Filter) SubscriptionOption
WithSubscriptionFilter adds a filter to the subscription.
func WithSubscriptionThreshold ¶ added in v0.8.0
func WithSubscriptionThreshold(threshold float32) SubscriptionOption
WithSubscriptionThreshold sets the minimum similarity threshold for matches.
type TTLCleaner ¶ added in v0.9.0
type TTLCleaner struct {
// contains filtered or unexported fields
}
TTLCleaner periodically removes expired records from collections.
func (*TTLCleaner) Stop ¶ added in v0.9.0
func (tc *TTLCleaner) Stop()
Stop stops the TTL cleaner and waits for it to finish.
type TTLCleanerCallback ¶ added in v0.9.0
TTLCleanerCallback is called after each cleanup cycle with the number of deleted records.
type TimeRangeSnapshot ¶ added in v0.14.0
type TimeRangeSnapshot = storage.TimeRangeSnapshot
TimeRangeSnapshot is the serializable state of a time range.
type TraversalConfig ¶ added in v0.8.0
type TraversalConfig struct {
// MaxDepth is the maximum number of hops from start nodes.
MaxDepth int
// MaxNodes is the maximum number of nodes to visit.
MaxNodes int
// MinWeight is the minimum relationship weight to follow.
MinWeight float32
// RelationshipTypes limits traversal to specific relationship types.
// Empty means all types.
RelationshipTypes []string
// EntityTypes limits traversal to specific entity types.
// Empty means all types.
EntityTypes []string
// Direction controls traversal direction: "outgoing", "incoming", or "both".
Direction string
}
TraversalConfig configures graph traversal behavior.
type TraversalResult ¶ added in v0.8.0
type TraversalResult struct {
// Entities are the entities found during traversal.
Entities []*Entity
// Relationships are the relationships traversed.
Relationships []*Relationship
// Depths maps entity IDs to their depth from start nodes.
Depths map[string]int
}
TraversalResult contains the results of a graph traversal.
type VectorSpaceConfig ¶ added in v0.16.0
type VectorSpaceConfig struct {
// Name uniquely identifies the space within its collection. Required, and
// must not be the reserved DefaultVectorSpace name.
Name string
// Dimension fixes the vector length for this space. 0 auto-detects from the
// first inserted vector.
Dimension int
// Distance is the distance metric for this space. Empty defaults to cosine.
Distance DistanceType
// Modality is an optional free-form hint such as "text", "image", or "audio".
Modality string
// Provider and Model record the embedding source. They are advisory metadata
// used by embedding-profile compatibility checks.
Provider string
Model string
// HNSW, when non-nil, enables an HNSW index for this space. Nil uses
// brute-force search.
HNSW *HNSWConfig
// Profile, when non-nil, attaches a first-class embedding profile to the
// space and enables compatibility validation on insert.
Profile *EmbeddingProfile
}
VectorSpaceConfig declares a named vector space on a collection.
A vector space is an independent index over one named embedding per record: its own dimension, distance metric, and optional HNSW index. One logical record (see RecordInput) can carry vectors in several spaces at once, for example a "text" embedding and an "image" embedding for the same item.
Apps still own embedding generation; VecLite only stores and searches the vectors the app provides for each space.
type VectorSpaceInfo ¶ added in v0.16.0
type VectorSpaceInfo struct {
// Name is the vector-space name ("default" for the implicit space).
Name string
// Dimension is the configured/auto-detected dimension (0 if not yet known).
Dimension int
// Distance is the distance metric used by this space.
Distance DistanceType
// Modality is the optional modality hint.
Modality string
// Provider and Model record the embedding source, if declared.
Provider string
Model string
// IndexType is "none" or "hnsw".
IndexType string
// VectorCount is the number of records carrying a vector in this space.
VectorCount int
// Profile is the space-level embedding profile, if any.
Profile *EmbeddingProfile
}
VectorSpaceInfo is a read-only view of a vector space's configuration and current state, returned by Collection.VectorSpaces and VectorSpace.
Source Files
¶
- bm25.go
- cleanup.go
- collection.go
- collection_spaces.go
- config.go
- consolidation.go
- conversation.go
- decay.go
- doc.go
- embedder.go
- embedder_factory.go
- episode.go
- errors.go
- explain.go
- filter.go
- fusion.go
- graph.go
- index.go
- index_hnsw.go
- insert_options.go
- logger.go
- metrics.go
- options.go
- record.go
- search.go
- storage.go
- subscription.go
- veclite.go
- vectorspace.go
- wal.go
Directories
¶
| Path | Synopsis |
|---|---|
|
Package client provides a thin Go client for a VecLite HTTP server.
|
Package client provides a thin Go client for a VecLite HTTP server. |
|
cmd
|
|
|
veclite
command
Command veclite provides a CLI for interacting with VecLite databases.
|
Command veclite provides a CLI for interacting with VecLite databases. |
|
embed
|
|
|
common
Package common provides shared utilities for embedder implementations.
|
Package common provides shared utilities for embedder implementations. |
|
ollama
Package ollama provides an Ollama-based embedding system for veclite.
|
Package ollama provides an Ollama-based embedding system for veclite. |
|
openai
Package openai provides an OpenAI-based embedding system for veclite.
|
Package openai provides an OpenAI-based embedding system for veclite. |
|
examples
|
|
|
basic
command
Example basic demonstrates the core VecLite operations: open a database, insert vectors, search, and close.
|
Example basic demonstrates the core VecLite operations: open a database, insert vectors, search, and close. |
|
batch
command
Example batch demonstrates batch operations, upsert, and iteration.
|
Example batch demonstrates batch operations, upsert, and iteration. |
|
filtering
command
Example filtering demonstrates VecLite's rich filter expressions.
|
Example filtering demonstrates VecLite's rich filter expressions. |
|
hnsw
command
Example hnsw demonstrates HNSW index configuration and performance.
|
Example hnsw demonstrates HNSW index configuration and performance. |
|
http-client
command
Example http-client demonstrates using VecLite's HTTP API.
|
Example http-client demonstrates using VecLite's HTTP API. |
|
internal
|
|
|
floats
Package floats provides optimized floating-point vector operations.
|
Package floats provides optimized floating-point vector operations. |
|
hnsw
Package hnsw implements the Hierarchical Navigable Small World graph algorithm for approximate nearest neighbor search.
|
Package hnsw implements the Hierarchical Navigable Small World graph algorithm for approximate nearest neighbor search. |
|
storage
Package storage provides persistence backends for VecLite databases.
|
Package storage provides persistence backends for VecLite databases. |
|
Package session provides a lazy dual-handle database session that resolves the multi-process lock contention problem with veclite's file-based storage.
|
Package session provides a lazy dual-handle database session that resolves the multi-process lock contention problem with veclite's file-based storage. |