praixis-go
Go SDK for PraixisEngine — a FastAPI backend for local LLM chat sessions and RAG (retrieval-augmented generation) over your own documents.
Zero external dependencies. Stdlib only.
Installation
go get github.com/mettjs/praixis-go
Requires Go 1.22+.
Quick start
import (
praixis "github.com/mettjs/praixis-go"
"github.com/mettjs/praixis-go/chat"
"github.com/mettjs/praixis-go/rag"
)
client := praixis.New("http://localhost:8080", "praixis_your_key")
The client exposes two sub-clients:
| Field |
Package |
Purpose |
client.Chat |
chat |
Chat sessions, file summarization, session history |
client.RAG |
rag |
Vector collections, document Q&A, upload, embed |
Chat
Every generative endpoint comes in two flavors: a buffered method that returns
the whole result in one call, and a streaming method (suffixed Stream) that
yields tokens as they arrive.
Buffered chat
resp, err := client.Chat.Send(ctx, chat.Request{
Prompt: "Explain Go interfaces in two sentences.",
})
if err != nil { /* handle */ }
fmt.Println(resp.SessionID) // assigned by the server
fmt.Println(resp.Content) // the full reply
Set ResponseFormat: "json" to ask the model to return JSON; Content is still a
string (the model's raw JSON text), which you then unmarshal yourself.
Streaming chat
stream, err := client.Chat.Stream(ctx, chat.Request{
Prompt: "Explain Go interfaces in two sentences.",
})
if err != nil { /* handle */ }
defer stream.Close()
fmt.Println(stream.SessionID()) // assigned by the server
for stream.Next() {
fmt.Print(stream.Token())
}
if err := stream.Err(); err != nil { /* handle */ }
Continue an existing session by passing SessionID to either method:
resp, err := client.Chat.Send(ctx, chat.Request{
Prompt: "Follow-up question",
SessionID: previousSessionID,
})
Summarize a file
f, _ := os.Open("report.pdf")
defer f.Close()
// Buffered:
summary, err := client.Chat.SummarizeFile(ctx, "report.pdf", f, nil)
fmt.Println(summary.Filename) // "report.pdf"
fmt.Println(summary.Content) // the full summary
// Streaming, with options:
stream, err := client.Chat.SummarizeFileStream(ctx, "report.pdf", f,
&chat.FileSummaryOptions{Task: "Extract action items", Tone: "Bullet points"},
)
defer stream.Close()
fmt.Println(stream.Filename()) // "report.pdf"
fmt.Println(stream.Progress()) // e.g. "reducing 4 chunks"
for stream.Next() {
fmt.Print(stream.Token())
}
if stream.BackendError() != "" {
// server-side error (e.g. GPU busy) reported inside the stream
}
Session management
// Fetch message history
history, err := client.Chat.History(ctx, sessionID)
for _, msg := range history.History {
fmt.Printf("[%s] %s\n", msg.Role, msg.Content)
}
// List all active sessions
sessions, err := client.Chat.ListSessions(ctx)
// Per-session token usage. Counts the streamed answers (chat and RAG), RAG
// query reformulation, and compaction calls; counters expire with the session.
// EstimatedContextTokens shows how close the session is to auto-compacting.
usage, err := client.Chat.Usage(ctx, sessionID)
fmt.Println(usage.TotalTokens, usage.EstimatedContextTokens)
// Compact a session on demand: fold older exchanges into an LLM-written summary
// (the server also does this automatically near its context budget). Returns a
// 400-wrapped APIError when there's nothing to fold yet.
res, err := client.Chat.Compact(ctx, sessionID)
fmt.Printf("%d -> %d messages\n", res.MessagesBefore, res.MessagesAfter)
// Undo the last exchange: removes the most recent user message and the
// assistant reply that followed it, so it can be retried or regenerated.
// Compaction summaries are kept. Returns a 400-wrapped APIError when there's
// no user message left to undo.
undo, err := client.Chat.UndoLastExchange(ctx, sessionID)
fmt.Println(undo.UndonePrompt, undo.MessagesRemaining) // UndonePrompt = the removed user message
// Delete a session
err = client.Chat.DeleteSession(ctx, sessionID)
RAG
Ask a question
// Buffered:
resp, err := client.RAG.Ask(ctx, rag.QuestionRequest{
CollectionName: "hr-docs",
Question: "How many vacation days do employees get?",
})
if err != nil { /* handle */ }
fmt.Println(resp.SessionID) // session for follow-up questions
fmt.Println(resp.SearchQuery) // reformulated query used for retrieval
fmt.Println(resp.Sources) // []string{"policy.pdf", "hr-handbook.docx"}
fmt.Println(resp.Content) // the full answer
// Streaming:
stream, err := client.RAG.AskStream(ctx, rag.QuestionRequest{
CollectionName: "hr-docs",
Question: "How many vacation days do employees get?",
})
if err != nil { /* handle */ }
defer stream.Close()
fmt.Println(stream.SessionID()) // session for follow-up questions
fmt.Println(stream.SearchQuery()) // reformulated query used for retrieval
for stream.Next() {
fmt.Print(stream.Token())
}
if err := stream.Err(); err != nil { /* handle */ }
fmt.Println(stream.Sources()) // []string{"policy.pdf", "hr-handbook.docx"}
Restrict retrieval to a single source document with MetadataFilter. The only
honored key is source; any other keys are ignored (not an error):
rag.QuestionRequest{
CollectionName: "hr-docs",
Question: "What is the notice period?",
MetadataFilter: map[string]any{"source": "policy.pdf"},
}
Search (retrieval only)
Search returns the ranked raw chunks that match a query — no LLM call, no synthesis.
Use it when you want the evidence and its scores to reason over yourself (e.g. fusing
these chunks with another source) instead of a finished answer. Unlike Ask it does not
reformulate the query, so pass a standalone one.
resp, err := client.RAG.Search(ctx, rag.SearchRequest{
CollectionName: "hr-docs",
Query: "vacation accrual rate",
NResults: 5, // optional, 1–20; server default 5
})
if err != nil { /* handle */ }
for _, hit := range resp.Results {
fmt.Printf("%s (%.4f): %s\n", hit.Source, hit.Score, hit.Text)
}
// resp.ScoreType is "rrf" (hybrid pgvector backend) or "similarity" (dense Chroma backend),
// telling you how to read each hit's Score.
Returns a 404-wrapped APIError if the collection does not exist.
Upload documents
f, _ := os.Open("policy.pdf")
defer f.Close()
resp, err := client.RAG.Upload(ctx,
[]rag.FileUpload{{Filename: "policy.pdf", Data: f}},
&rag.UploadOptions{
CollectionName: "hr-docs",
ChunkingStrategy: "semantic", // or "character" for fixed-size splits
ChunkSize: praixis.Ptr(2000),
ChunkOverlap: praixis.Ptr(150), // only used when ChunkingStrategy is "character"
ImprovedSearch: true, // background hypothetical-question indexing for better natural-language search
},
)
// resp.Succeeded, resp.Results[i].Status
Pass nil for UploadOptions to use server defaults (collection "main", semantic chunking, chunk size 2000).
Filename is required — it is the document's stored identity and the server's primary format signal, so prefer a .pdf/.docx/.txt extension. For extension-less names the server falls back to the part's Content-Type (filled from the extension when ContentType is empty), then to the file's magic bytes.
ImprovedSearch enables hypothetical-question indexing: questions are generated in the background after the upload returns, so the document is searchable immediately and natural-language matching improves once generation finishes.
Upload raw text
UploadText ingests text directly — no file wrapping — for callers that already hold the content (scraped pages, database records). It runs the same pipeline as Upload after text extraction, so the stored document works with every other endpoint. Filename becomes the document's stored identity; re-using one replaces the prior document.
resp, err := client.RAG.UploadText(ctx, rag.TextUploadRequest{
Text: "full document text…",
Filename: "faq-2026.txt",
CollectionName: "hr-docs", // optional; default "main"
ImprovedSearch: true, // optional; same semantics as Upload
})
// resp.ChunksStored
Inspect chunks
Chunks returns a document's stored chunks in order — exactly what retrieval sees, useful for debugging why a search returned what it did. Returns a 404-wrapped APIError if the file does not exist.
chunks, err := client.RAG.Chunks(ctx, "hr-docs", "policy.pdf")
for _, c := range chunks.Chunks {
fmt.Printf("#%d: %s\n", c.ChunkIndex, c.Content)
}
Question-index management
ImprovedSearch is no longer locked in at upload time. RegenerateQuestions backfills or rebuilds the hypothetical-question index for an already-stored document (generation runs in the background; the existing questions are replaced only once the new pass succeeds), and QuestionStatus reports its progress. Returns a 409-wrapped APIError while a pass is already running, a 400-wrapped one when indexing is disabled server-side.
sched, err := client.RAG.RegenerateQuestions(ctx, "hr-docs", "policy.pdf")
// sched.Status == "scheduled", sched.Chunks
status, err := client.RAG.QuestionStatus(ctx, "hr-docs", "policy.pdf")
// status.QuestionsStored, status.GenerationPending — poll until pending is false
Collections
// List all collections
list, err := client.RAG.ListCollections(ctx)
// list.ActiveCollections []string, list.TotalDocuments int
// Files in a collection
files, err := client.RAG.ListFiles(ctx, "hr-docs")
// Delete a file from a collection
err = client.RAG.DeleteFile(ctx, "hr-docs", "old-policy.pdf")
// Delete an entire collection
err = client.RAG.DeleteCollection(ctx, "hr-docs")
Summarize / compare files
// LLM summary of a stored file (buffered)
summary, err := client.RAG.SummarizeDocument(ctx, "hr-docs", "policy.pdf", nil)
fmt.Println(summary.Content)
// Bullet-point diff between two files (buffered)
cmp, err := client.RAG.Compare(ctx, "hr-docs", "policy-v1.pdf", "policy-v2.pdf", nil)
fmt.Println(cmp.Content)
Both take an optional final argument for ResponseFormat ("text" or "json")
and have streaming counterparts SummarizeDocumentStream / CompareStream:
stream, err := client.RAG.CompareStream(ctx, "hr-docs", "v1.pdf", "v2.pdf",
&rag.CompareOptions{ResponseFormat: "json"})
defer stream.Close()
for stream.Next() {
fmt.Print(stream.Token())
}
Embedding
resp, err := client.RAG.Embed(ctx, "some text to embed")
// resp.Embedding []float64 (384 dimensions), resp.Dimensions int
Error handling
_, err := client.Chat.Stream(ctx, chat.Request{Prompt: "hi"})
switch {
case praixis.IsGPUBusy(err):
// 503 — server is processing another request; retry after a moment
case praixis.IsRateLimit(err):
// 429 — slow down
case praixis.IsUnauthorized(err):
// 401/403 — check your API key
case praixis.IsNotFound(err):
// 404 — collection, file, or session does not exist
}
The helpers live on the top-level praixis package. For full detail, extract the typed error with errors.As:
var apiErr *praixis.APIError
if errors.As(err, &apiErr) {
fmt.Println(apiErr.StatusCode, apiErr.Message)
}
Configuration
client := praixis.New(baseURL, apiKey,
praixis.WithTimeout(60*time.Second), // default 30s; does not affect streams
)
Streams are not subject to the HTTP timeout — cancel them via context.Context:
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
defer cancel()
stream, err := client.Chat.Stream(ctx, ...)
Examples
Runnable examples are in _examples/:
| Directory |
Description |
_examples/chat_stream |
Stream a single chat turn |
_examples/rag_ask |
Ask a question over a collection |
_examples/rag_upload |
Upload a file into a collection |
PRAIXIS_URL=http://localhost:8080 PRAIXIS_API_KEY=praixis_xxx go run ./_examples/chat_stream
Package layout
praixis-go/
├── praixis.go # New(), Ptr(), error helpers (APIError, IsGPUBusy, …), Option types
├── chat/
│ ├── types.go # Request, ChatResponse, FileSummary, History, FileSummaryOptions, …
│ ├── stream.go # ChatStream, FileSummaryStream
│ └── client.go # Send, Stream, SummarizeFile, SummarizeFileStream, History, …
├── rag/
│ ├── types.go # QuestionRequest, AskResponse, FileUpload, UploadOptions, …
│ ├── stream.go # AskStream, SummaryStream, CompareStream
│ └── client.go # Ask, AskStream, Search, Upload, Embed, SummarizeDocument, Compare, ListCollections, …
└── internal/
├── httpclient.go # DoJSON, DoStream, DoMultipart, APIError, …
└── stream.go # StreamReader (metadata + token iterator)
License
MIT — see LICENSE.