README
ΒΆ
pdf2md-tui
High-speed PDF β Markdown ingestion engine for multimodal RAG pipelines. Extracts structured text + isolated images so downstream chunkers, LlamaIndex, and VLM agents get context that actually works.
Why?
PDFs break AI pipelines. Binary encoding, embedded fonts, and layout metadata inflate token counts β but the real damage is structural. Legacy parsers flatten documents into raw text, destroying tables, orphaning image references, and producing vector embeddings that are functionally useless for retrieval.
pdf2md-tui solves this by extracting structured Markdown with preserved table layouts alongside an isolated ./images/ directory β the exact format that modern RAG frameworks (LlamaIndex, LangChain) and Vision-Language Models expect.
| Document | Naive VLM / Image Processing (est. API cost) | Clean Markdown (est. token cost) | Context Quality |
|---|---|---|---|
| 50-page technical spec | ~42,500 tokens (Image API) | ~15,000 tokens (Clean text) | Noise-free, structured |
| 200-page legal contract | ~170,000 tokens (Image API) | ~60,000 tokens (Clean text) | Noise-free, structured |
| Research paper (12 pages) | ~10,200 tokens (Image API) | ~3,500 tokens (Clean text) | Noise-free, structured |
Estimates reflect common API token costs for Vision vs. Text ingestion. Actual savings depend on document density and VLM provider.
Built for the "Look Twice" methodology β extract text + isolate images at ingestion time (Phase 1), then let your downstream VLM pipeline handle deep visual reasoning at retrieval time (Phase 2).
Use cases:
- Preprocessing document archives for RAG pipelines (LlamaIndex
SimpleDirectoryReadercompatible) - Building multimodal knowledge bases with text + image vector stores
- Feeding autonomous AI agents via MCP with structured, parseable document context
- Reducing token costs when processing large document sets via API
Features
- Chunking-safe tables β Basic table reconstruction exists; robust chunking-safe table blocks are planned. Positional text analysis detects column alignment and emits GFM pipe tables.
- Two-path extraction β positional extraction first (preserves structure); falls back to plain-text for edge cases
- Worker pool β concurrent conversion using
runtime.NumCPU()workers by default; configurable via--workers - Live TUI β Bubble Tea dashboard for batch conversion with live stats, recent activity, and a completion menu
- Interactive Menu β run
pdf2md-tuiwithout arguments in a terminal to launch a guided configuration wizard - Graceful OCR detection β scanned/image-only PDFs are detected and skipped cleanly, reported in the summary (no empty output files)
--strip-noiseβ aggressively removes page numbers, repeated headers/footers, and excess whitespace for maximum token density- Smart Overwrites β prompts before overwrite in a terminal; non-interactive and
--quietruns fail unless--forceis set - Automation-friendly modes β
--quietemits a JSON summary; non-TTY non-quiet runs fall back to a plain-text summary without alt-screen UI - Date-stamped outputs β
report_2026-05-06.mdso you always know which version was processed - Single static binary β no runtime dependencies;
CGO_ENABLED=0, pure Go
π Recent Quality Improvements
We've recently overhauled the core extraction engine to move beyond simple text scraping toward semantic document reconstruction. These changes ensure that the output Markdown is truly RAG-ready and "chunking-safe."
ποΈ Advanced Extraction Engine (v1.2.7+)
- Adaptive Spacing Heuristics: Uses line-level gap statistics (Mean/StdDev) to correctly coalesce intentionally tracked headings (e.g.,
D E F I N I T I V E) into single words while maintaining natural word separation. - Double-Letter Preservation: Disabled aggressive character deduplication in favor of a content-first strategy. Legitimate double letters (e.g.,
across,ebooks,www) are now perfectly preserved across all document sources. - Block-Aware Optimization: Refactored the
--strip-noisepipeline. Whitespace collapsing is now scoped to individual blocks, strictly preserving paragraph boundaries (\n\n) and table structures. - Standardized Unicode: Integrated
norm.NFKCfor industry-standard normalization. All extracted text now features consistent resolution of ligatures (fi,fl, etc.) and PUA characters.
β Automated QA Validation
To ensure long-term scalability and prevent regressions, we implemented a formal QA Test Plan with an automated validation suite (pkg/service/qa_validation_test.go).
Core Assertions Verified:
- Content Fidelity: Zero collapse of legitimate double letters.
- Structural Integrity: Guaranteed minimum paragraph density in optimized output.
- Semantic Coherence: Coalescing of styled headings and specialized technical symbols (
β,β’). - Cleanliness: Zero-tolerance policy for replacement characters (
U+FFFD) and "garbage" encoding sequences.
π¦ Installation
Homebrew (macOS and Linux)
brew tap nawodyaishan/tap
brew install pdf2md-tui
Go install
go install github.com/nawodyaishan/pdf2md-tui/cmd/pdf2md-tui@latest
Binary download
Download a pre-built binary for your platform from the latest release.
Platforms: Linux, macOS, Windows Γ amd64 / arm64. Packages: .tar.gz, .zip, .deb, .rpm.
Usage
Interactive Mode
Run the tool without arguments in a terminal to launch the guided interactive wizard. It prompts for the target directory and common flags, and currently preselects Extract images in the multiselect:
pdf2md-tui
The zero-argument menu is only used when both stdin and stdout are interactive terminals. Otherwise the command defaults to converting the current directory.
CLI Mode
# Convert all PDFs in ./docs, write Markdown to ./docs/md/
pdf2md-tui convert ./docs
# Recurse into subdirectories and strip layout noise for LLM ingestion
pdf2md-tui convert ./archive --recursive --strip-noise
# Use 8 workers, custom output directory, no date suffix, and force overwrite existing files
pdf2md-tui convert ./papers --workers 8 --output out --date-format none --force
# Extract embedded images and inject markdown links into the output
pdf2md-tui convert ./docs --extract-images
# CI-friendly mode: emit only JSON to stdout
pdf2md-tui convert ./docs --quiet
# Print version and build info
pdf2md-tui version
Go Library Usage
Since the refactoring to Clean Architecture, you can embed the conversion engine directly into your own Go applications.
[!NOTE] The core logic resides in the
pkg/directory, making it importable as a standard Go module.
import (
"github.com/nawodyaishan/pdf2md-tui/pkg/domain"
"github.com/nawodyaishan/pdf2md-tui/pkg/repository/pdf"
"github.com/nawodyaishan/pdf2md-tui/pkg/repository/storage"
"github.com/nawodyaishan/pdf2md-tui/pkg/service"
)
func main() {
// 1. Initialize configuration
cfg := domain.NewConfig()
cfg.ExtractImages = true
cfg.StripNoise = true
// 2. Initialize dependencies (Clean Architecture)
store := storage.NewStorage()
parser := pdf.NewParser()
// 3. Initialize the service
conv := service.NewConverterService(cfg, store, parser)
// 4. Run conversion
// Convert(pdfPath, outDir) returns a domain.Result
res := conv.Convert("input.pdf", "output_dir")
if res.Err != nil {
fmt.Printf("Conversion failed: %v\n", res.Err)
return
}
fmt.Printf("Successfully converted %s to %s (Saved %d bytes)\n",
res.InputPath, res.OutputPath, res.InputBytes - res.OutputBytes)
}
Flags
| Flag | Short | Default | Description |
|---|---|---|---|
--output |
-o |
md/ |
Output subdirectory (relative to the target directory) |
--recursive |
-r |
false |
Scan subdirectories |
--workers |
-w |
NumCPU |
Number of concurrent conversion workers |
--date-format |
2006-01-02 |
Date suffix format (Go reference time); none disables the suffix |
|
--force |
-f |
false |
Overwrite existing output files without prompting |
--strip-noise |
false |
Aggressively remove page numbers, headers/footers, and excess whitespace | |
--extract-images |
false |
Extract embedded images and inject markdown references | |
--quiet |
-q |
false |
Suppress the TUI and emit a JSON summary to stdout |
--verbose |
-v |
false |
Print per-file errors to stderr |
--log-file |
pdf2md.log |
Path for detailed conversion logs |
Runtime Modes
- Interactive terminal: shows the pterm banner/discovery phase, then launches the Bubble Tea dashboard during conversion. When the batch completes, the dashboard offers
Open Output Directory,View Detailed Logwhen failures occurred, andExit. - Non-interactive without
--quiet: skips the Bubble Tea alt-screen and prints a concise text summary after conversion finishes. --quiet: prints only the JSON summary to stdout. If any conversions fail, the command still exits non-zero.
Overwrite and Log Behavior
- Existing output files require confirmation only in an interactive terminal.
- Non-interactive runs and
--quietrefuse to overwrite existing outputs unless--forceis set. - When failures occur, inspect the configured
--log-filepath for details. The completion menu exposesView Detailed Login the dashboard, and non-interactive summaries print the log path directly.
Output structure
./docs/
βββ report.pdf
βββ spec.pdf
βββ md/
βββ images/
β βββ report/
β β βββ report_1_5.png
β β βββ report_2_11.jpg
β βββ spec/
β βββ spec_1_8.png
βββ report_2026-05-06.md
βββ spec_2026-05-06.md
How it works
- Discovery β scans the target directory (optionally recursive) for
.pdffiles. - Worker pool β distributes files across
Ngoroutines via a buffered channel. - Extraction β for each page, attempts positional character extraction to preserve table structure; falls back to plain-text if the page is image-only or yields no content.
- Table detection β identifies column-aligned rows (β₯3 columns, β₯50 pt gaps, appearing in β₯40% of rows) and renders them as GFM pipe tables.
- Output β writes one
.mdfile per PDF to the output directory.
See docs/ARCHITECTURE.md for a detailed breakdown of the extraction pipeline and concurrency model.
Roadmap Status
- π‘οΈ Graceful OCR Detection β Detect and skip scanned PDFs without failing.
- πΌοΈ Image Extraction Pipeline β Extract raw images for "Look Twice" VLM workflows.
- β‘ Zero-Arg Usability β Run
pdf2md-tuiin any folder with no arguments. - ποΈ Clean Architecture β Decoupled domain/service/repository structure for scaling.
- π§± Chunking-Safe Tables β Basic GFM pipe table support (v1.0). Robust indivisible blocks planned.
- π CI-Friendly Quiet Mode β Non-interactive JSON output for automation.
- π MCP Server Wrapper β Native tool support for Model Context Protocol agents.
- βοΈ VLM Cloud Integration β High-accuracy Markdown generation via GPT-4o/Claude.
See ROADMAP.md for the full strategic vision.
Building from source
git clone https://github.com/nawodyaishan/pdf2md-tui.git
cd pdf2md-tui
make build # β bin/pdf2md-tui
make test # run tests with race detector
make lint # golangci-lint
make check # fmt + vet + lint + test
make help # list all targets
Git hooks
This repository uses Lefthook as a Go-friendly alternative to Husky.
The hook split follows the usual Go workflow:
pre-commit: format staged Go files withgofmt -wand rungo vet ./...pre-push: rungolangci-lint run ./...,go test -race ./..., andgo build ./cmd/pdf2md-tui
Install Lefthook and wire the hooks into .git/hooks:
brew install lefthook
make hooks-install
Official Go-based install is also supported:
go install github.com/evilmartians/lefthook/v2@v2.1.6
make hooks-install
You can also run the hook suites manually:
make hooks-run-pre-commit
make hooks-run-pre-push
Roadmap
The roadmap is organized around maximizing context quality for downstream AI pipelines:
- Near-term (v0.x) β Image extraction pipeline (
pdfcpu), graceful OCR detection, chunking-safe table output,--quietJSON mode for CI/MCP - Mid-term (v1.x) β "Look Twice" VLM pipeline (cloud vision providers), MCP server prototype,
.docx/.txtingestion - Long-term (v2.x) β Pluggable post-processors, full MCP server + REST API
See ROADMAP.md for the full vision, strategic goals, and how to contribute.
Contributing
Bug reports and feature requests: open an issue using the provided templates.
For code contributions:
- Check ROADMAP.md and open issues for
help wanteditems. - Fork, branch, implement, add tests, and open a PR.
- PRs must pass
make check(fmt + vet + lint + test with race detector).
License
MIT β see LICENSE.