pdf2md-tui

module
v1.2.7 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: May 9, 2026 License: MIT

README ΒΆ

pdf2md-tui

High-speed PDF β†’ Markdown ingestion engine for multimodal RAG pipelines. Extracts structured text + isolated images so downstream chunkers, LlamaIndex, and VLM agents get context that actually works.

CI Go Version Latest Release


pdf2md-tui banner

Why?

PDFs break AI pipelines. Binary encoding, embedded fonts, and layout metadata inflate token counts β€” but the real damage is structural. Legacy parsers flatten documents into raw text, destroying tables, orphaning image references, and producing vector embeddings that are functionally useless for retrieval.

pdf2md-tui solves this by extracting structured Markdown with preserved table layouts alongside an isolated ./images/ directory β€” the exact format that modern RAG frameworks (LlamaIndex, LangChain) and Vision-Language Models expect.

Document Naive VLM / Image Processing (est. API cost) Clean Markdown (est. token cost) Context Quality
50-page technical spec ~42,500 tokens (Image API) ~15,000 tokens (Clean text) Noise-free, structured
200-page legal contract ~170,000 tokens (Image API) ~60,000 tokens (Clean text) Noise-free, structured
Research paper (12 pages) ~10,200 tokens (Image API) ~3,500 tokens (Clean text) Noise-free, structured

Estimates reflect common API token costs for Vision vs. Text ingestion. Actual savings depend on document density and VLM provider.

Built for the "Look Twice" methodology β€” extract text + isolate images at ingestion time (Phase 1), then let your downstream VLM pipeline handle deep visual reasoning at retrieval time (Phase 2).

Use cases:

  • Preprocessing document archives for RAG pipelines (LlamaIndex SimpleDirectoryReader compatible)
  • Building multimodal knowledge bases with text + image vector stores
  • Feeding autonomous AI agents via MCP with structured, parseable document context
  • Reducing token costs when processing large document sets via API

Features

  • Chunking-safe tables β€” Basic table reconstruction exists; robust chunking-safe table blocks are planned. Positional text analysis detects column alignment and emits GFM pipe tables.
  • Two-path extraction β€” positional extraction first (preserves structure); falls back to plain-text for edge cases
  • Worker pool β€” concurrent conversion using runtime.NumCPU() workers by default; configurable via --workers
  • Live TUI β€” Bubble Tea dashboard for batch conversion with live stats, recent activity, and a completion menu
  • Interactive Menu β€” run pdf2md-tui without arguments in a terminal to launch a guided configuration wizard
  • Graceful OCR detection β€” scanned/image-only PDFs are detected and skipped cleanly, reported in the summary (no empty output files)
  • --strip-noise β€” aggressively removes page numbers, repeated headers/footers, and excess whitespace for maximum token density
  • Smart Overwrites β€” prompts before overwrite in a terminal; non-interactive and --quiet runs fail unless --force is set
  • Automation-friendly modes β€” --quiet emits a JSON summary; non-TTY non-quiet runs fall back to a plain-text summary without alt-screen UI
  • Date-stamped outputs β€” report_2026-05-06.md so you always know which version was processed
  • Single static binary β€” no runtime dependencies; CGO_ENABLED=0, pure Go

πŸš€ Recent Quality Improvements

We've recently overhauled the core extraction engine to move beyond simple text scraping toward semantic document reconstruction. These changes ensure that the output Markdown is truly RAG-ready and "chunking-safe."

πŸ—οΈ Advanced Extraction Engine (v1.2.7+)
  • Adaptive Spacing Heuristics: Uses line-level gap statistics (Mean/StdDev) to correctly coalesce intentionally tracked headings (e.g., D E F I N I T I V E) into single words while maintaining natural word separation.
  • Double-Letter Preservation: Disabled aggressive character deduplication in favor of a content-first strategy. Legitimate double letters (e.g., across, ebooks, www) are now perfectly preserved across all document sources.
  • Block-Aware Optimization: Refactored the --strip-noise pipeline. Whitespace collapsing is now scoped to individual blocks, strictly preserving paragraph boundaries (\n\n) and table structures.
  • Standardized Unicode: Integrated norm.NFKC for industry-standard normalization. All extracted text now features consistent resolution of ligatures (fi, fl, etc.) and PUA characters.
βœ… Automated QA Validation

To ensure long-term scalability and prevent regressions, we implemented a formal QA Test Plan with an automated validation suite (pkg/service/qa_validation_test.go).

Core Assertions Verified:

  • Content Fidelity: Zero collapse of legitimate double letters.
  • Structural Integrity: Guaranteed minimum paragraph density in optimized output.
  • Semantic Coherence: Coalescing of styled headings and specialized technical symbols (β†’, β€’).
  • Cleanliness: Zero-tolerance policy for replacement characters (U+FFFD) and "garbage" encoding sequences.

πŸ“¦ Installation

Homebrew (macOS and Linux)
brew tap nawodyaishan/tap
brew install pdf2md-tui
Go install
go install github.com/nawodyaishan/pdf2md-tui/cmd/pdf2md-tui@latest
Binary download

Download a pre-built binary for your platform from the latest release.

Platforms: Linux, macOS, Windows Γ— amd64 / arm64. Packages: .tar.gz, .zip, .deb, .rpm.


Usage

Interactive Mode

Run the tool without arguments in a terminal to launch the guided interactive wizard. It prompts for the target directory and common flags, and currently preselects Extract images in the multiselect:

pdf2md-tui

The zero-argument menu is only used when both stdin and stdout are interactive terminals. Otherwise the command defaults to converting the current directory.

CLI Mode
# Convert all PDFs in ./docs, write Markdown to ./docs/md/
pdf2md-tui convert ./docs

# Recurse into subdirectories and strip layout noise for LLM ingestion
pdf2md-tui convert ./archive --recursive --strip-noise

# Use 8 workers, custom output directory, no date suffix, and force overwrite existing files
pdf2md-tui convert ./papers --workers 8 --output out --date-format none --force

# Extract embedded images and inject markdown links into the output
pdf2md-tui convert ./docs --extract-images

# CI-friendly mode: emit only JSON to stdout
pdf2md-tui convert ./docs --quiet

# Print version and build info
pdf2md-tui version
Go Library Usage

Since the refactoring to Clean Architecture, you can embed the conversion engine directly into your own Go applications.

[!NOTE] The core logic resides in the pkg/ directory, making it importable as a standard Go module.

import (
	"github.com/nawodyaishan/pdf2md-tui/pkg/domain"
	"github.com/nawodyaishan/pdf2md-tui/pkg/repository/pdf"
	"github.com/nawodyaishan/pdf2md-tui/pkg/repository/storage"
	"github.com/nawodyaishan/pdf2md-tui/pkg/service"
)

func main() {
	// 1. Initialize configuration
	cfg := domain.NewConfig()
	cfg.ExtractImages = true
	cfg.StripNoise = true

	// 2. Initialize dependencies (Clean Architecture)
	store := storage.NewStorage()
	parser := pdf.NewParser()

	// 3. Initialize the service
	conv := service.NewConverterService(cfg, store, parser)

	// 4. Run conversion
	// Convert(pdfPath, outDir) returns a domain.Result
	res := conv.Convert("input.pdf", "output_dir")

	if res.Err != nil {
		fmt.Printf("Conversion failed: %v\n", res.Err)
		return
	}

	fmt.Printf("Successfully converted %s to %s (Saved %d bytes)\n", 
		res.InputPath, res.OutputPath, res.InputBytes - res.OutputBytes)
}
Flags
Flag Short Default Description
--output -o md/ Output subdirectory (relative to the target directory)
--recursive -r false Scan subdirectories
--workers -w NumCPU Number of concurrent conversion workers
--date-format 2006-01-02 Date suffix format (Go reference time); none disables the suffix
--force -f false Overwrite existing output files without prompting
--strip-noise false Aggressively remove page numbers, headers/footers, and excess whitespace
--extract-images false Extract embedded images and inject markdown references
--quiet -q false Suppress the TUI and emit a JSON summary to stdout
--verbose -v false Print per-file errors to stderr
--log-file pdf2md.log Path for detailed conversion logs
Runtime Modes
  • Interactive terminal: shows the pterm banner/discovery phase, then launches the Bubble Tea dashboard during conversion. When the batch completes, the dashboard offers Open Output Directory, View Detailed Log when failures occurred, and Exit.
  • Non-interactive without --quiet: skips the Bubble Tea alt-screen and prints a concise text summary after conversion finishes.
  • --quiet: prints only the JSON summary to stdout. If any conversions fail, the command still exits non-zero.
Overwrite and Log Behavior
  • Existing output files require confirmation only in an interactive terminal.
  • Non-interactive runs and --quiet refuse to overwrite existing outputs unless --force is set.
  • When failures occur, inspect the configured --log-file path for details. The completion menu exposes View Detailed Log in the dashboard, and non-interactive summaries print the log path directly.
Output structure
./docs/
β”œβ”€β”€ report.pdf
β”œβ”€β”€ spec.pdf
└── md/
    β”œβ”€β”€ images/
    β”‚   β”œβ”€β”€ report/
    β”‚   β”‚   β”œβ”€β”€ report_1_5.png
    β”‚   β”‚   └── report_2_11.jpg
    β”‚   └── spec/
    β”‚       └── spec_1_8.png
    β”œβ”€β”€ report_2026-05-06.md
    └── spec_2026-05-06.md

How it works

  1. Discovery β€” scans the target directory (optionally recursive) for .pdf files.
  2. Worker pool β€” distributes files across N goroutines via a buffered channel.
  3. Extraction β€” for each page, attempts positional character extraction to preserve table structure; falls back to plain-text if the page is image-only or yields no content.
  4. Table detection β€” identifies column-aligned rows (β‰₯3 columns, β‰₯50 pt gaps, appearing in β‰₯40% of rows) and renders them as GFM pipe tables.
  5. Output β€” writes one .md file per PDF to the output directory.

See docs/ARCHITECTURE.md for a detailed breakdown of the extraction pipeline and concurrency model.


Roadmap Status

  • πŸ›‘οΈ Graceful OCR Detection β€” Detect and skip scanned PDFs without failing.
  • πŸ–ΌοΈ Image Extraction Pipeline β€” Extract raw images for "Look Twice" VLM workflows.
  • ⚑ Zero-Arg Usability β€” Run pdf2md-tui in any folder with no arguments.
  • πŸ—οΈ Clean Architecture β€” Decoupled domain/service/repository structure for scaling.
  • 🧱 Chunking-Safe Tables β€” Basic GFM pipe table support (v1.0). Robust indivisible blocks planned.
  • πŸ”‡ CI-Friendly Quiet Mode β€” Non-interactive JSON output for automation.
  • πŸ”Œ MCP Server Wrapper β€” Native tool support for Model Context Protocol agents.
  • ☁️ VLM Cloud Integration β€” High-accuracy Markdown generation via GPT-4o/Claude.

See ROADMAP.md for the full strategic vision.


Building from source

git clone https://github.com/nawodyaishan/pdf2md-tui.git
cd pdf2md-tui
make build          # β†’ bin/pdf2md-tui

make test           # run tests with race detector
make lint           # golangci-lint
make check          # fmt + vet + lint + test
make help           # list all targets
Git hooks

This repository uses Lefthook as a Go-friendly alternative to Husky.

The hook split follows the usual Go workflow:

  • pre-commit: format staged Go files with gofmt -w and run go vet ./...
  • pre-push: run golangci-lint run ./..., go test -race ./..., and go build ./cmd/pdf2md-tui

Install Lefthook and wire the hooks into .git/hooks:

brew install lefthook
make hooks-install

Official Go-based install is also supported:

go install github.com/evilmartians/lefthook/v2@v2.1.6
make hooks-install

You can also run the hook suites manually:

make hooks-run-pre-commit
make hooks-run-pre-push

Roadmap

The roadmap is organized around maximizing context quality for downstream AI pipelines:

  • Near-term (v0.x) β€” Image extraction pipeline (pdfcpu), graceful OCR detection, chunking-safe table output, --quiet JSON mode for CI/MCP
  • Mid-term (v1.x) β€” "Look Twice" VLM pipeline (cloud vision providers), MCP server prototype, .docx/.txt ingestion
  • Long-term (v2.x) β€” Pluggable post-processors, full MCP server + REST API

See ROADMAP.md for the full vision, strategic goals, and how to contribute.


Contributing

Bug reports and feature requests: open an issue using the provided templates.

For code contributions:

  1. Check ROADMAP.md and open issues for help wanted items.
  2. Fork, branch, implement, add tests, and open a PR.
  3. PRs must pass make check (fmt + vet + lint + test with race detector).

License

MIT β€” see LICENSE.

Directories ΒΆ

Path Synopsis
cmd
debug-pdf command
pdf2md-tui command
internal
pkg

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL