# liteparse
liteparse is [[LlamaIndex]]'s open-source (Apache 2.0), **local-first document parser**: PDFs and Office documents in, Markdown/JSON/plain text out, no cloud, no LLM required. It does spatial text extraction with bounding boxes via PDFium, which makes the output position-aware rather than a text dump.
It exists as the free floor under LlamaParse, [[LlamaIndex]]'s paid cloud parser. The division of labor is explicit: liteparse for straightforward documents, fast and local; LlamaParse when the document is hard (dense tables, multi-column layouts, charts, handwriting, scans).
## The useful parts
- **Complexity detection** — scores a document *before* parsing so a pipeline can route: simple → parse locally, complex → escalate to OCR or a cloud parser. That routing step is the design idea worth stealing.
- **Markdown reconstruction** — headings, tables, lists, links, and images rebuilt from spatial layout, not just extracted strings.
- **Pluggable OCR** — bundled [[Tesseract]], an HTTP OCR server, or your own implementation.
- **Screenshot generation** — page images for agents that want to look at the document, not just read it.
- **Runs everywhere** — Rust core with Python, Node.js/TypeScript, and browser (WASM) bindings. Office formats (Word, PowerPoint, spreadsheets) via LibreOffice conversion; images natively.
## Positioning
Same niche as [[MarkItDown]] (Microsoft's convert-anything-to-Markdown tool), but with bounding boxes, complexity routing, and WASM as differentiators. For scanned or messy documents, a dedicated OCR model like [[Mistral OCR]] still wins. [[pdf-inspector]] (Firecrawl) covers the classification half of this problem with the same routing philosophy.
## References
- [liteparse on GitHub](https://github.com/run-llama/liteparse)
## Related
- [[MarkItDown]] — closest alternative
- [[pdf-inspector]] — classify-then-route sibling from Firecrawl
- [[Mistral OCR]] — where to escalate scanned/complex documents
- [[Retrieval-Augmented Generation (RAG)]] — the pipeline this usually feeds
- [[Rust]]
- [[LlamaIndex]] — the framework it belongs to
- [[Tesseract]] — its default OCR engine