# liteparse liteparse is [[LlamaIndex]]'s open-source (Apache 2.0), **local-first document parser**: PDFs and Office documents in, Markdown/JSON/plain text out, no cloud, no LLM required. It does spatial text extraction with bounding boxes via PDFium, which makes the output position-aware rather than a text dump. It exists as the free floor under LlamaParse, [[LlamaIndex]]'s paid cloud parser. The division of labor is explicit: liteparse for straightforward documents, fast and local; LlamaParse when the document is hard (dense tables, multi-column layouts, charts, handwriting, scans). ## The useful parts - **Complexity detection** — scores a document *before* parsing so a pipeline can route: simple → parse locally, complex → escalate to OCR or a cloud parser. That routing step is the design idea worth stealing. - **Markdown reconstruction** — headings, tables, lists, links, and images rebuilt from spatial layout, not just extracted strings. - **Pluggable OCR** — bundled [[Tesseract]], an HTTP OCR server, or your own implementation. - **Screenshot generation** — page images for agents that want to look at the document, not just read it. - **Runs everywhere** — Rust core with Python, Node.js/TypeScript, and browser (WASM) bindings. Office formats (Word, PowerPoint, spreadsheets) via LibreOffice conversion; images natively. ## Positioning Same niche as [[MarkItDown]] (Microsoft's convert-anything-to-Markdown tool), but with bounding boxes, complexity routing, and WASM as differentiators. For scanned or messy documents, a dedicated OCR model like [[Mistral OCR]] still wins. [[pdf-inspector]] (Firecrawl) covers the classification half of this problem with the same routing philosophy. ## References - [liteparse on GitHub](https://github.com/run-llama/liteparse) ## Related - [[MarkItDown]] — closest alternative - [[pdf-inspector]] — classify-then-route sibling from Firecrawl - [[Mistral OCR]] — where to escalate scanned/complex documents - [[Retrieval-Augmented Generation (RAG)]] — the pipeline this usually feeds - [[Rust]] - [[LlamaIndex]] — the framework it belongs to - [[Tesseract]] — its default OCR engine