# pdf-inspector pdf-inspector is [[Firecrawl]]'s open-source (MIT) [[Rust]] library that answers one question fast: **does this PDF need OCR?** It classifies a document as TextBased, Scanned, ImageBased, or Mixed in 10-50ms with a confidence score — by scanning content streams for text operators instead of loading the whole file — then extracts text with positions and fonts from the pages that don't need OCR. The economics are the point. Most PDFs in a pipeline are text-based; running them through an OCR service wastes money and time. [[Firecrawl]] built this to handle that majority locally in milliseconds and send only the genuinely scanned pages to OCR. Classification and extraction share a single parse, so routing costs almost nothing extra. ## Capabilities - **Per-page OCR routing** — only pages that need OCR get sent out - **Markdown conversion** — headings, lists, code blocks, and tables detected from layout - **Table detection** — rectangle-based plus heuristic approaches - **Multi-column and RTL layouts** handled - **Bindings** for Rust, Python, Node.js, and WASM ## Positioning The classification-first design mirrors [[liteparse]]'s complexity detection: both encode the same lesson, that document pipelines should **route before they parse**. pdf-inspector is narrower (PDFs only) and sharper at the triage step; liteparse covers more formats and does fuller Markdown reconstruction. Escalation target for the scanned remainder: an OCR model like [[Mistral OCR]]. ## References - [pdf-inspector on GitHub](https://github.com/firecrawl/pdf-inspector) ## Related - [[liteparse]] — broader local parsing with the same routing idea - [[Mistral OCR]] — the expensive step this library helps you avoid - [[Retrieval-Augmented Generation (RAG)]] - [[Rust]] - [[Firecrawl]] — the maker; web-scraping counterpart of the same cost logic