firecrawl/pdf-inspector
Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.
Rust17K stars1.1K forks+1.8K today
Topics
#markdown#nodejs#pdf#pdf-extraction#pdf-parser#python#rust#text-extraction#ocr-routing#pdf-classification
What it does
pdf-inspector is a fast Rust library for PDF inspection and text extraction, intelligently distinguishing between scanned and text-based documents to optimize processing. It's crucial for content creators needing efficient PDF handling without OCR.
Star history
Not enough history yet — 1 day(s) recorded. The daily snapshot builds this up.
Tracking
- Last trending
- 2026-08-03
Creator kit
Hook
Revolutionize your PDF workflow with pdf-inspector, a Rust library that detects text-based vs scanned documents in under 50ms. #PDFHandling
Content angles
- How to Use pdf-inspector for Efficient Document Processing
- Benchmarking pdf-inspector Against Competitors: A Deep Dive
- Building Web Apps with pdf-inspector's Browser WebAssembly
Who should care
Developers, content creators, and anyone needing efficient PDF processing without OCR services.