Open-source repositories gaining traction right now.
12 repositories
#ocr
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

A community-supported supercharged document management system: scan, index and archive all your documents

ShareX is a free and open-source application that enables users to capture or record any area of their screen with a single keystroke. It also supports uploading images, text, and various file types to a wide range of destinations.

PDF Parser for AI-ready data. Automate PDF accessibility. Open-source.

A fast, helpful, and open-source document parser

OCR model that handles complex tables, forms, handwriting with full layout.

GLM-OCR: Accurate × Fast × Comprehensive

RuVector is a High Performance, Real-Time, Self-Learning, Vector Graph Neural Network, and Database built in Rust.

Sponsor Star icereed / paperless-gpt Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI

批量OCR工具,底层用了Ollama和GLM -OCR、PP-DocLayoutV3(检测模型),速度极快,正常PDF书籍平均0.5秒每页(4090显卡笔记本)

Paste a github.com URL. Submissions are reviewed before they are tracked.