Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Why this repo matters
•Converts PDFs and images into structured data specifically optimized for LLM ingestion and RAG pipelines.
•Supports over 100 languages with lightweight models suitable for diverse hardware including CPU and edge devices.
•Provides comprehensive document parsing capabilities like table extraction and layout analysis beyond standard text recognition.