Overview
An open-source OCR engine for converting text in images into machine-readable text. The repository is maintained under tesseract-ocr on GitHub. Its primary language is C++. See the [project README](https://github.com/tesseract-ocr/tesseract#readme) for the supported workflows.
Key features
- An open-source OCR engine for converting text in images into machine-readable text.
Requirements, installation and quick start
[Read the upstream installation and quickstart instructions](https://github.com/tesseract-ocr/tesseract#readme).
Usage
[Usage examples and configuration reference](https://github.com/tesseract-ocr/tesseract#readme).
Model compatibility and use cases
Model compatibility is not stated in the repository metadata.
License and risk notes
GitHub reports Apache-2.0 for this repository. Review the upstream license file; model weights, datasets and dependencies may have separate terms.
Source review 2026-09-05: GitHub search metadata and repository README. No runtime benchmark performed. License metadata: Apache-2.0.
Release and maintenance
[View upstream releases](https://github.com/tesseract-ocr/tesseract/releases).