Vision par ordinateur

GLM-OCR: document recognition with Markdown and JSON output

zai-org/GLM-OCR

An OCR project for tables, formulas and complex layouts, with hosted and self-hosted pipelines, a CLI and a Python SDK.

★ 7,5KÉtoiles
⑂ 671Forks
48Problèmes ouverts
PythonLangue
Apache-2.0Licence
Q84Score éditorial

Vue d’ensemble

GLM-OCR combines layout analysis with recognition. Its complete pipeline uses PP-DocLayoutV3 to identify document regions before recognition. The SDK can forward images and PDF documents to a hosted service or connect to self-hosted inference. In a knowledge-base workflow, it sits before indexing: convert pages into inspectable text and structure, then chunk and index those results while preserving page relationships. The OCR model has approximately 0.9B parameters; the complete service also includes layout processing.

Fonctionnalités clés

  • Handles document text, tables and formulas.
  • Includes layout analysis in the complete pipeline.
  • Hosted mode returns Markdown and JSON layout details.
  • Provides the glmocr parse command-line interface.
  • The Python API accepts individual images and multiple pages.
  • Offers hosted, vLLM and SGLang deployment paths.

Prérequis, installation et démarrage rapide

1. For hosted use, run pip install glmocr.
2. Configure pipeline.maas.enabled and an API key according to the official example, keeping credentials in local configuration or deployment secrets.
3. For the self-hosted layout pipeline, install pip install "glmocr[selfhosted]" and configure the model endpoint.
4. Run glmocr parse image.png --output ./results/ on a sample.
5. In Python, call parse("image.png") and result.save(output_dir="./results").
6. Inspect the extracted text and structure before processing directories or multi-page documents.

Utilisation

To process a collection of manuals, start with one text page, one table page and one formula page. Check reading order, table cells and mathematical symbols before running a batch. Store the original image, page identifier and output together. A list of images passed to the Python API is treated as pages of one document, so callers should explicitly preserve boundaries between separate documents.

Implementation notes
Blur, skew, stamps and dense tables can introduce recognition errors. Sample-check numbers and fields that affect downstream retrieval or analysis, and validate extracted values against application rules.

Compatibilité des modèles et cas d’usage

Hosted mode needs credentials and network access and sends documents to the provider. Self-hosting requires a compatible inference environment. A GPU-free client does not remove the compute requirements of the server.

Licence et notes sur les risques

Repository code uses Apache-2.0, the GLM-OCR model uses MIT, and PP-DocLayoutV3 in the full pipeline uses Apache-2.0.

opencv

opencv/opencv

★ 90,7KC++

tesseract

tesseract-ocr/tesseract

★ 76,3KC++

yolov5

ultralytics/yolov5

★ 58KPython

faceswap

deepfakes/faceswap

★ 57,5KPython