Resumen
GLM-OCR combines layout analysis with recognition. Its complete pipeline uses PP-DocLayoutV3 to identify document regions before recognition. The SDK can forward images and PDF documents to a hosted service or connect to self-hosted inference. In a knowledge-base workflow, it sits before indexing: convert pages into inspectable text and structure, then chunk and index those results while preserving page relationships. The OCR model has approximately 0.9B parameters; the complete service also includes layout processing.
Características principales
- Handles document text, tables and formulas.
- Includes layout analysis in the complete pipeline.
- Hosted mode returns Markdown and JSON layout details.
- Provides the glmocr parse command-line interface.
- The Python API accepts individual images and multiple pages.
- Offers hosted, vLLM and SGLang deployment paths.
Requisitos, instalación y guía rápida
2. Configure pipeline.maas.enabled and an API key according to the official example, keeping credentials in local configuration or deployment secrets.
3. For the self-hosted layout pipeline, install pip install "glmocr[selfhosted]" and configure the model endpoint.
4. Run glmocr parse image.png --output ./results/ on a sample.
5. In Python, call parse("image.png") and result.save(output_dir="./results").
6. Inspect the extracted text and structure before processing directories or multi-page documents.
Uso
Implementation notes
Blur, skew, stamps and dense tables can introduce recognition errors. Sample-check numbers and fields that affect downstream retrieval or analysis, and validate extracted values against application rules.
Compatibilidad de modelos y casos de uso
Hosted mode needs credentials and network access and sends documents to the provider. Self-hosting requires a compatible inference environment. A GPU-free client does not remove the compute requirements of the server.
Notas sobre la licencia y los riesgos
Repository code uses Apache-2.0, the GLM-OCR model uses MIT, and PP-DocLayoutV3 in the full pipeline uses Apache-2.0.