whisper — Speech recognition and transcription
openai/whisper
Robust Speech Recognition via Large-Scale Weak Supervision
ToolAI.io · Kênh GitHub
Chỉ mục dựa trên dữ kiện về các dự án phát triển LLM, tác nhân, MCP, RAG và AI, kèm giấy phép, thiết lập, tải xuống, ảnh chụp màn hình và tài nguyên liên quan cho từng mục.
Dữ liệu kho lưu trữ công khai
openai/whisper
Robust Speech Recognition via Large-Scale Weak Supervision
huggingface/transformers
A model-definition framework for state-of-the-art machine learning models across text, vision, audio, and multimodal domains, supporting both inference and training.
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages
deepseek-ai/DeepSeek-OCR
Contexts Optical Compression
QwenLM/Qwen3-VL
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud
deepseek-ai/Janus
Janus-Series: Unified Multimodal Understanding and Generation Models