whisper — Speech recognition and transcription
openai/whisper
Robust Speech Recognition via Large-Scale Weak Supervision
ToolAI.io · Kênh GitHub
Chỉ mục dựa trên dữ kiện về các dự án phát triển LLM, tác nhân, MCP, RAG và AI, kèm giấy phép, thiết lập, tải xuống, ảnh chụp màn hình và tài nguyên liên quan cho từng mục.
Dữ liệu kho lưu trữ công khai
openai/whisper
Robust Speech Recognition via Large-Scale Weak Supervision
huggingface/transformers
A model-definition framework for state-of-the-art machine learning models across text, vision, audio, and multimodal domains, supporting both inference and training.
mintplex-labs/anything-llm
An all-in-one, local-first AI application for chatting with documents, building AI agents, and running a private, multi-user ChatGPT-like experience with zero setup friction.
nvidia-nemo/speech
A scalable generative AI framework built for researchers and PyTorch developers working on Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Speech LLMs.
lancedb/lancedb
An open-source, developer-friendly embedded retrieval library and multimodal AI lakehouse designed for fast, scalable, and production-ready vector search, built on the Lance columnar format.
xorbitsai/inference
A powerful and versatile library designed to serve language, speech recognition, and multimodal models. It allows users to swap GPT for any LLM by changing a single line of code and run models on cloud, on-prem, or locally via a unified, production-ready inference API.
genkit-ai/genkit
Genkit is an open-source framework built and used in production by Google's Firebase for developing full-stack, AI-powered applications. It provides cross-language SDKs and unified APIs for integrating various AI models to build chatbots, automations, and recommendation systems.
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages
deepseek-ai/DeepSeek-OCR
Contexts Optical Compression
QwenLM/Qwen3-VL
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud
deepseek-ai/Janus
Janus-Series: Unified Multimodal Understanding and Generation Models