whisper — Speech recognition and transcription
openai/whisper
Robust Speech Recognition via Large-Scale Weak Supervision
ToolAI.io · GitHub चैनल
LLM, एजेंट, MCP, RAG और AI डेवलपमेंट प्रोजेक्ट्स की तथ्यों पर आधारित इंडेक्स, जिसमें हर प्रविष्टि के लिए लाइसेंस, सेटअप, डाउनलोड, स्क्रीनशॉट और संबंधित संसाधन शामिल हैं।
सार्वजनिक रिपॉज़िटरी का डेटा
openai/whisper
Robust Speech Recognition via Large-Scale Weak Supervision
huggingface/transformers
A model-definition framework for state-of-the-art machine learning models across text, vision, audio, and multimodal domains, supporting both inference and training.
mintplex-labs/anything-llm
An all-in-one, local-first AI application for chatting with documents, building AI agents, and running a private, multi-user ChatGPT-like experience with zero setup friction.
sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
nvidia-nemo/speech
A scalable generative AI framework built for researchers and PyTorch developers working on Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Speech LLMs.
lancedb/lancedb
An open-source, developer-friendly embedded retrieval library and multimodal AI lakehouse designed for fast, scalable, and production-ready vector search, built on the Lance columnar format.
xorbitsai/inference
A powerful and versatile library designed to serve language, speech recognition, and multimodal models. It allows users to swap GPT for any LLM by changing a single line of code and run models on cloud, on-prem, or locally via a unified, production-ready inference API.
genkit-ai/genkit
Genkit is an open-source framework built and used in production by Google's Firebase for developing full-stack, AI-powered applications. It provides cross-language SDKs and unified APIs for integrating various AI models to build chatbots, automations, and recommendation systems.
Tencent-Hunyuan/UniRL
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
harry0703/MoneyPrinterTurbo
An AI-assisted workflow for generating short videos from a topic or keywords.
zai-org/GLM-V
GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages
पृष्ठ 1 / 2 · 21 प्रोजेक्ट