whisper — Speech recognition and transcription
openai/whisper
Robust Speech Recognition via Large-Scale Weak Supervision
ToolAI.io · GitHub چینل
LLM، ایجنٹ، MCP، RAG اور AI ڈویلپمنٹ پروجیکٹس کی حقائق پر مبنی فہرست، جس میں ہر اندراج کے لیے لائسنس، سیٹ اپ، ڈاؤن لوڈ، اسکرین شاٹس اور متعلقہ وسائل شامل ہیں۔
عوامی ریپوزٹری کا ڈیٹا
openai/whisper
Robust Speech Recognition via Large-Scale Weak Supervision
Tencent-Hunyuan/UniRL
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
harry0703/MoneyPrinterTurbo
An AI-assisted workflow for generating short videos from a topic or keywords.
zai-org/GLM-V
GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
PaddlePaddle/PaddleOCR
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages
deepseek-ai/DeepSeek-OCR
Contexts Optical Compression
QwenLM/Qwen3-VL
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud
deepseek-ai/Janus
Janus-Series: Unified Multimodal Understanding and Generation Models
Tencent-Hunyuan/HY-World-2.0
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
AUTOMATIC1111/stable-diffusion-webui
A browser interface for Stable Diffusion image generation and extensions.
Tencent-Hunyuan/HunyuanWorld-1.0
Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
Tencent-Hunyuan/HunyuanImage-3.0
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
صفحہ 1 / 2 · 14 پروجیکٹس