DeepSeek-OCR — Document OCR and visual understanding
deepseek-ai/DeepSeek-OCR
Contexts Optical Compression
ToolAI.io · Kênh GitHub
Duyệt toàn bộ thư mục ToolAI gồm các dự án AI mã nguồn mở trên GitHub, được sắp xếp theo chủ đề, ngôn ngữ và giấy phép.
Dữ liệu kho lưu trữ công khai
deepseek-ai/DeepSeek-OCR
Contexts Optical Compression
QwenLM/Qwen3-VL
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud
deepseek-ai/Janus
Janus-Series: Unified Multimodal Understanding and Generation Models
Tencent-Hunyuan/HY-World-2.0
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
AUTOMATIC1111/stable-diffusion-webui
A browser interface for Stable Diffusion image generation and extensions.
Tencent-Hunyuan/HunyuanWorld-1.0
Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
Tencent-Hunyuan/HunyuanImage-3.0
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
zai-org/CogVideo
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
Tencent-Hunyuan/Hunyuan3D-2
High-Resolution 3D Assets Generation with Large Scale Hunyuan3D Diffusion Models.
Trang 2 / 2 · 21 dự án