DeepSeek-OCR — Document OCR and visual understanding
deepseek-ai/DeepSeek-OCR
Contexts Optical Compression
ToolAI.io · Canal de GitHub
Un índice basado en hechos de proyectos de desarrollo de LLM, agentes, MCP, RAG y AI, con licencia, configuración, descarga, capturas de pantalla y recursos relacionados para cada entrada.
Datos del repositorio público
deepseek-ai/DeepSeek-OCR
Contexts Optical Compression
QwenLM/Qwen3-VL
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud
deepseek-ai/Janus
Janus-Series: Unified Multimodal Understanding and Generation Models
Tencent-Hunyuan/HY-World-2.0
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
AUTOMATIC1111/stable-diffusion-webui
A browser interface for Stable Diffusion image generation and extensions.
Tencent-Hunyuan/HunyuanWorld-1.0
Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
Tencent-Hunyuan/HunyuanImage-3.0
HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
zai-org/CogVideo
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
Tencent-Hunyuan/Hunyuan3D-2
High-Resolution 3D Assets Generation with Large Scale Hunyuan3D Diffusion Models.
Página 2 / 2 · 21 proyectos