hal0
hal0ai/hal0
An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.
ToolAI.io · GitHub चैनल
LLM, एजेंट, MCP, RAG और कोडिंग विषयों में सक्रिय ओपन-सोर्स AI रिपॉज़िटरीज़ को स्टार्स और सत्यापित प्रोजेक्ट जानकारी के साथ ट्रैक करें।
सार्वजनिक रिपॉज़िटरी का डेटा
hal0ai/hal0
An open-source, self-hosted home AI inference platform designed to turn a Linux box into an OpenAI-compatible inference appliance, with native optimization for AMD Strix Halo hardware.
xllm-ai/xllm
xLLM is a C++ inference framework for LLM, VLM, DiT and REC models. It targets high-throughput, low-latency deployment on several AI accelerator families and separates service-layer scheduling and availability from engine-layer computation.
vllm-project/vllm-ascend
A community-maintained hardware plugin that enables vLLM to run model inference and serving workloads on supported Ascend NPUs.
dstackai/dstack
A vendor-agnostic control plane for provisioning and orchestrating training, inference, development, and agentic workloads across GPU clouds, Kubernetes, and on-premises infrastructure.
avifenesh/memra
A from-scratch LLM inference engine built in Rust and CUDA, specifically optimized for RTX 5090 (sm_120a) and H100 (sm_90a) architectures, delivering exactness-gated performance without relying on frameworks like ggml.
mozilla-ai/any-llm
any-llm is a Python SDK for communicating with multiple LLM providers through one interface. It supports switching providers and models with minimal code changes while using official provider SDKs when available.
openai/tiktoken
tiktoken is a fast BPE tokeniser for use with OpenAI's models
xorbitsai/inference
A powerful and versatile library designed to serve language, speech recognition, and multimodal models. It allows users to swap GPT for any LLM by changing a single line of code and run models on cloud, on-prem, or locally via a unified, production-ready inference API.
katanemo/plano
Plano is an AI-native proxy server and data plane built in Rust that centralizes LLM routing, agent orchestration, observability, and guardrails, allowing developers to focus on core agent logic rather than infrastructure plumbing.
huggingface/transformers
A model-definition framework for state-of-the-art machine learning models across text, vision, audio, and multimodal domains, supporting both inference and training.
opencsgs/csghub
CSGHub is an open-source, on-premise platform for managing the full lifecycle of Large Language Model assets, including models, datasets, spaces, and code, offering functionality comparable to a private Hugging Face.
lemonade-sdk/lemonade
Lemonade is an open-source local AI server that enables users to run optimized Large Language Models (LLMs), speech, and image generation models directly on their own GPUs and NPUs, providing a free and private alternative to cloud APIs.
पृष्ठ 2 / 4 · 41 प्रोजेक्ट