추론, 배포 및 런타임

NVIDIA Container Toolkit: NVIDIA GPU access for containers

NVIDIA/nvidia-container-toolkit

Configures NVIDIA GPU access for container runtimes such as Docker, supporting containerized inference, training and CUDA workloads.

★ 4.6K별점
⑂ 589포크 수
57미해결 이슈
Go언어
Apache-2.0라이선스
Q84편집 점수

개요

NVIDIA Container Toolkit connects the host driver and container runtime so containers can access GPU devices and required driver components. It is infrastructure for GPU-enabled images rather than an AI model. Validate the host, runtime and a simple test container before introducing a model-serving image; this separates driver issues from runtime configuration and application dependencies.

주요 기능

  • Provides runtime integration for GPU containers.
  • Includes the nvidia-ctk configuration utility.
  • Documents Docker runtime configuration.
  • Documents paths for runtimes including containerd and CRI-O.
  • Supports device access configuration for different deployment modes.
  • Works with CUDA application images for model workloads.

요구 사항, 설치 및 빠른 시작

1. Install a compatible NVIDIA driver on the Linux host and confirm that nvidia-smi detects the devices.
2. Configure NVIDIA’s package repository following the guide for your distribution.
3. Install the documented Container Toolkit packages, pinning versions when required.
4. For Docker, run sudo nvidia-ctk runtime configure --runtime=docker.
5. Apply the configuration with sudo systemctl restart docker during an appropriate maintenance window.
6. Follow the official sample workload with --gpus all and verify GPU visibility inside the container before starting a model image.

사용 정보

When deploying inference, validate the host driver, then GPU access in a minimal container, then mount model storage and start the application. Record the driver, image tag, Toolkit version and device allocation. If the host sees the GPU but the container does not, investigate runtime configuration first. If GPU access works but the model fails, inspect CUDA compatibility, memory and application dependencies.

Implementation notes
Restarting Docker affects running workloads, so schedule the configuration change. Rootless Docker, containerd and CRI-O require their own setup steps rather than the complete standard Docker command sequence.

모델 호환성 및 사용 사례

Requires a supported Linux environment, an NVIDIA GPU driver and a container runtime. The host does not need the CUDA Toolkit installed, but it does need the NVIDIA driver. The container’s CUDA requirements must still be compatible with that driver.

라이선스 및 위험 참고 사항

The repository uses Apache-2.0. GPU drivers, CUDA images and deployed models may have different terms.

vllm

vllm-project/vllm

★ 89.5KPython

sglang

sgl-project/sglang

★ 32.2KPython

OpenVINO

openvinotoolkit/openvino

★ 10.6KC++

GPUStack

gpustack/gpustack

★ 5.4KPython