ロボティクス&エッジ AI

Isaac GR00T: vision-language-action models for robotics

NVIDIA/Isaac-GR00T

NVIDIA’s robot models and reference code offer multimodal action prediction, demonstration-data adaptation, fine-tuning and deployment tooling.

★ 8.1Kスター
⑂ 1.5Kフォーク数
303未解決の問題
Python言語
Apache-2.0ライセンス
Q92編集部スコア

概要

Isaac GR00T N1.7 combines a vision-language backbone with a network that predicts continuous actions. Inputs include language, camera observations and robot state. The repository provides checkpoints, data-format guidance, inference examples, fine-tuning and deployment tooling. Adapting it to a robot requires mapping the expected observations and actions to the actual sensors and controller.

主な機能

  • Uses multimodal inputs including vision, language and state.
  • Provides N1.7 base and task-adapted checkpoints.
  • Uses a LeRobot-style dataset with additional modality metadata.
  • Includes open-loop inference on demonstration data.
  • Supports post-training for specific embodiments and tasks.
  • Includes server-client inference and export tooling.

要件、インストール、クイックスタート

1. Prepare the platform-specific GPU, CUDA and Python environment; the default dGPU path uses Python 3.12.
2. Run git clone --recurse-submodules https://github.com/NVIDIA/Isaac-GR00T.
3. Install platform dependencies and run uv sync --python 3.12.
4. Obtain Hugging Face access to the gated Cosmos-Reason2-2B backbone and authenticate.
5. Follow the standalone_inference_script.py example using nvidia/GR00T-N1.7-3B, demo_data/droid_sample and the matching embodiment tag.
6. Inspect predicted sample trajectories before converting your own data and fine-tuning.

使用方法

For an arm-manipulation task, align camera views, joint or end-effector state, action coordinates and sampling frequency. Prepare demonstrations with modality.json, compare predictions on offline trajectories, then evaluate in simulation. Physical deployment should begin with controlled motions under the existing controller’s speed and workspace limits. Match the embodiment tag to both the checkpoint and data.

Implementation notes
Good open-loop predictions do not establish closed-loop task success. Sensor latency, control frequency and action units affect execution, so record offline, simulation and physical evaluations separately.

モデルの互換性とユースケース

Upstream lists 16GB+ VRAM for inference and recommends 40GB+ for fine-tuning. Platforms have different CUDA and dependency combinations. The default dGPU video stack requires a supported FFmpeg version, and the gated backbone needs separate access approval.

ライセンスとリスクに関する注意事項

The N1.7 main project is released under Apache-2.0. Check the access conditions and terms for downloaded backbones, datasets and dependencies.

mediapipe

google-ai-edge/mediapipe

★ 36.9KC++

cosmos

NVIDIA/cosmos

★ 11.2KJupyter Notebook