Robótica y Edge AI

Isaac GR00T: vision-language-action models for robotics

NVIDIA/Isaac-GR00T

NVIDIA’s robot models and reference code offer multimodal action prediction, demonstration-data adaptation, fine-tuning and deployment tooling.

★ 8,1KEstrellas
⑂ 1,5KBifurcaciones
303Problemas abiertos
PythonIdioma
Apache-2.0Licencia
Q92Puntuación editorial

Resumen

Isaac GR00T N1.7 combines a vision-language backbone with a network that predicts continuous actions. Inputs include language, camera observations and robot state. The repository provides checkpoints, data-format guidance, inference examples, fine-tuning and deployment tooling. Adapting it to a robot requires mapping the expected observations and actions to the actual sensors and controller.

Características principales

  • Uses multimodal inputs including vision, language and state.
  • Provides N1.7 base and task-adapted checkpoints.
  • Uses a LeRobot-style dataset with additional modality metadata.
  • Includes open-loop inference on demonstration data.
  • Supports post-training for specific embodiments and tasks.
  • Includes server-client inference and export tooling.

Requisitos, instalación y guía rápida

1. Prepare the platform-specific GPU, CUDA and Python environment; the default dGPU path uses Python 3.12.
2. Run git clone --recurse-submodules https://github.com/NVIDIA/Isaac-GR00T.
3. Install platform dependencies and run uv sync --python 3.12.
4. Obtain Hugging Face access to the gated Cosmos-Reason2-2B backbone and authenticate.
5. Follow the standalone_inference_script.py example using nvidia/GR00T-N1.7-3B, demo_data/droid_sample and the matching embodiment tag.
6. Inspect predicted sample trajectories before converting your own data and fine-tuning.

Uso

For an arm-manipulation task, align camera views, joint or end-effector state, action coordinates and sampling frequency. Prepare demonstrations with modality.json, compare predictions on offline trajectories, then evaluate in simulation. Physical deployment should begin with controlled motions under the existing controller’s speed and workspace limits. Match the embodiment tag to both the checkpoint and data.

Implementation notes
Good open-loop predictions do not establish closed-loop task success. Sensor latency, control frequency and action units affect execution, so record offline, simulation and physical evaluations separately.

Compatibilidad de modelos y casos de uso

Upstream lists 16GB+ VRAM for inference and recommends 40GB+ for fine-tuning. Platforms have different CUDA and dependency combinations. The default dGPU video stack requires a supported FFmpeg version, and the gated backbone needs separate access approval.

Notas sobre la licencia y los riesgos

The N1.7 main project is released under Apache-2.0. Check the access conditions and terms for downloaded backbones, datasets and dependencies.

mediapipe

google-ai-edge/mediapipe

★ 36,9KC++

cosmos

NVIDIA/cosmos

★ 11,2KJupyter Notebook