Übersicht
DINOv2 encodes images into visual features that can support downstream tasks. A common starting point is to freeze a pretrained backbone, extract embeddings and train a small classifier or build a similarity index. The repository also includes resources for tasks such as depth estimation and semantic segmentation. It is useful for establishing a vision baseline on an existing image collection, with performance assessed on representative application data.
Wichtige Funktionen
- Learns reusable visual representations through self-supervision.
- Offers ViT-S, B, L and g backbone options.
- Includes model variants with registers.
- Loads pretrained models through PyTorch Hub.
- Supports classification, similarity and feature-extraction workflows.
- Includes resources for depth and semantic segmentation.
Voraussetzungen, Installation und Schnellstart
2. Start with PyTorch Hub for simple feature extraction.
3. Load the small backbone with torch.hub.load("facebookresearch/dinov2", "dinov2_vits14").
4. Follow the documented preprocessing, using consistent image sizes and normalization, and switch to evaluation mode.
5. Extract features for a small batch with gradients disabled.
6. Install the repository’s task-specific dependencies when moving to training or specialized heads.
Nutzung
Implementation notes
Record crop settings, normalization, backbone version and embedding dimensions so indexing and queries use the same pipeline. Medical and cellular extensions have separate model documentation and should not inherit assumptions about the general-purpose models.
Modellkompatibilität und Anwendungsfälle
Requires PyTorch, consistent image preprocessing and model weights. Larger backbones need more resources. Visual similarity does not necessarily match business semantics, so compare models using task-specific retrieval or classification metrics.
Lizenz- und Risikohinweise
The main DINOv2 repository uses Apache-2.0. The separately listed XRay-DINO weights use the FAIR Noncommercial Research License; the main repository license should not be assumed to cover those weights.