Computer Vision

DINOv2: visual features for classification and retrieval

facebookresearch/dinov2

Meta’s self-supervised vision models provide pretrained backbones and task heads for feature extraction, visual retrieval and downstream adaptation.

★ 13.3KStars
⑂ 1.3KForks
300Open issues
PythonLanguage
Apache-2.0License
Q84Editorial score

Overview

DINOv2 encodes images into visual features that can support downstream tasks. A common starting point is to freeze a pretrained backbone, extract embeddings and train a small classifier or build a similarity index. The repository also includes resources for tasks such as depth estimation and semantic segmentation. It is useful for establishing a vision baseline on an existing image collection, with performance assessed on representative application data.

Key features

  • Learns reusable visual representations through self-supervision.
  • Offers ViT-S, B, L and g backbone options.
  • Includes model variants with registers.
  • Loads pretrained models through PyTorch Hub.
  • Supports classification, similarity and feature-extraction workflows.
  • Includes resources for depth and semantic segmentation.

Requirements, installation and quick start

1. Create a compatible Python/PyTorch environment for your device.
2. Start with PyTorch Hub for simple feature extraction.
3. Load the small backbone with torch.hub.load("facebookresearch/dinov2", "dinov2_vits14").
4. Follow the documented preprocessing, using consistent image sizes and normalization, and switch to evaluation mode.
5. Extract features for a small batch with gradients disabled.
6. Install the repository’s task-specific dependencies when moving to training or specialized heads.

Usage

For product-image retrieval, associate each image with a product identifier, extract embeddings and apply consistent normalization before indexing. Test matching products from different angles as well as different products with similar backgrounds. This reveals whether similarity is driven by the item or the scene. For business-specific categories, train a small classifier over frozen features and evaluate on held-out examples.

Implementation notes
Record crop settings, normalization, backbone version and embedding dimensions so indexing and queries use the same pipeline. Medical and cellular extensions have separate model documentation and should not inherit assumptions about the general-purpose models.

Model compatibility and use cases

Requires PyTorch, consistent image preprocessing and model weights. Larger backbones need more resources. Visual similarity does not necessarily match business semantics, so compare models using task-specific retrieval or classification metrics.

License and risk notes

The main DINOv2 repository uses Apache-2.0. The separately listed XRay-DINO weights use the FAIR Noncommercial Research License; the main repository license should not be assumed to cover those weights.

opencv

opencv/opencv

★ 90.7KC++

tesseract

tesseract-ocr/tesseract

★ 76.3KC++

yolov5

ultralytics/yolov5

★ 58KPython

faceswap

deepfakes/faceswap

★ 57.5KPython