Visi komputer

Segment Anything: Generate Image Segmentation Masks with Points and Boxes

facebookresearch/segment-anything

Meta’s SAM image segmentation project supports prompt inputs such as points and boxes, and can also automatically generate candidate masks for an entire image. It is suitable for interactive annotation and image processing workflows.

★ 54,8KBintang
⑂ 6,4KFork
596Isu terbuka
PythonBahasa
Apache-2.0Lisensi
Q84Skor editorial

Ringkasan

The core output of Segment Anything is a mask, which marks which pixels in an image belong to a certain region. Users can provide prompt points or bounding boxes on an object, allowing the model to generate candidate regions, and then select and adjust them according to the actual use case. Another entry point automatically searches the entire image for candidate regions, making it suitable for generating an initial annotation draft. In product images, asset organization, and data annotation, SAM can handle the region-selection step. The application continues to process subsequent category naming, background replacement, quality checks, and data export. Designing these steps separately makes it easier to determine whether an issue lies in target selection, edge quality, or output format.

Fitur utama

  • SamPredictor accepts prompts such as points and boxes to generate segmentation masks for the current image.
  • SamAutomaticMaskGenerator generates multiple candidate regions for an entire image.
  • Provides entry points for corresponding model checkpoints such as ViT-B, ViT-L, and ViT-H.
  • Retains image embeddings, making it convenient to repeatedly adjust prompts for the same image.
  • Provides examples for Notebooks, command-line batch processing, and mask output.
  • Supports ONNX export of a lightweight mask decoder and provides a browser demo.

Persyaratan, instalasi, dan mulai cepat

Prepare compatible PyTorch and TorchVision in an independent Python environment, then run pip install git+https://github.com/facebookresearch/segment-anything.git. Download the selected checkpoint according to the repository’s model list, and save the model type and file path.

First use the predictor_example Notebook, load a clear image, and provide a target prompt. Install the corresponding optional dependencies only when image reading, COCO-format processing, or ONNX export is needed. The model type must match the checkpoint, and the GPU configuration should also correspond to the PyTorch installation method.

Penggunaan

Practical case: creating region annotations for a batch of product images.

1. Select a small number of samples with different backgrounds and object sizes, and agree on whether the entire product or a particular component should be retained.
2. After loading an image, call set_image and add positive prompt points for the target; if the mask includes the background, add prompts to exclude the background.
3. Compare the edges, holes, and small components of the candidate masks, and select the result that matches the annotation objective.
4. Add the manually confirmed category and object ID to the selected region.
5. Export a binary mask or annotation format, and record the original image dimensions and coordinate system.
6. Before batch processing, read back the exported files and overlay them on the original images to check whether scaling, orientation, and IDs are consistent.

Interactive annotation is suitable for confirming targets one by one; automatic masks are suitable for first obtaining a candidate set and then filtering it according to area, overlap, and task rules.

How it works
The image encoder first converts the image into a feature representation. The prompt encoder processes points, boxes, or existing masks, and the mask decoder combines both to predict regions. When adjusting prompts for the same image, the image features can be reused. The automatic mask workflow samples prompts in the image and filters candidate results, after which the application organizes the output according to its own annotation rules.

Who it is for
Computer vision developers, data annotation teams, and product teams that need to add interactive region-selection functionality to image editors.

Environment and inputs
Python, PyTorch, TorchVision, model checkpoints, and image-reading tools. A GPU can be used for batch processing and large models; before running, allocate GPU memory according to the input resolution and model selection.

Practical use cases
• Annotation acceleration: generate region drafts first, then manually confirm categories and edges.
• Image editing: pass selected regions to background replacement, local enhancement, or cropping steps.
• Data organization: export region areas, positions, and masks to prepare inputs for subsequent vision tasks.

Implementation notes
Small objects, transparent objects, thin lines, and complex occlusions require category-by-category inspection. Masks represent region extents; category semantics need to be supplemented by another model or by humans. The ONNX export example exports the mask decoder; a browser application also needs to arrange the generation and transmission of image features.

Common questions
Q: Can it directly produce object names?
The primary result of SAM is a region mask. Category names must be supplemented by your annotation workflow or another recognition component.

Q: Why do multiple masks sometimes appear for the same target?
A prompt may correspond to regions at different granularities. Selection can be based on the target definition, candidate scores, and edge inspection.

Related projects and workflow ideas
facebookresearch/sam2: Suitable for comparing the image and video segmentation capabilities of the later model. The model files and interfaces of the two projects should be configured separately.

google-ai-edge/mediapipe: Suitable for developers with existing on-device vision tasks to compare deployment methods and task interfaces. Choose the execution path according to the specific platform.

Kompatibilitas model dan kasus penggunaan

Use the model architecture and checkpoint corresponding to SAM. SAM 2 has a separate repository and interface; for video tasks, consult the SAM 2 documentation before deciding whether to migrate the workflow.

Catatan lisensi dan risiko

The model and project use the Apache-2.0 license. The SA-1B dataset uses a separate dataset research license; when downloading the data, follow the instructions on the dataset page.

opencv

opencv/opencv

★ 90,7KC++

tesseract

tesseract-ocr/tesseract

★ 76,3KC++

yolov5

ultralytics/yolov5

★ 58KPython

faceswap

deepfakes/faceswap

★ 57,5KPython