In 2026, most computer vision projects come down to two choices. You pick the model family you build on, and then you decide how much glue code you are willing to write around it. The libraries below cover both halves. A few of them shipped major releases this year: Ultralytics put out YOLO26 in January, OpenCV hit 5.0 in June, and Hugging Face Transformers moved to its v5 line. Some older favourites have quietly stalled, and a stalled library is a real cost when you are the one maintaining the pipeline.
We ranked each library on four things: how actively it is maintained, how good the docs are, how clear the licence is for commercial use, and how short the path is from a trained model to a running deployment. Affiliate deals had no effect on the order. Most of these libraries are free and open source anyway.
1. Hugging Face Transformers: 9.2/10
Hugging Face Transformers is where most new vision research models show up first. That covers DETR-family detectors, DINOv2 and SigLIP backbones, and a steady stream of vision-language models. The v5 line cleaned up years of drift, so multimodal models now take a single processor object, and a lot of legacy APIs are gone. Fine-tuning a transformer detector on custom data takes a few dozen lines with the Trainer, and the Hub gives you thousands of checkpoints to start from. The weak spots are speed and churn. For real-time detection on a CPU or a Jetson, a YOLO model will beat most Transformers checkpoints on latency. Minor releases also arrive often enough that you should pin your version. Check each model card for the licence, because the library is Apache 2.0 but the weights can come with anything.
- Best for: using recent research models, VLM-based vision tasks, and transfer learning from strong pretrained backbones.
- Pricing: the library is free under Apache 2.0. Hub hosting and Inference Endpoints are billed separately.
2. OpenCV 5: 8.8/10
Opencv is the library every other entry on this list quietly depends on for camera I/O, colour conversion, resizing and geometry. Version 5.0.0 shipped on 6 June 2026, and it is the biggest update in years. The DNN module has been rewritten, ONNX operator coverage is now above 80%, and there is built-in support for running LLMs and vision-language models, along with a much better 3D vision toolkit. The hardware layer has tuned paths for Intel IPP, Arm KleidiCV, Qualcomm FastCV and RISC-V vector extensions, which is why OpenCV is still the default for embedded work. The upgrade costs something too. C++17 is now the minimum, and the legacy C API is gone, so older codebases that still call cvLoadImage-era functions will need porting. As a model runner the DNN module works well enough, but it is not where you train anything.
- Best for: preprocessing, classical vision, calibration and 3D work, and running ONNX models on constrained devices.
- Pricing: free under the Apache 2.0 licence.
3. Roboflow Supervision: 8.6/10
Supervision fixes the dull half of every vision project, which is everything that happens after the model returns its predictions. It turns Ultralytics, Transformers, SAM and other model outputs into one Detections object. From there you get box and mask annotators, ByteTrack tracking, zone and line-crossing counters, and conversion between COCO, YOLO and Pascal VOC dataset formats. The package is MIT-licensed and model-agnostic, and it cuts a surprising amount of hand-written drawing and counting code. It will not train or run models for you. Its value depends on pairing it with one of the entries above.
- Best for: video analytics, object counting, annotation and format conversion around any detector.
- Pricing: the library is free under MIT. The Roboflow platform has a free Public plan with $60 a month in credits, and Core costs $79 a month billed annually or $99 billed monthly. Enterprise is custom.
4. Meta SAM 3
Sam 3 (Segment Anything Model 3) came out on 19 November 2025 and changed what zero-shot segmentation can do. SAM 1 and 2 segmented one object per click or box. SAM 3 takes a short noun phrase such as 'yellow school bus', or an example image, and returns a mask and an ID for every matching instance in an image or video. Meta also released the SA-Co benchmark with it, covering more than 200,000 concepts, and the later SAM 3.1 update made video tracking faster. The most practical use today is auto-labelling. You can pre-annotate a dataset with SAM 3, fix the misses by hand, and then train a small YOLO model for production. SAM 3 is too heavy to be that production model on most edge hardware. The custom SAM License allows commercial use but rules out military, ITAR, nuclear and weapons applications, so defence-adjacent teams should read it first.
- Best for: open-vocabulary segmentation, dataset labelling, and video object tracking where GPU budget is not tight.
- Pricing: free code and checkpoints under the SAM License, with the use restrictions listed above.
5. torchvision: 7.8/10
torchvision ships with PyTorch and is the sensible base for a custom training loop. The v2 transforms apply augmentations to images, bounding boxes and masks together, which removes a whole class of silent label bugs. It also includes datasets, ops such as NMS and RoIAlign, and pretrained ResNet, ViT and ConvNeXt backbones. The detection model zoo (Faster R-CNN, Mask R-CNN, RetinaNet) now looks dated next to YOLO26 or current DETR variants, so treat torchvision as a toolkit and look elsewhere for the latest detectors.
- Best for: research code, custom architectures, and teams who want full control of the training loop.
- Pricing: free under BSD-3-Clause.
6. Ultralytics YOLO (YOLO26): 7.6/10
Ultralytics is still the fastest way to go from a labelled dataset to a working detector. YOLO26 launched on 14 January 2026 and removes non-maximum suppression. The model is trained to output one box per object, so the export no longer carries a post-processing step that breaks differently on every runtime. Ultralytics claims up to 43% faster CPU inference, and in practice the gain matters most on small edge boxes and Raspberry Pi-class hardware. One pip install covers detection, segmentation, pose, oriented boxes and classification, and a single export call writes ONNX, TensorRT, CoreML or TFLite. The catch is the licence. The code and weights are AGPL-3.0, so if you ship YOLO inside a closed-source product you need an Enterprise License, and Ultralytics does not publish a price for it. The high-level API also hides a lot of detail, and you will end up reading the source once you need to change the loss or the data loader.
- Best for: real-time detection and segmentation on edge hardware, and teams that want a trained model this week.
- Pricing: free under AGPL-3.0. The Enterprise License for closed-source commercial use is quoted per customer. The hosted Ultralytics Platform bills training on credits.
7. Kornia: 7.4/10
Kornia reimplements classical computer vision as differentiable PyTorch operations, including geometry, filtering, colour spaces and homographies. It also runs GPU-batched augmentations and includes feature matchers such as LoFTR and LightGlue. When you need gradients to flow through a warp, or want augmentation off the CPU so the data loader stops holding back your GPU, nothing else does the job as cleanly. It is a narrow tool, the docs are thinner than OpenCV's, and most application projects will never need it.
- Best for: differentiable geometry, GPU augmentation pipelines, and feature matching research.
- Pricing: free under Apache 2.0.
What we left out
Detectron2 and MMDetection shaped the field for years, but releases for both have slowed sharply and their dependency pins fight with current PyTorch builds. They are fine for reproducing an older paper. For a new project, start with one of the seven above.
Comparison table
| Library | Score | Licence | Best for | Cost |
|---|---|---|---|---|
| Ultralytics YOLO26 | – | AGPL-3.0 / Enterprise | Real-time edge detection | Free; Enterprise quoted |
| OpenCV 5 | 8.8 | Apache 2.0 | Preprocessing, embedded, 3D | Free |
| Hugging Face Transformers | 9.2 | Apache 2.0 (weights vary) | Research models, VLMs | Free; hosting extra |
| Meta SAM 3 | – | SAM License | Zero-shot segmentation, labelling | Free |
| Roboflow Supervision | 8.6 | MIT | Tracking, counting, annotation | Free; platform from $79/mo |
| torchvision | 7.8 | BSD-3-Clause | Custom training loops | Free |
| Kornia | 7.4 | Apache 2.0 | Differentiable CV, GPU augmentation | Free |
Final picks
For a production detector on edge hardware, use Ultralytics with YOLO26, and budget for the Enterprise License if your product is closed source. Put Opencv underneath it for camera input and preprocessing, and Supervision on top for tracking and overlays. That three-library stack covers most commercial vision work we see.
If you have no labelled data yet, run Sam 3 over your raw images to build a first dataset, then train a small model from it. If AGPL is off the table and you need a permissive licence end to end, use a DETR-family model from Hugging Face Transformers and accept some extra latency. Keep torchvision and Kornia for the cases where you are writing the model yourself.