Development ceased to be purely conceptual and became more engineering-oriented:
🟢 Review of YOLO v4-v6
➡️ YOLO v4: Model turned into an engineering encyclopedia (2020)
YOLO v4 became a "BIBLE" for improving architectures. It packed in as many tricks as possible without killing FPS.
GOLDEN FEATURE: In new version, mosaic augmentation was introduced. It collects training picture from several different ones, which improves model's performance. As a result, the quality was improved by +6% mAP compared to YOLOv3, while maintaining a speed of 60 FPS.
OTHER CHANGES: Pyramidal architecture (CSPDarknet-53 + PANet + SPP). Instead of simply cutting out pieces from picture, a multi-scale approach was implemented at level of network itself. Network itself extracted features of different scales and recognized contexts.
TRICKS & AUGMENTATIONS. Architecture integrated such developments as Mish-activation, DropBlock, and CloU loss. Together with mosaic augmentation, they improved model's quality by 10% without drastically changing it.
DOWNSIDES of YOLO v4 include difficulty of integrating model and manual hyperparameters left over from previous versions.
There are no more problems to fix, so developers focused on improvements.
➡️ YOLO v5: "Ugly Duckling" and mass adoption (2020-2021)
YOLO v5 was released four months after v4 - version was nicknamed "Ugly Duckling", because there were no architectural breakthroughs in it.
GOLDEN FEATURE: YOLO v5 was rewritten in PyTorch and made it more user-friendly. Everyone could integrate it into their project and retrain it for their own tasks. PyTorch soon gained popularity and dominated the DL field, which led to mass adoption of YOLO.
There weren't many other features - they were released to promote article about the new version. But there were a lot of problems:
📎 Version didn't work due to bugs. For first two months, the buggy implementation simply didn't allow to use model. Memory leaks, incorrectly specified areas for three candidates.
📎 Version didn't add anything new. Each new YOLO either solved an engineering problem or an idea problem. Fifth model was considered a rewrite of what already existed - just on a different framework. Community didn't like this approach.
📎 Version was developed by Ultralytics. Community was wary of it: previously, YOLO was developed by a CIS superstar in the CV field - Bachkovsky and now it's some no-names. So developers were worried about fate of beloved model.
📎 Version never got an article. Company promised to release it within a few months. But it's been four years - Article hasn't appeared. They just released a couple of technical reports on archive.
Fortunately, Ultralytics didn't abandon model and kept improving and enhancing it. Thanks to PyTorch and support from developers, YOLO v5 is widely used as a component of a comprehensive solution.
➡️ YOLO v6: Model was made more convenient for deployment (2022)
Company focused on developing most convenient real-time deployment for frameworks like TensorRT and Edge devices.
GOLDEN FEATURE: An Anchor-Free Head was introduced. Instead of predicting shifts for candidates, Model searches for exact center of object. It's faster and more accurate.
OTHER INNOVATIONS: New architecture. EfficientRep, an analogue of EfficientNet, was chosen as the backbone. They also abandoned DarkNet backbone - it was outdated.
HIGH SPEED. Model became super-lightweight and demonstrated 120 FPS on a T4 at a resolution of 640x640. Therefore, it was used in tasks related to thermal imagers and Edge computing.
There were no obvious downsides or problems with the model. Except for the accuracy compared to v5 and v7. But v6 is best for Edge devices.
In next post, we'll discuss at Why YOLO v8 became the most popular model in the family? and
How commercialization turned the project into a conveyor?
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
