Launching a series of posts about the evolution of one of the most popular architectures in computer vision.
We'll break down:
Before 2015, the task of detection was solved by searching for the most likely regions. There were two-stage approaches, such as Faster R-CNN.
🟢 How did the YOLO architecture evolve from v1 to v3?
📝 First, they searched for candidate regions, and then used a refine process to refine the classes and coordinates.
PROBLEM: The process was very slow. Imagine the task of tracking a tennis ball on the court during a match. Old networks would have taken 5 minutes to process a video, even on a good GPU. Players would have had to stand and wait for the VAR system.
A real-time approach was needed, where speed was more important than perfect results. Thus, YOLO was born.
➡️ YOLO v1: A model that looks at the entire scene (2015)
The idea was to turn detection from a region-searching task into a regression problem. Combine all stages into a single network that directly "spits out" coordinates.
How it was implemented technically?
📝 They made an architecture similar to GoogLeNet. Two fully connected and 24 convolutional layers. Although it was large, it detected bounding boxes and immediately determined the coordinates.
📝 All images were divided into a 7x7 grid. Each cell predicted 2 bounding boxes and 20 classes. The input was a 448x448 image, which was further divided into 64x64.
PROBLEM: YOLO v1 couldn't handle other resolutions. To work with detection on large images, they resized them to 448x448 or cut them into patches. Due to the extra operations, the main advantage over Faster R-CNN — speed — was lost.
➡️ YOLO v2 / YOLO9000: Scale and anchors (2016–2017)
To level the complex LOSS, multi-scale was added to the new version: YOLO9000 simultaneously detects more than 9,000 classes without full annotation — hence the name.
What new features were added?
📝 Anchor Boxes: Instead of directly predicting coordinates, they switched to predicting shifts relative to the X and Y axes for candidates. This maximized object capture.
📝 Skip Connections: They introduced pass-through layers and added batch normalization, which solved the problem of gradient fading.
PROBLEM: The accuracy of detections became heavily dependent on anchor boxes. The anchors were manually selected, and if they were poorly chosen for the dataset, the model's metrics suffered.
➡️ YOLO v3: Victory over other models (2018)
Thanks to the update, YOLO v3 became a foundation in ML. It surpassed Faster R-CNN in popularity and became a favorite of many developers.
What was added new?
📝 Multiscale detection. It removed noise when detecting small objects and stopped ignoring them.
📝 The "third eye". The network immediately outputted three candidates at different resolutions — large, smaller, and the smallest.
PROBLEM: The version became slower. Due to the complexity of the architecture, v3 became heavier than its predecessors. The anchors were still manually selected, which also slowed down the detection process.
Model continued to evolve, but no longer in the hands of its original author, Joseph Redmon: he left ML and handed over project to a large company.
In next post, we'll break down why YOLO v4 is called the "engineer's constitution" and YOLO v5 is a "ugly duckling"?.
••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
