Explicitly Modeling Pre-Cortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness
See where this sits in the topic map →Summary AI-generated
- TL;DR
- Convolutional neural networks struggle with corrupted images in the real world, but taking inspiration from human biology can make them more robust.
- Problem
- While convolutional neural networks excel at clean image classification, they struggle to classify images with common corruptions, which limits their real-world applicability.
- Method
- The authors introduce two novel biologically-inspired CNN model families: RetinaNet, which combines a pre-cortical front-end with a standard back-end, and EVNet, which adds a V1 visual cortex block after the front-end.
- Results
- RetinaNet achieved a relative robustness improvement of 12.3% compared to standard models, while EVNet achieved an 18.5% relative gain across all corruption categories, though with a small decrease in clean image accuracy.
- Contributions
- The work expands on prior V1-inspired approaches by designing a new front-end block that simulates pre-cortical visual processing.
- Limitations
- The approach is accompanied by a small decrease in clean image accuracy.
- Takeaways
- Simulating multiple stages of early visual processing in the early layers of CNNs provides cumulative benefits for model robustness.
- Applications
- Not specified in the abstract.
- Topics
- Computer Vision; Deep Learning; Neural Network Robustness; Bio-inspired AI
- For industry
- Not specified in the abstract.
- Why it matters
- Not specified in the abstract.
Abstract
While convolutional neural networks (CNNs) excel at clean image classification, they struggle to classify images corrupted with different common corruptions, limiting their real-world applicability. Recent work has shown that incorporating a CNN front-end block that simulates some features of the primate primary visual cortex (V1) can improve overall model robustness. Here, we expand on this approach by introducing two novel biologically-inspired CNN model families that incorporate a new front-end block designed to simulate pre-cortical visual processing. RetinaNet, a hybrid architecture containing the novel front-end followed by a standard CNN back-end, shows a relative robustness improvement of 12.3% when compared to the standard model; and EVNet, which further adds a V1 block after the pre-cortical front-end, shows a relative gain of 18.5%. The improvement in robustness was observed for all the different corruption categories, though accompanied by a small decrease in clean image accuracy, and generalized to a different back-end architecture. These findings show that simulating multiple stages of early visual processing in CNN early layers provides cumulative benefits for model robustness.