preprint · arXiv (Cornell University) · 2024

Explicitly Modeling Pre-Cortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness

Lisa Piper, Arlindo L. Oliveira, Tiago Marques · 0 citations

View original publication →

See where this sits in the topic map →

Summary AI-generated

TL;DR
Convolutional neural networks struggle with corrupted images in the real world, but taking inspiration from human biology can make them more robust.
Problem
While convolutional neural networks excel at clean image classification, they struggle to classify images with common corruptions, which limits their real-world applicability.
Method
The authors introduce two novel biologically-inspired CNN model families: RetinaNet, which combines a pre-cortical front-end with a standard back-end, and EVNet, which adds a V1 visual cortex block after the front-end.
Results
RetinaNet achieved a relative robustness improvement of 12.3% compared to standard models, while EVNet achieved an 18.5% relative gain across all corruption categories, though with a small decrease in clean image accuracy.
Contributions
The work expands on prior V1-inspired approaches by designing a new front-end block that simulates pre-cortical visual processing.
Limitations
The approach is accompanied by a small decrease in clean image accuracy.
Takeaways
Simulating multiple stages of early visual processing in the early layers of CNNs provides cumulative benefits for model robustness.
Applications
Not specified in the abstract.
Topics
Computer Vision; Deep Learning; Neural Network Robustness; Bio-inspired AI
For industry
Not specified in the abstract.
Why it matters
Not specified in the abstract.

Abstract

While convolutional neural networks (CNNs) excel at clean image classification, they struggle to classify images corrupted with different common corruptions, limiting their real-world applicability. Recent work has shown that incorporating a CNN front-end block that simulates some features of the primate primary visual cortex (V1) can improve overall model robustness. Here, we expand on this approach by introducing two novel biologically-inspired CNN model families that incorporate a new front-end block designed to simulate pre-cortical visual processing. RetinaNet, a hybrid architecture containing the novel front-end followed by a standard CNN back-end, shows a relative robustness improvement of 12.3% when compared to the standard model; and EVNet, which further adds a V1 block after the pre-cortical front-end, shows a relative gain of 18.5%. The improvement in robustness was observed for all the different corruption categories, though accompanied by a small decrease in clean image accuracy, and generalized to a different back-end architecture. These findings show that simulating multiple stages of early visual processing in CNN early layers provides cumulative benefits for model robustness.

← All publications