Contribution of V1 Receptive Field Properties to Corruption Robustness in CNNs
See where this sits in the topic map →Summary AI-generated
- TL;DR
- Simulating early primate visual processing up to the primary visual cortex (V1) at the front of convolutional neural networks (CNNs) improves their robustness to image corruptions.
- Problem
- It remains unclear whether this robustness requires precisely matching the biological receptive field (RF) properties of V1 neurons, or if only certain aspects are sufficient.
- Method
- We build several variants of a CNN model with a V1 front-end—using a Gabor Filter Bank (GFB) and cell nonlinearities—with varying levels of biological detail in how their receptive field properties are sampled.
- Results
- The model variant sampling parameters according to empirical biological distributions was 8.72% more robust to image corruptions than a variant sampling them uniformly and independently. However, a uniform variant that also captured correlations between certain GFB parameters achieved the same performance as the biological sampling variant.
- Contributions
- Demonstrating that capturing observed correlations between receptive field properties is necessary when modeling V1, while fully replicating their exact empirical distributions is not required.
- Limitations
- Not specified in the abstract.
- Takeaways
- Approximating V1 requires including observed parameter correlations rather than just the right class of function and parameter range, but fully replicating empirical distributions is not necessary for robustness.
- Applications
- Not specified in the abstract.
- Topics
- Computer Vision; Convolutional Neural Networks; Biological Vision; Robustness
- For industry
- Not specified in the abstract.
- Why it matters
- Not specified in the abstract.
Abstract
Recently, it has been shown that simulating computations in early primate visual areas, up to the primary visual cortex (V1), at the front of convolutional neural networks (CNNs) leads to improvements in robustness to image corruptions. However, it remains unclear whether this improvement requires precisely matching the receptive field (RF) properties of V1 neurons or if some aspects are sufficient. Here, we explore this question by building several variants of a CNN model with a front-end modeling the primate V1 using a classical neuroscientific model, a Gabor Filter Bank (GFB) followed by simple- and complex- cell nonlinearities. Each model variant had varying levels of biological detail according to how the RF properties were sampled. The model variant sampling these parameters according to empirical biological distributions was considerably more robust to image corruptions than the variant sampling the parameters uniformly and independently (relative difference of 8.72%). However, a uniform variant capturing correlations between some GFB parameters obtained the same performance as the biological sampling variant. Our results show that it is not sufficient to approximate V1 with only the right class of function and parameter range, as it is also required to include the observed correlations between RF properties. However, it is not necessary to fully replicate the empirical distributions of V1 RF properties to obtain the desired improvement in robustness.