journal · Discover Applied Sciences · 2024

Training environmental sound classification models for real-world deployment in edge devices

Manuel Goulão, L. Bandeira, Bruno Martins, Arlindo L. Oliveira · 12 citations

View original publication →

See where this sits in the topic map →

Summary AI-generated

TL;DR
As smart city technologies grow, extracting useful information from urban sensor data—such as detecting vehicles and sirens using sound sensors—has become a major research challenge.
Problem
Classifying sounds in street environments is complex because sound quality is heavily affected by changing weather, traffic volume, and microphone variations.
Method
We present a deep learning model for multi-label sound classification designed for real-world deployment on edge devices, utilizing a training methodology that includes knowledge distillation pre-training alongside careful data collection and preparation.
Results
When benchmarked on the ESC-50 dataset, our model achieves an 85.4% accuracy, performing comparably to state-of-the-art models that require significantly more computational resources.
Contributions
We describe two key components for real-world edge deployment: data collection and preparation, and a training methodology incorporating knowledge distillation pre-training.
Limitations
Real-world evaluations using early edge-integrated luminaire prototypes show that while the approach works well for most vehicles, it faces significant limitations in classifying 'person' and 'bicycle' sounds.
Takeaways
All results demonstrate the great benefits of pretraining the model using knowledge distillation for edge-device deployment.
Applications
Smart city sound monitoring using edge devices integrated into street luminaires to detect vehicles and sirens.
Topics
Smart cities, deep learning, environmental sound classification, edge devices, knowledge distillation
For industry
Smart city infrastructure and urban sensor technology
Why it matters
Not specified in the abstract.

Abstract

Abstract The interest in smart city technologies has grown in recent years, and a major challenge is to develop methods that can extract useful information from data collected by sensors in the city. One possible scenario is the use of sound sensors to detect passing vehicles, sirens, and other sounds on the streets. However, classifying sounds in a street environment is a complex task due to various factors that can affect sound quality, such as weather, traffic volume, and microphone quality. This paper presents a deep learning model for multi-label sound classification that can be deployed in the real world on edge devices. We describe two key components, namely data collection and preparation, and the methodology to train the model including a pre-train using knowledge distillation. We benchmark our models on the ESC-50 dataset and show an accuracy of 85.4%, comparable to similar state-of-the-art models requiring significantly more computational resources. We also evaluated the model using data collected in the real world by early prototypes of luminaires integrating edge devices, with results showing that the approach works well for most vehicles but has significant limitations for the classes “person” and “bicycle”. Given the difference between the benchmarking and the real-world results, we claim that the quality and quantity of public and private data for this type of task is the main limitation. Finally, all results show great benefits in pretraining the model using knowledge distillation.

← All publications