Environmental sound analysis with mixup based multitask learning and cross-task fusion

03/30/2021
by   Weiping Zheng, et al.
0

Environmental sound analysis is currently getting more and more attentions. In the domain, acoustic scene classification and acoustic event classification are two closely related tasks. In this letter, a two-stage method is proposed for the above tasks. In the first stage, a mixup based MTL solution is proposed to classify both tasks in one single convolutional neural network. Artificial multi-label samples are used in the training of the MTL model, which are mixed up using existing single-task datasets. The multi-task model obtained can effectively recognize both the acoustic scenes and events. Compared with other methods such as re-annotation or synthesis, the mixup based MTL is low-cost, flexible and effective. In the second stage, the MTL model is modified into a single-task model which is fine-tuned using the original dataset corresponding to the specific task. By controlling the frozen layers carefully, the task-specific high level features are fused and the performance of the single classification task is further improved. The proposed method has confirmed the complementary characteristics of acoustic scene and acoustic event classifications. Finally, enhanced by ensemble learning, a satisfactory accuracy of 84.5 percent on TUT acoustic scene 2017 dataset and an accuracy of 77.5 percent on ESC-50 dataset are achieved respectively.

READ FULL TEXT

page 1

page 2

research
04/05/2022

How Information on Acoustic Scenes and Sound Events Mutually Benefits Event Detection and Scene Classification Tasks

Acoustic scene classification (ASC) and sound event detection (SED) are ...
research
09/05/2018

CNNs-based Acoustic Scene Classification using Multi-Spectrogram Fusion and Label Expansions

Spectrograms have been widely used in Convolutional Neural Networks base...
research
06/21/2022

Joint Analysis of Acoustic Scenes and Sound Events Based on Multitask Learning with Dynamic Weight Adaptation

Acoustic scene classification (ASC) and sound event detection (SED) are ...
research
02/28/2023

Incremental Learning of Acoustic Scenes and Sound Events

In this paper, we propose a method for incremental learning of two disti...
research
05/02/2019

City classification from multiple real-world sound scenes

The majority of sound scene analysis work focuses on one of two clearly ...
research
11/26/2018

Scene Recognition Through Visual and Acoustic Cues Using K-Means

We propose a K-Means based prediction system, nicknamed SERVANT (Scene R...
research
05/21/2020

Team Neuro at SemEval-2020 Task 8: Multi-Modal Fine Grain Emotion Classification of Memes using Multitask Learning

In this article, we describe the system that we used for the memotion an...

Please sign up or login with your details

Forgot password? Click here to reset