Multi-scale aggregation of phase information for reducing computational cost of CNN based DOA estimation

by   Soumitro Chakrabarty, et al.

In a recent work on direction-of-arrival (DOA) estimation of multiple speakers with convolutional neural networks (CNNs), the phase component of short-time Fourier transform (STFT) coefficients of the microphone signal is given as input and small filters are used to learn the phase relations between neighboring microphones. Due to this chosen filter size, M-1 convolution layers are required to achieve the best performance for a microphone array with M microphones. For arrays with large number of microphones, this requirement leads to a high computational cost making the method practically infeasible. In this work, we propose to use systematic dilations of the convolution filters in each of the convolution layers of the previously proposed CNN for expansion of the receptive field of the filters to reduce the computational cost of the method. Different strategies for expansion of the receptive field of the filters for a specific microphone array are explored. With experimental analysis of the different strategies, it is shown that an aggressive expansion strategy results in a considerable reduction in computational cost while a relatively gradual expansion of the receptive field exhibits the best DOA estimation performance along with reduction in the computational cost.


page 1

page 2

page 3

page 4


Accelerating Large-Kernel Convolution Using Summed-Area Tables

Expanding the receptive field to capture large-scale context is key to o...

ASCNet: Adaptive-Scale Convolutional Neural Networks for Multi-Scale Feature Learning

Extracting multi-scale information is key to semantic segmentation. Howe...

PSDNet and DPDNet: Efficient channel expansion, Depthwise-Pointwise-Depthwise Inverted Bottleneck Block

In many real-time applications, the deployment of deep neural networks i...

Cascaded multi-scale and multi-dimension convolutional neural network for stereo matching

Convolutional neural networks(CNN) have been shown to perform better tha...

Broadband DOA estimation using Convolutional neural networks trained with noise signals

A convolution neural network (CNN) based classification method for broad...

Should You Go Deeper? Optimizing Convolutional Neural Network Architectures without Training by Receptive Field Analysis

Applying artificial neural networks (ANN) to specific tasks, researchers...

Direction of Arrival Estimation of Sound Sources Using Icosahedral CNNs

In this paper, we present a new model for Direction of Arrival (DOA) est...

Please sign up or login with your details

Forgot password? Click here to reset