ConvRNN-T: Convolutional augmented recurrent neural network transducers for streaming speech recognition

Martin Radfar; Rohit Barnwal; Rupak Vignesh Swaminathan; Feng-Ju (Claire) Chang; Grant Strimel; Nathan Susanj; Thanasis Mouchtaris

Publication

ConvRNN-T: Convolutional augmented recurrent neural network transducers for streaming speech recognition

By Martin Radfar, Rohit Barnwal, Rupak Vignesh Swaminathan, Feng-Ju (Claire) Chang, Grant Strimel, Nathan Susanj, Thanasis Mouchtaris

2022

Download Copy BibTeX

Share

Download

Copy BibTeX

Share

The recurrent neural network transducer (RNN-T) is a prominent streaming end-to-end (E2E) ASR technology. In RNN-T, the acoustic encoder commonly consists of stacks of LSTMs. Very recently, as an alternative to LSTM layers, the Conformer architecture was introduced where the encoder of RNN-T is replaced with a modified Transformer encoder composed of convolutional layers at the frontend and between attention layers. In this paper, we introduce a new streaming ASR model, Convolutional Augmented Recurrent Neural Network Transducers (ConvRNN-T) in which we augment the LSTM-based RNNT with a novel convolutional front end consisting of local and global context CNN encoders. ConvRNN-T takes advantage of causal 1-D convolutional layers, squeeze-and-excitation, dilation, and residual blocks to provide both global and local audio context representation to LSTM layers. We show ConvRNNT outperforms RNN-T, Conformer, and ContextNet on Librispeech and in-house data. In addition, ConvRNN-T offers less computational complexity compared to Conformer. ConvRNNT’s superior accuracy along with its low footprint make it a promising candidate for on-device streaming ASR technologies.

ConvRNN-T: Convolutional augmented recurrent neural network transducers for streaming speech recognition

Latest news

Work with us