Spatial Audio Coding Through Relative Room Impulse Response Estimation

Article | Audio examples | Slides | Bibtex | Acknowledgement

Abstract
Immersive virtual listening relies on spatial audio technologies such as Higher-Order Ambisonics (HOA), which represent sound scenes as multichannel signals. As the desired spatial resolution increases, so does the number of channels, making efficient compression essential for transmission over bandwidth-limited networks. Moreover, to facilitate deployment by network operators, the target bitrate for immersive audio coding should ideally remain close to the 25 kbps currently allocated to VoLTE audio services. State-of-the-art parametric codecs, such as the recently standardized Immersive Voice and Audio Services (IVAS) codec, achieve compression by transmitting spatial metadata together with a reduced number of transport channels. However, recent studies have shown that IVAS performance degrades on reverberant content, particularly at low bitrates, a limitation that suggests its inability to accurately model room acoustics. In this paper, we propose a novel HOA coding scheme based on the explicit and blind estimation of the Relative Spatial Room Impulse Response (ReSRIR), using a beamformed version of the HOA signal as a reference signal. By exploiting the structure and sparsity of the estimated ReSRIR, we derive an efficient parametric representation for immersive audio coding. Experimental evaluations show that the proposed method achieves higher compression than IVAS in the single-transport-channel regime, while maintaining comparable to slightly better quality.
Block diagram of the proposed coding method.

Block diagram of the proposed coding method (ReSRIR Codec)

Audio examples
Below are examples of input and reconstructed 3rd-order Higher-Order Ambisonics (HOA) audio signals rendered binaurally for headphone playback. Our codec is compared against IVAS at two bitrates: 32 kbps and 64 kbps. All signals are synthesized from 40-second long Spatial Room Impulse Responses (SRIRs), which are truncated versions of the original SRIRs from each dataset. Please use headphones for optimal spatial audio perception.

Condition SRIR from Aalto SRIR from IKSLabIR SRIR from Motus SRIR from RSoANU SRIR from 50_Lebedev
Original
ReSRIRCodec @ 27 kbps (ours)
IVAS @ 32 kbps
IVAS @ 64 kbps

Mean QASTANet spatial audio quality scores on tested models (IVAS and proposed model) and on various testsets (each containing 50 samples). The vertical lines denote the standard deviation. Bar width is proportional to bitrate.