High-quality resynthesis from coarse neural audio codec tokens

Establish a method for high-quality audio resynthesis from coarse tokens produced by neural audio codecs based on Residual Vector Quantization, thereby improving the fidelity attainable by systems that generate such tokens.

Background

Neural audio codecs based on Residual Vector Quantization represent audio using hierarchical discrete codebook tokens, with early layers encoding coarse structure and later layers encoding progressively finer detail. Resynthesis from only coarse codec tokens requires recovering the missing residual information and is therefore a bottleneck for the fidelity of audio-generation systems that operate over codec tokens.

The paper identifies this problem as unresolved and proposes geometric iterative retrieval as one approach to resynthesis. The method predicts continuous codebook vectors layer by layer, using the RVQ hierarchy as the iterative structure, but the explicit open-problem statement concerns the broader challenge of achieving high-quality resynthesis from coarse neural audio codec tokens.

References

Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every system that generates them.

Geometric Iterative Retrieval for Neural Audio Codec Resynthesis  (2608.19141 - Schmidt-Traub et al., 19 Aug 2026) in Abstract