Coherent Audio Generation Beyond the Training Duration Range

Generate coherent audio scenes with MiDashengLM-Gen beyond the 1–20-second duration range represented in its training data.

Background

MiDashengLM-Gen supports variable-length autoregressive mixed-audio generation, but its training examples are limited to durations between 1 and 20 seconds. The authors explicitly identify the generation of coherent audio beyond this duration range as unresolved, motivating future work on scaling the system to longer durations.

References

Variable-length generation is bounded by the training data distribution (1--20 seconds); generating coherent audio beyond this range remains open.

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching  (2608.11804 - Sun et al., 12 Aug 2026) in Section Discussion and Conclusion, subsection “Limitations”