Symbolic Planning Lifts YuE2’s Full-Song Musicality
A shared autoregressive–non-autoregressive model writes lead sheets before audio tokens, raising expert overall-quality preferences to 49.3% versus 34.6% without planning.
Underlying Paper
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning. A single AR-NAR Mixture-of-Transformers (MoT) first writes a readable score specifying melody and harmony, expands it into semantic music tokens, and realizes it as full-song audio. In comparisons using the same checkpoint, experts prefer symbolic planning for overall quality and musicality, with 49.3% of overall preferences versus 34.6% without planning. Experts also favor the unified model over a separate language model and diffusion Transformer. On WildSongBench, YuE2 scores 6.73 on SongBench Global Avg, exceeding all evaluated public baselines. Selecting from eight candidates (best-of-8), YuE2 reaches 6.96, the highest observed mean among all evaluated systems. Expert listening further establishes its competitiveness with proprietary song generators, favoring best-of-8 over Suno v4.5 and yielding nearly balanced preferences against Suno v5. To learn this generation process from recordings without aligned scores, we introduce MERT2 and SheetSage2 to supply semantic and symbolic supervision. MERT2 sets a new state of the art in music representation learning, surpassing previous best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in our lead-sheet transcription comparison. The same checkpoint follows score edits while largely preserving unedited musical content and generates zero-shot covers without cover-specific training. Its readable score also enables agentic music editing, with external language models translating user feedback into revisions of the composition.
Audio music generators can produce finished recordings, but their musical decisions are usually buried in latent representations; symbolic systems expose melody, harmony, rhythm, and form, yet stop short of the recording. YuE2 attempts to connect those stages in one model: it first creates an editable score-like plan, then turns that plan into semantic tokens and stereo audio. The paper’s central empirical case is that exposing the composition stage improves what listeners hear, rather than merely making generation easier to inspect.
Core Contribution
YuE2 combines symbolic composition and audio realization in a single checkpoint. Its symbolic plan specifies melody and harmony in a readable score representation, then conditions generation of 25-Hz semantic music tokens and continuous acoustic latents before 48-kHz stereo VAE decoding. That differs from a pipeline with a separate language model for planning and a separate diffusion Transformer for audio, and from audio-only systems whose compositions must be inferred after the fact.
The practical consequence is editing. A user can modify the symbolic representation directly, or have an external language model translate feedback into score revisions, while YuE2 regenerates the recording. The paper also applies the same model to zero-shot covers: it extracts a symbolic representation from a reference recording, accepts edits to that representation, and renders audio again. This is a credible design response to a real failure mode in audio generation: changing a chord progression or melodic phrase without asking a model to rediscover the entire song.
Technical Approach
The model is an AR-NAR Mixture-of-Transformers. Within each of 28 layers, autoregressive and non-autoregressive experts retain separate normalization, projections, and MLPs while sharing attention. The AR stream takes style and lyric text and predicts the symbolic plan and semantic sequence causally. The NAR stream predicts flow velocity for acoustic latents with bidirectional attention. This division assigns discrete compositional choices and continuous audio realization to different generation regimes without splitting them into independently trained generators.
Figure 2 makes the three operating modes explicit: creation generates a score from text, covering derives one from a recording, and editing starts from a user-modified score. All feed the same semantic and acoustic stages.
The supporting supervision stack matters because full recordings rarely arrive with aligned lead sheets. MERT2 learns audio representations from four residual-vector-quantized, 25-Hz code streams jointly derived from frozen MuQ and Qwen2-Audio-Instruct views. Its training proceeds from masked-audio pretraining to full-song adaptation and causal adaptation. SheetSage2 then maps full-context MERT2 features into grammar-constrained event sequences for timing, meter, structure, harmony, and melody, which a deterministic builder converts to ABC notation. The paper therefore does not assume a large corpus of manually aligned recordings and scores.
Results and Analysis
On WildSongBench’s 192 prompts, YuE2 reports a 6.73 SongBench Global Avg, above the evaluated public baselines; selecting the best of eight candidates raises that mean to 6.96, the highest observed among the systems considered. That result supports strong performance on the benchmark, but best-of-8 is a different operating point from a single generation: it trades extra sampling and selection for quality.
The paper’s more diagnostic comparison holds the checkpoint fixed and changes whether it uses symbolic planning. Experts preferred planning on overall quality in 49.3% of judgments, versus 34.6% for generation without planning. Figure 7 further attributes the listener advantage to perceived song quality, melody, and chord progression. The paired score example is useful here because it gives a mechanism for the preference result: planning develops and recalls a melodic idea across verse and chorus, whereas the unplanned counterpart repeats a local pattern with weaker harmonic continuity.
The representation and transcription components also show broad benchmark coverage. MERT2 exceeds the previous best result on 14 of 15 MARBLE metrics, while SheetSage2 leads 12 of 15 benchmark-metric pairs in the lead-sheet transcription comparison. These results make the symbolic interface more than a hand-authored control layer. Still, the paper’s headline quality measurements combine automatic indices, expert preferences, and selected-candidate evaluation; they establish a persuasive case for the evaluated prompts and systems, not a general guarantee that symbolic planning will dominate for every genre, listener group, or deployment budget.
Evidence Box
strongKey Claims
- •Symbolic planning improves full-song quality and musical coherence
- •One AR-NAR model can support creation, covering, and score-based editing
- •MERT2 and SheetSage2 provide semantic and symbolic supervision from recordings
- •YuE2 is competitive with evaluated proprietary song generators
Key Results
- •49.3% expert preference for symbolic planning versus 34.6% without planning
- •6.73 SongBench Global Avg on WildSongBench’s 192 prompts, above evaluated public baselines
- •6.96 SongBench Global Avg with best-of-8 candidate selection
- •MERT2 surpasses prior best results on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 transcription benchmark-metric pairs
Limitations & Caveats
- •Best-of-8 quality uses selection among 8 candidates rather than a one-shot generation
- •WildSongBench evaluation covers 192 prompts and the listed comparison systems
- •Human listening evidence is based on expert preferences rather than broad consumer sampling
- •Editing preservation is described as largely retaining unedited content without a quantitative preservation result in the supplied material