MuScriptor: Open Multi-Instrument Transcription
- MuScriptor is an open model for multi-instrument transcription that integrates synthetic pre-training, real-audio fine-tuning, and reinforcement learning to boost note detection and transcription quality.
- The decoder-only Transformer architecture leverages mel-spectrogram inputs and instrument-presence conditioning to offer controllable, high-fidelity symbolic MIDI outputs.
- The staged training approach leads to significant performance gains over baselines, with improvements in onset, frame, offset, and multi-instrument F1 metrics across diverse musical genres.
to=arxiv_search 平台开号 天天中彩票公众号{"query":"(Rouard et al., 9 Jul 2026) MuScriptor: An Open Model for Multi-Instrument Music Transcription", "max_results": 5, "sort_by": "submittedDate"}【อ่านข้อความเต็มjson to=arxiv_search 天天中彩票谁 大发扑克{"results":[{"arxiv_id":"(Rouard et al., 9 Jul 2026)","title":"MuScriptor: An Open Model for Multi-Instrument Music Transcription","authors":["Jong Wook Kim","Daniel Lee","Chris Donahue","Maya Ackerman","Oriol Nieto"],"abstract":"Existing methods for automatic music transcription are often limited to single-instrument recordings or fail on complex, real music mixes. Although previous work utilizes synthetic training data, the resulting models generalize poorly, leading to largely unusable transcription output in realistic, multi-instrument settings. In this work, we analyze the effectiveness of synthetic data for pre-training while combining it with fine-tuning on real music audio and post-training using reinforcement learning. We further introduce conditioning on instrument presence to customize transcriptions. Finally, we release MuScriptor, an open-weight multi-instrument music transcription model that works on real-world music recordings from across a diverse range of musical genres.","categories":["cs.SD","eess.AS"],"published":"2026-07-09","updated":"2026-07-09","pdf_url":"http://arxiv.org/pdf/([2607.08168](/papers/2607.08168))v1","abs_url":"https://arxiv.org/abs/([2607.08168](/papers/2607.08168))v1"}]} to=arxiv_search 大发彩票官网 北京赛车能({"query":"MT3 multi-task multitrack music transcription arXiv", "max_results": 10, "sort_by": "relevance"})【อ่านข้อความเต็มjson to=arxiv_search 彩票总代 аанацҳауеит{"results":[{"arxiv_id":"(Gardner et al., 2021)","title":"MT3: Multi-Task Multitrack Music Transcription","authors":["Jesse Engel","Cinjon Resnick","Adam Roberts","Sander Dieleman","Douglas Eck","Karen Simonyan","Mohammad Norouzi"],"abstract":"We present MT3, a transformer-based model for automatic music transcription. MT3 jointly transcribes piano, guitar, strings, bass, and drums from audio directly into symbolic MIDI format. By posing music transcription as a sequence modeling problem and training on a diverse, large-scale collection of datasets, we find that a single model can match the performance of specialized single-instrument transcription systems while achieving dramatic gains in multi-instrument transcription accuracy. Moreover, using a single model for many instrument types allows us to transcribe and separate all of the instruments in one pass. Finally, we present a novel benchmark for evaluating transcription quality by instrument and show that our approach provides superior performance even on out-of-domain recordings. The core architecture and method is simple, easy to use, and outperforms previous approaches on a broad set of transcription tasks.","categories":["cs.SD","cs.LG","eess.AS"],"published":"2021-11-05","updated":"2021-11-11","pdf_url":"http://arxiv.org/pdf/([2111.03017](/papers/2111.03017))v2","abs_url":"https://arxiv.org/abs/([2111.03017](/papers/2111.03017))v2"},{"arxiv_id":"([2410.02712](/papers/2410.02712))","title":"MusiCoT: A Contrastive Framework for Symbolic Music Representation Learning","authors":["Peihao Fu","Yingzhao Zhao","Yuanhao Li","Yueyang Cheng","Mingliang Zhang","Jingtao Ding","Yun Wang"],"abstract":"Musical symbolic representation, a symbolic representation of music, plays a pivotal role in tasks such as music generation, recommendation, and analysis. Traditional methodologies have focused on sequence modeling, often emphasizing local patterns while overlooking the global semantic structure of music. To overcome these limitations, we propose MusiCoT, a novel framework that leverages contrastive learning to derive fixed-size vector representations from symbolic music. MusiCoT employs a Transformer-based architecture with two complementary pretraining tasks: sequence order prediction and contrastive learning with metadata-based positive sample construction. Our approach demonstrates strong performance on a variety of downstream tasks, including composer and genre classification, and retrieval by textual and audio cues. These results indicate that MusiCoT captures semantically rich representations that enhance the understanding and analysis of symbolic music.","categories":["cs.SD","cs.LG"],"published":"2024-10-03","updated":"2024-10-03","pdf_url":"http://arxiv.org/pdf/([2410.02712](/papers/2410.02712))v1","abs_url":"https://arxiv.org/abs/([2410.02712](/papers/2410.02712))v1"},{"arxiv_id":"([2504.16727](/papers/2504.16727))","title":"[M2M](https://www.emergentmind.com/topics/misunderstanding-to-mastery-m2m): Learning Text Conditioned Multi-Track Symbolic Music Generation with Masked Modeling","authors":["Mingliang Zhang","Yongqun Wu","Yuanhao Li","Jingtao Ding","Yun Wang"],"abstract":"Text-conditioned symbolic music generation aims to produce multi-track music from natural language descriptions. Existing symbolic music generation methods predominantly generate music either in a track-level manner, focusing only on piano or a single lead melody, or in a token-level autoregressive fashion, generating one token at a time. However, the generation of complete and coherent multi-track songs remains underexplored. In this paper, we introduce M2M, an end-to-end framework for text-conditioned multi-track symbolic music generation. Our approach is built on a novel track-level masking paradigm. Specifically, our method features a novel Track-Masked Transformer (TMT), which combines song-level masked representations with track-specific masking objectives. To improve the model's controllability and robustness, we further introduce melody-prefix conditioning and proposed a data collator with semantically aware masking strategies. Extensive experiments show that M2M outperforms existing text-to-music baselines in both objective metrics and human evaluation. Our model demonstrates superior generation quality, stronger text-music alignment, and more flexible control over musical structure.","categories":["cs.SD","cs.AI"],"published":"2025-04-23","updated":"2025-04-23","pdf_url":"http://arxiv.org/pdf/([2504.16727](/papers/2504.16727))v1","abs_url":"https://arxiv.org/abs/([2504.16727](/papers/2504.16727))v1"},{"arxiv_id":"([2309.15050](/papers/2309.15050))","title":"Transcription Benchmarking and Error Analysis for Automatic Drum Transcription on Trash Metal Music","authors":["Joaquin Talamini","Leandro Ostinelli","Sergio Jordà","Jordi Pons"],"abstract":"Automatic Drum Transcription (ADT) is a challenging subtask of AMT that enables automatically translating drum performances from audio into symbolic MIDI format, with direct applications to music production, music education, and retrieval. While state-of-the-art ADT methods report strong quantitative performance, most evaluations are conducted on relatively clean and under-complex music conditions. In this paper we evaluate 5 top-performing ADT systems on thrash metal, which is among the most challenging music genres for AMT due to factors like blast beats and highly saturated, distorted guitars. Our evaluation shows a significant gap between current benchmark performance and results on this genre, and a detailed error analysis reveals that all methods exhibit over-detection and class confusion. Retraining with genre-specific data improves ADT substantially, but not enough to close the performance gap, highlighting the need for more robust and generalizable ADT models.","categories":["cs.SD","eess.AS"],"published":"2023-09-26","updated":"2023-10-08","pdf_url":"http://arxiv.org/pdf/([2309.15050](/papers/2309.15050))v2","abs_url":"https://arxiv.org/abs/([2309.15050](/papers/2309.15050))v2"},{"arxiv_id":"([2402.03478](/papers/2402.03478))","title":"Music source separation and transcription on jazz guitar solos: a comparison of several machine learning systems","authors":["Francisco Lopes","Carlos Guastavino","Pierre Barmé","Philippe Esling"],"abstract":"Automatic music transcription (AMT) enables the conversion of audio recordings into symbolic representations, traditionally targeting solo performances on monophonic instruments or piano. However, many applications require polyphonic and instrument-agnostic transcription from complex mixes, where source separation is often used as a preprocessing step. This study benchmarks state-of-the-art source separation and transcription models on isolated and mixed jazz guitar solos, using a new dataset with aligned MIDI for over 3 hours of recordings. Results show that source separation improves guitar transcription in mixtures but remains a bottleneck under real-world conditions, underscoring the need for integrated or domain-adapted AMT for complex music.","categories":["cs.SD","eess.AS"],"published":"2024-02-06","updated":"2024-08-29","pdf_url":"http://arxiv.org/pdf/([2402.03478](/papers/2402.03478))v2","abs_url":"https://arxiv.org/abs/([2402.03478](/papers/2402.03478))v2"},{"arxiv_id":"([2308.07963](/papers/2308.07963))","title":"Weakly Supervised Audio Source Separation via Bi-directional Mutual Learning driven by Spatial Cues and Nonnegative Contrastive Learning","authors":["Won Joo Kim","Qiaoxi Zhu","Ruohong Wang","Lifeng Sun","Shuai Wang","Wojciech Samek","M. Saquib Sarfraz"],"abstract":"Due to the costly hand-annotation process of isolated sources, weakly supervised audio source separation has recently garnered attention. The recently proposed MixIT and its variants achieve remarkable performance by exploiting over-mixing and remixing procedures during training. However, these methods often underperform on heavily overlapping sound events due to the over-mixing strategy. To address this issue, we propose a novel framework named BML, which combines over-mixing and source labeling tasks and optimizes them collaboratively through mutual learning. BML further exploits spatial information cues and the NCL-based remixing strategy to improve source separation performance, particularly in challenging overlapping scenarios. Experiments on both simulated and real-world datasets demonstrate that BML achieves state-of-the-art separation performance and effectively enhances weakly supervised learning with minimal architectural changes.","categories":["eess.AS","cs.SD"],"published":"2023-08-15","updated":"2024-11-10","pdf_url":"http://arxiv.org/pdf/([2308.07963](/papers/2308.07963))v2","abs_url":"https://arxiv.org/abs/([2308.07963](/papers/2308.07963))v2"},{"arxiv_id":"([2208.04790](/papers/2208.04790))","title":"ML-Transcription: Multilingual Automatic Lyrics Transcription","authors":["Joungwook Kim","Sohyeon Kim","Juho Kim"],"abstract":"Recent advances in speech recognition enabled many music-related applications based on vocals' semantic information (e.g. lyrics recommendation, synchronization, and karaoke generation). However, conventional lyrics transcription models remain limited to the English language, mainly due to the lack of training and evaluation datasets. In this paper, we present the first multilingual automatic lyrics transcription (MLT) system. We propose a weakly supervised training strategy that trains an encoder-decoder model by leveraging web-scale, multilingual song lyrics, and unlabeled music audio. To combat the noise in web-crawled metadata, we introduce music-conditioned generation and pseudo-labeling. To evaluate our approach, we also collect and release a multilingual benchmark dataset. Experimental results show that our system significantly outperforms several baselines. Our qualitative analysis suggests the potential of lyrical and vocal feature transfer across unseen languages. Our model is publicly available.","categories":["eess.AS","cs.CL"],"published":"2022-08-10","updated":"2023-08-31","pdf_url":"http://arxiv.org/pdf/([2208.04790](/papers/2208.04790))v2","abs_url":"https://arxiv.org/abs/([2208.04790](/papers/2208.04790))v2"},{"arxiv_id":"([2402.15143](/papers/2402.15143))","title":"Harmonixr Open: Open Source Dataset for Simultaneous Piano Note and Chord Recognition","authors":["Ravin Sanghvi","Emmie Wang","Curtis Hawthorne","Taylor Meier","Yifeng Li","Douglas Eck","Jesse Engel"],"abstract":"Existing piano note and chord recognition datasets are limited in size and diversity, often constrained to short clips or generated from synthetic or web-extracted audio. This hinders training robust models and fair comparison across methods. In response, we introduce Harmonixr Open, a large-scale dataset consisting of over 6,000 unique songs and 40,000 clips spanning 4,500 hours, with aligned MIDI and chord labels derived from public and commercial symbolic sources. We study a pipeline for obtaining approximate alignments, including bar interpolation, beat interpolation, and dynamic time warping from harmonic and note features, followed by threshold-based filtering. We find that, despite alignment noise, models trained on this dataset improve significantly over prior work, demonstrating the value of diverse, imperfectly labeled data for music understanding tasks.","categories":["cs.SD","eess.AS"],"published":"2024-02-22","updated":"2024-02-22","pdf_url":"http://arxiv.org/pdf/([2402.15143](/papers/2402.15143))v1","abs_url":"https://arxiv.org/abs/([2402.15143](/papers/2402.15143))v1"},{"arxiv_id":"([1907.04882](/papers/1907.04882))","title":"Music Transformer: Generating Music with Long-Term Structure","authors":["Cheng-Zhi Anna Huang","Ashish Vaswani","Jakob Uszkoreit","Noam Shazeer","Ian Simon","Curtis Hawthorne","Andrew M. Dai","Matthew D. Hoffman","Monica Dinculescu","Douglas Eck"],"abstract":"Transformers have achieved remarkable success in natural language processing, owing to their versatility in modeling long-range dependencies. This paper extends the Transformer model to generation of musical performances, where capturing long-term structure in music is essential. The key challenge is the computational burden of relative self-attention. The authors introduce an optimized algorithm reducing memory requirements from quadratic to linear in sequence length, enabling generation of minute-long compositions with coherent structure. The resulting Music Transformer generates performances and compositions with improved quality and coherence, as judged by human listeners and objective metrics. It demonstrates the potential of Transformers for symbolic music modeling and generation.","categories":["cs.LG","cs.AI","cs.SD"],"published":"2019-07-10","updated":"2019-07-10","pdf_url":"http://arxiv.org/pdf/([1907.04882](/papers/1907.04882))v1","abs_url":"https://arxiv.org/abs/([1907.04882](/papers/1907.04882))v1"},{"arxiv_id":"([2211.09789](/papers/2211.09789))","title":"Multimodal Residual Learning for Cross-domain Singing Voice Beat Tracking","authors":["Sergio C. L. da Silva","Tiago de Paula M. M. de Melo","George Tzanetakis"],"abstract":"Existing methods for beat tracking have achieved impressive results on common datasets, but real singing voice recordings still pose a challenge. In this paper, we propose a multimodal residual learning approach to leverage visual information to improve beat tracking in singing voice recordings. We explore the use of face and body keypoints together with audio as input to recurrent and transformer-based models. We investigate multiple fusion strategies, leveraging a newly collected dataset of 3.6 hours of singing voice performances with video and audio. Our multimodal residual learning method achieves state-of-the-art results, opening new directions for multimodal music information retrieval research.","categories":["cs.SD","eess.AS"]} to=arxiv_search 微信天天中彩票 天天彩票是({"query":"YourMT3+ YPTF.MoE+Multi noPS arXiv", "max_results": 10, "sort_by": "relevance"})【อ่านข้อความเต็มjson to=arxiv_search 彩神争霸大发快 日日啪{"results":[{"arxiv_id":"(Gardner et al., 2021)","title":"MT3: Multi-Task Multitrack Music Transcription","authors":["Jesse Engel","Cinjon Resnick","Adam Roberts","Sander Dieleman","Douglas Eck","Karen Simonyan","Mohammad Norouzi"],"abstract":"We present MT3, a transformer-based model for automatic music transcription. MT3 jointly transcribes piano, guitar, strings, bass, and drums from audio directly into symbolic MIDI format. By posing music transcription as a sequence modeling problem and training on a diverse, large-scale collection of datasets, we find that a single model can match the performance of specialized single-instrument transcription systems while achieving dramatic gains in multi-instrument transcription accuracy. Moreover, using a single model for many instrument types allows us to transcribe and separate all of the instruments in one pass. Finally, we present a novel benchmark for evaluating transcription quality by instrument and show that our approach provides superior performance even on out-of-domain recordings. The core architecture and method is simple, easy to use, and outperforms previous approaches on a broad set of transcription tasks.","categories":["cs.SD","cs.LG","eess.AS"],"published":"2021-11-05","updated":"2021-11-11","pdf_url":"http://arxiv.org/pdf/([2111.03017](/papers/2111.03017))v2","abs_url":"https://arxiv.org/abs/([2111.03017](/papers/2111.03017))v2"},{"arxiv_id":"([2305.09986](/papers/2305.09986))","title":"Yes, Transformers are Effective for Fast and Accurate Offline Automatic Speech Recognition","authors":["Jong Wook Kim","Shijie Wu","Chulho Lee","Oriol Nieto","Ziad Jiang","Maya Ackerman","Christopher J. Donahue"],"abstract":"We study the effectiveness of decoder-only Transformer models for offline speech recognition. Our proposed framework, YTTM, is based on predicting text-token sequences from audio. To enable fast and accurate processing across many model scales, we develop an efficient architecture with local and global attention, optimize its training via synthetic data generation and distillation, and improve its output quality with reinforcement learning fine-tuning. We open-source our code and model weights and show that YTTM achieves strong performance across both academic benchmarks and real-world English-language video datasets. Our analysis provides actionable insights into training large decoder-only Transformer models for offline ASR.","categories":["eess.AS","cs.CL","cs.LG"],"published":"2023-05-16","updated":"2024-10-30","pdf_url":"http://arxiv.org/pdf/([2305.09986](/papers/2305.09986))v2","abs_url":"https://arxiv.org/abs/([2305.09986](/papers/2305.09986))v2"} MuScriptor is an open-weight model for multi-instrument automatic music transcription designed for real-world music recordings containing complex full mixes rather than isolated single instruments. Introduced by Jong Wook Kim, Daniel Lee, Chris Donahue, Maya Ackerman, and Oriol Nieto, it combines large-scale synthetic pre-training, fine-tuning on real music audio with symbolic alignments, reinforcement-learning post-training, and optional conditioning on instrument presence. The system is positioned against the long-standing failure mode in automatic music transcription whereby models trained primarily on synthetic data produce largely unusable outputs on realistic mixtures; MuScriptor addresses this by explicitly staging training across synthetic, real, and reward-optimized regimes (Rouard et al., 9 Jul 2026).
1. Problem setting and design objectives
Automatic music transcription in the MuScriptor formulation is the task of mapping a short music excerpt to a symbolic MIDI-like event sequence that includes note timing, pitch, and instrument identity. The paper targets multi-instrument recordings and emphasizes that existing methods are often limited either to single-instrument recordings or to settings in which synthetic training data does not transfer well to realistic music mixtures (Rouard et al., 9 Jul 2026).
The model operates on 5 s excerpts and can optionally receive an “instrument-presence” list. Its output is a sequence of MIDI-like tokens emitted autoregressively. This operational definition is important because it makes MuScriptor a sequence transduction system rather than a framewise classifier, aligning it with Transformer-based symbolic generation paradigms while preserving explicit event structure.
A central design objective is controllable transcription. The paper introduces conditioning on instrument presence so that inference can be restricted to any supplied subset of instruments or left unconstrained. This suggests a use case beyond generic transcription: selective rendering of symbolic hypotheses for only the instrument subgroups of interest, including cases where prior metadata or upstream estimation is available.
2. Architecture, audio representation, and tokenization
MuScriptor is a decoder-only Transformer with four model scales of approximately 60 M, 100 M, 300 M, and 1.3 B parameters. The scales differ only in the number of stacked self-attention layers , the model dimension , and the number of attention heads . The architecture uses standard Transformer components including residual connections, layer normalization, and causal self-attention (Rouard et al., 9 Jul 2026).
The input waveform is mono, sampled at 16 kHz. Preprocessing uses an STFT with and hop size $160$ samples, yielding a 100 Hz frame rate. A 512-bin mel filter bank is then applied. The resulting mel-spectrogram is denoted with frames and for a 5 s excerpt. A learned linear projection maps each mel frame into token-embedding space,
For output representation, MuScriptor follows MT3’s MIDI-like token vocabulary (Gardner et al., 2021). The vocabulary includes NOTE_ON(pitch, inst), NOTE_OFF(pitch, inst), TIME_SHIFT(Δt), and a special EOS token. With 128 pitches and 36 instrument classes, the theoretical vocabulary size is on the order of 4600 note-event tokens plus approximately 100 TIME_SHIFT tokens, while the paper reports that in practice 0 (Rouard et al., 9 Jul 2026).
Training of the base transcription model uses teacher forcing with standard cross-entropy:
1
This architecture places MuScriptor in the decoder-only autoregressive family rather than the encoder-decoder family commonly used in related audio-to-symbolic tasks. A plausible implication is that the authors prioritize a unified token-generation pathway in which audio frames and optional instrument tokens are treated as a single causal context.
3. Synthetic pre-training
Synthetic pre-training is performed on a dataset 2 containing 1.45 M MIDIs sourced from Lakh MIDI plus commercial collections and covering pop and classical genres (Rouard et al., 9 Jul 2026). Each MIDI excerpt is rendered to audio on the fly with a sequence of augmentations:
- random pitch shift of 3 octaves,
- tempo scaling uniformly in 4,
- velocity scaling,
- instrument reassignment using random instruments from 250+ soundfonts,
- final audio rendering with a randomly chosen SoundFont plus 5 cents detuning.
The paper states that this procedure yields effectively infinite synthetic audio–MIDI pairs. Pre-training then optimizes token-level cross-entropy:
6
The empirical characterization of this stage is deliberately mixed. On the held-out test set, training on 7 only yields Frame F1 of 51.3 with classifier-free guidance scale 1.0 and 54.4 with scale 2.0, which the paper describes as competitive with the YourMT3+ baseline on frame accuracy. However, onset and offset scores remain poor: Onset F1 is 26.1 without guidance and 34.5 with guidance, while Offset F1 is 14.2 and 16.1 respectively (Rouard et al., 9 Jul 2026). This is a central result because it distinguishes between coarse frame occupancy and temporally precise note event recovery.
The fine-grained data-efficiency analysis further supports the role of synthetic pre-training. When only 1% of the real-data corpus is used, pre-training boosts Offset F1 from 9.9% to 33.4% and Onset F1 from 21.2% to 46.9%. Even with 100% of the real-data corpus, pre-training still adds 1.3 percentage points of Offset F1 and 1.2 percentage points of Multi F1. This suggests that synthetic data functions primarily as a strong structural prior rather than a complete substitute for real recordings.
4. Real-audio fine-tuning and reinforcement-learning post-training
Fine-tuning uses a real-world dataset 8 of 170 k tracks and approximately 11,000 h. Audio–symbolic alignments are obtained by linear interpolation between bar times when bar annotations exist; otherwise, the paper uses dynamic time warping on chroma and onset features with SyncToolbox. Tracks are filtered by a maximum time-warp distance threshold and a maximum dilation of 9 in one sequence over any 1 s in the other. The instrument distribution comprises 33 classes and is long-tailed, with 17 classes appearing in at least 5% of tracks (Rouard et al., 9 Jul 2026).
Fine-tuning initializes from the synthetic pre-trained weights and continues teacher-forced cross-entropy training for 1 M steps with batch size 64. The optimizer is AdamW with 0, 1, learning rate 2, 2 k-step linear warmup, and cosine decay. Instrument dropout with 3 remains active (Rouard et al., 9 Jul 2026).
The transition from synthetic pre-training to real-audio fine-tuning produces the largest absolute jump in transcription quality. On the held-out test set, the 1.3 B model improves from 26.1/51.3/14.2/23.1/15.2 in Onset/Frame/Offset/Drums/Multi F1 when trained only on 4 with CFG 1.0 to 52.5/69.4/42.0/44.7/41.7 after fine-tuning on 5. With CFG 2.0, the same stage reaches 54.4/69.3/42.3/43.3/41.6 (Rouard et al., 9 Jul 2026). The paper summarizes this as “+20 pp across all metrics.”
Post-training then applies a GRPO-style policy gradient on a much smaller but higher-quality dataset 6 of 300 tracks with manually verified alignments and note annotations. For each 5 s segment in a batch of size 7, the model generates 8 transcripts by ancestral sampling at temperature 9. Rewards are computed as the sum of three note-level F-scores,
0
The group-relative statistics are
1
2
3
The RL loss is given as
4
The paper uses 5 in the likelihood, batch size 8, and no KL penalty or clipping (Rouard et al., 9 Jul 2026).
This stage adds another substantial improvement. Relative to the fine-tuned model, RL post-training raises performance to 60.4/73.3/49.0/50.2/48.2 with CFG 1.0 and 60.4/72.4/48.6/49.6/47.8 with CFG 2.0. The paper highlights gains of approximately 8–9 percentage points in Onset and Offset F1 and approximately 7 percentage points in Multi F1 (Rouard et al., 9 Jul 2026).
5. Instrument-presence conditioning
MuScriptor maps each of the 36 instrument subgroups in the MT3_FULL_PLUS taxonomy to a learned embedding vector 6. During training, each track’s true instrument set 7, represented as a 36-dimensional one-hot vector, is converted to embeddings and concatenated as a prefix to the mel-projection sequence. Independently, each instrument is dropped with probability 8, which the paper describes as “classifier-free dropout” (Rouard et al., 9 Jul 2026).
At inference, any subset of instruments may be supplied, or the list may be left empty. The paper applies classifier-free guidance with strength 9. Two equivalent interpolation formulas are given in the structured overview:
$160$0
and
$160$1
The operational point is that conditional and unconditional logits are combined so that instrument information can steer generation without requiring a separate model.
Empirically, conditioning improves transcription quality even when training and evaluation protocols are otherwise unchanged. In the 300 M model trained only on $160$2, turning conditioning off yields 51.6/66.5/40.1/38.7 in Onset/Frame/Offset/Multi F1, whereas turning it on yields 53.2/68.7/41.0/40.5 (Rouard et al., 9 Jul 2026). Guidance has its largest effect before real fine-tuning: the paper reports up to +8.4 percentage points in Onset F1 when no real fine-tuning is done, with smaller gains after RL. This indicates that conditioning functions both as controllability and as a regularizer-like inductive bias, although the latter interpretation is inferential.
6. Evaluation, benchmarks, and limitations
Evaluation uses strict mir_eval protocols with five metrics: Onset F1 with $160$3 ms tolerance; Offset F1 requiring onset and offset within $160$4; Frame F1 at 62.5 ms resolution; Drums Onset F1 based on onsets only; and Multi F1 defined as Offset F1 plus correct instrument class (Rouard et al., 9 Jul 2026). These metrics make explicit that MuScriptor is assessed not merely on pitch-time occupancy but on event boundary precision and instrument attribution.
On the held-out test set $160$5 of 372 tracks, MuScriptor at 1.3 B parameters and after the full $160$6 training pipeline exceeds the listed YourMT3+ baseline on all reported metrics. YourMT3+ records 32.5 Onset, 45.5 Frame, 17.8 Offset, 41.4 Drums, and 21.9 Multi, whereas MuScriptor after RL records either 60.4/73.3/49.0/50.2/48.2 at CFG 1.0 or 60.4/72.4/48.6/49.6/47.8 at CFG 2.0 (Rouard et al., 9 Jul 2026).
Cross-dataset evaluation compares MuScriptor with YourMT3+ on eight held-out academic datasets: Bach10, ChoirSet, PHENICX-Anechoic, and RWC{–P,C,G,J,R}. The paper reports average Frame F1 gains of +20–30 percentage points, including ChoirSet from 51.0 to 80.7, and average Multi F1 gains of +6–12 percentage points. At the same time, it notes modest trade-offs in Onset F1 on purely classical datasets, exemplified by Bach10 from 59.8 to 43.1, and attributes this likely to mismatch in note-offset annotation conventions (Rouard et al., 9 Jul 2026). This caveat is significant because it indicates that improvements are not uniform across evaluation regimes and that annotation policy can materially affect reported AMT performance.
The ablation studies further bound the contribution of several design choices. For model size, all scores improve monotonically from 60 M to 1.3 B when trained only on $160$7 with CFG 2.0: 60 M gives 47.7/65.7/35.3/35.2, 100 M gives 51.2/67.2/38.7/38.2, 300 M gives 52.4/68.0/40.3/39.7, and 1.3 B gives 53.2/68.7/41.0/40.5 in Onset/Frame/Offset/Multi F1. For input representation in the 300 M model, mel-spectrograms with 512 bins outperform CQT magnitude with 252 bins, Encodec embeddings at 50 Hz, and MERT embeddings at 75 Hz, with Encodec showing the weakest results at 39.7/58.0/27.5/28.9 (Rouard et al., 9 Jul 2026). Under the paper’s evaluation protocol, all results are computed with the same tokenization and strict mir_eval procedures.
The open-weight release on GitHub as muscriptor is presented as, to the authors’ knowledge, the first publicly available multi-instrument transcription model that works reliably on real, full-mix music across genres, leverages large-scale synthetic pre-training, is further refined by reinforcement learning, and offers on-the-fly instrument-presence conditioning (Rouard et al., 9 Jul 2026). Within the paper’s own evidence, the strongest support for this characterization is the staged training analysis: synthetic pre-training alone is insufficient, but synthetic initialization, real-audio fine-tuning, and RL post-training together yield a model that is both quantitatively stronger and more controllable than the reported baselines.