Somax 2: Real-Time Co-creative AI Music
- Somax 2 is a real-time co-improvisation system that uses corpus-based memory and context-sensitive recombination to achieve musical coherence.
- It recombines learned musical fragments rather than generating from scratch, enabling controlled stochastic creativity within stylistic boundaries.
- The system continuously listens to live performances, adapting its output in real time to balance semi-autonomous creativity with high human control.
Somax 2 is a real-time system for machine co-improvisation with human musicians that operates through corpus-based memory, machine listening, and context-sensitive recombination rather than generation “ex nihilo.” In a thematic review of AI music systems, it is presented as the clearest example of a system at the “high coherence / high human control / real-time” end of the spectrum of structured uncertainty, where stochasticity remains bounded by stylistic memory and live musical context (Browne, 29 Sep 2025).
1. Conceptual placement within structured uncertainty
In the review’s framing, Somax 2 is a paradigmatic instance of structured uncertainty. The central claim is that randomness in computational music creativity need not be eliminated, but can be shaped so that novelty remains musically legible. Somax 2 exemplifies this by producing surprise through constrained transformation of learned material rather than through unbounded random invention.
The review treats this positioning as analytically significant. Among the six systems compared—Musika, MIDI-DDSP, Melody RNN, RAVE, Wekinator, and Somax 2—Somax 2 is the strongest case of a real-time, co-improvisational AI music system that achieves novelty without sacrificing musical sense. A plausible implication is that the system functions as a model case for how stochastic procedures can be embedded inside a musically meaningful interaction loop rather than deployed as autonomous generation alone (Browne, 29 Sep 2025).
2. Corpus-based memory and generative substrate
Somax 2 is described as “a state-of-the-art system for machine co-improvisation with human musicians.” Its core representation is a corpus-based memory learned from musical material supplied by the user, such as recordings or MIDI files in a given style. This corpus grounding is not incidental; it defines the boundaries of the system’s possible output.
The review emphasizes that Somax 2 does not generate music from scratch. Instead, it “recombines and transforms what it has learned.” Its novelty therefore derives from fragment recombination rather than arbitrary note generation. The system uses statistical modeling to create “a memory structure it can navigate to form new, stylistically coherent musical sequences.” This suggests that the operative generative space is deliberately narrow: the source corpus constrains stylistic possibility, while recombination provides local variation and freshness within that space (Browne, 29 Sep 2025).
3. Real-time listening, adaptation, and execution in performance
Somax 2 is not characterized as a one-shot generator. As performance unfolds, “Somax’s machine listening modules analyse it and update the model in real time.” It “listens continuously to live musicians and adapts its behavior accordingly,” and its generative algorithms must be “fast and incremental.” These features place it firmly in the category of live improvisation systems rather than offline composition tools.
The review links this behavior to the broader idea of “Live Algorithms” and participatory sense-making. In that account, uncertainty becomes musically useful when it is embedded in reciprocal exchange. Somax 2 is therefore presented not merely as a playback or accompaniment engine, but as an interactive improvisation partner whose outputs are conditioned as the performance happens. This suggests that the system’s temporal coupling to human action is a primary source of its co-creative salience (Browne, 29 Sep 2025).
4. Randomness, contextual prediction, and musical coherence
The review states that randomness in Somax 2 is highly regulated rather than unstructured. “Randomness enters in selecting which learned fragment to play next, or how to vary it,” but this stochasticity is “highly context-sensitive.” The system ensures that any such choice “aligns with the live input’s harmony, timing, and style.”
The resulting coherence is attributed to four interlocking constraints: corpus grounding, fragment recombination, real-time listening, and context-sensitive selection. The review explicitly states that this “ensures a high degree of musical coherence, particularly important in real-time co-improvisation scenarios.” In comparative terms, the table places Somax 2 as Real-time: Yes, Randomness mechanism: Contextual prediction, Randomness type: Structured, and Level of human control: High. The significance of these labels is that the system’s uncertainty is not free-floating; it is filtered through harmony, timing, style, and live interaction, so novelty remains within “meaningful boundaries” (Browne, 29 Sep 2025).
5. User control, semi-autonomy, and shared agency
Somax 2 is not described as a system of direct note-by-note specification. Instead, control is distributed across several higher-level mechanisms. The user prepares the training corpus, adjusts behavioral parameters such as “how adventurous the agents should be,” and shapes the system during performance through actual playing.
The review presents this as a distinctive co-creative design. The performer is not a passive consumer, but neither are they the sole author of every output detail. Somax is described as a “semi-autonomous partner” that “contributes its own creative flavor while remaining grounded in the musical language defined by the user.” This is the basis for the review’s claim that Somax 2 supports shared agency: the human introduces motifs, shapes the musical environment, and guides style, while the machine responds and elaborates in ways that are not fully predictable but remain musically legible. A common misconception is that strong user control in AI music must mean low system autonomy; Somax 2 is presented as evidence that high human control can coexist with semi-autonomous contribution when control is exerted by framing the improvisational space rather than by micromanaging output events (Browne, 29 Sep 2025).
6. Comparative position among AI music systems
The review situates Somax 2 by contrasting it with several adjacent systems. Against Musika, Somax 2 is far more constrained: Musika is described as more autonomous and more “unstructured,” relying on random latent vectors and broad stylistic plausibility from training data, whereas Somax 2 uses source material directly and adapts to a live performer. Against Melody RNN, Somax 2 is more strongly framed by live interaction and stylistic continuity in performance, while Melody RNN is described as “semi-structured” through next-note sampling and temperature control.
Relative to MIDI-DDSP, Somax 2 shares an emphasis on preserving coherence through structure but differs in being a live improvising partner rather than mainly an assistive offline composition or synthesis tool. Relative to RAVE, Somax 2 is more explicitly tied to musical syntax and corpus memory, while RAVE is presented as an instrument for timbral exploration in latent space. Relative to Wekinator, both systems have high co-creative potential, but Wekinator externalizes control to the user through mappings and is not itself a generative improviser, whereas Somax 2 functions as an autonomous responding agent. The comparative pattern is therefore consistent: Somax 2 occupies the region in which novelty is produced through contextual recombination, coherence through corpus memory and live context, and co-creativity through responsive improvisational dialogue (Browne, 29 Sep 2025).
7. Terminological scope and name overlap
Within AI music research, Somax 2 refers to the co-improvisational music system associated in the review with Borg, 2019. Separately, the name “Somax” is also used for a composable, Optax-native stack for curvature-aware training in JAX. That system is a planned, composable second-order optimization stack with first-class modules for curvature operators, estimators, linear solvers, preconditioners, damping policies, and telemetry, and it is explicitly characterized as a systems framework for second-order optimization rather than a single optimizer (Korbit et al., 26 Mar 2026).
The shared name can invite confusion across domains, but the two usages are distinct in object, method, and application area. In the musical context, Somax 2 denotes a real-time co-improvisation system built from corpus memory and live machine listening. In the optimization context, Somax denotes a single planned, JIT-compiled step interface for curvature-aware training. The overlap is nominal rather than conceptual (Korbit et al., 26 Mar 2026).