Comment by hazenut

9 hours ago

As a musician, I have a feeling how this will go on:

Most of the existing music creation heavily relies on daw and audio plugins, there are two major approaches among many: 1. recording with acoustic instruments, then using plugins to process 2.using a lot of virtual instruments directly, which is more often used in scoring.

Some of the process is already heavily ai assisted, like the postprocessing and mastering. For scoring, more film scorers are using the Noteperformer which is directly note to audio but quality is far from Suno level. The very popular virtual singers(vocaloid, etc) evolved from audio concatenation to full neural network generations. But so far, none of these have encroached into the essential part of the music creation.

Things will get more interesting if this get to the next level, when every single part of the music creation process (in the digital chain) is analyzable and synthesizable by ai. YuE2 appears to me can do note level analyze but stil not synthesize. To be able to synthesize, you get Noteperformer on steroid. Then, the daw and ai music creator will become one. This is somehow unappealing but paradoxical situation: You, as a creator, now can vastly do more. The ai tools gives you back the creative process. The ai music on the internet will no longer be pure slops because lots of them will be finely crafted by human musicians, but at the same time it all becomes super pointless because ai's capability is much stronger and do all these things itself, the only difference is whether you choose to do it yourself or let ai automate some of the parts.

Therefore I'm not sure all the bad qualities of the ai music (slop, lacking soul, or whatever you name) is a sin or blessing. The condemnable part of the ai music is exactly what currently saves human musician.

There are other side effects: the current virtual instrument, audio plugin market will mostly collapse. AI based audio generation will be a hegemony which reduces the diversity. You know that limitation of the process is a major source of curiosity and creativity, and itself is fun part. Wait until we rediscover the appeal of limitation, like the pixel art genre. We will see more of chiptunes, concatenative-sampling punks, 2010-digital vaporwaves and so on. But now the most important part is, genuineness will be a requirement: you have to show the process - you fake it with ai, you lose. Music will be a much more performative art (as it always was in the past) and community centric. Music streaming in its currently state of affair will be abandoned by all but the most casual listeners.

Another trend will be gamifications and participatory music. There is still one part of music tech that is currently underexplored. The physical modeling. The synthesizer tech has hit a wall: it still struggle to generate natural and interesting sound, which is also why sample based virtual instrument still thrives. The physical modeling is promising but extremely compute intensive. AI accelerated physical modelling can help with this. This can be actually an antidote to the generative AI. Physical modeling can be enhanced by more visceral visualizations or hardware analogues (materialization of virtual process). There is joy of participating and watching the corporeal and physical process of making music. While elite artist can create very elaborate process(think of the eurorack and experimental synths) and visualizations and show them to the audience, the process can be simplified and gamified to promote collective creative performance, all without dumbing down, as the physical process itself is open ended, and we as human beings, have stornger connection to the physical sounds than the abstract sound of the previous gen synthesizers.

To talk about participatory music and community, we can get glimpse of it from the Vocaloid scene. These crowd has very interesting take on ai music which I feel can be very enlightening to those uninitiated. They are enamored by the new AI soundbanks (which is the progression from the old contatenative technology) while very opposed to (generative) ai music. How is it so? They see the music by the Vocaloid producers a token of love poured into the community. The quality of music is not deciding factor, it is whether the producer know their community and pay respect to the virtual idol they loved. Their idol is a collective creation, it is not directed by any single entity. Derivative work also plays central role in the scene. Diaglogue is weighted a lot more than pure output. This is quite different from some of the AI idols though many uninitiated tend to conflate the two together. The AI industry may see Vocaloid as the predecessors to the AI idols, and you see some of the singing synth engine and soundbank makers are also giving nods to the AI music industry (ace studio, traditionally a singing synth company, is dipping toes in prompt based generations, and looks like is behind YuE2), but the two have very different spiritual cores. The conflation of AI music and Virtual Singers has caused further damage to the scene. One example is some streaming platform categorizes all virtual singer songs into AI music, which completely denies the hard work of the producers, and monetization and visibility possibilities, despite that some virtual singer still uses concatenative technology and have absolutely zero AI involved. But just for the AI soundbanks many has seen their drawbacks, now many consider they are too realistic. They want to keep the essence of the mechanical sound of their beloved idols, so you see the latest Miku soundbank delibrately not pursuing the realism like many do, e.g. SynthV, and those keep their idol unique characters are more successful in retaining their audience.

"The physical modeling is promising but extremely compute intensive. " This is not true, at least for the makers of pianoteq. Their simulation runs fine on a raspberry 5 with 1 gb.

btw. thanks for your perspective.

  • Interestingly, I'm long time user of Pianoteq, and other physically modeled instruments like SWAM. In my eyes, pianoteq has somewhat plateaued in terms of realism in recent years, and still not comparable to sampled piano in the most critical listening, and other instruments still have long way to go. And these are traditionally instruments. For more avant-garde, musik-concrete and foley style sounds, physical modeling is still almost unexplored at all. There is major limitation in compute, and it will remain this way unless we go in the AI accelerated approach.