Comment by dimmuborgir
4 years ago
AI is bad at music also. Even the state of the art transformer models can't produce more than a few seconds of coherent melodic phrases.
4 years ago
AI is bad at music also. Even the state of the art transformer models can't produce more than a few seconds of coherent melodic phrases.
I think if we replaced "AI" with "taking averages over subsets of historical examples", then there'd be no mystery for when "AI" will be good or bad at anything.
Would we expect a discrete melodic structure to be expressible as averages of prior music? No.
Have you heard the piano continuations of AudioLM?
https://google-research.github.io/seanet/audiolm/examples/
Pretty sure the first continuation is a famous piece with a few notes messed up. Can't remember the name. Honestly it only sounds marginally better than the old markov chain continuations.
Yep, Moonlight Sonata (mov. 3) no less. Talk about over-fitting!
Isn’t that as good as it gets? The whole point of the continuations is that given a short leading prompt from a real piece that it should continue it realistically.
It didn’t get to train on the test set, if that’s what you’re implying, and I find it hard to believe the assertion that continuations are copies of the train set (if that’s your claim).
3 replies →
Indeed, there is lots of denial or ignorance in this thread (ignorance in the technical sense). AudioLM already produced impressive results and it's a tiny fraction of what is already possible because performance simply improves with scale. One can probably solve music generation today with a ~$1B budget for most purposes like film or game music, or personalized soundtracks. This is not science fiction.
I don't see a lot of progress in AudioLM compared to results from 2018: https://storage.googleapis.com/magentadata/papers/maestro/in...
What's more interesting and concerning - listen carefully to the first piano continuation example from AudioLM, notice the similarity of the last 7 seconds to Moonlight sonata: https://youtu.be/4Tr0otuiQuU?t=516
I'm afraid we will see a lot of this with music generation models in the near future.
6 replies →
Which is extra funny, because GOFAI models (e.g. David Cope's work) were doing a pretty OK job back in the 1990s!
It doesn't surprise me that an AI model for language can't grok maths or music. I can't see how a language model can map to maths. Hell, I don't even know how to describe music in words. It's possible to articulate some maths in words, but that often involves using words with unexpected definitions.
AI can be quite good at music,
but yes there is not yet at on-demand button rendering from a text prompt of bitstreams encoding composed performed and mastered music.
AI is bad at Audio. AI can do MIDI fine.
MIDI is extraordinarily expressive and is likely used to sequence a large majority of music produced within the last three decades. A lot of the instruments you hear are synthesizers or samplers running directly from MIDI. There is a lot more to what MIDI can do, and is used for, than the conception most people have from "canyon.mid" or old website background music. If an AI can do MIDI just fine then it's an extremely small leap to doing audio just fine.
If an AI can do MIDI just fine then it's an extremely small leap to doing audio just fine.
Unfortunately this is not true. It takes a huge amount of human effort to make MIDI encoded music sound good. The difference between MIDI and raw audio music generation is the same as the difference between drawing a cartoon and producing a photograph.
To clarify, yes MIDI can be expressive, but what's being generated when people say "AI generates MIDI music" is basically a piano roll.
3 replies →
Which is a real shame. AI-powered restoration of poor-quality audio would be highly useful.
That particular niche has had some pretty amazing successes already. It's coming.
We can't produce arbitrary media streams with many "stack layers" of meaning and detail yet, but we can do a lot of specific instrumental transformations...
Vaguely relevant: https://koe.ai/recast/
That's wrong, and shows how ignorant you are of SOTA techniques for music generation. They are far ahead of that.
That’s what a musician does. They make short loops and loop them.
This reads like someone who knows sheet music and theory but does not listen to music. It’s repetition of short phrases over and over.
I’m not really sure what people expect of general AI trained on human generated outputs. It can’t make up anything anything “net new” only compose based upon what we feed it.
I like to think AI is just showing us how simple minded we really are and how our habit of sharing vain fairy tales about history makes us believe we’re masters of the universe.
Those models are not trained on short loops. They are trained on whole songs just like image generation models are trained on whole images. And yet they struggle to repeat sections, modulate to a different key, create bridges, intros and outros. After a few seconds of hallucinating a melodic line they simply abandon the idea and migrate to another one. There is no global structure whatsoever.
Maybe that's the problem.
We're trying to train a full composer AI without allowing to learn about different instrument sections independently at first. The human composer will have a good idea of the different parts and know how to merge them in harmony.
I think we might get better results training separate AI systems on percussions, strings, vocals etc. then somehow create connections between them so they learn together. A band AI if you will.
We could try a BERT for each, with the generator learning to output logical sequences of sounds instead of words.
Musicians don’t spit out an album in one sitting and they’re highly trained in theory. They get bored and tired of a process and take breaks. They come up with an album of loops composed together over time.
AIs state will forever be constrained to the limits of human cognition and behavior as that’s what it’s trained on.
I read published research all year. Circular reasoning. Tautology. It’s all over PhD thesis.
There’s no “global structure” to humanity. Relativity is a bitch.
Seeing the world through the vacuum of embedded inner monologue ignores the constraints of the physical one. It’s exhausting dealing with the mentality some clean room idea we imagine in a hammock can actually exist in a universe being ripped asunder by entropy.
It’s living in memory of what we were sold; some ideal state. Very akin to religious and nation state idealism.
4 replies →