Comment by Genego

1 day ago

Whenever I see the new releases around video generation (and image) generation models, I get goosebumps, because it just feels so fun to work with them. But then I remember that I spend upwards of $10k on inference generating well over 50k images for storyboards, training models; and probably creating almost an hour of video (I assume). Yeah, I get that things can be economic if you don't use the latest models (ran some case studies on this), but the latest models are the most fun to work with. It doesn't scale as well as "vibe coding" stuff together on the weekend. And when things work really well its almost as if you're seeing an zoopraxiscope come to life for the first time; and you just want to keep going.

I got a few offers to work with some startups in this space, but it also seems that many startups work on stuff that just doesn't seem to be very worthwhile (like creating masses of spam for YT or TikTok shorts), or even straight out morally/ethically wrong (cloning/deepfakes, etc). But seeing advances in this space; and coming from a filmmakers background, I might just end up being naturally drawn to this space on an engineering level and figuring something out along the way. As you can see I worked on a lot of stuff just for the fun of it, and documenting the process: https://edwin.genego.io/blog (but I stopped at the beginning of the year .... might.. just pick it up again.

I went to Art Center for film. Loved it. But ended up writing software instead of shooting movies (while still also handling a lot of visual art direction, graphics work, UI, 3D animation, etc). Now I feel like we're starting to be roughly in the same boat as far as using prompts.

What bothers me is that every piece of content generated this way helps flood an already saturated market for content, while slowly degrading the expectations of what people see, to the point that no one will bother with shooting or animating anything anymore. Even if it's 50% worse, it's 90% cheaper, so the economics argur against producing any new physically made content. Simultaneously, it's cannibalizing all existing content. This points toward a feedback loop, like a snake eating its own tail. And even though Hollywood blockbusters have followed that pattern for a couple decades, it's demoralizing to me to see it enshrined as the future of film (or to hear from someone who makes films that it would be a preferred mode of creation).

  • > Even if it's 50% worse, it's 90% cheaper, so the economics argue against producing any new physically made content.

    That is an unfortunate pattern I fear is become applicable to a lot of domains (film, software, food, clothing, built environment, electronics, physical goods et al). AI is just accelerating that in a few.

  • I feel these tools are destined to be mostly used as content generators - I'm happy to be proven otherwise but they've been around a while now and aside from a few pop videos I've not seen stuff that seems to be infused with the outer edge of quality art direction - perhaps because the model data limits it or perhaps it's the way they're used.

MiniMax H3 is going to release weights. You can locally run it with definitely less than $10k (and possibly faster than Seedance's queue), and it's fun to train it for whatever you need.

  • H3 results are underwhelming compared to LTX 2.3 so far:

    https://www.reddit.com/r/StableDiffusion/comments/1vciy35/lt...

    Perhaps refined ComfyUI workflows will squeeze more quality out of it, but it's definitely not in the realm of Seedance, and Lightricks is training LTX 2.5/"LTX-Next".

    • Note: Those results are a little misleading the good LTX results were generated with a good workflow in ComfyUI, and a prompt expanding local model.

      Whilst the H3 result was generated with the raw api using his raw prompt.

      If you ask ChatGPT or other AI model to "improve" your prompt (with cinematic, good lighting) generally, you will get also very good result from H3 also.

      My own H3 test show that H3 (API edition) is undoubtedly better than LTX in prompt adherence ! The real comparison will be with the edition of H3 we get to run locally.

  • Awesome! Have been a bit in the dark of the latest models, will have a look. I am due an upgrade for my local machine and GPU, so that excites me as well.

Curious what you think of sentienttube.com

I’ve built it as a side project, and the cost to produce one 20-40 minute video is in the 5$ range.

My friends/family have watched some of the better videos. All one shotted, and the storylines come out surprisingly well.

Currently working on a new engine that generates video, but it’s expensivee

Open to collabing

I've been dabbling in this space on the application layer and have spent a few hundred myself experimenting with video models.

It can be really entertaining/addicting to build with them because you're essentially pulling the slot machine and having TikTok/Marvel/YouTube come out of it. I think that was the bet with Sora but the problem is mostly that the novelty wears off quick, and most people want to just consume content without typing in what content they want to see, or sifting through mass-generated spam "content" with nothing behind it to make it worthwhile (a lot of people engage with content parasocially)

Once the tools for creators to steer and integrate models in this space get better, it will explode. We've been working on what I think will be one of the first use case for integrating these models, because I think we're approaching a middle ground where they can be integrated in experiences to provide entertainment/engagement/fun experiences without feeling like slop.

Btw, I'm impressed with some of the AI content on your site but I think you might want to pare down the non-demo pages because it has a different impression me (can I trust that this text is true? / I'm reading a lot of words but not really learning about this person) than you might have intended). I'm a bit of a hypocrite here but also speaking from experience.

  • most people want to just consume content without typing in what content they want to see

    I think that's always going to be true, but making it easier to produce will definitely increase the number of people who otherwise wouldn't bother.

  • Thanks for the input! And no I agree with your assessment of the website, I have some plans of a much simpler redesign soon, and I definitely value that critique. I think my "digital garden" has been through at least a few dozen of iterations, and sometimes I just get too carried away with it. The good part being, the next iteration always starts out better than the previous one (or at least I hope so :)).

IS that a real person's website? It looks entirely AI generated, even the copy and sample projects.