← Back to context

Comment by jjcm

3 days ago

Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this.

Original images: https://image.non.io/neonRamenDesigns.webp

Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7

Opus 5 build for comparison: https://html.non.io/neonRamen

Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM price wise, which is Grok 4.6: https://html.non.io/neonRamenGrok4.6 . I thought Gemini would blow Grok out of the water (it generally has in the past), but Grok has really caught up.

Other thoughts: I really think Google has fallen behind here. Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf... .

It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.

  • Can you help me understand how it is hard to get an API key from Google? You just head on over to http://aistudio.google.com/api-keys and create a key... not any different from platform.openai.com?

    Disclaimer: I work in Google so it might be that this link is not publicly well known

    • Disclaimer that I haven't tried this since January, so things may have changed in the last 7mo, but this was my experience at that time: https://x.com/pwnies/status/2010523020629274723

      At a high level though, as a rule of thumb Google assumes that they're serving companies at Google scale first, and at a human scale second. For other companies it's the opposite. Generally what that means is the first experience you get with a Google product will route you through 8 different dashboards to set up ACLs before you've hired your 2nd employee.

      33 replies →

    • Personally I went to https://console.cloud.google.com since I already had some GCP projects. Then I searched for Gemini API Key. It brought me to https://console.cloud.google.com/agent-platform/studio/setti.... Then, there was a banner saying "Enable APIs to access full platform capabilities." Then, I did that, which took quite a while (minutes). Finally, I was able to see the way to create an API key.

      The fact that there's two ways to get keys is also very confusing.

      5 replies →

    • You need to create a Google Cloud project to create an api key and when you try to create one you very often get error messages like:

      “Failed to create project, The request is suspicious. Please try again” or “ You do not have permission to create a key in this project”. You can then navigate multiple screens in GCP to make it work but it’s a hassle compared to any other provider (OAI/Ant/OpenRouter or any of the Chinese labs).

      2 replies →

    • We have probably 10+ years old account with Google cloud etc. We recently had a production deployment, I went over to AI studio to get new keys and it kept failing saying "Failed to generate API key, The request is suspicious. Please try again" - It was through my standard browser, same geo-ip. And it just worked after 2 days.

    • > Can you help me understand how it is hard to get an API key from Google?

      Using Google products in general is an effing nightmare as soon as you have to give them money.

      The one thing you want in a business is to remove friction when people want to give you money, a concept Google has never been able to understand.

      2 replies →

    • Maybe things have changed but it was a big mess trying to getting an API key from Google as an individual a few years ago. Way too much conflicting documentation.

      Eventually I gave up and run a few hundred million tokens (edit a few billion) through openrouter.ai using Gemini Flash 1.5 to Flash 2.5

      Every since price increases on Flash 3.0 I've stopped using Gemini, too expensive for basic classification, sentiment detection, ocr etc.

      As other posters said Google assumes you are some bigcorp trying to use their products. The Vertex versus AI studio confusion did not help.

    • I tried to use Gemini for one of my projects a month ago. Immediately after signing up and paying for credits, I got an email saying, “Action required: your billing account {redacted} is past due or has invalid payment information.” I have no idea why it says this. My credit card on file works. My balance updated with a new amount from that card. 11 days later, my account was terminated. I still don’t understand what went wrong or how to fix it.

      I do pay for OpenAI, Anthropic, and ElevenLabs keys.

      2 replies →

    • I have a Google account that I had tied to a domain I originally purchased from google (when they had the .dev offerings), and now that it's been purchased by square space my ability to use AI through that account is in a bizarre state. Even my free gmail account has more AI offerings, and I'm not allowed to pay for improved AI offerings on my custom domain & google workspace account.

      I know this is probably a pretty small edge case, but it is a bit frustrating. Any other provider lets you sign up with an email and give them a payment processor/card, but because google wants me to only use their unified workspace for signing up, I'm completely locked out now.

    • it worked for me okay when I needed it for myself in my personal account. But when I tried setup this for a company I spent almost a day solving lot of small puzzles in GCE like how to tell CEO that he have to connect billing account created for other purposes (and he not even remember at time that it exist) to new project and all other things that others talking about.

    • As someone who has been running Gemini models in production for a year, recently (last 2 months), I have been actively moving away from it.

      The primary reason for me has been that Google autonomously decides to downgrade usage tiers and then upgrade them again - and does this incorrectly.

      Over the last week itself, in the span of two days, our account for first downgraded and then upgraded. This is despite matching the criteria to remain at the tier we operate at throughout.

      Google Support (when you finally get to a human) has accepted that these are potentially bugs, but the first time it happened, we were rate limited so severely for ~4 hours that I find it really difficult to continue trusting Google.

    • Simple question - can i use Gemini 3.7 flash with a subscription in my own harness and not in agy client ? You're from Google so the question.

      By the way - I love Gemini's personality . it's phenomenal to work with

      The reason i want my own harness - is the custom tools that i provide vs the low tier tools that come with the custom harnesses.

      You'll would really benefit , if we could use Gemini in our own harness and not be forced to use it via agy . i've tried using gemini to circumvent - but not been successful.

      If you see this - please reply here

    • On top of being the hardest website to navigate, Google console a) doesn't have real time billing (!) b) doesn't allow you to set a budget limit.

      Sorry but it's not worth waking up with a 100k$ bill, fix your platform first.

      4 replies →

    • Same experience. It took minutes to find the API key from the console.

      Now following up with Google support team without luck to find the logs. Prompts send to the model and the responses including the thinking was available in the ai studio. But it’s unclear where to find the same in console.

      To make matters worse there is vertex api and rebranded to Gemini something and making it very confusing.

    • i use gemini api in the third-party platform like openrouter and evolink.ai , easier to get key and manage my bill

    • That is easy but I’ve also found myself in account setup dashboards that were obviously geared toward enterprise trying to set up access to tinker with something AI related. It might have been TTS but it’s been a little while and I can’t quite remember.

    • The problem is that this doesn’t work for enterprise. The rate limits of that is super low. Then you need to migrate to Vertex and that is just a pain. Who ever thought of using JSON instead of an api key…

    • I even know the link existed but forgot what it was specifically and couldn’t remember what AI* property it was offered under and it took me a long time to figure it out.

    • I haven’t tried in about a year, but I could never do any meaningful work outside 1P Google apps (antigravity) due to such fast throttling.

    • I just asked Gemini how to do it and that's exactly what it sent me. Was up and running in a few minutes.

  • I use LLMs rarely, and only for digging into subjects which I can't find enough information using search engines. I only tried Claude and Gemini, but Gemini both returns faster and higher quality information which I can use for more targeted digging myself.

    Google being Google, their models tend to be better at finding, organizing and presenting information, from my experience.

    • Yep, I use Gemini for this too and it’s great - very fast and high quality.

      I’d be very willing to try it out as an API, but it’s far too complicated to set up payment, and I don’t want to risk taking a wrong step and being locked out of other Google services. So Anthropic and Mistral get my money instead.

  • Sol on Cerebras is going to be expensive AF

    • Is it? I think waferscale might actually be cheaper per-token, it's just so many more tokens, and of course right now it's not a full buildout so the availability is limited as well. I'd imagine they'll be migrating to whichever inference method is least expensive, and I expect asics to be the ultimate answer.

      1 reply →

    • Moving from either frontier intelligence or frontier latency to a single model that does both at the same time is potentially a game changer in certain industries. I can easily see e.g. hedge funds dropping tons of money on this, because it means they can now do the same thing as their competitors, but much faster. That's basically a license to print money.

      8 replies →

  • I like using 3.5-flash-lite for doing cheap PDF and Image data extraction stuff. I don't think there is a better bang / buck model right now (3.1 is cheaper but a lot worse).

  • The API key you are mentioning is just ridiculous. Onboarding your company or personal account is a trap. I ended up getting assigned to sales guy just to test their Vertex API because I used a company email.

    Of course, we just used OpenRouter for testing and never touched a Gemini model anymore.

  • > but I just don't know what situation I'd reach for 3.7 Flash

    You reach for it every time you do a Google search

    • > You reach for it every time you do a Google search

      [my self-important Kagi shtick awakens, pokes at it's restraints]

  • Depends on the definition of friction. If someone is in the Google ecosystem, why would they reach out of it.

    • I already use GCP and Google for work, and getting an API key was so annoying that even I couldn't be bothered after a while of looking around.

      Maybe things there have improved some, but when I was looking it was a huge runaround.

      2 replies →

  • It's much cheaper tho. Junie says Fable is 5-10x more than default model (Gemini 3 Flash Preview).

  • > especially given how hard it is to get an API key from them

    What does this mean? Anybody can get an API key

These look exactly like all of the low budget bodega signs near me. They also look like a bunch of cheap ads for parties that I keep seeing. The sameness of style is uncanny.

(I don't have the Bodega signs, but I'm thinking of shit like this, from a quick google: https://linkstub.com/en/wet-wild-foam-party)

  • Sigh. You might be interested in this piece that came to mind from Doctorow:

      “Let me explain: on average, illustrators don't make any money. They are already one of the most immiserated, precarized groups of workers out there. They suffer from a pathology called "vocational awe." That's a term coined by the librarian Fobazi Ettarh, and it refers to workers who are vulnerable to workplace exploitation because they actually care about their jobs – nurses, librarians, teachers, and artists.
    
      If AI image generators put every illustrator working today out of a job, the resulting wage-bill savings would be undetectable as a proportion of all the costs associated with training and operating image-generators. The total wage bill for commercial illustrators is less than the kombucha bill for the company cafeteria at just one of Open AI's campuses.
    
      The purpose of AI art – and the story of AI art as a death-knell for artists – is to convince the broad public that AI is amazing and will do amazing things. It's to create buzz. Which is not to say that it's not disgusting that former OpenAI CTO Mira Murati told a conference audience that "some creative jobs shouldn't have been there in the first place," and that it's not especially disgusting that she and her colleagues boast about using the work of artists to ruin those artists' livelihoods.”
    

    https://pluralistic.net/2025/12/05/pop-that-bubble/

    OK so after all that, THIS (bad foam party posters) is what we get!

    • I like this concept of "vocational awe", and I do think it's why, for example, there is so much sexual and financial abuse in the movie, music and game industry: those people are willing to be treated poorly so they can work their craft - do the thing they love. Basically they are trading away good working conditions in exchange for actually doing work you like or find meaningful, the opposite of someone who does something they hate or find meaningless but pays a high salary and where you are treated really well.

      I don't know how well this actually translates to AI, though. We all understand that AI art has a pretty low quality, but sometimes low quality is enough. If I want an image for my D&D character that only I and my DM will likely ever see, I am fine with a 7 cent AI-generated image, but I'm not willing to pay 150$ for an artist to do it - not that I don't value their time, I just don't value the image that much. Before AI this was the same, I'd just have used an image from Pinterest and thought "Well, this isn't exactly a good match, but I can't find anything better".

      But I assume real illustrators do things like illustrations for children's books? I'd like to believe that those are still done by actual people, not AI.

      1 reply →

  • All AI "art" is proving out to be basically crap, and the good thing is that it has started to become a very good filter, as in whoever uses AI art most definitely is doing a shitty job with the rest of his/her business so it's better not to bother with said business.

    The thing is though that most of the nerds here are oblivious to all that, so it will take a while for them to fully acknowledge what's happening. In a positive note, I'm here on a AI thread discussing AI-image generation and I haven't yet seen any link to the dreaded pelican on a bicycle thing, so hopefully we're getting into the right direction.

I think both outputs are really good. I don't see a lot of differences. So what exactly should be looking at and notice that one model did worse or better than the other one.

EDIT: OKAY I see it's mostly the "image" generation, not so much the HTML... Noticeable in the food photos and the foodtruck/cart photo

  • clicking the add buttons and scrolling the menu is just much better in Opus 5. It feels like an actual website vs a simulation of one.

    • i do feel like opus here is the most natural. theres something off about 3.7 and grok while its an improvement feels flat and not complete

      i do wonder why gpt sol was not compared here but honestly it's not really known to be the best at UI

      a fable 5 comparison would've been also interesting and likely the best.

There is some irony being a developer and reading along the lines of: "oh look at the comparison between these models executing a task for a few cents on a job i'd be charging 1k minimum"

  • FYI, developers are rarely given such a rich UX mock.

    • I don't think I'd say rarely. Companies rarely allocate the design resources to produce that, but the companies that do are typically much larger, so the actual number of individual developers that get rich mocks is probably closer to 40-50%.

      1 reply →

    • Depends who you work with, what's the intention, budget, etc. I'd agree this is a really good one.

      I'm used to incremental Figma wireframe -> final product and working together with a designer.

  • I know what you're saying but this is the most tedious, soul crushing dev work there is

  • meh I have no interest in that type of software dev anyway

    • Unfortunately, as soon as they can’t find work, everybody interested in the more easily automated dev work will suddenly become very interested in up-skilling into all other kinds of dev work. So then you have an excess supply, which means little job security and littler salaries. Developers were in the cool kid club in SV because it was more painful to fill dev roles than to treat developers with kid gloves. Without high labor demand, there is no leverage for developers. Increasingly, management has the leverage. Oh well.

How are you doing this with Opus. Clearly I’m missing something. I always turn to ChatGPT when I need images because Opus typically refuses. I’ve tried Claude Code and Claude online in the past. I’m pretty sure neither created images for me and I thought this was because Anthropic was focused on code.

I guess I need to try harder. :)

  • Images and build step were generated with my own tool (https://news.ycombinator.com/item?id=48995754 - it's why I'm often running these img->html tests).

    Opus can't generate images since A\ doesn't have a diffusion model.

    • This is pretty interesting take on design and your trick to make the surface maps is pretty slick!

      I did have one question about the tool, is it possible to set how many variations you want per step? I would rather be able to guide it manually at some steps where maybe I know pretty well what I want or just need minor tweaks, and then let it loose on others when really trying to experiment with an idea.

  • I believe they are testing giving it an image, which you can do in Claude code by dragging/dropping into the terminal or copy/pasting, and asking it to build the html equivalent.

One of my favorite image tests with AI models is schematic analysis...I build and repair tube amps for a living, and use AI for such work a LOT.

so far, IMHO, the best has been opus and fable\mythos.

  • Nice! I test if models know which tubes I can use for an amp given the power and number of pins.

I'm curious how much the harness plays into this. I'm somewhat surprised by the gemini and grok results, they seem to have strongly deviated from the original images. I'm thinking maybe the harness has a big effect? It's possible to proxy in different models to claude code, if you're curious you might find it interesting to test!

  • Harness could be a part of it, but worth noting both the Opus and Gemini 3.7 flash tests were both ran through opencode.

    The grok test was ran through the cursor cli agent however.

There's one specific aspect I like better with the grok version: The prices are above the fold. ON the Opus and Gemini versions, I have to scroll to see the full menu item showcase.

I'm not sure what prompt you put in but did Gemini replace the all of the images in the original with its own? That would be really weird behavior unprompted.

  • The prompt is a build step generated by my tool for image->html conversion, which includes APIs the model can call to generate images/patterns/svgs.

    https://image.non.io/12275ee8-71e9-4941-823b-e51fec157b4d.we...

    The agent is told to generate assets as part of the buildout. It gets to decide what the prompt is for them / whether to do postprocessing like background removal / what type of asset to generate.

How are you prompting it with the original images to create such sites?

Was the original concept generated by Claude somehow? It gives me Claude UI vibes with all the extraneous small-caps text elements.

I'm not sure what you consider good design, but if it's subjective, then I see it differently from your examples.

Gemini 3.7 looks the best. Opus 5 looks almost as good as Gemini. Grok 4.6 looks pretty terrible.

They both have horizontal scroll on mobile...

  • Oh yea, as a disclaimer the models didn't have any instructions to do a mobile version. I haven't tested them on mobile at all.

This has so much less character than the pelican smdh..

Plus is the ramen in HK even any good?

  • I think they're both testing very different things. The pelican test is testing if a LLM can come up with visuals on its own via writing bezier curves directly.

    This is testing if it can match visuals that have already been established, and represent them with all the tools available to a web developer. The ramen example was chosen in particular because there are a lot of things that aren't easy to do with CSS, and require creative strategies: 45deg button cuts, angular repeating pattern elements, blending of raster art and svgs, microglyphs, low contrast subtle elements, etc.

    Don't ask yourself whether it's a good design, as yourself whether it's a good test.