← Back to context

Comment by bluegatty

5 hours ago

This is a failure of the AI foundries; if we have to use totally different prompting techniques for every model, this wont work.

AI is rapidly saturating it's ability to be useful and these products need to start to mature.

It's not 'fun' to manage 50 different broken MCPs and their variety of ways in which they are broken.

It was 'fun' at the start, now it's just 'broken technology'.

Astra and Opus 5.5 are the 'starting point' for the next era of AI where we expect robust tooling.

All LLMs understand natural language. All LLMs understand examples. That's honestly more compatibility than you get nearly anywhere, in anything.

The reason why advanced prompting is a moving target is that a lot of prompting is "use extra instructions to compensate for specific ways in which the target LLM is weak or prone to errors". And guess what? LLMs get better over time - obsoleting your advanced prompting.

"Tune a prompt to death for the specific task and the specific model" gets you better performance in the moment, but "trust LLM to be smart" ages a lot more gracefully.

  • "but "trust LLM to be smart" ages a lot more gracefully."

    That it doesn't even work now.

    The word 'smart' there is actually doing a lot of heavy lifting, it's entirely contextualized.

    • So much AI discourse tacitly assumes that there's an objective quality that corresponds to "being smart", rather than a chaotic patchwork of extremely contextual social practices and expectations.

      1 reply →

  •     And guess what? LLMs get better over time - obsoleting your advanced prompting.
    
    

    It's nowhere near that simple. For instance, models used to be WAY better at writing, until the labs decided that coding ability was a better thing to focus on, and trained successor models accordingly.

Agreed.

I'm genuinely worried about all our short term investment in mitigating the failure modes of models that may only be SOTA for a few months.

It's very possible people being 'late' adopting AI may end up with a leg up, not only because they spent more time polishing personal skills during this time, but also because they don't bring all the baggage of 'AI competence' that is becoming irrelevant at breakneck speed.

  • Suppose you use LLMs in a more straightforward way, like asking coding agents to make specific changes rather than attempting to build a software factory?

    That could be seen either as early adoption that’s overfitted to current capabilities or as late adoption of LLM’s more advanced capabilities.

  • The people who are adopting LLMs later are also slower to adopt new technology in general. The samples of people who adopt early and adopt late have different characteristics.

  • This is “second mover advantage.” There are several dynamics that can make it better to wait and move later. Framed in terms of firms, moving late is advantaged when:

    - the product category is long lived

    - switching costs are low for buyers

    - there are objective standards of quality

    - product imitation costs are low

    https://insight.kellogg.northwestern.edu/article/the_second_...

    Let’s consider those criteria for an individual competing in the labor market with AI. The category should be long lived, AI is here to stay. Switching costs (here, hiring/firing by employers/clients) are low. Objective quality standards fails; technical labor is notoriously difficult to quantify. Imitation costs (can you copy someone else’s good ideas) are moderate but decreasing. That’s where model and tooling improvement shows up.

    Based on this analysis, I agree that late movers are well positioned IF the market leaders continue to improve models and tooling to integrate best practices that were previously individual skills.

    Early movers should exploit the lack of objective standards. Use your experience with the first generation of tools as marketing to win and retain clients. Continue to invest in soft skills like communication.

> It was 'fun' at the start, now it's just 'broken technology'.

It was even more 'broken' at the start. We overcame some of the issues by 'prompt engineering', which is needed less in the newer, smarter models.

  • Of course - what I mean to say is that we did not perceive it as broken.

    The first combustion engine was a miracle. It only becomes 'broken' when we evaluate in some kind of applicable context.

I really want to know if I am doing something wrong so let me know

I don't use MCPs, agents.md, skills.md, plugins, nothing. I just open a DeepSeek Harness workspace and start a brainstorming session with a request for an architecture.md prompt.md and plan.md files, then I go prepare coffee while it does all it needs asking questions along the way and writing them in decisions.md so it understands why we took that route

Minutes later a fully functioning product that I run, check it complies with the initial plan and then ask for minor cosmetic changes

I've been doing it for six months now while I see posts and posts about people making their harnesses do things I don't see the need for. Why so complicated?

No special prompts, no rehearsed inputs, just a simple "Hello my friend, today we are going to create an app for transportation, ask all the questions you may have and at the end write an architecture.md ..."

It works, it is simple, it is enjoyable, like a friend of mine and as such we treat each other

  • "Minutes later a fully functioning product that I run"

    There are very few people who operate in this kind of environment aka 'small new product from scratch, move on'.

    Like if that's what dev was, this would be easy.

    Also FYI is no such thing as a 100x developer, other than some very senior architects who's wisdom and guidance affects the outcome of gigantic projects.

  • Are your projects standalone? Because why I need tooling is to carry over the learning and knowledge from one project to next.

    • Most of them are but when I need it to learn something from one already done I just point it to that workspace, like "For the invoices app take a look at inventory workspace and learn about X and Y" or "For Prince of Persia find all available code and references online and propose architecture changes to adapt to Swift on MacOS" so yes, it learns from both local or online sources

You'll find similar documentation anytime a language or framework or other systems software ships a new major version. It doesn't seem like the way to prompt Opus has changed all that much. Certainly not enough to require a "totally different prompting technique."

Counterpoint, the differentiation is maturity. If all models are simply interchangeable commodities, what's the payoff for Anthropic or OpenAI?

Vastly different ways of interacting with each provider is another story, but really we are pretty spoiled here. Slightly different prompting techniques is not really a big deal. If anything it shows the user has some nuance and appreciation for what each model provides.

Fow what it's worth, I am super happy with Opus 5.5. Less verbose than 5 and just gets work done. The progress has been astounding, and if I have to coax it out a bit differently on Opus 5.5 vs Astra 6, I am happy to pay that small price.