← Back to context

Comment by paimapi

3 days ago

there's a wide array of assessments when it comes to reading comprehension. the one you refer to, the GRA, sets the 'sixth grade level' as whether or not a reader understands the author's main points, is able to answer conceptual questions related to the text, and then apply those to relevant situations. beyond this level is the ability to essentially be skeptical of a text and to know how to critically analyze it. so if your comprehension level stops before this you get 'big words in complex sentence structure sounds smart and right so it is smart and right' even if the reasoning and process is poor

it makes me think about how people engage with movies and television - as passive, plot-and-character driven consumption (eg I hope Walter White survives) with no critical analysis of how and why the writers added ABC thematic element (eg Walter White as a motif of a toxically masculine narcissist with specialized knowledge as a larger critique how mass media tends to valorize their male leads in the same vein as many other prestige shows at the time like Mad Men), and the larger, downstream sociocultural impact that piece of media has on how people see the world (eg people who now have the Heisenberg tattoo, unironically)

there's been some musings on why this the case like Hofstadter's Anti-Intellectualism in American Life - the valorization of obedience and trust in hierarchy and the state are net wins if you're an institution that seeks to increase it's power, whether religious or governmental. I was talking about this with a few friends the other day and it's a dismal future reality where not only did we make anti-intellectualism normalized and politically legitimate in the USA (eg Fox News, clickbait articles, and all the other forms of yellow journalism that have emerged), we now have tools by which individuals can even further remove themselves from having to critically engage with thoughts, feelings. I heard a story about how someone scanned a group activity at a baby shower into ChatGPT and had it answer for them instead of, well, socially interacting with the other guests and forming a memory of the moment with their friends

the counterargument to that might be that Claude/ChatGPT/etc have more epistemic rigor than your average American (sure) but the sycophancy of modern day LLMs is an actual danger that enables more harm than good. it does seem as if Claude is the only one interested in guarding against some small amount of it (though to the detriment of people just trying to get work done. as an aside, I get the feeling Mythos was intended to be the bespoke enterprise solution without the guardrails but the Anthropic marketing department or some power-hungry department lead made it about how dangerous/effective it was from a security perspective which threw a wrench in things). but then I think about people like my parents asking ChatGPT which specific house to buy in their retirement only to later find out the house was sold weeks ago, or just in bad condition, or in a neighborhood where the housing value has already reached equilibrium, it makes me think about how it's not enough and the future is bleak

I'll also say that I think Claude sounds the way that it does because it, like many other LLMs, are RLHF trained largely by lowly paid gig-workers, many of them ESL speakers. if their trainers were, for example, dedicated and highly trained academics, scientists, and other researchers, you'd likely see a lot more concise and more importantly skeptical reasoning and responses. but that won't happen in our current reality of capitalist-driven development so we get encoded solutions like MoE that still largely depend on the messy, imprecise RLHF training at baseline

in the right hands, I do think AI is a wonderful tool. one of the first things I did with it was to create a research skill that reviews white papers from the lens of someone who knows how to read/interpret research methodology, is aware of things like p-hacking, and deterministically assigns weight according to the hierarchy of evidence. even still, I'll still read the studies because there's so often nuance that's missed if the sub-agent read only a search snippet but that takes effort, time, and the practiced knowledge of critical analysis to even want to do it

I’m inherently skeptical of big walls of text like this these days.

(So here’s a big wall of text of my own!)

However, a lot of what is written here makes sense.

And particularly “if your comprehension level stops [here] you get 'big words in complex sentence structure sounds smart and right so it is smart and right' even if the reasoning and process is poor”

This is exactly the problem.

And another point you make:

> but the sycophancy of modern day LLMs is an actual danger that enables more harm than good

I don’t think it is necessarily the sycophancy that is the biggest problem (though that is definitely a problem) but rather the combination of authoritative sounding text plus “complete answers” which sound wholly believable but are deeply flawed unless you have domain expertise.

I moderate a forum that deals with people who face a relatively common but somewhat complex (and nuanced) set of legal problems.

The purpose of the forum is peer support, shared experience (“lived experience”) and community.

It’s not legal advice, though moderators will sometimes step in to highlight relevant legal resources (e.g. case law/precedent or primary legislation/instruments).

Prior to AI infecting the forum someone would post their problem, people would respond with their often incomplete or poorly communicated thoughts, the OP would ask more questions - or argue - and a dialogue would occur. That created a community and people would post updates and ask more questions and find common shared experience. Many of them became correspondents with each other and some became actual friends.

In the past 12-18 months the discourse has changed from “here is my personal experience and here is what I did” to “here’s a bunch of stuff an AI says and I’m pretending it is me giving advice”.

Almost without exception the person who has started the thread will react positively to the AI generated content, even when it is egregiously incorrect - but won’t ask questions.

More problematically, these AI posters will often argue specific incontestable points of law “because I asked ChatGPT/Grok/Claude and it says this” and ChatGPT clearly cannot be wrong. And the border of precedence seems to be ChatGPT, Grok and then Claude some way behind.

I’m slowly seeing a pushback from people as “normies” begin to spot AI. But it’s ruined a community because the advice sounds so authoritative and complete that people won’t argue or ask questions.

As a result we have banned AI generated posts and remove repeat infringers.

That’s significantly reduced the volume of posting (below what it was pre-AI) but has significantly increased the value the members are getting.

  • I do appreciate the thoughtful response to a really long wall of text lol. and yes, I agree - I think that'll be the lesson that society is going to take probably far too long to learn, to not see everything as a nail that AI can hammer at. a lot of tech companies are in essentially a 'fuck around and find out' phase with AI taking over code review, testing, etc. combined with the expectation of shipping 3X the amount of code, we've enshittified the entire SDLC. and so we have near-daily incidents, data leaks, etc, something that I was able to measure and report on at my old place of work to, well, no avail

    it's the old tortoise vs hare parable, I think. go fast, make a bunch of mistakes, get too arrogant, and you lose out. your forum might be slightly lower engagement now while people are caught up in the latest fad but your rules are proactive for a future where average people hopefully realize that you can't trust an LLM that has zero context, no real harness and determinstic tests to speak of, and a propensity towards probabilistic rabbit holes that result in hallucinations. at least that's the kind of space I'd look for now and largely why I've given up on a lot of other forums

    • That’s quite encouraging to hear, because it aligns with what we are trying to do.

      Which is basically weather the AI storm and come out the other side with something that is essentially purely human.

      And then we might - where appropriate - use AI to help surface or explain relevant external content. “Idiots guides” but human reviewed.

  • You are fighting a good fight! Props.

    • Good fight / entirely thankless fight maybe.

      I can’t help feeling like this is the last gasp of the old internet. Those tiny corners of expertise can so easily be eliminated through a few months of “AI! SHINY!” and there’s no coming back. I’ve seen a couple of other communities decimated by AI. The participants start posting AI slop and then remarkably quickly everyone else just stops commenting. It’s awful.

  • The problem is one of expertise, sometimes general, sometimes specific.

    If you don't know better, you don't know better to question what the AI says.

    I've seen this in the work environment with a coworker who insisted that I implement my side of the control system using the control law ChatGPT recommended instead of building off the empirically tuned control law. I eventually sectioned off a part of the codebase for him to work on independently.

    Needless to say he didn't get a whole lot farther.

    Later characterization of the entire system end-to-end showed the existing system was already close to the theoretical limits and ChatGPT's tearup would have bought us precisely nothing except for more work to tune the new control loop.

    And I see this in everything that requires expertise. You need to know enough to know when it's bullshitting you, and it's hard to be enough of an expert in everything to tell when it's bullshitting you for something you aren't enough of an expert in.

    • I see it as a problem of context which maps to my theoretical understanding of LLMs. models trained on large data sets will probabilistically veer towards the median in all aspects - reasoning, assumptions, environments, etc. specialized context about your specific codebase's solutions don't exist unless you add them in, either in the prompt, as a skill or rule, or more generally in the harness via memories, tests, etc (though ideally a combination of all of the above). without that the LLM will suggest the median solution for the median codebase according to some ephemeral, unqualifiably trained understanding of best practices

      it makes me wonder if the solution that businesses/users need to implement is just the same solution to everything since the beginning of time ie standardization. skills/harnesses/agents.md/etc maintained by codeowning teams that must be invoked for AI-assisted code changes on ABC part of the codebase, these existing as replacement for the bevy of other documentation required for the days of hand-written code. a company-wide orchestration skill knows how to search and pull down the relevant .mds, cleans it as cruft at the end of a session, every merge with a short changelog saved to a corpus somewhere with a TOC + appendix that an LLM can navigate to and read for context, major changes in the logic documented in the working skill doc, all of it generally automated but requiring HITL vetting

      this wouldn't fully solve the problem of subject matter expertise but it seems like it would remove a lot of the friction for new employees and other teams with dependencies on your work or with whom you have dependencies

> it makes me think about how people engage with movies and television - as passive, plot-and-character driven consumption (eg I hope Walter White survives) with no critical analysis of how and why the writers added ABC thematic element (eg Walter White as a motif of a toxically masculine narcissist with specialized knowledge as a larger critique how mass media tends to valorize their male leads in the same vein as many other prestige shows at the time like Mad Men), and the larger, downstream sociocultural impact that piece of media has on how people see the world (eg people who now have the Heisenberg tattoo, unironically)

That's not the only smart-person way to read that show. And even if a character has flaws, or even if it's an outright villain, people can still like the character. If I tattoo Scar on me from the Lion King, does it mean I didn't understand that he's not a positive character? I can still think he's cool. I'm sure people also put Darth Vader tattoos on them. Also you're using phrases of political ideology that one doesn't have to subscribe to in order to enjoy the series.

  • sure, that's very much the 'just let people like things' argument where literal white supremacists can enjoy Rage Against the Machine in spite of the music literally being in total opposition of their ideology

    everyone's free to enjoy media however they want, with whatever level of interpretation they like. I provided the BB references as short examples, they aren't meant to represent the definitive diagetic experience of the show. if you have a different view, great. if it was triggering for you to hear 'toxically masculine', also fine but... might be something worth self-examination on given that it's very much also a clinical and academic term [0] as much as it is one steeped in the artificially manufactured culture wars by people who don't want to change their anti-social behaviors

    I would also say that understanding Scar as a villain is the sixth grade reading level understanding of the character. and you're free to stay at that level of understanding. someone who wants to engage more critically might map the character to Claudius, comparing and contrasting how they're characterized given the context of the audience for Disney and Shakespeare, and appreciate the character that way, as a standardized trope utilized throughout all other forms of media. they may even get a tattoo of Scar, symbolizing their dive into the analysis

    my point is not that people should or shouldn't engage critically. it's that this practice of critical engagement, of being skeptical and analytical provides you with the skills to not be a total sucker who falls for the latest manufactured fad that someone with a strong theoretical understanding of semiotics and social capital created (ie most modern marketers). the pertinent example being how people engage with AI - seeing it either as a specialized tool with a set of flaws that need to be accounted for and checked against or as just some kind of authoritative voice because it sounds smart and so-called smart people like Elon are terrified of it and AGI. which, again, if you prefer the latter engagement all power to you but the chances of you taking some really bad advice forward is not negligible

    [0] https://www.wi.edu/news-Shifting-the-Conversation-From-Toxic...

    • You're the live stereotype of that middle-of-bell-curve meme thing. Using thesaurus words is no longer impressive. And that academic world you refer to is a navel gazing self referential nothingburger.

      It's like a jumbled up string. You pull on the two ends and it turns out to be just a loop, it resolves to a big null.

      5 replies →

> I'll also say that I think Claude sounds the way that it does because it, like many other LLMs, are RLHF trained largely by lowly paid gig-workers, many of them ESL speakers. if their trainers were, for example, dedicated and highly trained academics, scientists, and other researchers, you'd likely see a lot more concise and more importantly skeptical reasoning and responses. but that won't happen in our current reality of capitalist-driven development so we get encoded solutions like MoE that still largely depend on the messy, imprecise RLHF training at baseline

No really, that's not particularly accurate, they use so much gig work because no-one else wants to work for them not because they would be unwilling to pay a little extra, or only want the absolute cheapest labor they can get on the planet.

They want senior white collar professionals and scientists and researchers especially since these companies already on some level believe their models are as good as any senior employee in any field (it's probably the models generating text saying that, but that's besides the point). But who's going to work on contract for a company that wants to automate them out of a job? Realistically no-one unless they get some shares in the thing that will destroy their future earnings potential and ability to control their own destiny if it works out.

But they can find enough educated white collar professionals on unemployment or in unstable academic employment that will take an extra job on even if it's only 50 $/h or 70 $/h and compromise on any solitary they might have but the work output you get from that is only going to be as good as what you ask for, if they had better respect for the professions they want to automate, it would be better.

Like is that an acceptable wage in the US for difficult skilled work, not particularly but it's not rock bottom exactly, and it's not bad for other English speaking countries, working conditions and stated mission are more of an issue than being cheap.

Training pipeline on a modern LLM is also going to be quite indirect during the long tail of post training, and heavy on automated RL, the human feedback might end up getting used in the form of automated grading guidelines like what you did for research, with the same issues as that, compounded by the input being LLM generated and models being biased towards model output by default. It's more of a feedback on the loop rather than in the loop.

  • so why don't people want to work for them? they don't get paid enough? what if they were paid more? what if they were FTE with all the benefits? what if AI projects were nationalized and trainers were funded by grants? what if we increased the NIH budget, made peer review and journals far less exploitative of researcher's time, and got rid of academic middle management, focusing mostly on paying more towards actual research and academia?

    definitely a utopian vision that is not likely to happen in our current reality but I like to imagine better worlds that are possible. as LeGuin once said, "We live in capitalism, its power seems inescapable — but then, so did the divine right of kings. Any human power can be resisted and changed by human beings."