Can a MUD evaluate LLMs? A $99 proof of concept

6 days ago (cruciblebench.ai)

I'm the author of a paper my friends and I wrote after we were curious if a MUD, text games originating in the 1970s, could be used to evaluate LLMs. We've spent the last several months on nights and weekends running this experiment and writing the paper on just our personal computers with about $99 in API credits.

Our experiment did have an interesting leaderboard but even more surprising was the measurements of each LLM. We scored each on four behavioral dimensions, two of which lean heavily on an LLM classifier. When we removed those two, one of the frontier models fell six positions. When we then checked the classifier against a second judge, the per-model agreement between them ranged from 85% to 22%. The aggregate kappa (0.04 on probe detection) indicated the instrument was noisy without saying which models the noise was hitting. The most affected model shared a model family with the classifier. This isn't proof of bias, just one observation we recorded.

We realize LLM judges can be unreliable, and while it wasn't our original intent to test this, it ended up being the most interesting finding. The divergence between the two judges is the finding we think generalizes to other judge-based benchmarks.

We emphasize this is just a proof of concept and not a validated benchmark. We prepared a thorough limitations section in the paper, including just 50 runs per model, overlapping CIs among the top models, no human raters, a tiny environment, etc.

Everything we did is publicly available, the paper and data are CC BY 4.0, while the code is MIT. The paper, transcripts, code, and complete API billing export can be found at https://doi.org/10.5281/zenodo.21386663

If you find issues, please let us know, that's why we're sharing this. We're currently designing Phase 2 and want it to be as robust as possible. We're looking at human baselines, multiple judges, more objectives, a larger environment, etc.

A little irrelevant, but I need to vent about something MUD adjacent.

I really wish MUDs were still a thing. The text-based roleplaying communities were largely swallowed by Discord, where as before they existed scattered over proboards, jcink, and a few others.

Discord, being the imperfect text platform, limits the text and leaves most interaction limited by awkward formatting. To boot, even when a better ecosystem is invented, players refuse to give it a shot even if directly invited, because they got comfortable with Discord's terribleness.

It's so frustrating, and I don't know how to voice it, or even what can be done.

  • MUDs never disappeared. People stopped finding them engaging enough to use. Blaming Discord reverses the causality. MUDs lost engagement long before disco was created. People congregate on disco because it offers a more compelling social interface, even if it lacks the depth of a persistent game world. The opportunity I'd draw from this isn’t to convince people to return to older and less engaging interfaces. It’s to build persistent roleplaying systems around the interaction modes that have proven to be more engaging already. Potentially combine voice, speech recognition, text, tts, and synthesized characters with a proper world simulation underneath. If roleplaying communities already gather in disco, then deliver the experience where they already are. The interface was always irrelevant to what made MUDs fun anyway. Imagination and connecting with other people was always the thing.

    • Is your autocomplete passing you disco instead of Discord and you're just going with it? Or are you implying that before the creation of disco music people had grown tired of MUDs?

      1 reply →

    • I specifically outlined non-MUD role playing environments because I know MUDs lost popularity long before Discord. I am outlining that any form of tech, besides Discord, is sidelines -- MUDs and a hopeful resurgence as well.

  • I agree with you. You've also got more / easier competition. IIRC, MUDs peaked in the mid-late 90s. MMOs were really the death knell, then the explosion of other types of games, mobile games, then to your point things like Discord.

    One adjacent project we're exploring for the future is if you took all of the current understanding of game design and mechanics with modern AI functionality, knowledge bases, etc, and built a MUD, what would it look like and would it be enough to bring back players from other games?

    • There is an ongoing project doing this (that launches this weekend I believe) to revive a MUD called Shadows of Isildur. I think a lot of effort has been put into the new changes to the codebase to bring QoL features that a modern audience would expect while keeping the bones of the game the same. I'm not certain how it will fare as, to be honest, MUDs take an inordinate amount of time to really groove into and immerse yourself in - but I am excited to see efforts like this invested on as we older MUDers find ourselves finally reaching the point of having free time again.

      4 replies →

    • Would be kind of a cool experiment for a mud-style interface, where each npc/avatar was backed by a lower-cost LLM... similar rules for the npcs in the game, but a background feed, and history of interactions with other players as background.

      Maybe limiting npc's to only a certain number of moves that aren't a response to other users per day...

      It could be a lot of fun.

      10 replies →

    • I honestly worry that the relevance of text-based systems is largely tied with literacy and how it interacts with the player's imagination -- something largely in decline, at least in the US.

      4 replies →

  • Simutronics is has been running and developing Gemstone and DragonRealms for 30+ years and counting. The communities are about 400-1000 players each. They've embraced microtransactions (in the form of quarterly events), but both games are entirely playable without participating in that aspect.

    With modern technology, training in both games is very automatable. But developing and optimizing your training routines is (in my opinion) incredibly fun. Actual roleplaying is less common it was in the 90s, but there are still nightly events where dozens of players will get together to socialize and explore the worlds in-character.

    • I highly recommend it, and for those who don't have the time to grind, you can join the shattered server where 24/7 scripting is allowed and encouraged. F2P subscriptions don't work on that server, unfortunately, but it has been a great hobby for someone like me who used to play back in the day but could never play it without scripting due to time constraints.

  • Technically MUDs still are a thing, there's a good number still running. Though sounds like you don't really just want a MUD where you want to go out mobbing but a more social thing (but not IRC?).

    I suspect MUDs may have a bit of a resurgence soon, I've seen more and more interest in them as of late simply because they seem like an excellent pairing with LLMs which easily consume text. Though I think for your needs someone would need to work on the barrier to entry to getting on a MUD as well as the modern UI/UX expected today.

    • Having stepped away from MUDs ourselves for years, it's been interesting to see what the community has been cooking up in terms of UI/UX, Evennia, MUDlet, TUIs, etc.

  • >I really wish MUDs were still a thing.

    They aren't as popular as they once were, but they are still out there. One of the hardest-core OG muds, Armageddon Mud (which was based on the Dark Sun world) recently went down, but the founders have been working on a new reboot with a totally new code base. Looks interesting (I played the original but have no connection with the MUD itself). They are aiming to launch BETA in the relatively near future if anyone wants to check it out.

    https://www.zalanthas.org/

    There are also some OG muds still chugging along after 30+ years, like the hardcore PVP mud Duris:Land of Bloodlust. The player base has dwindled mightily over the years but people still play and it is a load of fun to explore, if MUDs are your thing, as it incredibly expansive having been in active development for decades.

    https://www.durismud.com/

  • MUDs are still kicking around and if you'd like to play in one you'll find a few solid communities with good engagement that would love to include you. They're in an odd spot though since MUDs existed in two sorts of categories 1. I really want to adventure and text is the only way to do that (these sorts died off to MMOs) and 2. I want to roleplay - those survive but the mechanical portions of those games are becoming more and more irrelevant as new games are constantly supplying better systems and mechanisms than MUDs were able to.

    This leaves MUDs in a weird place since things like MUSHes generally have most of the roleplaying tools without as much weird code for attacks and PVP interactions being involved so I've seen those thrive more recently.

  • > I really wish MUDs were still a thing.

    About 15 years ago, for a while I ran a MUD over AX.25, because MUDs aren't nerdy enough on their own, and amateur radio - even packet radio - isn't nerdy enough on its own.

    All three of us that used it found it pretty entertaining.

  • The MUD I play (Aardwolf) averages about 200 players online at any time, but it's effectively dead as a MUD. There's very little player interaction, which is what made MUDs special originally. Part of this is that nearly all multiplayer features have been abandoned, and the rest of the features are centered around loot box equivalent features. Despite this, there are a couple hundred people who seem to enjoy the gamification. The MUD is highly optimized for scripting, so it scratches the side project programming itch for me.

  • > even what can be done.

    I think this could be solved with bridge bots to make your MUD multiplatform. You could have a dedicated web app, an IRC bridge, a Discord bridge, etc.

    Discord-brained users can still use your MUD and contribute to network effects, and you can provide an off-ramp ("I just want to play VarelionScape, maybe I'll just open the dedicated website on my phone instead of going on Discord and digging through the server list").

  • Fully agree here. I love MUDs. The shear level of creativity that you can have in an MUD is practically infinite, and MUDs are just so fun to play in general. Especially when you start adding in triggers and sound packs and all that

  • > Discord, being the imperfect text platform, limits the text and leaves most interaction limited by awkward formatting.

    I don't think I understand your meaning here. I'd much rather use Markdown than BBCode or related.

  • I have a pile of ideas and not enough time to get to them all, but one of them is the opposite direction. Embrace Discord, and build a MUD whose social systems are funneled through it.

  • You would probably like my project https://github.com/timbran-project/moor

    • Very cool! Why reimplement the "moo code", though? Is it "just" for backward compatibility, or did you determine that none of the existing languages can be easily modified to live in a MOO environment?

      I started playing on an LP MUD in the late '90s. Over the past 20 years, I tried a few times to implement a similar environment, only with everyone being a "wizard". LPC was also a prototype-based OO language with multiple inheritance, and it supported live development, with code for objects stored in files and dynamically reloadable. I would always hit some kind of blocker, no matter the language I chose for the implementation. It looks like MOOs did (and do) what I wanted, but I just wasn't aware of their existence. :( Maybe if I used MOOs as a model, I'd get better results... but I'm not sure if I want to go all in on DB. I tried looking up how the code is loaded into the DB in the book (https://timbran.org/book), but couldn't find it. Is there an explanation somewhere of how MOOs are bootstrapped from sources, how you can maintain sources in files on disk, and how new definitions are pulled into the DB at runtime?

      2 replies →

  • Mine is still running

    > telnet playlom.com 4000

    Owe my programming skills to it, and also dropping college, only to start a company later.

> the measurements of each LLM

I think you have the correct idea.

If you ask an LLM to solve a multiplication problem using reasoning without code tools, depending on the model, it will get 15 digit (15D) * 15 digit multiplication correct (123456789012345 x 998765432109876). It will take between 4000 and 8000 tokens if Sonnet. Eventually with enough digits, it will start only solving the problem 80% ... then 70% of the time until it has so many digits it will never solve. There will be a certain number of digits where it will not converge on a solution nor will it stop working thinking it can solve it. That is very, very expensive. It takes a lot of runs to determine the probability it will solve it.

What you can do now, this is likely the most important thing, is change the prompt and evaluate how many tokens and at what speed it takes to accomplish the task. Sure the measurements of each LLM are important! Nonetheless, if you can say to a company that you have a technique to tune prompts and evaluate them so instead of spending $1 X 100,000 times a day, they can instead spend $0.90 X 100,000 times a day, you will make a ton of money.

I am having a hell of a lot of fun letting agents play and understand a (still alive, human populated) MUD, which I have also played for the last 30 years.

It’s mostly my way to play with local llm inference (m5 64gb, gwen3.6 27).

It’s amazing. They build maps, classify events (building a grammar for a parser), run experiments (to verify the grammar). They are now (given the correct tools/infrastructure) trying to fine-train a 3b model for fighting (where you need a decision for 5 seconds rounds). Basically autonomously!

Overall, a MUD does prove a great constrained sandbox for them to play in.

What started as an experiment to test local inference landed in a sweet spot for seeing models strength/weaknesses/tradeoffs. And it’s really fun.

Only problem is that Claude gets really jealous when I ask him to code their po harness running local inference. Weird world.

  • How is the agent interfacing with it? Are you just manually copy pasting game output and doing the response or something more integrated?

    • in a container a daemon is managing the telnet connection/logins and throws into a parser, that takes grammaries (basically, lists of regexes) and produces events, which are appended into simple files as json blobs (with types: rooms, npc enters/exists, unknown, etc)

      agents can tail live logs or parse in any sense: i built agents that try to classify (via grammar file edits) "what that unknown event is", build experiments and verify against live runs

      surprisingly, a local 27b is quite smart

      ultimately you need to run fast inference for combats (5 second rounds), for that i built a system to prepare fine-tuning runs over 3b models

      overall, using container primitives for everything seems to be the biggest gain: in the end they mostly are coding models, so intuitively know how to interact without needing too many explanations

      example: mud commands are small bash commands, the "game help" is provided via man for those commands, and hence when there are attempts at reasoning (which is done for knowledge building via grammar files) agents are great at using apropos to find out more about a specific topic

      another example: setting a "pick all money from body" is a tail + grep, once the agent is shown it he picked up usage quite fast

      to run agents i use pi and simple setup where a main 27b agent acts as an orchestrator and delegates runs to agents (orchestrator also build prompts), usually as "experiments" "if i classify event x than what should happen in-game"?

      i do not really keep it running all time, it's mostly about me putting up challenges, see what happens, and using that as a learning experience

I was a bit afraid that MUD was an acronym for something else, I clicked in about 10% hope that this will be for multi user dungeons and behold, thank you. Interesting read.

That brought back memories when we were trying to find free terminal in few departments at our uni to get access to one that has internet connectivity.

I tried MUD a few times as a teenager, but for some reason I didn't like it. I prefered playing RPG on a long dead phpBB forum (LdC), and to this day I remember the very simple mechanic (for both players and gamemaster): an asterisk for character speech, and (#) for everything else.

A single post might look like this:

    # By hearing those absurdities, memories rush to the mind
    of Vicent Panclast — memories he'd like to avoid. He gets
    physically agitated and raises his pistol high in the air.
    * They shall not pass! We must do anything we can to stop
    these mf nazis!!

From what I remember there was no dice at all — the gm just did whatever he'd like. Your character sheet was just a story of your character plus some description. We had a fantastic gamemaster, so it made for great rpg. Good times.

I wonder what something like that could look like with the right use of AI, today. I'd like to think kids are exploring this kind of thing nowadays - I hope they are, because it's a ton of fun (even without AI, kids).

Speaking of MUD, has anyone tried attaching an AI to the Hitchhikers MUD[1]?

[1]: https://www.bbc.co.uk/programmes/articles/1g84m0sXpnNCv84GpN...

  • This is a text-based game, but it is not a MUD. The MU in MUD stands for multi-user.

  • Not Hitchhiker's specifically, we'd considered another existing one but as soon as we started digging into IP law related to MUDs, we decided if we're going to do it, we'll just start our own from scratch.

    • > we decided if we're going to do it, we'll just start our own from scratch

      And because scripts and information about famous text based adventure games can be part of what the LLM ingested during training (so unusable for fresh reasoning).

      Anyway: very, very, very, very good idea.

  • Speaking of Douglas Adams, Starship Titanic would be the perfect game to try reimplementing with a modern LLM. It already runs on a primitive chat AI system.

  • that's funny. the very first thing i had an llm do was to ask it to pretend to be the HHGTTG infocom game. Can't remember how it did.. not amazingly but was interesting still.

+1 to using MUD as a learning tool. Learn about agent loops, prompting iteration, etc.

I vibe coded a few room proof of concept MUD. I added a tools called `oracle` (aka me) that it could ask me questions to help along its way; and the ability to interject in the loop with a hint. Just as I might want an llm to stop to ask me for help instead of plodding along...

Interesting to limit the number of turns and see where it gets stuck or how quickly it can finish.

It's interesting that you chose to measure how well each LLM did in talking to other NPCs, and having each NPC also use an LLM to react to each input. Why not have the LLM fight NPCs and loot items (both a significant part of the MUD experience, neither requiring LLMs on the server side), then measure character progression in experience and stats?

I'm actually building a MUD powered in part by LLMs so this is very timely. It never occurred to me to use it to evaluate LLMs, which is a really neat idea. I'll be keeping an eye on this project.

Withered technology - that's a nice turn of phrase. I think I'll start using that to refer to some of the 2000s era MS tech we're saddled with that gets a coat of paint and some KBs to cover CVEs but ultimately is rather withered...

That’s a fun design that brings back some old memories.

  • Thanks! Yeah, it started with a friend's idea to take a MUD we played in the 90s / early 2000s that had been open sourced and adapt it to mobile games. He's a middle school teacher and wanted to introduce the next generations to MUDs.

    We then decided to pivot and thought, what if this would be a good eval environment for LLMs from the simplistic standpoint of, text is their native environment.

"We did not choose a MUD because it is charming. We chose it because its constraints make behavior measurable." I so fucking hate these LLMisms

  • That's fair. We leaned on LLMs to help draft parts of the paper but disclosed it in the acknowledgements section. We didn't appreciate the amount of work this "hobby" experiment would require, and so looked for ways to help get this out the door.

I tried to build an AI agent product recently, frankly people are not so fancy about AI, they still just want a simple product as old time

Well now, you've given me an excellent idea for how to speed up development in my vibecode MUD reboot.

I don't understand this article. In fact I even read the entire whitepaper and I don't understand it.

You didn't use "a MUD"; you used a very limited "MUD-style environment". This is not a MUD. Your headline, your article, your whitepaper is a lie. Your "MUD" didn't originate in the 1970s; you coulnd't even be arsed to include its source code!

I was at first mystified when you couldn't be clear about what genre or species of MUD software you're using. There are many types of MUDs out there, and you used exactly none of them. You used some sort of bespoke, vibe-coded, constrained environment that is not a MUD.

Shame on you and your research. Shame on you for falsely capitalizing on the "MUD" term. Flagging this post.