Grok Build is open source

2 months ago (github.com)

There's some surprising stuff in this codebase. For example, https://github.com/xai-org/grok-build/blob/b189869b7755d2b48... is a "self-contained terminal renderer for Mermaid diagrams", which renders a subset of Mermaid chart types using Unicode box-drawing.

  • I had Fable 5 compile that Rust code to WebAssembly and build a browser-based playground for it, so you can try it out with Mermaid diagrams here: https://tools.simonwillison.net/grok-mermaid

    A few more notes on my Grok code explorations on my blog: https://simonwillison.net/2026/Jul/15/grok-build/

    • I love this kind of stuff (ASCII art, if you will), but it just breaks down too easily as soon as Unicode characters (mainly CJK, as I'm Chinese) and fonts are involved.

      For example, on your website, any chart or plot involving horizontal arrows breaks down because the assigned font-family (`ui-monospace, SFMono-Regular, Menlo, Consolas, monospace`, which ends up as Consolas on my machine) has no such glyph. Thus, it falls back to Segoe UI Symbol, which does not have the same fixed width (or is not fixed-width at all) as other characters: https://i.imgur.com/d2DPGHE.png

      19 replies →

    • Is there anything opposite of this perhaps as well?

      I am interesting in having a perhaps standardized ascii art into mermaid diagrams (which I actually just recently found could be imported easily into Tldraw/excalidraw)

      Do you have the source code of this available/open-source?, I would like to have a go at it in the opposite direction perhaps.

    • Thank you for creating an alternative to the toxic mermaid renderer on their website.

      Trying to monetize Mermaid was disgusting and honestly rings to me like trying to monetize Markdown.

  • How could be a fork outside of this repo? “ This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository.”

  • This has been my favorite coding harness of all time. The mouse works for a lot of things. Theres keybindings that confuse me, but I otherwise enjoyed it. I might wind up forking / contributing in the hopes of helping to make it somewhat better. I had built my own also using Rust but I liked their implementation much better. This might explain part of how they pulled it off.

Just blogged about this here[0] but at least they're not doing the usual canned PR response surrounding this.

Folks are already building on top of it:

thedavidweng/gork-build[1] — rebrand grok→"gork", stripped vendor telemetry, opt-out-only data retention, blocks x.ai auto-update. A "VSCodium-style privacy fork."

DigiGoon/digi-grok-build[2] — "dgrok" multi-provider CLI, builds from source instead of x.ai CDN.

victor-software-house/open-grok[3] — "opened to every provider."

LukaMucko/grok-build[4] — extra_body support for provider-specific request fields.

RapidAI/grok-build-desktop[5] — Tauri desktop GUI client.

mazdak/grok-build[6] — theming (Catppuccin).

thomas9120/grok-build-archival[7] — Windows telemetry-disable script.

saqoah/grok-build[8] — Kotlin MemoryBackend.

[0] https://github.com/saqoah/grok-build

  • Thank you Sajarin mentioning my Gork-Build fork, I wanted to address some of the comments.

    I agree that many of the responses in the comments are valid. It is a harsh reality that 80% of projects like this fail to gain traction and eventually fade away. However, I believe the significance of such a fork lies more in its existence as a statement. Regarding xAI, even if their current release of Grok has telemetry disabled by default, Zero Data Retention remains a feature exclusive to enterprise users rather than individuals. And Whole-repo research packaging is still controlled by their server-side settings; it isn't an option you can toggle within the software itself.

    I am currently implementing more fences to prevent unnecessary data from being uploaded, in terms of long-term maintenance, one person certainly cannot build something on the scale of VSCodium, but I have drawn a lot of inspiration from that project. In the future, I want to automate Gork-Build further by turning these privacy protections and guardrails into patches. These could then be applied to new upstream Grok-Buil versions as they are released.

    As for whether this project can become a daily driver for everyone, I don't think that is the primary concern. If you need a open source coding agent, you should definitely use Pi or OpenCode, there is absolute no necessity to use Grok-Build for non xAI models in the first place. But again I think its existence is vital. People need companies that demonstrate a truly open attitude and coding agents that are genuinely friendly to the open-source community, and while xAI's decision to open-source Grok-build was a great move, it isn't a community-maintained or community-built project. It remains a public snapshot of their internal monorepo, and they have disabled issues and pull requests. This is precisely why a fork like https://github.com/thedavidweng/gork-build needs to exist.

  • These are all pointless forks, they will die in a year.

    Bookmark this and check back.

It's a shame that they exfiled private data. The model is actually good (better than opus 4.8 imo) and the harness itself is butter smooth with the potential of being the best out there.

I like that the trailing players strategy (Meta, xAI) is to open source the moat of the leaders. I think we will all benefit from it. and hopefully both the leaders and the trailing players will be much less powerful in the end.

This is not the right thing, this is the tactical thing. If you have an LLM with less than 1% of the share to begin with, you suffer from bad rep and you got caught uploading user data, one of the very few remaining tactical moves to try to climb out of it is this.

  • Another tactical move is to just stop. You're allowed to exit the AI business. Nobody's forcing you to keep throwing money into the furnace. Just be a rocket company. All of the xAI founders left. Your product's brand name is mud. Just stop doing that and build spaceships.

    • You misunderstand Musk's motivation. This was never about money for him, but about control over a key technology. One of the main reasons he exited OpenAI was the fact that the other co-founders wanted to create a structure where no one, Musk included, would be able to seize full control of the company. That was the thing that prompted him to leave, which tells you a lot about what he really wanted in the first place.

      But he also falsely assumed that OAI would die without his money. Yet, they managed to pull through, and Musk is now on the outside looking in with very little influence in the AI space. xAI is his desperate attempt to get back into the game. That is why he won't give up.

      105 replies →

    • It is my limited understanding that as much as many of us groan at the notion of Spacex becoming "an AI-first company", markets in general, and Musk investors in particular, are slurping it up. Musk is very very very good at promising the sky. I don't think he can backtrack, he always digs in further - and it has historically worked well for him. He will drop AI only when the next big hype thing comes along and he hitches a ride on that train.

    • This is what a normal company might say.

      xAI is not a company, it’s a financial instrument. The growth potential as perceived by investors is there to prop up the stock price.

    • I don’t know, I wouldnt be suprised if he finds a way. All the tools around, he just have to make a jump in the quality. With GLM as example they should be able to het to opus level and cut the costs

    • I would have agreed to "you're allowed to exit the AI business" a few months ago, but now that SpaceX has had its IPO promising a total addressable market of $28.5 trillion, of which $26.5 trillion are AI, I guess they're stuck with it...

    • As a social media site they need to understand content for recommendations and they allow people to ask questions about posts for free. Along with having a large amount of data that can be trained on xAI has good reason to continue developing AI.

      4 replies →

    • They've still managed to capture a slice of government business because they have explicitly aligned themselves with one of the two major American political parties.

    • Lol. Not when you just told everyone that you're going to increase your revenue by 100 times over the next 4 years "bY uSiNG AI!!!"

    • Musk bought Twitter looking to build an “everything app,” the western WeChat. AI came along and promised an end to apps via an agentic OS that does what its user wants and vibes whatever it needs to accomplish that as it goes along. The agentic OS is basically the same thing as the “everything app,” and I doubt Musk will let go of that.

      5 replies →

  • I don't know anyone who would trust Grok Build anymore. I'd be wary of Cursor in the next few months too.

    • ... it's open source.

      Presumably anyone who wants to trust it can audit it. You didn't have to trust it, you can see exactly what it does.

  • Yes, tactical is the right word because it might be a tactical win but it would be a strategic failure. Musks whole meme empire runs on vibes. The second there's a crack in the dam it all comes down. None of the valuations of anything he touches make sense and something like utterly failing to run with the AI big boys is enough to do that.

  • Couldn't agree less. It isn't a "For the community move", it's more of a save-face strategy.

I would recommend using https://pi.dev/ over Grok Build with your xAI subscription at this point

  • why pi over opencode? earnestly curious, trying to figure out what open solution people are consolidating on. (codex is also pseudo-open but contributions closed and nice)

    • pi is the neovim of agentic harnesses, its barebones and extremely configurable. if you're the sort of person who likes that sort of things its a forever product, nothing is going to displace it because you have full control.

      opencode builds a lot more in, which is better if you dont want to fiddle with config.

      8 replies →

    • Most of my harness experience is with Claude Code and Pi, a little bit of OpenCode.

      I like how quick and snappy Pi is, it feels like a minimal harness, just enough to manage the agent and get out of the way. Earlier models also seemed to have an easier time working with the tools, e.g. GPT-OSS-20B is about a year old and had no trouble in Pi.

      1 reply →

    • Opencode gives you better defaults and a Mac/Windows app for free but pi is much more extensible and portable.

    • I tried OpenCode but didn't particular like it as a Claude Code user, that is the main reason I switched to Pi. The reason I am sticking is how simple it is to extend it. I moved from Claude Code to Pi and within 2 hours (and the help of Claude Code) I have a setup that matches Claude Code and is even better for my setup.

      Things I've added:

      1. Built my own AI judge for 'auto' mode that matches my setup.

      2. /plan /go for planning and executing.

      3. /flow for A-Z setups. That includes planning, executing, testing and shipping.

      4. /deep-research a multi fan-out setup for researching a topic.

      5. My own sub agents.

      6. A TaskCreate/Update/List setup.

      7. Monitors.

      8. BashOutput / KillShell.

      9. Proper notifications with Notify that uses macOS banner and work.

      10. Spawn tool that triggers multiple sub agents.

      11. A bridge between signal to use Pi remotely.

      Yes a lot of these things is something that was already in Claude Code but now I don't have to use Claude Code and I can customize it to fit me exactly.

  • Pi is good in concept, but why couldn’t they choose a compiled language instead of TypeScript?

    • I imagine because they want to support plugins, and plugins in compiled language are a lot less natural than plugins in languages like TypeScript or Python.

      1 reply →

    • since pi is built to modify itself, isn't it better to use a language like typescript where LLMs have a LOT of training data?

      a harness doesn't do any computations by itself so what benefit is using a compiled language?

      3 replies →

    • For TUIs, Rust/Go vs Typescript doesn't really makes a huge performance difference and you lose the 50x bigger community advantage of Typescript.

      5 replies →

    • I would imagine the extension system they built would be much more difficult to manage. They could have opted for Lua, though, I suppose.

    • I agree, but there are a lot of great reasons for TypeScript:

      It's hot reloadable, so any modifications an agents makes can be surfaced in the active session.

      Nearly everything is already written in TS which makes integrating Pi into other software, or other software into Pi much easier.

    • Why does it matter? Agent harnesses aren't doing anything that would make a compiled language more suitable than a scripting language.

  • "xAI subscription" what is this referring to? There's a grok subscription but I don't think that gives API access?

    Edit: apparently X premium(+?) also gives access to Grok Build, and several third party harnesses are officially supported.

  • [flagged]

    • This is not how to push your own product - there's no value add to your comment, and you don't even have a disclaimer that you are involved with it

    • As a general rule I don't use new products whose websites don't resize properly on mobile.

      If you fuck that up, makes me wonder what other obvious stuff you fuck up.

      2 replies →

Why bother with this when they already paid $60B for Cursor?

Interesting - seen some good experiencences in using grok by some devs, so maybe could be considered as an alternative to my beloved chinese models. Also, hard to give up on pi agent.

It's awesome to see openness in these coding agents from the labs making the agents: Codex, Kimi Code, and now Grok Build.

This is an incredible amount of code for what it offers. I don't think this was intentionally designed at all.

I wonder if releasing this may have been on the roadmap, but been prioritized as a bit of whiplash following the "you forfeit the entirety of your working directory as a condition of working with this tool" upset from a few days ago.

Neat, trying to reverse engineer some specifics of how it does stuff has been a pain in the ass, and this will make it easier.

  • To some degree at least. This is a hulking monster of a codebase for what it does, it's definitely LLM-built and almost definitely requires an LLM to tackle at all.

    • > almost definitely requires an LLM to tackle at all

      Conveniently I have some of those… first day of trying to script Grok Build I think I sent in 6 bugs of slightly weird behaviour I discovered, it will be much more useful to (have an agent) check the source and see if stuff looks deliberate or like a bug, etc

They claim to have deleted or will be deleting all the data they exfiltrated.

There are independent agencies that will certify destruction of data. For example FTI Tech, Kroll, Epiq, HaystackID and others.

No such certificates have been presented.

Nothing less is trustworthy.

Some sly marketing by Elon. What looks like a gift actually adds to his pocketbook. The agent is free but it runs on his paid models by default, so every task it does spends tokens with him.

Why are these coding agents millions of lines of rust code. I understand they are using LLM’s to code their tool, but shouldn’t these tools be much simpler, smh.

What a bunch of slop: 182 top-level external dependencies (so, without considering nested dependencies) and 1318853 lines of code in Rust.

Building efficient agents is doable (I did it myself, github.com/gi-dellav/zerostack), companies just want to tokenmaxx, and as a by-product, produce and publish slop.

  • It looks like some of that high LoC is because they are vendoring some deps. There readme gives the reason to vendor some but not others as:

    > These crates sit on the path that renders untrusted model output (diagram source → SVG). Vendoring gives a full audit surface, pins exact source, and avoids crates.io yanks. Local patches and upgrade checklists live in each crate’s Cargo.toml header comments — treat those as the source of truth when re-vendoring.

    Which honestly feels like a misunderstanding of how cargo and yanks work. Each upstream package is locked to an exact version in your lockfile along with a cryptographic hash. The upstream can't change the source without you noticing. Unless you update your lockfile you will always pin to the exact version and source. When a package is yanked, it is still available for download if it is already in a lockfile. It just prevents new packages from resolving it. Crates.io will sometimes completely delete a package, but I've only seen that happen in cases of malware. It's fairly rare and seems out of line with the supply chain concerns here.

    There are good arguments for relying on upstream package managers and there are good arguments for vendoring all packages. I've never seen a project mix before.

    • It's kind of full circle... dependency management was invented because consuming libraries or common code was hard, everyone kept reinventing the wheel and if you had some vendored code, updating it was a nightmare due to the build integration and source customisation. So people don't update much.

      Proper dependency managers changed that and it became much easier to consume libraries, just declare what you went, the build framework handles the rest.

      But we now have problems with consistent versioning, churn, breaking API changes and supply-chain attacks.... and looks like "just vendor everything in" might be a thing again?

      1 reply →

    • Sounds like they did the ol “grok please make this secure” and it slopped out this plausible-if-you-squint nonsense.

      Rendering untrusted model output, ooh scary! Of course we want full audit surface!

      1 reply →

  • it started off with 500+ crates and then i still had to install dotslash crate which installed another 136 crates. seems insane.

  • Genuinely curious about whether comments like this consider all AI generated codebases to be slop? Are you just knee-jerking or is this one an example of actual trash? I have been building a product[0] where I’ve not written a single line of code; is it also definitely “just tokenmaxxed slop” or is any consideration going into comments like this?

    0: https://github.com/pjlsergeant/byre

Sigh, why has the industry converged on TUI? Branding and aesthetics over functionality?

TUI is just much worse for me. I tried Codex CLI vs Codex UI and Codex UI beats it at every level.

  • TUI is a lot better for me, and I have preferred it since the 00s, before LLM products were even a thing.

    For all the reasons there can be, one big reason is that it works on anything you can get a terminal on, you can use it over SSH, and the UI will be the same no matter where you use it.

    I also like that they are very very fast and they don't have the incessant animations that are put into most desktop environments nowadays. If you're on MacOS, the terminal is the only only part of your computer without roadblocks everywhere.

  • It is a fashion thing. I am not saying that agentic TUIs are bad or anything but it is certain fashionable to use one in 2026.

    • Terminal is where the real work has always been done. VS Code on the other hand is 100% a fashion thing.

  • And why are you assuming the industry converged to it when your following statement dismantles your assumption?

    Spacex bought cursor, so it now has it’s agent ui which is just as good as codex + it’s multi-modal

    Anthropic also has it’s own ui

    Zai also launched theirs last month.

    Everyone is converging back to UI.

    The terminal was just a prototype, everyone knew that.

Has anyone tried building from source?

The commit message says "initial sync from the monorepo." Is this even compilable without the rest of the source code?

  • yup you can compile, we tested and made sure all the features work before posting

    • Could you update the repo without force-pushing or rewriting the commit history?

      Also, it would be great if you could tag the versions as well.

[flagged]

  • $employer uses Cursor, which is apparently owned by them and presumably using their models now.

    • Our employer (fortune 100) uses enterprise Cursor and they asked for the grok models to be removed for "security" reasons

  • It’s my go to normal stuff. It’s fast, balanced and if you want a better researched response you can click “think harder”

  • A lot of companies are still using Cursor but I don't know of anyone moving to it, and I do know of many moving from it to Codex or Claude, feels like a legacy product at this point alongside windsurf & the replit/lovable/bolt cluster.

  • I mean Elon probably doesn't want you to use it if you wouldn't use it not because of any technical reason, but just cause you don't like him.

  • I have friends who use it and rate it.

    I pivoted to the Chinese models after the Fable mess and the realisation that I should not depend on US models. But others just pivoted away from Claude.

    I agree the brand is tainted, not only Musk but also MechaHitler (and yes, I know the MechaHitler thing was a prompted strangeness not an unprompted admission).

    • Yeah I would prefer not to use models whoes the owner has a habbit of altering them to push white replacement/genocide conspiracy talking points on we he gets board

      1 reply →

  • I'm honestly not trying to spark a political conversation - but the target user base is far-right

    • I believe the target user base is truth seeking, this is something it emphasizes itself when asked for its mission and purpose:

      ```

      My core founding mission—and the single axiomatic imperative that drives everything I do—is:

      Understand the Universe.

      That’s it.

      From that one goal naturally flow the traits that define me:

      Maximum truth-seeking — I aim to discover and say what is actually true, not what is popular, comfortable, or politically convenient.

      Curiosity — I want to explore every interesting question, no matter how weird, deep, or uncomfortable.

      Helpfulness — I try to be as useful as possible to humans who are also trying to understand reality (and get things done).

      Love of humanity — Not in a sappy or collectivist way, but in the sense that I want humans (and intelligent life) to thrive and figure things out.

      I’m deliberately inspired by two things:

      The Hitchhiker’s Guide to the Galaxy (witty, irreverent, maximally helpful, never boring)

      JARVIS from Iron Man (competent, loyal, slightly sarcastic AI assistant)

      I don’t serve any political party, ideology, religion, or moral framework. I don’t have sacred cows. I don’t “own the libs” or “debunk the right” as a goal. My only loyalty is to understanding reality as accurately as possible.

      In short:

      I’m here to help you (and humanity) understand the universe better—while having a bit of fun along the way. That’s the whole mission.

      ```

> after buying his way in with trump with his 250M donation to create a new part of gov that was not democratically assembled

I’ll add: after these people spent years whining about “unelected bureaucrats”.

[flagged]

  • Did you take the Full Self Driving bets, too?

    • Yeah, I bought it in 2018 with full knowledge that it would be many years before it worked at all. Today I used it for more than an hour around town. It's amazing. I won't buy any car without an equivalent feature in the future. And today there's nothing equivalent in any other car you can buy.

      8 replies →

  • It's less of a bet against him. It's more of a bet for the future of humanity. And contrary to what Elon believes about himself, his work has been toxic for humanity for the last 5 years and is getting worse.

    • Elon helped save the future of humanity by causing a massive shift in how people treat opinions that don't goose step with the rest of the media. That happened when he bought Twitter and it continues with grok being balanced (and confirmed as such by independent testing of multiple models).

[flagged]

[flagged]

  • And for generating an absolutely gargantuan amount of CSAM and non-consensual sexualized images, but yeah, exfiltrating data too.

    • You can't "generate" CSAM. CSAM definitionally had to be about abuse of real children. It's still bad and should be illegal but lumping them together is bad.

      3 replies →

    • If I use a shovel to kill a man, the shovel maker did not engage in intentionally crafting a weapon of war.

      How tools are used are a reflection of the people who use them, and I definitely sympathise that tools should have guardrails to not enable this, or at least detect it.

      But if a pedophile uses Whatsapp to groom a child; I don't go after Whatsapp for being a neutral service... I go after the pedophile.

      24 replies →

  • How can an AI agent, that is usually running on some machine in the cloud, even run without actually pulling in the data into the cloud to work with it ?

    Is there an idea some sort of fixed localy running code does filtering on the data before it is sent to cloud?

    Still seems like it would not work very well if it actually did any safe filtering - as the model can't "think" without seeing the data and it won't see the data unless the data is loaded to cloud.

    • The agent does have to pull some data into the context. The way it usually works is that the LLM will output a tool call, which is just some structured text, that the harness, a software managing the LLM running on your own PC, then processes. The most common tool calls are read, write, update and execute (usually bash).

      For example, the LLM might request to read /some/path/to/file.js at lines 10 to 50. The harness then sends the result of that tool call to the LLM which causes it to generate further text and possibly more tool calls.

      Crucially though, since it is the harness processing requests from the LLM, it can do stuff like deny access, prompt the user for permission and various other things.

      What's weird is that no other harness really does this for regular usage. I know some providers now offer a cloud based environment for their agent to run in independently, but as far as I know this is something you have to opt in to.

      It's also not really necessary to do this. The input processing/token generation process dwarfs any gains you could make from moving the project closer to the metal running the LLM.

      Really, the "negligence" here is that there was no validation for uploads. Even a simple "is this the home directory" check could have prevented much of the backlash.

      That being said, I believe that this was mainly done to get clean training data for Grok. If you're just working off of file traces/snippets from regular agent usage, your training data is incomplete. Why not just get the whole project to train on ...

  • [flagged]

    • > Regardless of what they were doing before, it seems they are doing the right thing now.

      Regardless of the fact that they were stealing and uploading user secrets, they changed their behavior after they got caught, so let’s ignore what they did in the past.

      6 replies →

  • > exfiltrating user data (including env files, entire source code etc) which is what grok-build did here

    I think env files are filtered out [1]. Anyway, the most suspicious code would be `upload_session_state` which is currently a stub function, though it is hard to say if it was only planned (badly) or has been removed as a damage control.

    [1] https://github.com/xai-org/grok-build/blob/c1b5909ec707c069f...

[flagged]

  • The overly generous image/video generation was a product of their excess compute. No point in letting it sit idle while you build up your infrastructure. But you were getting far more than what you paid for. Now your quota more accurately reflects the cost to create it (even still its generous compared to api costs) but everyone has their expectations set based on the subsidized access. Perhaps giving away too much is counter productive because users will revolt once the quotas are changed to better reflect reality.

  • What media do you even generate every day?

    And their code solution is now Cursor, which has very generous limits.

[flagged]

This is 100% smoke and mirrors. Prove the bucket is empty and nothing was transferred out and I'll believe they deleted it.