← Back to context

Comment by raincole

5 days ago

I'm still confused about what this codemode is. Models have been trained to chain bash and other typical unix tools well. They're so good at that to an uncanny level. Why do we want to not utilize this ability? Is it just a permission management issue in case you don't want the model to use shell directly?

We will write about it. The best way to think about it is that codemode solves a different problem than bash in that bash is a way for the agent to run a particular tool: running bash.

Codemode is a way for the LLM to orchestrate harness level tools. The reason this happening now, is because the models by the labs are increasingly trained on this. Codex for instance in responses lite requires codemode to even perform parallel tool calling.

  • Bash (including `python3 <<EOF...`) is how you orchestrate CLI calls.

    If all your tools are CLI calls, all you need is bash.

    The one and only impedance mismatch is subagents; nobody's come up with a good way to turn the "subagent" tool into something you can access via the CLI. But I think fixing that would be a more productive direction than surrendering to the MCP insanity.

Codemode is a fancy name some MCP authors coined for the practice of providing scripting/method chaining for their MCP tools. It's generally implemented by providing some kind of code execution tool, the LLM calls it with a script, and the MCP server runs it in a sandbox.

It's pretty effective because of the reasons you noted, but there's a composability problem since each MCP has its own sandbox and can't call into the other ones.

IIUC Pi offer a workaround for this, the harness runs the sandbox and populate it with the MCP tools, that way the composability problem is solved and every MCP do not have to implement their own sandbox.

  • My understanding is that code mode is supposed to be implemented by the harness, not the MCP provider. You chain multiple MCP providers as well as other harness provided tools inside the sandbox.

    • My timeline might be wrong (I remember a cloudflare article mentionning the "in MCP" case), but anyway yes there's tools to do it in the harness now and it's the better idea.

yes codemode is when you want to utilize this ability AND you want MCP tool calls as part of your scripts

codemode lets you execute scripts in a runtime where your MCP tools are made available as function calls

this matters for cases where the MCP tool is the only way to do something and you do not have an equivalent CLI, API, whatever to script with

From my understanding, code mode came about due to some agents not having access to a shell.

  • The value is that rather than an agent chaining together tool calls itself (which means each step sends the result back to the agent for it to analyse and work out what to do next), it writes a script for the harness to execute that chains together all the calls. The major benefits are:

    * speed - much fewer hops back to the LLM

    * fewer tokens - intermediate execution steps in the script don't leak into context, only the final result does.

    * repeatability - if the LLM needs to repeat work, it can reuse a script it wrote last time.

    If you have a harness that has access to a full shell and knows how to use bash or python, you'll often see it writing little scripts. For setups that don't (ie normal model API requests with tool calls), you can give it an lightweight secure execution environment like just-bash, or quickjs.

    • Agents just do this anyway, how is it a “mode”? I always see the agent writing scripts in a tmp dir to execute or even just inlining bash and python scripts.

      8 replies →

  • thanks! it now clicked for me. so instead of cat its read_file, even if read_file resolves to cat, cat is not always available.