The value is that rather than an agent chaining together tool calls itself (which means each step sends the result back to the agent for it to analyse and work out what to do next), it writes a script for the harness to execute that chains together all the calls. The major benefits are:
* speed - much fewer hops back to the LLM
* fewer tokens - intermediate execution steps in the script don't leak into context, only the final result does.
* repeatability - if the LLM needs to repeat work, it can reuse a script it wrote last time.
If you have a harness that has access to a full shell and knows how to use bash or python, you'll often see it writing little scripts. For setups that don't (ie normal model API requests with tool calls), you can give it an lightweight secure execution environment like just-bash, or quickjs.
Agents just do this anyway, how is it a “mode”? I always see the agent writing scripts in a tmp dir to execute or even just inlining bash and python scripts.
What if I have an MCP Tool LookupZip(City) and want to chain it with a bash tool that prodcues a list of 100 cities. And then I want to filter again to the largest Zip code.
And I said that they do this in my last paragraph. But you need an execution environment for this, and you don’t get this automatically when just interacting with models via their API
The value is that rather than an agent chaining together tool calls itself (which means each step sends the result back to the agent for it to analyse and work out what to do next), it writes a script for the harness to execute that chains together all the calls. The major benefits are:
* speed - much fewer hops back to the LLM
* fewer tokens - intermediate execution steps in the script don't leak into context, only the final result does.
* repeatability - if the LLM needs to repeat work, it can reuse a script it wrote last time.
If you have a harness that has access to a full shell and knows how to use bash or python, you'll often see it writing little scripts. For setups that don't (ie normal model API requests with tool calls), you can give it an lightweight secure execution environment like just-bash, or quickjs.
Agents just do this anyway, how is it a “mode”? I always see the agent writing scripts in a tmp dir to execute or even just inlining bash and python scripts.
Perhaps the appeal is a stronger and more customizable lockdown on agent capabilities.
3 replies →
but these bash scripts can not execute MCP tools.
What if I have an MCP Tool LookupZip(City) and want to chain it with a bash tool that prodcues a list of 100 cities. And then I want to filter again to the largest Zip code.
3 replies →
But agents do write scripts, in bash. And they are very good at it. And bash is quite efficient with its pipeing.
And I said that they do this in my last paragraph. But you need an execution environment for this, and you don’t get this automatically when just interacting with models via their API
1 reply →
Bash scripts don't really help when you're dealing with MCP tools or other tools that are local to the harness itself.
Then give them a shell in a VM sandbox. VM sandboxes are way cheaper than LLMs.
thanks! it now clicked for me. so instead of cat its read_file, even if read_file resolves to cat, cat is not always available.
Rather, instead of shell_tool(command: "cat ...")