Comment by kamranjon
4 hours ago
New coding harness that seems to have some novel concepts and one of the pretty cool things on their landing page for it here: https://deepseek.com/harness/en/ is the Every Run is Traceable view:
"Everything the model sees is recorded in an append-only session log: system prompts, reasoning, tool calls and results, subagent scheduling, and every context injection. In the Trajectory view, you can inspect these records by source. Resume, fork, search, and replay all operate on the same event stream."
Seems pretty helpful - have sort of wanted something similar (I use Pi).
They also released this research paper that backs their whole plugin composability system that seems pretty cool: https://github.com/cordiverse/paper
I'm glad they're doing this also and that more people are adopting it. Event sourcing [0] is the right way to represent informaiton like tool calls, user interactions, etc. --- it makes it easy to fork conversations and maintain a cohesive conversation stream and stable message history that does not break the cache.
[0] https://www.dreamcoder.ai -> scroll down to the event graph.
I promise this isn't meant to be snarky, but is that not just...logs?
It's only logging if an obsolete human does it.
But the future is here and thus it's called "Agentic causality's reified temporal traceability."
"Temporally Reified Agentic Causality Traceability Report" - TRACTR.
It's logs of activity the US models hide from you in fear of them being used by competitors.
Which are unavailable with the leading American models. You can't look at the complete traces of OpenAI or Anthropic model agents, as they are encrypted (there's been discussion of a couple of different ways to expose those, but that violates terms of service, and well, you shouldn't have to find complicated ways unencrypt your own usage logs).
Logs that aren't missing anything out of the box. I'd say it's pretty underused concept in time of 8TB consumer SSD drives.
Don't those cost 1-2k?
It's useful logs which i think is an important distinction.
is it just for coding? the docs don't mention code, just "agents"
Seems like Agentsview, but built in and likely less features(at least, as of now): https://github.com/kenn-io/agentsview