Terminal-Bench-Science: Evaluating AI agents on scientific research workflows 2 hours ago (terminal-bench-science.ai) 5 comments matt_d Reply Add to library akshay_akula 1 hour ago Evals on actual research workflows is the right direction, most agent benches are toy tasks. rubslopes 2 hours ago I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style. mlmonkey 24 minutes ago Sad to see no mention of Gemini ... vatsachak 1 hour ago Damn. These things aren't AGI... but I don't care.Luna is good enough for me to give a parser spec and have it write one.
akshay_akula 1 hour ago Evals on actual research workflows is the right direction, most agent benches are toy tasks.
rubslopes 2 hours ago I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.
vatsachak 1 hour ago Damn. These things aren't AGI... but I don't care.Luna is good enough for me to give a parser spec and have it write one.
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.
Sad to see no mention of Gemini ...
Damn. These things aren't AGI... but I don't care.
Luna is good enough for me to give a parser spec and have it write one.