Comment by dominotw 3 days ago they are not really optimized for 'wide variety of models' . what optimization did pi do for glm 5.3? 2 comments dominotw Reply rpdillon 2 days ago Tool calling success rate, in the case of omp. rpdillon 2 days ago Can't edit my post anymore, but here's the omp blog from February talking about improved tool calling rates across 15 models, with only the harness being tweaked to get the improvements.https://stencil.so/blog/the-harness-problemThree GLM models are mentioned, but so is Deepseek, Grok, Minimax, Kimi, and Gemini.
rpdillon 2 days ago Tool calling success rate, in the case of omp. rpdillon 2 days ago Can't edit my post anymore, but here's the omp blog from February talking about improved tool calling rates across 15 models, with only the harness being tweaked to get the improvements.https://stencil.so/blog/the-harness-problemThree GLM models are mentioned, but so is Deepseek, Grok, Minimax, Kimi, and Gemini.
rpdillon 2 days ago Can't edit my post anymore, but here's the omp blog from February talking about improved tool calling rates across 15 models, with only the harness being tweaked to get the improvements.https://stencil.so/blog/the-harness-problemThree GLM models are mentioned, but so is Deepseek, Grok, Minimax, Kimi, and Gemini.
Tool calling success rate, in the case of omp.
Can't edit my post anymore, but here's the omp blog from February talking about improved tool calling rates across 15 models, with only the harness being tweaked to get the improvements.
https://stencil.so/blog/the-harness-problem
Three GLM models are mentioned, but so is Deepseek, Grok, Minimax, Kimi, and Gemini.