Orca-Bench: How Ready Are Language Model Agents for Oncall?

2 hours ago (arxiv.org)

Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.

  • Attackers advantage in the iterative fast feedback loop?

    It’s harder to have a loop to ensure you are defending all possible attacks?

    I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.

    Finding all possible attacks and patching them against yourself is inherently more expensive?