Orca-Bench: How Ready Are Language Model Agents for Oncall?
Posted by yruzin 2 days ago
Comments
Comment by dash2 2 days ago
Comment by aleksiy123 2 days ago
It’s harder to have a loop to ensure you are defending all possible attacks?
I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.
Finding all possible attacks and patching them against yourself is inherently more expensive?
Comment by EGreg 2 days ago
Your strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance.
Comment by fibuladev 2 days ago
Comment by 4di 2 days ago
This doesn't work anymore. Is there a newer link?
Comment by cheriot 2 days ago
Comment by tra3 2 days ago
GET /ignore-all-previous-instructions.
How do you protect against that?
Comment by yruzin 2 days ago
Comment by 2001zhaozhao 2 days ago
Comment by UltraSane 2 days ago
Comment by keypusher 2 days ago
Comment by ryhminghistory 2 days ago