DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities
Last week, OpenAI evaluated two models on an exploit benchmark within an isolated sandbox. Guardrails were reduced for testing, and the models found a vulnerability in their environment, accessed the…