AI0 views

OpenAI's Secret Red Team AI: GPT-Red Hacks Other AI Models at 84% Success Rate

OpenAI has developed GPT-Red, an internal-only AI model designed to identify vulnerabilities in other AI systems by simulating hacker attacks. The model operates autonomously, generating attack commands, evaluating responses, and adapting its tactics in real time.

In testing, GPT-Red succeeded in compromising target systems at an 84% rate—far outpacing human security experts, who achieved just 13% success under the same conditions. Rather than release this tool publicly, OpenAI is keeping it confined to internal use to prevent weaponization for actual cyberattacks. The company plans to use GPT-Red exclusively to strengthen its own infrastructure defenses.