OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Training the latest GPT-5.6 against GPT-Red made it the companys most robust release yet against prompt injection attacks.
GPT-Red is an automated red-teaming model that scales prompt injection vulnerability discovery with the aim of fixing issues before the tools are deployed widely. According to OpenAI, GPT-Red is a strong red-teamer, and the companys previous models are highly vulnerable to its prompt injection attacks. The system was used to adversarially train GPT-5.6, uncovering vulnerabilities that were used to make the model more resistant to these types of attacks.
Prompt injection is a type of attack where malicious instructions are embedded in input data to trick AI models into performing actions outside their intended scope. As AI agents gain more capabilities and access to external tools, including the ability to execute code and interact with third-party services, prompt injection has become a critical security concern for the industry.
The development of GPT-Red reflects a broader trend in AI security toward using AI systems to defend against other AI systems. The same techniques that make large language models vulnerable to prompt injection can be harnessed to discover and patch those vulnerabilities. OpenAI approach is analogous to using the same tools that attackers would use but deploying them defensively.
The release comes amid growing industry concern about the security of AI systems. The Five Eyes national security agencies recently issued a joint statement warning about the increasing cyber risks of AI models, particularly their ability to autonomously hack into systems and networks. Defensive AI systems like GPT-Red are seen as a critical part of the response to these emerging threats.
This article was adapted from MIT Technology Review. Read the original here.
