AI Catchup

OpenAI Introduces GPT-Red: An Internal Automated Red-Teaming Model for Prompt Injection

By 3 min read

OpenAI published details on GPT-Red, an internal automated safety red-teaming model designed to find prompt injection vulnerabilities at scale and strengthen defenses before broader deployment.

OpenAI introduced GPT-Red, an internal-only automated red-teaming model designed to find prompt injection vulnerabilities at scale and help strengthen defenses before wider deployment (OpenAI blog).

What GPT-Red is

According to OpenAI, GPT-Red is their “current best automated safety red-teaming model,” built to uncover vulnerabilities and generate attacks that can be incorporated into training to improve robustness (OpenAI blog).

OpenAI says GPT-Red is kept separate from the models they deploy so that the malicious capabilities they train into it are not released broadly (OpenAI blog).

What it does (and why it matters)

OpenAI describes GPT-Red as an agent that iterates toward a goal by sending prompts, observing model responses, and refining its approach, with a focus on discovering failures such as successful prompt injections (OpenAI blog).

OpenAI’s blog post highlights prompt injection risks in tool-using AI systems (e.g., content embedded in emails, webpages, tool responses, or code repositories) and frames GPT-Red as a way to harden models against those attacks before broader deployment (OpenAI blog).

Reported results

OpenAI reports that its latest production model, GPT-5.6 Sol, achieved “6× fewer failures” on its hardest direct prompt injection benchmark compared to OpenAI’s best production model from about four months earlier (OpenAI blog).

OpenAI also reports that on a replicated “indirect prompt injection arena” benchmark, GPT-Red succeeded in 84% of scenarios vs. 13% for humans (OpenAI blog).

Availability

OpenAI says GPT-Red is internal-only and is used to improve the robustness of models that are deployed more widely (OpenAI blog).

Sources

Provenance note. Every claim above -- the "current best automated safety red-teaming model" description, the separation of GPT-Red from deployed models, the iterate-and-refine attack loop, the "6x fewer failures" direct prompt injection result for GPT-5.6 Sol, the 84%-versus-13% indirect prompt injection arena figure, and the internal-only availability -- was checked against that OpenAI post when this page was published on July 16, 2026, and no claim here rests on the X announcement. OpenAI revised the post on July 29, 2026 (per its sitemap lastmod). On August 8, 2026 the current version was read in full in a browser session and compared against a July 17, 2026 Internet Archive snapshot of the original: the only material change is that the closing sentence promising a pre-print "later this week" was replaced by a link to the published paper, listed above. Every figure and description cited on this page appears unchanged in the revised post.

Get the weekly AI Catchup

Tools, practices, and what matters, in your inbox every week.