OpenAI unveils GPT-Red to boost AI safety with internal red-teaming
OpenAI introduced GPT-Red, an internal red-team model that finds prompt-injection flaws at scale. Trained via self-play, it outperformed humans in attacks and has already reduced failure rates in newer GPT models through adversarial training.
🗞 #openai @testingcatalog
Post #8716
1.24K