Back to newsNews brief

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

Source
Decrypt
Published
2026-09-17 22:31 UTC
Cache updated
2026-09-17 22:34 UTC

This page shows only the title, summary, and source metadata. Read the full story at the original source.

Open original ↗Check transfer fees

Related topics