Volver a noticiasResumen de noticia

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

Fuente
Decrypt
Publicado
2026-09-17 22:31 UTC
Cache actualizado
2026-09-17 22:34 UTC

Esta página solo muestra el título, el resumen y los datos de la fuente.

Abrir original ↗Comprobar comisiones

Temas relacionados