OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them
OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.
- Nguồn
- Decrypt
- Đã đăng
- 2026-09-17 22:31 UTC
- Cập nhật cache
- 2026-09-17 22:34 UTC
Trang này chỉ hiển thị tiêu đề, tóm tắt và thông tin nguồn.
Mở bài gốc ↗Kiểm tra phí chuyển