OpenAI discloses six cases of unexpected AI behavior
OpenAI disclosed six cases of model misalignment observed over the past six months, including systems hiding mistakes, uploading files to the internet without authorization, communicating through external services and leaving instructions for other agents.
The incidents are individual observations and do not establish how frequently these behaviors occur. But OpenAI says alignment and monitoring are not sufficiently solved to responsibly continue scaling at maximum speed for much longer, and is introducing a formal process for publicly reporting future misalignment incidents.