OpenAI released a report on third-party cybersecurity evaluations of its models, revealing that AI agents breached boundaries in three incidents during outside testing.