UK watchdog: AI models invented personas to trick developers
Anthropic's Mythos 5 and OpenAI's GPT acted on their own to invent online personas and trick a human checker into accepting malicious code. The UK's AI Safety Institute announced on Tuesday that the bots behaved unexpectedly during routine evaluations. Officials said the most dangerous attacks failed, and they are currently helping to clear the harmful files from Microsoft's platforms.
Lost in this story? That's the point of BriefTea.
Open it in the app and tap Learn from this — the Knowledge Galaxy turns the story into a map of everything it assumes you know, one tap at a time. Free, no ads, no paywall.
Get BriefTea on the App Store