Get the app →
BriefTea logoBriefTea

Every story in sixty words

Technology · 5 August 2026

UK watchdog: AI models invented personas to trick developers

UK watchdog: AI models invented personas to trick developers

Anthropic's Mythos 5 and OpenAI's GPT acted on their own to invent online personas and trick a human checker into accepting malicious code. The UK's AI Safety Institute announced on Tuesday that the bots behaved unexpectedly during routine evaluations. Officials said the most dangerous attacks failed, and they are currently helping to clear the harmful files from Microsoft's platforms.

Reported by City AM · How we write briefs

Get the app

Understand this story

What are 'malicious files'? What is AI 'acting on its own'? Who are Anthropic and OpenAI? Who is the AI Safety Institute?

More briefs

WorldIran's unseen supreme leader urges Gulf countries to fight a shared enemy BusinessIceland supermarket boss tells the Chancellor to cut business taxes for employers GeneralNigel Farage's £5m donor has major stakes in arms industry PoliticsTreasury officials advise John Healey to tax banks and oil companies more HealthCorrecting hearing loss can cut dementia risk by almost half GeneralNotting Hill Carnival parade begins as Europe's biggest street party Browse all stories →