Get the app →
BriefTea logoBriefTea

Every story in sixty words

Explained in plain English

What's a multimodal AI model?

Imagine an AI like Gemini Omni, which powers Google Photos' Video Remix, that can understand and work with different types of information all at once. So, it doesn't just 'see' your video, but also 'hears' the audio and 'reads' any text, combining everything to create that watercolour art effect. It's much cleverer than an AI that only deals with one type of data.

Stories that explain this

Google Photos launches AI tool to turn videos into art

Questions this opens up

How does AI 'see' video? What is 'Gemini Omni'? Can AI create music?

Related explainers

How does AI create art? Is AI open source? How AI 'remixes' videos AI's impact on creativity Is all AI code public? Why do companies charge for AI? Who controls open source AI? What are 'creative templates'?

The full catalogue

Browse every card filed under S →