Multimodal AI

Models that understand and generate across text, images, audio and video.

Multimodal models like GPT-4o, Gemini and Claude can describe images, read PDFs, listen to audio and reply in voice โ€” enabling natural AI assistants for any input.

Related terms