Multi-Modal AI
AI models that can process and generate multiple types of content: text, images, audio, and video. Enables tasks like describing images or generating visuals from text.
Why it matters
Multi-modal AI matters because real life isn't only text. Being able to show an AI a photo, speak to it, or have it read a chart unlocks tasks that typing alone can't. It's why you can snap a picture of a broken part and ask what it is, or paste a screenshot instead of retyping everything. This is quickly becoming the default way people use AI.
A concrete example
Photograph the inside of your fridge and ask "What can I cook with this?" and a multi-modal model reads the ingredients from the image and suggests recipes. Or share a screenshot of a confusing error message and get a plain-English explanation. Because it handles images, audio, and text together, you can hand it whatever you have on hand instead of describing everything in words.
How to use it
The practical unlock is that you can stop transcribing things. Photograph a whiteboard and ask for the action items, screenshot an error and ask what it means, share a chart and ask what it shows. For anything where the detail matters — a handwritten figure, a serial number, a value read off an axis — treat the reading as a first pass and check it, because a confident misreading looks exactly like a correct one.
The common mistake
Assuming that because a model can describe an image well, it can read precise detail from it. Describing a scene and reliably transcribing a column of numbers are different tasks, and the second is much easier to get subtly wrong.
Related terms
Vision Language Model (VLM)
An AI model that can process both images and text, understanding visual content and answering questions about it. Used for image captioning, visual Q&A, document analysis, and multimodal reasoning tasks.
Multimodal Embedding
Vector representations that encode multiple data types — text, images, audio — into a shared mathematical space. Enables cross-modal search: find images using text queries, or retrieve documents using audio clips.
Hallucination
When an AI model generates information that sounds plausible but is factually incorrect or entirely fabricated. Common with statistics, citations, and historical claims.
Agentic AI
AI systems that operate autonomously over extended tasks — planning, executing, and self-correcting without step-by-step human guidance. Unlike chatbots, agentic AI sets sub-goals, uses tools, and adapts its strategy based on intermediate results.
AI Agent
An autonomous AI system that can perceive its environment, make decisions, and take actions to achieve goals — like managing your email, scheduling meetings, or monitoring data.
AI Disclosure
Stating that AI was used in producing a piece of work, where a policy, client or publisher requires it. Distinct from permission: some contexts allow AI use but require it to be declared.