← All explainers
AI in 3 · 3 min read

What "multimodal" actually changes for your product

Images, audio and documents as first-class inputs — and the new failure modes that come with them.

AI in 3Foundational

Multimodal means the model takes more than text: screenshots, photos, PDFs, audio, sometimes video. In product terms it removes a transcription step your users were doing by hand.

The biggest unlocks are unglamorous. Upload the invoice instead of typing the fields. Screenshot the error instead of describing it. Read the whiteboard photo instead of retyping it. Each removes a form.

The new risks are real: images carry hidden text, image tokens are expensive and count against the same context window, and accuracy on dense tables or handwriting is much lower than on clean text. Test with your users' actual photos, not stock samples.

Want this applied to your situation?

Sessions are direct and specific — you leave with a decision, not a reading list.

Book a session
More in AI in 3