On the multimodal-AI benchmark MMMU-Pro (multidisciplinary comprehension across images, charts, and figures), the top models now sit in the 80s. OpenAI reported 81.2% for GPT-5.5 (no tools) in April 2026, and Google reported 80.5% for Gemini 3.1 Pro in February. When the original MMMU launched in late 2023, GPT-4V scored about 56%. MMMU-Pro is a harder revision of MMMU (released September 2024), so the two are not the same yardstick — yet models now clear 80% on the harder one. The era of "text-only" AI is truly over.

Scores are not the only thing that changed. How these models are built has been shifting from "stitched" to "native". It used to be normal to chain separate models — one to transcribe speech, one to reason over the text, one to read the answer aloud (OpenAI describes ChatGPT's pre-GPT-4o Voice Mode as exactly this three-model pipeline). Today, designs where one model takes several kinds of input directly have spread, and Gemini and Qwen's Omni line handle all four — text, images, audio and video — in a single model. Not every vendor has gone that way, though. As of September 2026, OpenAI's flagship model and the Claude API accept only text and images as input, and OpenAI handles audio with separate voice models such as gpt-realtime (see the table in §4).

Let me get my take out front: multimodal has gone from "nice to have" to "not having it is a non-starter". Snap a photo of an error screen and have AI solve it on the spot, screenshot a PDF and pull out the key points, transcribe and summarize a YouTube video — these are now the baseline of 2026 AI fluency. This article covers the definition, the difference between stitched and native multimodal, what technically determines strength in each modality, what to judge on for each use case, benchmarks, and the limits. It is organized around how to decide rather than which model currently ranks first, so it stays useful as generations turn over.

MULTIMODAL AI · 2026

AI that handles text, images, audio and video

— how much of it one model can take in depends on the model

TEXT
Text
Prose, code, symbols
IMAGE
Image
Photos, charts, screenshots
AUDIO
Audio
Speech, music, ambience
VIDEO
Video
Time + visuals + audio

On MMMU-Pro, top models have reached the 80% range (vendor-reported figures and sources in §1).
The "image is a bonus" era is over. But only some models — Gemini and Qwen's Omni line among them — take all four in one model; OpenAI's flagship model and the Claude API stop at images (as of September 2026).

1. In 2026, AI Stopped Being "Text Only" — MMMU-Pro Crosses 80%

"Multimodal" started trending in 2024, but the models just before that could only read images as an afterthought: when MMMU (multidisciplinary multimodal understanding) launched in late 2023, even the top model, GPT-4V, scored only around 56%. The median human expert (82.6%) was far out of reach on image questions requiring specialist knowledge.

2026 looks entirely different. Published MMMU-Pro scores (the harder revision of MMMU):

"Crossing 80% means the benchmark is saturating" is the 2026 reality. Differentiation has moved to video understanding (Video-MMMU), OCR-dense documents, and joint audio-visual reasoning — harder territory. The public leaderboard at MMMU benchmark lets anyone compare.

2. What Is Multimodal AI? — AI That Takes More Than Text

Definition: "an AI model that can take inputs other than text (images, audio, video, etc.)." How far that goes varies by model — from text plus images up to audio and video handled by a single model (see the terminology box below).

Traditional AI was single-modality: GPT-3 handled text; Whisper handled speech-to-text only; Stable Diffusion handled text-to-image only. Combining them required a pipeline where the output of one model fed another, and information was lost at every handoff.

In multimodal AI, one model looks at several inputs at once and answers. For example: "Look at the error message in this screenshot (image) together with my question (text) and tell me the cause" — done in a single API call.

Terminology: LMM (Large Multimodal Model) = a large model with multimodal capability. VLM (Vision-Language Model) = text + image only. Omnimodal = next-generation models that unify 4+ modalities. The same "supports multimodal" label means very different things depending on whether a model is omnimodal or a VLM (text + image only). If you plan to work with audio or video, check that distinction rather than counting the ◎ marks in a support matrix — VLM-based models often have only limited audio and video handling.

3. Stitched vs Native — The Architectural Divide

Understanding the mechanical difference makes each model's strengths easier to read. There are broadly two ways to build a multimodal model.

Architecture comparison

Stitched vs Native

① Stitched
  • Transcription, reasoning and speech output run on separate models
  • Models hand off to each other as text
  • Tone, multiple speakers and background noise get lost on the way
  • Each extra stage adds latency
  • e.g., ChatGPT Voice Mode before GPT-4o (three chained models)
VS
② Native
  • One model takes several inputs directly
  • Input to output in the same network
  • The model can directly see tone and how picture and sound line up
  • Little information lost to hand-offs
  • e.g., GPT-4o (May 2024), Google's Gemini, Qwen's Omni line

Native models can "interpret a video's audio and picture together" inside one model.
Stitched setups insert a hand-off — such as transcribing speech to text first — and information is lost there.

Concrete example: "watch a YouTube cooking video and pull out the recipe." Stitched: audio → Whisper to text → GPT for the summary; video → pull frames → analyze separately. Many steps. A native model that accepts video directly (Gemini, for instance) can take the entire video file as input and return the recipe in one API call, matching the spoken explanation to the visible action inside the model.

4. What Determines Strength in Each Modality

"Which model is best" changes with every generation. The video leader and the best voice experience have both changed hands within six months. So the thing worth learning isn't the ranking itself — it's what technically determines strength in each modality. Once you have that, you can judge a new model on your own the day it ships.

🎬 What decides video understanding

① Frame sampling density (how many seconds between sampled frames) ② context long enough to hold the whole runtime ③ temporal causal reasoning. An hour of video is 3,600 frames even at 1 fps, which converts to hundreds of thousands of tokens. Video ability and context length are coupled — that's why models strong on video invariably have a big window.

🎙 What decides voice conversation

① Round-trip latency ② whether it's full duplex (can you cut in while it's still talking?) ③ retention of prosody and emotion. This is where the stitched/native divide from §3 shows up most: any setup that transcribes speech to text before processing it is structurally stuck with added delay and lost information.

📄 What decides document and UI parsing

① OCR accuracy ② layout retention (do tables, columns, and footnotes survive?) ③ whether it can return coordinates. The third is easy to overlook, but it is decisive for agents that operate a screen. Knowing that a button exists is useless if the model can't say where it is — you can't click it.

🔓 What decides the open-weight options

① Genuinely omnimodal, or a stitched-together set of per-modality models ② weight availability and licence. Plenty of models that call themselves "open" attach conditions to commercial use. If you plan to run it yourself, read the licence before you read the benchmarks.

The table below is our own rough rating as of May 2026, based on what each vendor had published about supported inputs — not benchmark measurements. Rankings move, so check the list of current models for what you can pick today, and a comparison piece such as GPT-5.6 Sol vs Gemini or each vendor's model card for the head-to-head. Read the table here as a worked example of how the four axes above actually show up as differences.

ModelTextImageAudioVideoStrength
GPT-5.5◎◎××Reads text + images; the API does not accept audio or video input
Gemini 3.1 Pro◎◎◎◎◎Strong at video understanding, including long-form video
Claude Opus 4.7◎◎××UI/document parsing; strong for agent workloads
Qwen3.5-Omni◎◎◎◎Omnimodal: one model takes all four inputs. Offered via API (weights unpublished as of September 2026)
DeepSeek V4-Pro◎×××Text-only and cheap (image input arrived with the Flash line from August 2026)

× = the API does not accept that input (app features that route through a separate model or transcription, such as voice chat, are not counted). Input support was checked against primary sources: OpenAI's GPT-5.5 model page, Google's Gemini 3.1 Pro model page, Anthropic's Claude Opus 4.7 model page, the Qwen3.5-Omni technical report, DeepSeek's API change log (checked September 2026).

What the table shows is that the four axes are each doing separate work:

  • The video gap opens up with length: Video-MME scores videos from 11 seconds to an hour in short, medium, and long buckets, and every model on its official leaderboard scores lower on the long bucket. Long footage is where axis ① frame density and axis ② context length become the bottleneck. If you only ever handle short clips, long-video rankings don't apply to you
  • Audio is settled by architecture: fast responses and emotional read-out are far easier to combine when speech is handled natively instead of being dropped to text first (a transcription step adds its own wait and strips intonation). A spec sheet can show "audio ◎" across the board and still describe two completely different experiences, if one of them routes through transcription
  • Document parsing splits on coordinates: the reason PDF and UI-screenshot reading gets praised in agent contexts like Cursor is less about raw accuracy than about whether the model can hand back a position
  • For open weights, omnimodal changes the value: one model that covers every modality is far lighter to operate than a rig that switches models per modality. Options at near-frontier quality are now arriving at dramatically lower cost

5. Benchmarks That Matter — MMMU / Video-MMMU / OCR / Audio

You'll choose the wrong model if you don't know what each benchmark actually tests. Four benchmarks to know in 2026:

Benchmarks × 4

What we measure multimodal AI by

① MMMU-Pro
Multidisciplinary understanding from images + figures + charts. Top models sit in the 80s, level with the median human expert (80.8%). Already weak as a differentiator.
② Video-MMMU
300 lecture-style expert videos + 900 questions. Tests whether a model can learn from a video and apply it, across perception, comprehension, and adaptation.
③ DocVQA / OCRBench
Document + in-image text. The gap shows in whether tables, columns and coordinates survive intact. For UI parsing, invoices and forms, look here rather than at MMMU.
④ AudioBench
Measures how well audio LLMs (AudioLLMs) understand what they hear: speech recognition and translation, audio scenes, and emotion and speaker traits in the voice (50+ datasets in the official repository). It does not measure latency or voice generation, so check the speed of voice conversation separately.

"High MMMU = good at everything" is wrong.
For video, check Video-MMMU; for documents, DocVQA; for audio, AudioBench — otherwise selection misses.

6. By Use Case — What to Judge On

Five common patterns, each framed as "what to check before you decide". No model names here — that section would be out of date within six months. Instead, these are steps you can run identically against whatever ships next.

  • ① Phone-photo Q&A / diagnosis (meal photo → nutrition, error screen → fix, product photo → search)
    → Almost anything will do. Try two on their free plans and keep whichever answers at the level of detail you like. This is the use case where models differ least, so it isn't worth much selection effort
  • ② PDF / document parsing (receipts, contracts, technical specs, papers)
    → Test with your own real files. Three things decide it: do tables survive intact, does it follow multi-column pages in the right reading order, and does it hold up on a badly scanned page. Throwing your single worst-quality page at it tells you more, faster, than any published OCR score
  • ③ Video transcription & summary (meetings, lectures, YouTube)
    → It branches on the length of video you actually handle. At around ten minutes the differences are small. If you're feeding in hour-long recordings whole, axes ① frame density and ② context length from §4 start to bite, so check long-form video benchmarks and the size of the window
  • ④ Voice conversation / interpreter / interview practice
    → First confirm whether audio is handled natively. "Voice supported" still means added latency and flattened intonation if it goes through transcription. The test is simple: try interrupting it mid-sentence
  • ⑤ Cost-first / bulk processing
    → Look past the unit price at batch discounts and at whether self-hosting is realistic for you. If you're going open-weight, read the licence's commercial terms before you look at performance
How to think about combining them: rather than committing to one, a primary model plus a second one that covers its weak spot is the more stable setup. Two $20/month plans come to $40, not far off a single top-tier subscription. Low-frequency needs like video can be pushed onto a free tier such as Google AI Studio as they come up. This way of allocating work holds no matter how many generations the models go through.

7. Hard Limits — Use, Don't Trust Blindly

Multimodal AI is strong, but three limits will bite you if ignored.

Limit ①: Don't read photo-derived "guesses" as facts

Asking "OCR the amount on this receipt" sounds simple, but if the image is low-resolution, dim, or skewed, AI fabricates plausible numbers. Even a model above 80% on MMMU-Pro still gets nearly one answer in five wrong. Amounts, dates, proper nouns — always have a human double-check. Especially in legal, finance, healthcare.

Limit ②: Video accuracy drops in the middle

Even for the model that tops the video benchmarks, retrieving information from the middle of a 1-hour video is hard — the same "Lost in the Middle" issue as the context-window problem. For key segments, specify timestamps: "analyze the 30:00–35:00 segment specifically" gets much better results.

Limit ③: Audio struggles with dialects and jargon

Standard English / Japanese speech is accurate, but regional dialects, specialist vocabulary, multi-speaker crosstalk, and noisy environments increase errors. For meeting records and other high-stakes use, pair with specialized tools (Otter.ai, Notta, etc.), or clean up audio first before sending to AI.

Summary

Recap:

  • 2026: GPT-5.5 and Gemini 3.1 Pro crossed 80% on MMMU-Pro (vendor-reported; figures and sources in §1). Multimodal AI has moved from "nice to have" to "must have"
  • Architecture has been shifting from stitched pipelines of separate models to native models that take several inputs directly. But taking all four in one model is limited to Gemini, Qwen's Omni line and a few others; OpenAI's flagship model and the Claude API stop at images (as of September 2026)
  • Four axes decide strength: video = frame density and context length / audio = native processing vs routed through transcription / documents = whether it returns coordinates / open weights = omnimodal or not, plus the licence. Rankings change hands; these axes don't
  • Benchmarks: MMMU-Pro / Video-MMMU / DocVQA / AudioBench — check all four axes before choosing
  • For per-use-case selection, testing on your own real material is the shortest path — with PDFs especially, your worst-scanned page beats any published spec. For subscriptions, one primary plus one that covers its weak spot is the stable arrangement
  • Three limits: low-quality image guesses / mid-video accuracy drop / dialect & jargon audio. Double-check critical outputs

In 2026, AI work that completes "in text alone" is shrinking fast. Phone photos, meeting recordings, YouTube videos, PDFs — handing them to AI is now routine (whether one model can take all of them depends on the model). Knowing how to use multimodal is no longer "a useful feature"; it is the floor of 2026 AI literacy. Start by feeding the AI one photo from your phone today — that's enough to begin.

FAQ

Q1. Can I try multimodal AI for free?

Yes. ChatGPT's free plan (image input OK), Google AI Studio (free tier, video included) and Claude.ai's free plan (images OK) all let you try. Note that in Google AI Studio and the Gemini API free tier, what you submit is used to improve Google products and may be read by human reviewers (Gemini API terms), so keep sensitive images and audio out. But the models assigned to free plans change with every generation, so rather than memorizing model names, check on each service's own screen which inputs it accepts (image, audio, video) and how much you can use per day. Some features, such as voice-conversation and long-video limits, expand on paid plans. See Free AI Tools Guide.

Q2. How is image-generation AI different from multimodal AI?

Different terms. Tools like Midjourney and Stable Diffusion specialize in generating images from text — a one-way text→image flow. Multimodal AI refers to understanding images (and other modalities) as inputs. Some services, such as ChatGPT and Gemini, handle both, but understanding and generating are separate abilities — being good at one doesn't mean being good at the other, so test each separately for your use case. See Image-Generation AI Tools Compared.

Q3. How do I send video over the API?

The Gemini API (the developer API on ai.google.dev) accepts video as-is. For long videos or ones you reuse, the recommended route is to upload with the Files API and pass the file reference; small videos can be embedded inline in the request. You can also pass a public YouTube URL, or register Cloud Storage files with the Files API (Google's developer docs). Through Google Cloud's Vertex AI, you put a Cloud Storage gs:// URI in fileData (Google Cloud docs). The OpenAI and Claude APIs do not accept video directly as of September 26, 2026 (OpenAI's model page for GPT-6 Astra lists Video as "Not supported"; Anthropic says all current models take text and image input). For both, the procedure is to extract frames from the video and send them as images — OpenAI has an example in its official Cookbook. See AI API Beginner Guide.

Q4. Is privacy okay?

Images, audio and video often contain sensitive data. Training use differs between consumer apps and API/business plans. On the consumer side, ChatGPT may use your conversations for training, and you can stop that by turning off "Improve the model for everyone" in settings (OpenAI's help center). Claude (Free, Pro, Max) uses your chats for training if you choose to allow it (Anthropic's Privacy Center). Gemini Apps use your chats — including through human review — to improve Google services while "Keep Activity" is on, and turning it off stops that (Gemini Apps Privacy Hub). By contrast, the OpenAI API and ChatGPT Business/Enterprise, the Anthropic API and business plans, and the paid tier of the Gemini API are not used for training by default (Anthropic's commercial FAQ, Gemini API terms). The exception: with the Gemini API free tier and Google AI Studio, your inputs are used to improve Google products. For faces, medical images or internal documents, pick a paid API or a business plan. For full secrecy, run an open-weight model on your own hardware (for an all-modality Qwen model, that means the previous-generation Qwen3-Omni; the latest Qwen3.5-Omni is API-only, with no published weights as of September 2026).

Q5. Is multimodal more expensive than text-only?

Images and videos bill by token conversion. One image ≈ a few hundred to ~1,000 tokens (resolution and model dependent); for video, the Gemini API counts about 100 tokens per second at the default setting and about 300 at high media resolution (Google's developer docs, checked September 2026). One hour at the default is 100 × 3,600 s ≈ 360,000 tokens. The cost techniques in AI Token Cost Saving (excerpt-only sending, caching) also work for video.