07 · TASK RANKING

Image & multimodal work

Images, documents, extraction, and mixed visual-text reasoning.

Official version
3
Public evidence
0
Data cutoff
Why these candidates
Include managed and open multimodal options, then test the exact image types, resolution, and output format you use.
What remains uncertain
The current public intelligence signal is not a dedicated vision benchmark and some candidates lack comparable data.
Starting candidate

Three official versions to test first

Public data narrows the field. It does not replace a controlled test on your own work. Select two or three candidates for a side-by-side public data comparison.

0 / 3
Starting candidate 1

Gemini 3.5 Flash

General availability
gemini-3.5-flash

GoogleNo comparable public result is currently mapped to this exact version.

Official website
Starting candidate 2

Command A Vision

General availability
command-a-vision-07-2025

CohereNo comparable public result is currently mapped to this exact version.

Official website
Starting candidate 3

Phi-4 Multimodal Instruct

General availability
microsoft/Phi-4-multimodal-instruct

MicrosoftNo comparable public result is currently mapped to this exact version.

Official website
All task rankings
Public evidenceThe models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

Gemini 3.5 Flash

No comparable public result is currently mapped to this exact version.

Command A Vision

No comparable public result is currently mapped to this exact version.

Phi-4 Multimodal Instruct

No comparable public result is currently mapped to this exact version.

A practical starting point, not a universal winnerSelect two or three candidates for a side-by-side public data comparison.

The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

The current public intelligence signal is not a dedicated vision benchmark and some candidates lack comparable data.

The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.