04 · TASK RANKING

Documents & grounded Q&A

Long files, summaries, citations, and answers bounded by supplied text.

Official version
3
Public evidence
1
Data cutoff:
Why these candidates
Compare models positioned for sustained context or enterprise retrieval, then test exact citation and refusal behavior.
What remains uncertain
Context-window claims do not establish retrieval quality or faithful use of a long document.
Public evidence
General capability
Data cutoff
Starting candidate

Three official versions to test first

Public data narrows the field. It does not replace a controlled test on your own work.

3 / 3
Starting candidate 1

Gemini 3.1 Pro

Preview
gemini-3.1-pro-preview

GoogleNo comparable public result is currently mapped to this exact version.

Official website
Starting candidate 2

Claude Fable 5

Stable
claude-fable-5

AnthropicGeneral capability · #1

Official website
Starting candidate 3

Command A+

Stable
command-a-plus-05-2026

CohereNo comparable public result is currently mapped to this exact version.

Official website
Public evidenceThe models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

Gemini 3.1 Pro

No comparable public result is currently mapped to this exact version.

Command A+

No comparable public result is currently mapped to this exact version.

A practical starting point, not a universal winnerSelect two or three candidates for a side-by-side public data comparison.

The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

Context-window claims do not establish retrieval quality or faithful use of a long document.

The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.