03 · TASK RANKING

Deep research

Multi-step analysis, evidence synthesis, and reasoned conclusions.

Official version
3
Public evidence
2
Data cutoff
Why these candidates
Start with models covered by current intelligence and objective-task sources, then test citation discipline and uncertainty.
What remains uncertain
Benchmark coverage does not prove source quality, browsing quality, or factual reliability in your domain.
Starting candidate

Three official versions to test first

Public data narrows the field. It does not replace a controlled test on your own work. Select two or three candidates for a side-by-side public data comparison.

0 / 3
Starting candidate 1

Claude Fable 5

General availability
claude-fable-5

AnthropicGeneral capability · 3

Official website
Starting candidate 2

GPT-5.6 Sol

General availability
gpt-5.6-sol

OpenAIGeneral capability · 5

Official website
Starting candidate 3

Kimi K3

General availability
kimi-k3

Moonshot AINo comparable public result is currently mapped to this exact version.

Official website
All task rankings
Public evidenceThe models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

Kimi K3

No comparable public result is currently mapped to this exact version.

A practical starting point, not a universal winnerSelect two or three candidates for a side-by-side public data comparison.

The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

Benchmark coverage does not prove source quality, browsing quality, or factual reliability in your domain.

The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.