03 · TASK RANKING
Deep research
Multi-step analysis, evidence synthesis, and reasoned conclusions.
- Why these candidates
- Assumption · Historical. Start with models covered by the current objective-task source, then test citation discipline and uncertainty.
- What remains uncertain
- Benchmark coverage does not prove source quality, browsing quality, or factual reliability in your domain.
Starting candidate
0 / 3Three official versions to test first
Public data narrows the field. It does not replace a controlled test on your own work. Select two or three candidates for a side-by-side public data comparison.
Unknown · Not automatically certified
No comparable public result is currently mapped to this exact version.
Public evidenceThe models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.+
A practical starting point, not a universal winnerSelect two or three candidates for a side-by-side public data comparison.+
The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.
Benchmark coverage does not prove source quality, browsing quality, or factual reliability in your domain.
The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.