01 · TASK RANKING

Everyday work

Drafting, analysis, planning, and mixed office tasks.

Official version
3
Public evidence
3
Data cutoff
Why these candidates
Start with broadly capable models that have current public preference evidence, then test them on your own recurring work.
What remains uncertain
A general preference rank cannot predict accuracy on a specialized workflow.
Starting candidate

Three official versions to test first

Public data narrows the field. It does not replace a controlled test on your own work. Select two or three candidates for a side-by-side public data comparison.

0 / 3
Starting candidate 1

Claude Fable 5

Stable
claude-fable-5

AnthropicUser preference · 1

Official website
Starting candidate 2

GPT-5.6 Sol

Stable
gpt-5.6-sol

OpenAIUser preference · 11

Official website
Starting candidate 3

Gemini 3.5 Flash

Stable
gemini-3.5-flash

GoogleUser preference · 17

Official website
Public evidenceThe models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.
A practical starting point, not a universal winnerSelect two or three candidates for a side-by-side public data comparison.

The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

A general preference rank cannot predict accuracy on a specialized workflow.

The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.

All task rankings