Starting candidate 1
StableClaude Fable 5
claude-fable-5AnthropicUser preference · 1
Drafting, analysis, planning, and mixed office tasks.
Public data narrows the field. It does not replace a controlled test on your own work. Select two or three candidates for a side-by-side public data comparison.
claude-fable-5AnthropicUser preference · 1
gpt-5.6-solOpenAIUser preference · 11
gemini-3.5-flashGoogleUser preference · 17
The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.
A general preference rank cannot predict accuracy on a specialized workflow.
The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.