Starting candidate 1
StableClaude Fable 5
claude-fable-5AnthropicUser preference · #1
Drafting, analysis, planning, and mixed office tasks.
Public data narrows the field. It does not replace a controlled test on your own work.
claude-fable-5AnthropicUser preference · #1
gpt-5.6-solOpenAIUser preference · #11
gemini-3.5-flashGoogleUser preference · #17
The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.
A general preference rank cannot predict accuracy on a specialized workflow.
The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.