Starting candidate 1
StableClaude Fable 5
claude-fable-5AnthropicGeneral capability · #1
Multi-step analysis, evidence synthesis, and reasoned conclusions.
Public data narrows the field. It does not replace a controlled test on your own work.
claude-fable-5AnthropicGeneral capability · #1
gpt-5.6-solOpenAIGeneral capability · #2
kimi-k3Moonshot AIGeneral capability · #4
The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.
Benchmark coverage does not prove source quality, browsing quality, or factual reliability in your domain.
The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.