Starting candidate 1
StableGrok 4.5
grok-4.5xAICost per successful task · #1
Reduce successful-task cost without choosing by price alone.
Public data narrows the field. It does not replace a controlled test on your own work.
grok-4.5xAICost per successful task · #1
gpt-5.6-lunaOpenAICost per successful task · #5
gemini-3.5-flashGoogleCost per successful task · #9
The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.
Published prices and workloads change; provider token accounting and your success criteria may differ.
The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.