08 · TASK RANKING

Cost-efficient production

Reduce successful-task cost without choosing by price alone.

Official version
3
Public evidence
3
Data cutoff:
Why these candidates
Use the published cost-per-success view to form a shortlist, then measure total retries, latency, and operational overhead.
What remains uncertain
Published prices and workloads change; provider token accounting and your success criteria may differ.
Public evidence
Cost per successful task
Data cutoff
Starting candidate

Three official versions to test first

Public data narrows the field. It does not replace a controlled test on your own work.

3 / 3
Starting candidate 1

Grok 4.5

Stable
grok-4.5

xAICost per successful task · #1

Official website
Starting candidate 2

GPT-5.6 Luna

Stable
gpt-5.6-luna

OpenAICost per successful task · #5

Official website
Starting candidate 3

Gemini 3.5 Flash

Stable
gemini-3.5-flash

GoogleCost per successful task · #9

Official website
Public evidenceThe models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

Grok 4.5

Source rank
#1
Source value
$0.128
Mapping reviewed

GPT-5.6 Luna

Source rank
#5
Source value
$0.202
Mapping reviewed

Gemini 3.5 Flash

Source rank
#9
Source value
$0.249
Mapping reviewed
A practical starting point, not a universal winnerSelect two or three candidates for a side-by-side public data comparison.

The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

Published prices and workloads change; provider token accounting and your success criteria may differ.

The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.