10 · TASK RANKING

Fast production workflows

High-volume classification, extraction, routing, and short responses.

Official version
3
Public evidence
1
Data cutoff:
Why these candidates
Start with official fast or lightweight variants, then measure latency, throughput, retries, and output validity in your region.
What remains uncertain
OpenGPT does not currently publish a comparable latency history for every provider.
Public evidence
Cost per successful task
Data cutoff
Starting candidate

Three official versions to test first

Public data narrows the field. It does not replace a controlled test on your own work.

3 / 3
Starting candidate 1

GPT-5.6 Luna

Stable
gpt-5.6-luna

OpenAICost per successful task · #5

Official website
Starting candidate 2

Gemini 3.1 Flash-Lite

Stable
gemini-3.1-flash-lite

GoogleNo comparable public result is currently mapped to this exact version.

Official website
Starting candidate 3

DeepSeek V4 Flash

Preview
deepseek-v4-flash

DeepSeekNo comparable public result is currently mapped to this exact version.

Official website
Public evidenceThe models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

GPT-5.6 Luna

Source rank
#5
Source value
$0.202
Mapping reviewed

Gemini 3.1 Flash-Lite

No comparable public result is currently mapped to this exact version.

DeepSeek V4 Flash

No comparable public result is currently mapped to this exact version.

A practical starting point, not a universal winnerSelect two or three candidates for a side-by-side public data comparison.

The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

OpenGPT does not currently publish a comparable latency history for every provider.

The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.