GPT-5.6 Luna
gpt-5.6-lunaOpenAICost per successful task · #5
High-volume classification, extraction, routing, and short responses.
Public data narrows the field. It does not replace a controlled test on your own work.
gpt-5.6-lunaOpenAICost per successful task · #5
gemini-3.1-flash-liteGoogleNo comparable public result is currently mapped to this exact version.
deepseek-v4-flashDeepSeekNo comparable public result is currently mapped to this exact version.
No comparable public result is currently mapped to this exact version.
No comparable public result is currently mapped to this exact version.
The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.
OpenGPT does not currently publish a comparable latency history for every provider.
The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.