05 · TASK RANKING

Customer support

Clear, safe, policy-aware replies to common requests.

Official version
3
Public evidence
1
Data cutoff:
Why these candidates
Begin with balanced and enterprise-oriented models, then score policy adherence, empathy, and escalation on your cases.
What remains uncertain
Public preference data does not test your policies, integrations, language mix, or regulated obligations.
Public evidence
User preference
Data cutoff
Starting candidate

Three official versions to test first

Public data narrows the field. It does not replace a controlled test on your own work.

3 / 3
Starting candidate 1

Claude Sonnet 5

Stable
claude-sonnet-5

AnthropicUser preference · #38

Official website
Starting candidate 2

GPT-5.6 Terra

Stable
gpt-5.6-terra

OpenAINo comparable public result is currently mapped to this exact version.

Official website
Starting candidate 3

Command A+

Stable
command-a-plus-05-2026

CohereNo comparable public result is currently mapped to this exact version.

Official website
Public evidenceThe models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

Claude Sonnet 5

Source rank
#38
Source value
1461
Mapping reviewed

GPT-5.6 Terra

No comparable public result is currently mapped to this exact version.

Command A+

No comparable public result is currently mapped to this exact version.

A practical starting point, not a universal winnerSelect two or three candidates for a side-by-side public data comparison.

The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

Public preference data does not test your policies, integrations, language mix, or regulated obligations.

The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.