06 · TASK RANKING

Chinese & multilingual work

Chinese-first writing, translation, localization, and mixed-language tasks.

Official version
3
Public evidence
2
Data cutoff:
Why these candidates
Start with providers that publish Chinese-focused or multilingual model lines, then evaluate terminology and regional language with your own samples.
What remains uncertain
The public overall ranking is not a Chinese-language benchmark; language quality must be tested directly.
Public evidence
User preference
Data cutoff
Starting candidate

Three official versions to test first

Public data narrows the field. It does not replace a controlled test on your own work.

3 / 3
Starting candidate 1

Qwen 3.7 Max

Stable
qwen3.7-max-2026-06-08

Alibaba QwenNo comparable public result is currently mapped to this exact version.

Official website
Starting candidate 2

Kimi K3

Stable
kimi-k3

Moonshot AIUser preference · #10

Official website
Starting candidate 3

DeepSeek V4 Pro

Preview
deepseek-v4-pro

DeepSeekUser preference · #46

Official website
Public evidenceThe models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

Qwen 3.7 Max

No comparable public result is currently mapped to this exact version.

Kimi K3

Source rank
#10
Source value
1486
Mapping reviewed

DeepSeek V4 Pro

Source rank
#46
Source value
1457
Mapping reviewed
A practical starting point, not a universal winnerSelect two or three candidates for a side-by-side public data comparison.

The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

The public overall ranking is not a Chinese-language benchmark; language quality must be tested directly.

The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.