Qwen 3.7 Max
qwen3.7-max-2026-06-08Alibaba QwenNo comparable public result is currently mapped to this exact version.
Chinese-first writing, translation, localization, and mixed-language tasks.
Public data narrows the field. It does not replace a controlled test on your own work. Select two or three candidates for a side-by-side public data comparison.
qwen3.7-max-2026-06-08Alibaba QwenNo comparable public result is currently mapped to this exact version.
kimi-k3Moonshot AINo comparable public result is currently mapped to this exact version.
deepseek-v4-proDeepSeekUser preference · 50
No comparable public result is currently mapped to this exact version.
No comparable public result is currently mapped to this exact version.
The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.
The public overall ranking is not a Chinese-language benchmark; language quality must be tested directly.
The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.