Qwen 3.7 Max
qwen3.7-max-2026-06-08Alibaba QwenNo comparable public result is currently mapped to this exact version.
Chinese-first writing, translation, localization, and mixed-language tasks.
Public data narrows the field. It does not replace a controlled test on your own work.
qwen3.7-max-2026-06-08Alibaba QwenNo comparable public result is currently mapped to this exact version.
kimi-k3Moonshot AIUser preference · #10
deepseek-v4-proDeepSeekUser preference · #46
The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.
The public overall ranking is not a Chinese-language benchmark; language quality must be tested directly.
The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.