09 · TASK RANKING

Private & local deployment

Open-weight candidates for controlled data paths and self-hosting.

Official version
3
Public evidence
0
Data cutoff
Why these candidates
Start with confirmed downloadable versions and open-weight public signals, then validate hardware, license, security, and serving cost.
What remains uncertain
Open weights do not automatically mean low cost, full commercial freedom, privacy, or easy operation.
Starting candidate

Three official versions to test first

Public data narrows the field. It does not replace a controlled test on your own work. Select two or three candidates for a side-by-side public data comparison.

0 / 3
Starting candidate 1

DeepSeek V4 Pro

General availability
deepseek-v4-pro

DeepSeekNo comparable public result is currently mapped to this exact version.

Official website
Starting candidate 2

Llama 4 Maverick 17B-128E

General availability
meta-llama/Llama-4-Maverick-17B-128E-Instruct

MetaNo comparable public result is currently mapped to this exact version.

Official website
Starting candidate 3

Qwen3 235B-A22B

General availability
Qwen/Qwen3-235B-A22B

Alibaba QwenNo comparable public result is currently mapped to this exact version.

Official website
All task rankings
Public evidenceThe models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

DeepSeek V4 Pro

No comparable public result is currently mapped to this exact version.

Llama 4 Maverick 17B-128E

No comparable public result is currently mapped to this exact version.

Qwen3 235B-A22B

No comparable public result is currently mapped to this exact version.

A practical starting point, not a universal winnerSelect two or three candidates for a side-by-side public data comparison.

The models use different public evidence scales. Missing evidence is not treated as zero, and OpenGPT does not combine unlike sources into one score.

Open weights do not automatically mean low cost, full commercial freedom, privacy, or easy operation.

The evaluation opens with this task template and the first two selected models. You can edit every task before collecting answers.