Three things that decide the answer.
What we deploy, and what each one needs.
| Model | Parameters | Weights in VRAM | Smallest node | Context | Licence | Published headline |
|---|---|---|---|---|---|---|
| Kimi K2.7Frontier quality, strongest for agent work | 1T total | ~1.0 TB at FP8 | 8x H200 | 256k | Modified MIT | Scores for this release are not published. The preceding K2.6 release led open models on aggregate intelligence indices. |
| GLM 5.2Best all-round open model on published indices | 753B / ~40B active | ~750 GB at FP8 ~380 GB at Q4 | 8x H200 8x H100 at Q4 | 1M | MIT | 91.2 GPQA Diamond, 81.0 Terminal-Bench 2.1, 62.1% SWE-bench Pro (vendor) |
| GLM 4.6The value option when 5.2 is more than you need | 355B | ~355 GB at FP8 ~180 GB at Q4 | 8x H100 4x H100 at Q4 | 200k | MIT | Superseded by 5.2. Vendor scores published on the model card. |
| DeepSeek V4 FlashSparse attention, small footprint for its class | 284B | ~284 GB at FP8 ~145 GB at Q4 | 4x H100 2x H100 at Q4 | 128k | MIT | Inherits the V4 architecture. The V4 Pro flagship reports 80.6% SWE-bench Verified and 90.1 GPQA Diamond. |
| DeepSeek V3.2Strong on mathematics and structured derivation | 685B / 37B active | ~685 GB at FP8 ~350 GB at Q4 | 8x H200 8x H100 at Q4 | 128k | MIT | 88.5 MMLU. Behind the V4 line on agentic coding. |
| Qwen3-CoderBuilt for code, used behind Senpiper Build | 480B / 35B active | ~480 GB at FP8 ~245 GB at Q4 | 8x H100 4x H100 at Q4 | 256k | Apache 2.0 | Vendor scores published on the model card. Qwen3-235B reports 80.6 MMLU-Pro and 69.5 LiveCodeBench. |
| GPT-OSS 120BThe one that fits on a single 80 GB card | 117B | ~63 GB at MXFP4 | 1x H100 | 128k | Apache 2.0 | 0.878 LiveCodeBench. Weak on agentic terminal work at 0.235 Terminal-Bench Hard. |
Memory figures are weights only, computed from the parameter count at the stated precision. Add the key-value cache for your context window and batch size, which the sizing calculator does for you. Benchmark scores are as published by the model vendors or independent harnesses at the date of the last review, and they move. Where a score is not published for a release, we say so rather than borrowing one from a different version.
The harnesses we run, so you can run them too.
Every score in the table above comes from one of these. They are all open source, and on your own box you can run them against your own prompts, which matters more than any public number.
You can pin a version and prove nothing left.
A hosted model can change under you between one procurement cycle and the next. A downloaded model cannot. For a regulated buyer that is usually worth more than the point or two a closed model leads by.