Models
18
Variants
34
sglang 33 · vllm 1
Optimized pairs
14
78% of models
Median uplift
+26.5%
+4.3% to +121%
Best uplift
+121%
mimo-v2.5
Perf reports
14
Hardware
4
H100 14 · any GPU 2 · B300 2 · H800 1

DeepSeek-V4-Flash

deepseek-v4-flash v1.0.0
+51%
vs baseline

DeepSeek-V4-Flash on 8 H100 GPUs. Configs use EAGLE speculative decoding.

  • sglang-tp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node, EAGLE speculative decoding. +51% on the tuning benchmark vs the baseline variant.
  • sglang-tp8-h100-baselinebaseline
    8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.

DeepSeek-V4-Flash-0731

deepseek-v4-flash-0731 v1.0.0
+58%
vs baseline

+58% at 50k + 1.5k tokens · +23.4% at 8k + 1k tokens

DeepSeek-V4-Flash-0731 on 8 H100 GPUs. Configs use DSPARK speculative decoding.

  • sglang-tp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node, DSPARK speculative decoding. Normalized throughput within the SLO vs the baseline variant: +58.0% (50k + 1.5k), +23.4% (8k + 1k).
  • sglang-tp8-h100-baselinebaseline
    8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.

DeepSeek-V4.1-Flash

deepseek-v4.1-flash v1.0.0
+89%
vs baseline

DeepSeek-V4.1-Flash on 8 H100 GPUs.

  • sglang-tp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node. +89% on the tuning benchmark vs the baseline variant.
  • sglang-tp8-h100-baselinebaseline
    8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.

GLM-5.3-Flash

glm-5.3-flash v1.0.0
+5%
vs baseline

GLM-5.3-Flash TP8 on eight H100 GPUs

  • sglang-tp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node. +5% on the tuning benchmark vs the baseline variant.
  • sglang-tp8-h100-baselinebaseline
    8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.

GLM 5.1 (Qwen3.6-35B-A3B)

glm5.1 v1.0.0

Qwen3.6-35B-A3B-793303-glm-5, served as "glm-5". Needs a merged MoE config directory on every node it can land on; SGLANG_MOE_CONFIG_DIR points at it.

  • sglang-tp2default
    2 × GPU · TP2 on two GPUs, single node.

GLM 5.3 (NVFP4)

glm5.3 v1.0.1

GLM-5.3 quantised to NVFP4, served as "glm-5.3" on four B300 GPUs. EAGLE speculative decoding over a 150 GiB hierarchical KV cache. The FP4 kernels are Blackwell-only, so there is no A100 or H100 variant to fall back to.

  • sglang-tp4-b300default
    4 × B300 · TP4 on four B300 GPUs. NVFP4 weights, EAGLE speculative decoding, hierarchical KV cache.

gpt-oss-120b

gpt-oss-120b v1.0.0
+44%
vs baseline

gpt-oss-120b on 8 H100 GPUs. Configs use fp8 KV cache.

  • sglang-tp2-dp4-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP2 DP4 on eight H100 GPUs, single node, fp8 KV cache, EAGLE3 speculative decoding (lmsys/EAGLE3-gpt-oss-120b-bf16). +44% on the tuning benchmark vs the baseline variant.
  • sglang-tp2-h100-baselinebaseline
    2 × H100 · Baseline: TP2 on two H100 GPUs, single node.

Hy3

hy3 v1.0.0
+64.2%
vs baseline

+64.2% at 50k + 1.5k tokens · +7.6% at 8k + 1k tokens

Hy3 on 8 H100 GPUs. Configs use NEXTN speculative decoding.

  • sglang-tp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node, NEXTN speculative decoding. Normalized throughput within the SLO vs the baseline variant: +64.2% (50k + 1.5k), +7.6% (8k + 1k).
  • sglang-tp8-h100-baselinebaseline
    8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.

Kimi K2.5

kimi-k2.5 v1.0.0

Does not fit on one node. A two-pod LeaderWorkerSet group, TP8 inside each node and PP2 across the InfiniBand fabric between them. Needs the LWS controller, rdma-shared-dev-plugin and rdma-injector.

  • sglang-pp2-lwsdefault
    8 × GPU × 2 nodes · TP8 per node, PP2 across two nodes over IB.

Kimi K3

kimi-k3 v1.0.2

Kimi K3 on a single eight-GPU B300 node -- TP8 with decode context parallel across the same eight cards, and an fp8 KV cache. DSPARK speculative decoding needs the Kimi-K3-DSpark draft weights present on every node it can land on, mounted beside the target model.

  • sglang-tp8-b300default
    8 × B300 · TP8 with DCP8 on one eight-GPU B300 node, fp8 KV cache, DSPARK speculative decoding.
  • vllm-tp8-b300
    8 × B300

Laguna-S-2.1

laguna-s-2.1 v1.0.0
+26%
vs baseline

Laguna-S-2.1 on 4 or 8 H100 GPUs. Configs use fp8 KV cache.

  • sglang-tp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node, fp8 KV cache. +26% on the tuning benchmark vs the baseline variant.
  • sglang-tp4-h100-baselinebaseline
    4 × H100 · Baseline: TP4 on four H100 GPUs, single node.

Ling-3.0-flash

ling-3.0-flash v1.0.0
+15.5%
vs baseline

+15.5% at 50k + 1.5k tokens · +21.3% at 8k + 1k tokens

Ling-3.0-flash on 8 H100 GPUs.

  • sglang-tp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node. Normalized throughput within the SLO vs the baseline variant: +15.5% (50k + 1.5k), +21.3% (8k + 1k).
  • sglang-tp8-h100-baselinebaseline
    8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.

Ling-3.0-flash-VL

ling-3.0-flash-vl v1.0.0
+15%
vs baseline

Ling-3.0-flash-VL on 8 H100 GPUs.

  • sglang-tp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node. +15% on the tuning benchmark vs the baseline variant.
  • sglang-tp8-h100-baselinebaseline
    8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.

MiMo-V2.5

mimo-v2.5 v1.0.0
+121%
vs baseline

MiMo-V2.5 on 8 H100 GPUs. Configs use fp8 KV cache, EAGLE speculative decoding.

  • sglang-tp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node, fp8 KV cache, EAGLE speculative decoding. +121% on the tuning benchmark vs the baseline variant.
  • sglang-tp8-h100-baselinebaseline
    8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.

Nemotron-3.5-Lightning-30B-A3B

nemotron-3.5-lightning-30b-a3b v1.0.0
+27%
vs baseline

Nemotron-3.5-Lightning-30B-A3B on 1 or 8 H100 GPUs.

  • sglang-tp1-dp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP1 DP8 on eight H100 GPUs, single node. +27% on the tuning benchmark vs the baseline variant.
  • sglang-tp1-h100-baselinebaseline
    1 × H100 · Baseline: TP1 on one H100 GPU, single node.

Nex-N2.5-Pro

nex-n2.5-pro v1.0.0
+24%
vs baseline

Nex-N2.5-Pro on 8 H100 GPUs. Configs use fp8 KV cache.

  • sglang-tp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node, fp8 KV cache. +24% on the tuning benchmark vs the baseline variant.
  • sglang-tp8-h100-baselinebaseline
    8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.

Qwen3.6 35B-A3B

qwen3.6-35b-a3b v1.1.0
+4.3%
vs baseline

+4.3% at 50k + 1.5k tokens · +10% at 8k + 1k tokens

Stock Qwen3.6-35B-A3B on two H100 or H800 GPUs or eight H100 GPUs, served as "qwen". The upstream weights, not one of the modelforge builds derived from them — those are published separately and served under other names.

  • sglang-tp2-h100default
    2 × H100 / H800 · TP2 on two H100 or H800 GPUs, single node.
  • sglang-tp8-h100-baselinebaseline
    8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.
  • sglang-tp8-h100-optimizedoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node. Normalized throughput within the SLO vs the baseline variant: +4.3% (50k + 1.5k), +10.0% (8k + 1k).

Qwen3.8-Flash-Next

qwen3.8-flash-next v1.0.0
+6%
vs baseline

Qwen3.8-Flash-Next on 8 H100 GPUs.

  • sglang-tp8-h100-optimizeddefaultoptimized
    8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node. +6% on the tuning benchmark vs the baseline variant.
  • sglang-tp8-h100-baselinebaseline
    8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.
Tags Variants (latest) per workload · in + out tokens Links
DeepSeek-V4-Flash
deepseek-v4-flash · v1.0.0
sglang-tp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp8-h100-baselinebaseline8 × H100
+51%
DeepSeek-V4-Flash-0731
deepseek-v4-flash-0731 · v1.0.0
sglang-tp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp8-h100-baselinebaseline8 × H100
50k + 1.5k+58%8k + 1k+23.4%
DeepSeek-V4.1-Flash
deepseek-v4.1-flash · v1.0.0
sglang-tp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp8-h100-baselinebaseline8 × H100
+89%
GLM-5.3-Flash
glm-5.3-flash · v1.0.0
sglang-tp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp8-h100-baselinebaseline8 × H100
+5%
GLM 5.1 (Qwen3.6-35B-A3B)
glm5.1 · v1.0.0
sglang-tp2default2 × GPU
— —
GLM 5.3 (NVFP4)
glm5.3 · v1.0.1
sglang-tp4-b300default4 × B300
— —
gpt-oss-120b
gpt-oss-120b · v1.0.0
sglang-tp2-dp4-h100-optimizeddefaultoptimized8 × H100
sglang-tp2-h100-baselinebaseline2 × H100
+44%
Hy3
hy3 · v1.0.0
sglang-tp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp8-h100-baselinebaseline8 × H100
50k + 1.5k+64.2%8k + 1k+7.6%
Kimi K2.5
kimi-k2.5 · v1.0.0
sglang-pp2-lwsdefault8 × GPU × 2 nodes
— —
Kimi K3
kimi-k3 · v1.0.2
sglang-tp8-b300default8 × B300
vllm-tp8-b3008 × B300
— —
Laguna-S-2.1
laguna-s-2.1 · v1.0.0
sglang-tp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp4-h100-baselinebaseline4 × H100
+26%
Ling-3.0-flash
ling-3.0-flash · v1.0.0
sglang-tp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp8-h100-baselinebaseline8 × H100
50k + 1.5k+15.5%8k + 1k+21.3%
Ling-3.0-flash-VL
ling-3.0-flash-vl · v1.0.0
sglang-tp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp8-h100-baselinebaseline8 × H100
+15%
MiMo-V2.5
mimo-v2.5 · v1.0.0
sglang-tp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp8-h100-baselinebaseline8 × H100
+121%
Nemotron-3.5-Lightning-30B-A3B
nemotron-3.5-lightning-30b-a3b · v1.0.0
sglang-tp1-dp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp1-h100-baselinebaseline1 × H100
+27%
Nex-N2.5-Pro
nex-n2.5-pro · v1.0.0
sglang-tp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp8-h100-baselinebaseline8 × H100
+24%
Qwen3.6 35B-A3B
qwen3.6-35b-a3b · v1.1.0
sglang-tp2-h100default2 × H100 / H800
sglang-tp8-h100-baselinebaseline8 × H100
sglang-tp8-h100-optimizedoptimized8 × H100
50k + 1.5k+4.3%8k + 1k+10%
Qwen3.8-Flash-Next
qwen3.8-flash-next · v1.0.0
sglang-tp8-h100-optimizeddefaultoptimized8 × H100
sglang-tp8-h100-baselinebaseline8 × H100
+6%