DeepSeek-V4-Flash on 8 H100 GPUs. Configs use EAGLE speculative decoding.
sglang-tp8-h100-optimizeddefaultoptimized
8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node, EAGLE speculative decoding. +51% on the tuning benchmark vs the baseline variant.
sglang-tp8-h100-baselinebaseline
8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.
+58% at 50k + 1.5k tokens·+23.4% at 8k + 1k tokens
DeepSeek-V4-Flash-0731 on 8 H100 GPUs. Configs use DSPARK speculative decoding.
sglang-tp8-h100-optimizeddefaultoptimized
8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node, DSPARK speculative decoding. Normalized throughput within the SLO vs the baseline variant: +58.0% (50k + 1.5k), +23.4% (8k + 1k).
sglang-tp8-h100-baselinebaseline
8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.
Qwen3.6-35B-A3B-793303-glm-5, served as "glm-5". Needs a merged MoE config directory on every node it can land on; SGLANG_MOE_CONFIG_DIR points at it.
sglang-tp2default
2 × GPU · TP2 on two GPUs, single node.
GLM 5.3 (NVFP4)
glm5.3v1.0.1
GLM-5.3 quantised to NVFP4, served as "glm-5.3" on four B300 GPUs. EAGLE speculative decoding over a 150 GiB hierarchical KV cache. The FP4 kernels are Blackwell-only, so there is no A100 or H100 variant to fall back to.
sglang-tp4-b300default
4 × B300 · TP4 on four B300 GPUs. NVFP4 weights, EAGLE speculative decoding, hierarchical KV cache.
gpt-oss-120b
gpt-oss-120bv1.0.0
+44%
vs baseline
gpt-oss-120b on 8 H100 GPUs. Configs use fp8 KV cache.
sglang-tp2-dp4-h100-optimizeddefaultoptimized
8 × H100 · Tuned by LLM AutoTune: TP2 DP4 on eight H100 GPUs, single node, fp8 KV cache, EAGLE3 speculative decoding (lmsys/EAGLE3-gpt-oss-120b-bf16). +44% on the tuning benchmark vs the baseline variant.
sglang-tp2-h100-baselinebaseline
2 × H100 · Baseline: TP2 on two H100 GPUs, single node.
+64.2% at 50k + 1.5k tokens·+7.6% at 8k + 1k tokens
Hy3 on 8 H100 GPUs. Configs use NEXTN speculative decoding.
sglang-tp8-h100-optimizeddefaultoptimized
8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node, NEXTN speculative decoding. Normalized throughput within the SLO vs the baseline variant: +64.2% (50k + 1.5k), +7.6% (8k + 1k).
sglang-tp8-h100-baselinebaseline
8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.
Does not fit on one node. A two-pod LeaderWorkerSet group, TP8 inside each node and PP2 across the InfiniBand fabric between them. Needs the LWS controller, rdma-shared-dev-plugin and rdma-injector.
sglang-pp2-lwsdefault
8 × GPU × 2 nodes · TP8 per node, PP2 across two nodes over IB.
Kimi K3
kimi-k3v1.0.2
Kimi K3 on a single eight-GPU B300 node -- TP8 with decode context parallel across the same eight cards, and an fp8 KV cache. DSPARK speculative decoding needs the Kimi-K3-DSpark draft weights present on every node it can land on, mounted beside the target model.
sglang-tp8-b300default
8 × B300 · TP8 with DCP8 on one eight-GPU B300 node, fp8 KV cache, DSPARK speculative decoding.
vllm-tp8-b300
8 × B300
Laguna-S-2.1
laguna-s-2.1v1.0.0
+26%
vs baseline
Laguna-S-2.1 on 4 or 8 H100 GPUs. Configs use fp8 KV cache.
sglang-tp8-h100-optimizeddefaultoptimized
8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node, fp8 KV cache. +26% on the tuning benchmark vs the baseline variant.
sglang-tp4-h100-baselinebaseline
4 × H100 · Baseline: TP4 on four H100 GPUs, single node.
+15.5% at 50k + 1.5k tokens·+21.3% at 8k + 1k tokens
Ling-3.0-flash on 8 H100 GPUs.
sglang-tp8-h100-optimizeddefaultoptimized
8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node. Normalized throughput within the SLO vs the baseline variant: +15.5% (50k + 1.5k), +21.3% (8k + 1k).
sglang-tp8-h100-baselinebaseline
8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.
MiMo-V2.5 on 8 H100 GPUs. Configs use fp8 KV cache, EAGLE speculative decoding.
sglang-tp8-h100-optimizeddefaultoptimized
8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node, fp8 KV cache, EAGLE speculative decoding. +121% on the tuning benchmark vs the baseline variant.
sglang-tp8-h100-baselinebaseline
8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.
Stock Qwen3.6-35B-A3B on two H100 or H800 GPUs or eight H100 GPUs, served as "qwen". The upstream weights, not one of the modelforge builds derived from them — those are published separately and served under other names.
sglang-tp2-h100default
2 × H100 / H800 · TP2 on two H100 or H800 GPUs, single node.
sglang-tp8-h100-baselinebaseline
8 × H100 · Baseline: TP8 on eight H100 GPUs, single node.
sglang-tp8-h100-optimizedoptimized
8 × H100 · Tuned by LLM AutoTune: TP8 on eight H100 GPUs, single node. Normalized throughput within the SLO vs the baseline variant: +4.3% (50k + 1.5k), +10.0% (8k + 1k).