ModelSphere Helm Charts
Helm charts for running large language models on Kubernetes.
helm repo add modelsphere https://modelsphere.github.io/helm-charts
helm repo update
helm search repo modelsphere
| Chart | Version | App version | Description |
|---|---|---|---|
modelsphere/autoconfig | 0.4.0 | 0.4.0 | k8s 后端发现 → 自动同步 openresty peers / cache-aware-router workers(ModelRoute controller) |
modelsphere/bodylog | 0.1.13 | 0.1.13 | Body-log listener for the OpenResty router: receives full request/response frames over TCP, writes them as hourly JSONL, and serves rollups over HTTP |
modelsphere/bodylog-exporter | 0.3.1 | 0.2.7 | Turns the body-log listener's detail records into Prometheus metrics, and polls the router's control-plane endpoints for live routing state |
modelsphere/cart | 0.2.2 | v0.6.4 | Cache-aware router: routes each request to the replica that already holds the longest matching prompt prefix, while keeping load balanced |
modelsphere/llm-slo-decision-gen | 0.3.4 | 0.7.0 | LLM and async-job SLO stack: decision-gen (recommendation service), optional slo-api (SLO storage HTTP/UI), and the LLMSLORequirement and JobSLORequirement CRDs under inference.modelsphere.dev. CRD ownership: this chart owns llmslorequirements.inference.modelsphere.dev and jobslorequirements.inference.modelsphere.dev (see crds/). Both CRDs carry helm.sh/resource-policy: keep so uninstall leaves them in the cluster — existing LLMSLORequirement / JobSLORequirement objects are not deleted with the release. Do not install a second chart that also ships these CRDs into the same cluster. LLMScaler CRD and the operator that reconciles it are provided by a separate chart (llmscaleoperator / llmscaleoperator-system). Deploy that one first if you need autoscaling, not just SLO storage. |
modelsphere/llmscaleoperator | 0.3.0 | 0.4.0 | Kubernetes operator that autoscales LLM inference workloads on LLM-specific signals -- KV-cache utilization, queue depth, TPM/capacity load -- rather than CPU and memory. An HPA specialized for token serving. Ships the LLMScaler CRD (autoscaling.modelsphere.dev/v1alpha1) and the controller that reconciles it. The SLO storage side lives in the llm-slo-decision-gen chart; the two are independent, and this one is what you need for autoscaling alone. Generated from the llm-operator repo (dist/chart, kubebuilder helm plugin) and maintained here, because this is where charts are published from. |
modelsphere/openresty | 0.1.20 | 0.1.20 | Session-affinity router for LLM inference backends: pins a conversation to the backend that already holds its prefix cache, with active health checking, adaptive concurrency and full body logging |
modelsphere/rdma-injector | 0.2.3 | 0.2.0 | Mutating admission webhook that spares workload manifests the RDMA boilerplate: pods labelled rdma-ib:"true" get the /etc/gpu-node mount and an NCCL_IB_HCA sourced from the node's healthy InfiniBand ports injected automatically. The webhook code is baked into the image; its serving certificate is self-signed by the chart at install time, so it installs into any namespace. |
modelsphere/sglang | 0.8.6 | 1.0.0 | SGLang inference deployment -- single-node or multi-node (LeaderWorkerSet), with optional cache-aware routing, autoscaling and hang detection |
modelsphere/vllm | 0.7.3 | 1.0.0 | vLLM inference deployment -- single-node or multi-node (LeaderWorkerSet), with optional cache-aware routing, autoscaling and hang detection |