ModelSphere Helm Charts

Helm charts for running large language models on Kubernetes.

helm repo add modelsphere https://modelsphere.github.io/helm-charts
helm repo update
helm search repo modelsphere
ChartVersionApp versionDescription
modelsphere/autoconfig0.4.00.4.0k8s 后端发现 → 自动同步 openresty peers / cache-aware-router workers(ModelRoute controller)
modelsphere/bodylog0.1.130.1.13Body-log listener for the OpenResty router: receives full request/response frames over TCP, writes them as hourly JSONL, and serves rollups over HTTP
modelsphere/bodylog-exporter0.3.10.2.7Turns the body-log listener's detail records into Prometheus metrics, and polls the router's control-plane endpoints for live routing state
modelsphere/cart0.2.2v0.6.4Cache-aware router: routes each request to the replica that already holds the longest matching prompt prefix, while keeping load balanced
modelsphere/llm-slo-decision-gen0.3.40.7.0LLM and async-job SLO stack: decision-gen (recommendation service), optional slo-api (SLO storage HTTP/UI), and the LLMSLORequirement and JobSLORequirement CRDs under inference.modelsphere.dev. CRD ownership: this chart owns llmslorequirements.inference.modelsphere.dev and jobslorequirements.inference.modelsphere.dev (see crds/). Both CRDs carry helm.sh/resource-policy: keep so uninstall leaves them in the cluster — existing LLMSLORequirement / JobSLORequirement objects are not deleted with the release. Do not install a second chart that also ships these CRDs into the same cluster. LLMScaler CRD and the operator that reconciles it are provided by a separate chart (llmscaleoperator / llmscaleoperator-system). Deploy that one first if you need autoscaling, not just SLO storage.
modelsphere/llmscaleoperator0.3.00.4.0Kubernetes operator that autoscales LLM inference workloads on LLM-specific signals -- KV-cache utilization, queue depth, TPM/capacity load -- rather than CPU and memory. An HPA specialized for token serving. Ships the LLMScaler CRD (autoscaling.modelsphere.dev/v1alpha1) and the controller that reconciles it. The SLO storage side lives in the llm-slo-decision-gen chart; the two are independent, and this one is what you need for autoscaling alone. Generated from the llm-operator repo (dist/chart, kubebuilder helm plugin) and maintained here, because this is where charts are published from.
modelsphere/openresty0.1.200.1.20Session-affinity router for LLM inference backends: pins a conversation to the backend that already holds its prefix cache, with active health checking, adaptive concurrency and full body logging
modelsphere/rdma-injector0.2.30.2.0Mutating admission webhook that spares workload manifests the RDMA boilerplate: pods labelled rdma-ib:"true" get the /etc/gpu-node mount and an NCCL_IB_HCA sourced from the node's healthy InfiniBand ports injected automatically. The webhook code is baked into the image; its serving certificate is self-signed by the chart at install time, so it installs into any namespace.
modelsphere/sglang0.8.61.0.0SGLang inference deployment -- single-node or multi-node (LeaderWorkerSet), with optional cache-aware routing, autoscaling and hang detection
modelsphere/vllm0.7.31.0.0vLLM inference deployment -- single-node or multi-node (LeaderWorkerSet), with optional cache-aware routing, autoscaling and hang detection