Running AI workloads on Kubernetes: GPU scheduling and cost without the pain
By VA2PT Team, . 9 min read

GPUs on Kubernetes fail in ways CPUs never do. A pod sits Pending for an hour because the scheduler cannot see the GPU. A 7B model that runs fine on a laptop gets OOM-killed on a 24 GB card because nobody set the memory fraction. The monthly bill triples because a data scientist's notebook held an A100 over a long weekend. None of these are Kubernetes problems; they are the result of treating a GPU like a CPU with a different name.
This article walks through the setup we use for teams running inference and batch AI jobs on EKS, AKS and GKE, with the cost maths that decides most of the choices.
The scenario
A Hyderabad healthtech runs two AI workloads on an EKS cluster in ap-south-1: a real-time report summariser (an 8B parameter model served with vLLM, roughly 2 requests per second at peak, latency budget 2 seconds) and a nightly batch that runs OCR and classification over 300,000 scanned pages. They started with one on-demand g5.2xlarge node running everything and a bill nobody could explain.
Step 1: a dedicated GPU node pool with taints
Never mix GPU and CPU workloads on the same node. Give GPU nodes a taint so only pods that ask for a GPU land there, and a label so schedulers and autoscalers can target them.
# EKS managed node group (eksctl)
managedNodeGroups:
- name: gpu-l4
instanceTypes: ["g6.xlarge"] # NVIDIA L4, 24 GB
amiFamily: AmazonLinux2023
minSize: 0
maxSize: 6
labels:
workload: gpu
nvidia.com/gpu.product: L4
taints:
- key: nvidia.com/gpu
value: "true"
effect: NoSchedule
The equivalents: on AKS, az aks nodepool add --node-taints sku=gpu:NoSchedule --node-vm-size Standard_NC8as_T4_v3; on GKE, gcloud container node-pools create ... --accelerator type=nvidia-l4,count=1 (GKE adds the nvidia.com/gpu taint automatically).
minSize: 0 matters. A GPU node that exists but runs nothing costs the same as one that is busy.
Step 2: the NVIDIA device plugin, and what "1 GPU" means
The device plugin advertises nvidia.com/gpu as a schedulable resource. Install it through the NVIDIA GPU Operator (which also manages drivers, the container toolkit and DCGM) or, on EKS with the GPU-optimised AMI, as a DaemonSet.
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm install gpu-operator nvidia/gpu-operator -n gpu-operator --create-namespace \
--set driver.enabled=false # EKS GPU AMI ships the driver; set true on plain images
A pod then requests a whole GPU:
resources:
limits:
nvidia.com/gpu: 1
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
nodeSelector:
workload: gpu
Note what is missing: there is no requests.nvidia.com/gpu distinct from limits. GPUs are not overcommittable by default; one pod, one card. That is why a single notebook can pin a card while using 3% of it.
Step 3: sharing a GPU with time-slicing or MIG
Two ways to run more than one pod per card:
- Time-slicing lets the device plugin advertise, say, 4 replicas of each GPU. Pods share the card by context switching. There is no memory isolation, so a misbehaving pod can OOM its neighbours. Good for notebooks, dev inference and small models.
- MIG (Multi-Instance GPU) partitions supported cards (A100, H100) into isolated slices with their own memory. Isolation is real; the cost is that slices are fixed sizes and only certain cards support it.
Time-slicing config for the device plugin:
apiVersion: v1
kind: ConfigMap
metadata:
name: time-slicing-config
namespace: gpu-operator
data:
any: |-
version: v1
sharing:
timeSlicing:
resources:
- name: nvidia.com/gpu
replicas: 4
Apply it to the ClusterPolicy (devicePlugin.config.name: time-slicing-config) and each L4 now shows nvidia.com/gpu: 4. For the healthtech, the dev namespace runs on time-sliced L4s; the production summariser gets whole cards.
Rule of thumb: production inference gets a dedicated GPU or a MIG slice; everything else can time-slice.
Step 4: autoscaling GPUs with Karpenter
The cluster autoscaler works with GPU node groups but is slow and rigid. Karpenter on EKS provisions the right instance for a pending pod in about a minute and can mix on-demand and spot.
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: gpu-inference
spec:
template:
metadata:
labels: { workload: gpu }
spec:
taints:
- key: nvidia.com/gpu
value: "true"
effect: NoSchedule
requirements:
- key: karpenter.k8s.aws/instance-family
operator: In
values: ["g6", "g5"]
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand"]
nodeClassRef:
group: karpenter.k8s.aws
kind: EC2NodeClass
name: gpu
limits:
nvidia.com/gpu: 8
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 5m
Create a second NodePool named gpu-batch with capacity-type: ["spot"] and a higher limit; the nightly OCR job targets it with a node selector on karpenter.sh/nodepool: gpu-batch. On AKS the analogue is the cluster autoscaler with a spot node pool; on GKE it is node auto-provisioning with --spot.
consolidateAfter: 5m is the setting that deletes idle GPU nodes. Check it exists. This is where the healthtech's bill was going.
Step 5: serving with vLLM (and when to use KServe)
vLLM is the practical default for serving open-weights LLMs: continuous batching and paged attention mean one L4 serves far more concurrent requests than a naive implementation. A production deployment:
apiVersion: apps/v1
kind: Deployment
metadata:
name: summariser
spec:
replicas: 2
template:
spec:
nodeSelector: { workload: gpu }
tolerations:
- { key: nvidia.com/gpu, operator: Exists, effect: NoSchedule }
containers:
- name: vllm
image: vllm/vllm-openai:latest # pin a version in production
args:
- --model=meta-llama/Llama-3.1-8B-Instruct
- --max-model-len=8192
- --gpu-memory-utilization=0.90
- --max-num-seqs=64
- --dtype=bfloat16
ports: [{ containerPort: 8000 }]
resources:
limits: { nvidia.com/gpu: 1, memory: 24Gi }
requests: { cpu: "4", memory: 16Gi }
readinessProbe:
httpGet: { path: /health, port: 8000 }
initialDelaySeconds: 60
volumeMounts:
- { name: hf-cache, mountPath: /root/.cache/huggingface }
volumes:
- name: hf-cache
persistentVolumeClaim: { claimName: hf-cache }
Three settings decide whether this works:
--max-model-lencaps context; an 8B model in bf16 needs about 16 GB for weights, leaving roughly 6 GB of a 24 GB L4 for KV cache. Longer context means fewer concurrent sequences.--gpu-memory-utilization=0.90tells vLLM how much of the card to claim. Leave headroom or the first burst OOMs.- The
hf-cachevolume is a shared PVC (EFS on EKS) so new replicas do not download 16 GB of weights from the internet on every scale-up. Without it, scale-up takes eight minutes instead of ninety seconds.
Use KServe when you need model versioning, canary traffic splitting and a standard inference protocol across many models; it can run vLLM underneath. For one or two models, a plain Deployment plus an HPA on a custom metric (queue depth from vLLM's /metrics) is simpler.
Step 6: queueing batch jobs with Kueue
The nightly OCR batch is 300,000 pages, run as 600 Jobs of 500 pages. Without a queue, all 600 pods go Pending, Karpenter tries to provision 600 GPU nodes, hits the limit, and the scheduler thrashes. Kueue gives you quotas and fair ordering.
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
name: gpu-batch
spec:
namespaceSelector: {}
resourceGroups:
- coveredResources: ["nvidia.com/gpu"]
flavors:
- name: spot-l4
resources:
- name: nvidia.com/gpu
nominalQuota: 6
---
apiVersion: kueue.x-k8s.io/v1beta1
kind: LocalQueue
metadata: { name: ocr, namespace: batch }
spec: { clusterQueue: gpu-batch }
Jobs carry the label kueue.x-k8s.io/queue-name: ocr and start suspended: true; Kueue admits six at a time as quota frees up. The overnight run finishes in a predictable window and never starves the production summariser.
Step 7: observability with DCGM
CPU utilisation tells you nothing about a GPU. The DCGM exporter (installed by the GPU Operator) exposes what matters:
| Metric | What it tells you | Alert when |
|---|---|---|
DCGM_FI_DEV_GPU_UTIL |
SM busy % | < 20% for 30 min on a production card (wasted money) |
DCGM_FI_DEV_FB_USED |
memory used | > 95% of FB_TOTAL (OOM imminent) |
DCGM_FI_DEV_GPU_TEMP |
temperature | > 85 °C sustained (throttling) |
DCGM_FI_DEV_XID_ERRORS |
driver errors | any non-zero (hardware or driver fault) |
Add vLLM's own metrics (vllm:num_requests_waiting, vllm:e2e_request_latency_seconds) and scale replicas on waiting requests, not on CPU.
The cost maths
Whether any of this is worth it comes down to cost per 1,000 tokens, so compute it.
An L4 instance (g6.xlarge) in Mumbai costs on the order of a dollar an hour on demand; check the current price. Suppose vLLM sustains 1,500 output tokens per second on the 8B model at your batch size (measure this; it varies widely with prompt length). That is 5.4 million tokens per hour, so the compute cost is well under a cent per 1,000 output tokens at full utilisation.
The catch is utilisation. At the healthtech's 2 requests per second averaging 300 output tokens, the card produces 600 tokens per second, 40% of its capacity, so the real cost per 1,000 tokens is 2.5 times the ideal figure. Still low, but the lesson is that self-hosting pays for itself only when you keep the cards busy. Three ways to raise utilisation: share the card with time-slicing for non-critical traffic, route batch work onto idle production cards at night, and scale to zero when there is genuinely nothing to do.
Compare that with a managed API's per-token price for a similar-sized model, and the crossover typically lands somewhere between one and a few hundred thousand requests a day. Below it, the API is cheaper once you count the engineer's time.
Spot GPUs and checkpointing
Spot GPU instances are often 60 to 70% cheaper but can be reclaimed with two minutes' notice. They are ideal for the batch job and wrong for the production summariser. Make batch work survive interruption:
- Split work into small units (500 pages per Job) so a lost pod loses minutes, not hours.
- Write progress to object storage after each unit, and make the Job idempotent so a restart skips completed units.
- Use
backoffLimitandpodFailurePolicyso a spot reclaim (SIGTERM) triggers a retry, not a failure. - Watch the AWS
spot-interruptionevents (Karpenter handles the node drain) and keep an on-demand fallback NodePool with a low weight for the last 5% of the batch.
For training or fine-tuning jobs, checkpoint to S3 every N steps and resume from the last checkpoint; frameworks such as PyTorch Lightning and Hugging Face Trainer support this with a few lines.
Common mistakes
- No taint on GPU nodes, so CPU pods land there and block GPU pods with memory requests.
- One shared GPU node running notebooks, dev and prod. Separate node pools, always.
- Downloading model weights on every pod start. Use a shared cache volume or bake them into the image.
- Scaling inference on CPU utilisation. Scale on queue depth or GPU utilisation.
- Leaving
consolidateAfterunset, so idle GPU nodes live forever. - Running production inference on spot.
What to do this week
- Add a taint and label to your GPU nodes and set the node group minimum to zero.
- Install the GPU Operator with DCGM and put GPU utilisation on a dashboard. Look at it for a week before changing anything.
- Move dev and notebook workloads to a time-sliced pool.
- Put a shared weights cache in front of your serving deployment and measure scale-up time.
- Run the cost-per-1,000-tokens calculation with your real throughput numbers and decide, per workload, whether self-hosting wins.
Most GPU bills are not high because GPUs are expensive. They are high because GPUs are idle.
- kubernetes
- gpu
- ai-infrastructure
- vllm
- karpenter
- eks
- finops