Running AI workloads on Kubernetes? The scheduling part works fine.
It's the GPU utilization and artifact management that breaks everything.
Typical inference workload uses 20-40% of GPU compute and maybe 30 GiB memory. But Kubernetes Device Plugin only reports integer device counts. So nvidia/gpu: 1 means you get the entire H100, and every other workload waits.
NVIDIA Time-Slicing lets multiple Pods share, but there's no VRAM isolation. One Pod OOMs, everything crashes.
MIG works but only on A100/H100, and you're stuck with fixed partition sizes.
Then there's the artifact problem.
Which model version is deployed? Does config.json match the weights? Can this move through CI/CD without manual file transfers?
I came across
@HAMiProject +
@Kit_Ops tutorial, and the pattern actually makes sense.
HAMi handles GPU virtualization through CUDA API interception. No driver changes, no app changes. It exposes nvidia/gpu, nvidia/gpumem, and nvidia/gpucores as schedulable resources.
One H100 becomes 10 vGPUs. Your Pod requests 30 GiB, gets scheduled, and can only see that allocation. Over-allocation returns OOM without crashing other workloads.
KitOps handles artifact packaging. Models, datasets, configs, code → stored as OCI artifacts in the same registry you use for containers.
The ModelKit gets pulled by an initContainer, unpacked into a flat directory, and the inference engine loads it from local disk. Registry-native supply chain, version controlled, reproducible.
The lab workflow:
→ Pull Qwen3-4B-Instruct ModelKit from Jozu Hub
→ Unpack with kitunpacker initContainer
→ Schedule Pod with HAMi GPU shares (30 GiB memory)
→ Serve with SGLang
→ Optionally co-locate vLLM on same physical GPU
Everything moves through standard K8s patterns. No one-off scripts, no S3 bucket coordination, no guessing which version is actually running.
Useful if you're running LLMs in production and trying to avoid building custom deployment infrastructure.