Operators make life easier, but they are not always an option. In this post, I'll walk through a practical way to deploy large language models on OpenShift without relying on the OpenShift AI or NVIDIA operators. The approach uses llama.cpp as a lightweight runtime engine and runs a quantized GGUF model to enable efficient inference with minimal dependencies.
https://medium.com/@ahmeddraz/deploy-llm-models-on-openshift-84ecb014f09a