How to Let Claude Code and Codex Deploy Self-Hosted Models Without Losing Control of Cost and Access
Blog post from Qovery
Platform teams can let AI coding agents such as Claude Code and Codex deploy self-hosted open-weight models without granting them cloud credentials by combining four distinct layers: a model-serving system such as vLLM or Ray Serve, LiteLLM as an OpenAI-compatible gateway for virtual keys, budgets, rate limits, routing, and logging, infrastructure provisioning through Terraform, OpenTofu, or an internal platform, and OpenCost or Kubecost for GPU-cost attribution. The recommended security model limits agents to opening pull requests or calling narrowly scoped platform APIs, while CI or the platform retains credentials, applies environment-specific RBAC, records audit trails, and enforces review gates for production. Cost control requires both gateway-level token and request caps and infrastructure-level safeguards including GPU quotas, node-pool size limits, auto-stop or scale-to-zero for non-production services, and daily anomaly alerts. Self-hosting can be economical under sustained, predictable utilization but becomes costly when GPUs are idle, making hosted APIs preferable for bursty workloads or tasks requiring frontier-model quality. A hybrid architecture can route routine, high-volume work to local models while sending complex reasoning to hosted providers through the same gateway. Kubernetes-native containers, manifests, and OpenAI-compatible endpoints support portability across cloud and on-premises GPU environments, though this flexibility requires more operational ownership than managed services.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 17 | No monthly metrics for this publish month. | |||
| Serverless | 10 | No monthly metrics for this publish month. | |||
| Platform Engineering | 5 | No monthly metrics for this publish month. | |||
| AI Agents | 2 | No monthly metrics for this publish month. | |||
| AI Coding Assistant | 2 | No monthly metrics for this publish month. | |||
| Developer Experience | 2 | No monthly metrics for this publish month. | |||
| LLM | 2 | No monthly metrics for this publish month. | |||
| AI Model Fine-tuning | 1 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.