Letting Claude Code and Codex Deploy Open Source Models In-House: The Stack That Keeps Costs and Access Under Control
Blog post from Qovery
Platform teams seeking to let coding agents such as Claude Code and Codex deploy self-hosted open-source models are advised to use a layered architecture rather than a single product: vLLM for high-throughput production inference, Kubernetes or KServe for deployment and scaling, LiteLLM for model routing, virtual keys, budgets, and rate limits, and an internal developer platform or policy-controlled infrastructure tooling for audited deployment automation. The central recommendation is to give agents a narrowly scoped deployment interface instead of broad cloud credentials, using environment-specific RBAC, pull-request workflows, OIDC-based ephemeral credentials, and deployment logs to limit risk. GPU costs should be controlled through both infrastructure measures, including scale-to-zero node pools, right-sizing, spot capacity, and automatic shutdown of nonproduction environments, and gateway-level budget caps and rate limits. Self-hosting can be economical for predictable, sustained, high-volume workloads or where data residency and private-network requirements matter, but hosted APIs are generally cheaper for intermittent or low-volume use, making a hybrid approach practical for many organizations.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 16 | No monthly metrics for this publish month. | |||
| Platform Engineering | 6 | No monthly metrics for this publish month. | |||
| AI Agents | 1 | No monthly metrics for this publish month. | |||
| AI Coding Assistant | 1 | No monthly metrics for this publish month. | |||
| Developer Experience | 1 | No monthly metrics for this publish month. | |||
| LLM | 1 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.