Home / Companies / Qovery / Blog / Post Details
Content Deep Dive

How to Let Claude Code and Codex Deploy Self-Hosted Models Without Losing Control of Cost and Access

Blog post from Qovery

Post Details
Company
Date Published
Author
-
Word Count
3,622
Company Posts That Month
43
Language
English
Hacker News Points
-
Post removed?
No
Summary

Platform teams can let AI coding agents such as Claude Code and Codex deploy self-hosted open-weight models without granting them cloud credentials by combining four distinct layers: a model-serving system such as vLLM or Ray Serve, LiteLLM as an OpenAI-compatible gateway for virtual keys, budgets, rate limits, routing, and logging, infrastructure provisioning through Terraform, OpenTofu, or an internal platform, and OpenCost or Kubecost for GPU-cost attribution. The recommended security model limits agents to opening pull requests or calling narrowly scoped platform APIs, while CI or the platform retains credentials, applies environment-specific RBAC, records audit trails, and enforces review gates for production. Cost control requires both gateway-level token and request caps and infrastructure-level safeguards including GPU quotas, node-pool size limits, auto-stop or scale-to-zero for non-production services, and daily anomaly alerts. Self-hosting can be economical under sustained, predictable utilization but becomes costly when GPUs are idle, making hosted APIs preferable for bursty workloads or tasks requiring frontier-model quality. A hybrid architecture can route routine, high-volume work to local models while sending complex reasoning to hosted providers through the same gateway. Kubernetes-native containers, manifests, and OpenAI-compatible endpoints support portability across cloud and on-premises GPU environments, though this flexibility requires more operational ownership than managed services.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Kubernetes 17 No monthly metrics for this publish month.
Serverless 10 No monthly metrics for this publish month.
Platform Engineering 5 No monthly metrics for this publish month.
AI Agents 2 No monthly metrics for this publish month.
AI Coding Assistant 2 No monthly metrics for this publish month.
Developer Experience 2 No monthly metrics for this publish month.
LLM 2 No monthly metrics for this publish month.
AI Model Fine-tuning 1 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.