How to Self-Host Mistral Large 4 for Claude Code and Codex: GPU Sizing, vLLM, and the Real Cost
Blog post from Qovery
Mistral Large 4 is presented as a roughly 1.05-trillion-parameter mixture-of-experts model with 49 billion active parameters per token, designed for frontier coding and multimodal use but, as of October 2026, not yet downloadable and without published final license terms for commercial self-hosting. Its full weights require substantial VRAM despite sparse activation, estimated at about 2.1 TB in bf16, 1.05 TB in FP8, or 525 GB at 4-bit quantization, making practical deployments dependent on multi-GPU H100 or H200 nodes and careful allowance for long-context KV cache memory. The recommended architecture uses vLLM or, for comparison, SGLang as the model-serving engine, with a LiteLLM or Portkey gateway providing authentication, budgets, logging, fallback routing, and protocol translation for Claude Code, while Codex can use an OpenAI-compatible gateway directly. The analysis argues that self-hosting is primarily justified by data residency, version control, rate-limit independence, and sufficiently steady demand rather than by cost alone, since an eight-H100 node can cost about $40,000 monthly when run continuously and hosted APIs generally remain cheaper at low utilization. It emphasizes scale-to-zero, discounted or spot capacity, quantization, cached weights, private networking, monitoring, and hybrid routing that sends routine coding tasks to self-hosted models while reserving difficult tasks for hosted frontier services. Different infrastructure tools are assigned distinct roles, including SkyPilot for capacity sourcing, KubeAI or KEDA for autoscaling, vLLM for inference, LiteLLM or Portkey for gateway governance, and Qovery or similar platforms for Kubernetes operations, access controls, and deployment workflows.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 15 | No monthly metrics for this publish month. | |||
| LLM | 6 | No monthly metrics for this publish month. | |||
| Observability | 3 | No monthly metrics for this publish month. | |||
| Developer Experience | 2 | No monthly metrics for this publish month. | |||
| Platform Engineering | 2 | No monthly metrics for this publish month. | |||
| Real-time | 2 | No monthly metrics for this publish month. | |||
| AI Coding Assistant | 1 | No monthly metrics for this publish month. | |||
| Harness engineering | 1 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.