Teaching a 9B model to investigate production alerts
Blog post from Datadog
Datadog describes fine-tuning the smaller Qwen3.5-9B model to perform agentic change attribution for production alerts, aiming to identify deployments, feature flags, or configuration changes that may have contributed to incidents at lower cost than frontier models. Using a data flywheel, the company generated successful investigation traces from GLM-5.3, retained traces matching proxy labels derived from Bits Investigation conclusions, and trained the student model with LoRA; the dataset grew from 100 to 186 internal-incident examples over two cycles. In an evaluation of 326 later production incidents, the fine-tuned model achieved a Recall@5 of 0.55 on internal incidents, or 87% of GLM-5.3’s 0.63 score, while estimated to cost $0.003 per investigation versus $0.06 for the teacher, and it also outperformed both the base model and a heuristics-based ranker. Fine-tuning shifted the model from broad searches toward more focused collection of log, span, and metric evidence, reducing token use relative to the untuned model, though it remained more resource-intensive than fixed rules. The reported metric measures agreement with changes referenced by Bits Investigation rather than independently established causality, and Datadog plans to explore reinforcement learning using production-derived feedback signals while acknowledging that these labels can be imperfect.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 11 | 139 | 28 | 14 | -75% |
| Reinforcement learning | 4 | 17 | 7 | 5 | -82% |
| MCP | 3 | 2,241 | 148 | 72 | -74% |
| LLM | 1 | 747 | 162 | 79 | -85% |
| Observability | 1 | 472 | 102 | 54 | -85% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.