Droid Shield 2.0: learned secret detection
Blog post from Factory
Factory’s Droid Shield 2.0 augments deterministic secret scanning for autonomous code commits with two fine-tuned language models designed to reduce false positives and catch secrets missed by fixed patterns. The risk model reviews suspicious lines that did not trigger the scanner and prioritizes recall, while the downgrade model evaluates scanner hits using masked context to determine whether they are likely false alarms without exposing credential values. Trained and evaluated primarily on public CredData benchmark material, supplemented by hand-labeled unknown samples and aggregate non-identifying production priors, the models use Qwen 3.6 35B A3B LoRA adapters and were tested with repository-level holdouts. Factory reports that the fine-tuned models achieved higher ROC-AUC scores than the base model and competitive or stronger results than GPT-5.5 and Opus 4.8 under selected operating thresholds, while aiming for lower cost and latency. The company acknowledges limitations involving benchmark representativeness, sparse real-secret prevalence, masked-context constraints, LLM-generated labels, threshold calibration, and comparison methods for closed models, and has released the model weights and supporting materials openly while offering Droid Shield 2.0 in research preview.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Secrets Management | 21 | 2,588 | 483 | 133 | +2% |
| AI Model Fine-tuning | 11 | 975 | 221 | 80 | +28% |
| Serverless | 7 | 775 | 251 | 99 | -24% |
| LLM | 3 | 7,655 | 1,347 | 245 | +22% |
| Data Pipeline | 1 | 530 | 192 | 77 | +1% |
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.