AutoSynthData: Generating Training Data for Enterprise Agents
Blog post from Hugging Face
AutoSynthData is a ServiceNow CoreAI pipeline for producing enterprise-specific synthetic training data by identifying where an agent fails, using stronger teacher models to characterize successful behavior, and generating new executable tasks that target those capability gaps. Each task includes a system specification, user prompt, and verifier, and is designed to be feasible in the environment, realistic for enterprise workflows, and difficult enough to provide learning value. The system creates core tasks and varied derivatives, validates reference solutions through positive and negative checks, repairs flawed candidates, and reviews batches for diversity, coverage, and redundancy. As models improve through supervised fine-tuning, the process shifts toward remaining weaknesses, creating an adaptive curriculum near the model’s capability boundary. In EnterpriseOps Gym experiments, fine-tuning Gemma-4-26B-A4B-it on roughly 2,000 generated samples improved Hybrid-domain Pass@1 by 7.2 percentage points, from 63.01% to 68.55% verifier success, and raised ITSM Pass@1 from 18.77% to 27.18%, suggesting that validated synthetic tasks can improve performance in stateful enterprise environments.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| AI Model Fine-tuning | 1 | No monthly metrics for this publish month. | |||
| Reinforcement learning | 1 | No monthly metrics for this publish month. | |||
| Voice AI | 1 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.