ServiceNow's AutoSynthData turns agent failures into new training tasks
ServiceNow's CoreAI team has described AutoSynthData, a pipeline for generating training data for AI agents that work inside enterprise systems, in a post on the Hugging Face blog. The idea is to train a model on exactly the kinds of tasks it currently gets wrong, rather than on generic examples.
The pipeline starts by running the target model and a stronger teacher model on diagnostic tasks in the environment. Where the target fails and the teacher succeeds, the team distils the pattern into a "capability specification card". The generator never sees the original evaluation prompts or trajectories, only the cards, and writes new tasks with different prompts, data and solution paths that exercise the same capability. The teacher then demonstrates a successful run of each task for supervised fine-tuning.
Most of the post is about quality control. Every generated task has a system specification, a user prompt and a verifier, and the authors set three requirements for the task (feasible, realistic, still difficult for the model) and three for the verifier (consistent, sound, complete). Each candidate passes a positive gate, where the reference solution is executed and must pass, and a negative gate, where parts of the expected outcome are mutated and must then fail, catching verifiers that award success too easily. Failed candidates go to a critic that proposes repairs with a fixed retry limit, and a batch-level review rebalances task families that are over-represented.
The reported results come from EnterpriseOps Gym. In its Hybrid domain, with Gemma-4-26B-A4B-it as the target and Qwen3.8-27B as the teacher, the pipeline generated 2,000 samples in about 18 hours; fine-tuning raised mean Pass@1 by 7.2 percentage points, a 35% relative gain, and verifier success from 63.01% to 68.55%, closing 59% of the gap to the reference model. In the ITSM domain, Pass@1 rose from 18.77% to 27.18%. The authors note these are results in the environment used for the experiments, and that they plan to test the approach with reinforcement learning.

Why it matters
The expensive part of training agents for a specific company is not the model but the tasks and the checks. The negative gate is the transferable idea: a verifier that cannot fail on a wrong answer silently teaches the model to be wrong.