AutoSynthData: Generating Training Data for Enterprise Agents
Back to Articles AutoSynthData: Generating Training Data for Enterprise Agents Enterprise Article Published October 2, 2026 Upvote - Esakkivel Esakkiraja esakkivel Follow ServiceNow-AI Shruthan Radhakrishna shruthan-r Follow ServiceNow-AI Denis Akhiyarov dtanow Follow ServiceNow-AI Sagar Davasam davasam Follow ServiceNow-AI Enterprises need agents that work well in their own environments.
The work they ask these agents to do is shaped by the systems they use, the rules they follow, and the state of their data.
A model may be broadly capable and still struggle with a particular environment: a workflow it handles poorly, a combination of tools it misuses, or a constraint it fails to respect.Those are the weaknesses an enterprise needs to improve.The difficulty is turning those weaknesses into training data.
An individual failure tells us something, but training a model requires many new tasks that exercise the same capability in different situations.
Those tasks must also be possible to complete in the environment, resemble work someone would actually request, and have a reliable way to check whether the agent succeeded.At ServiceNow CoreAI, we built AutoSynthData to turn those capability gaps into training data.
It uses a target model’s failures and a stronger teacher’s successes to decide what the model should learn next, then generates and validates new tasks that exercise those capabilities.As the model improves, the curriculum shifts toward what it still finds difficult.
We illustrate the pipeline with EnterpriseOps Gym (Malay et al., 2026), using the released dataset.We begin by describing the environment an agent operates in and what makes a task useful for training.What makes a useful agentic task?
An agentic environment defines the world in which an agent operates: the state it can observe and modify, the tools and APIs it can invoke, and the state transitions produced by its actions.A task is instantiated within this environment.
We use the following abstraction: task = (system specification, user prompt, verifier) System specification The system specification defines the constraints under which the agent operates, including system instructions, environment policies, and, when applicable, task-specific initialization such as a seeded database state or a set of knowledge articles.
The specification must be compatible with the environment’s tools, state, and supported actions.Its instructions should be clear and avoid arbitrary constraints introduced solely to manufacture difficulty.
Agent-facing task The user prompt specifies what the user wants the agent to accomplish, together with any user-level constraints.A generated task should satisfy three properties.Feasibility.
There should exist at least one trajectory in the current environment that satisfies the user prompt while respecting the system specification.This rules out tasks that depend on unavailable tools, inaccessible knowledge, impossible state transitions, or actions prohibited by policy.Realism.
The user prompt should resemble something a user would plausibly ask in the target environment.The space of executable behaviors is usually much larger than the space of realistic workflows.Difficulty.For training, the task should expose a weakness of the current agent.
Tasks that are already solved reliably provide little new training signal.The useful region is therefore tasks that are feasible and realistic, but not yet consistently solved.Verifier The verifier determines whether the resulting trajectory successfully completes the task.
It should satisfy three properties.Consistency.It should agree with the user prompt, the system specification, and the task-specific environment state.Soundness.It should reject trajectories that fail to satisfy the task or violate relevant constraints.Completeness.
It should accept valid solutions rather than encode one particular reference trajectory.These properties matter directly during training.A lax verifier can reward incorrect behavior, while an overly restrictive verifier can penalize valid solutions.
Overview Given an environment and a target model, AutoSynthData generates training tasks consisting of a system specification, user prompt, and verifier.The generated tasks are grounded in the environment and selected to provide useful training signal for the current model.
AutoSynthData first evaluates the target model in the environment using diagnostic tasks and identifies patterns in the tasks it struggles to complete.A stronger teacher helps characterize which of those tasks are solvable and what successful behavior looks like.
AutoSynthData turns the resulting capability gaps into new executable tasks, checks each task in the environment, and uses accepted samples for post-training.Evaluating the updated model reveals which gaps remain and can guide the next round of generation.
From model failures to a curriculum AutoSynthData uses evaluation runs in the target environment to identify what the model needs to learn next.In our EnterpriseOps Gym experiment, we run both the target model and a stronger teacher on the evaluation tasks.
We examine those runs to identify: the capability being tested; the tools and workflow structure involved; where the target model fails and how the teacher succeeds; the properties that a correct final state must satisfy; the dimensions that can vary while preserving the capability being tested.
We distill these findings into sanitized capability specification cards.The evaluation tasks guide what the model should learn, but the generator does not receive their original prompts, entities, trajectories, or verifier details.
It receives the cards and uses them to create new tasks with different prompts, states, and solution paths.Generating and scaling tasks Identifying a capability gap tells us what to teach, but training requires many varied tasks that exercise it.
AutoSynthData uses the specification card to generate those tasks.
Suppose the target model struggles with tasks that require the following workflow: The generator creates new tasks that exercise this workflow, varying the entities, initial environment state, workflow composition, tool combinations, wording, and difficulty.
The stronger teacher then demonstrates a successful trajectory for each task.For supervised fine-tuning (SFT), these demonstrations teach the target model how to apply the capability in new situations.
AutoSynthData builds the dataset in two phases: first generating and validating core samples, then expanding them into novel variants.Target The target phase creates the core set of training samples from the capability specifications.
Workers generate independent tasks in parallel, picking up a new target when they finish.Each candidate goes through validation, execution, solver evaluation, and repair before acceptance.The result is a batch of vetted examples built around what the target model needs to learn.
Multiply The multiply phase expands the dataset by creating novel variants of accepted target samples.Each variant has its own user request, environment state, entity configuration, reference trajectory, and verifier, and must pass the same validation and execution checks.
A multiplied sample cannot seed another multiplied sample.This anchors expansion to the vetted target set and limits drift across generations.Implementation details To support both phases, AutoSynthData separates generation control from environment-specific execution.
A shared controller coordinates generation, quality control, coverage, and dataset construction, while an adapter handles environment execution, task and state management, reference replay, deterministic verification, solver execution, and task profiling.
Together, parallel target generation and multiplication provide a path to training-scale datasets.Their usefulness depends on the checks applied to every candidate: the task must be executable, the solution must work, and the verifier must distinguish success from failure.
High-quality synthetic data needs more than generation Generating a plausible request is not enough to produce useful training data.A task may be impossible in the target environment, its reference solution may fail when executed, or its verifier may reward the wrong final state.
AutoSynthData checks these properties before accepting a task for training.AutoSynthData reviews quality at two levels: individual candidates must pass verification, and batches must provide useful coverage and diversity.
Sample-level verification and repair Each candidate must clear a quality-control loop before entering the training dataset.We begin with solver evaluation to measure difficulty.
In the configuration used here, we favor tasks the target model solves on no more than one of three trials and the stronger solver solves on at least two of three trials.Candidates also undergo positive and negative verification and a bounded repair process.
Positive verification The positive gate asks: Does the intended solution solve the generated task?The pipeline executes the reference trajectory in the target environment and checks the resulting state against the candidate’s verifier.
This reveals mismatches among the prompt, initial state, solution, and success criteria.Negative verification The negative gate asks: Do relevant incorrect outcomes fail?For example, it can mutate parts of the expected outcome and confirm that those states no longer pass verification.
This catches weak verifiers that award success without requiring the intended behavior.Critique and repair Failed candidates go to a critic before being discarded.
The critic examines the sample and its failure, looking for inconsistent state, impossible workflows, incorrect task construction, bad reference trajectories, weak verifier logic, or a mismatch with the intended capability.
The critic’s findings guide repairs, with a fixed limit on retries: candidate ↓ failure ↓ critique / diagnosis ↓ targeted repair ↓ run the gates again ↓ accept or retry A repaired task must pass the relevant checks again.
The diagnosis guides repairs to the existing candidate rather than requiring generation to start over.Passing these checks makes a sample eligible for training, but individually valid samples can still form a repetitive or unbalanced dataset.
AutoSynthData therefore also reviews generation at the batch level.Batch-level review A batch may overrepresent a few easy task families, miss a capability, or reflect too much generation effort spent on a low-yield pattern.
A meta-review examines accepted samples, rejected samples, and generation behavior across each batch.It asks: Which task families are overrepresented, and which capability dimensions are missing?Are the same kinds of examples appearing repeatedly?Do particular targets keep failing generation?
Are systematic problems appearing in critiques?What guidance should change for the next batch?The controller tracks coverage in the accepted dataset, reduces generation in overrepresented regions, and directs more work toward gaps.
When a region repeatedly produces poor candidates, critiques and meta-review guide changes to the generation strategy.These adjustments balance useful learning signal, task quality, coverage, diversity, and low redundancy within the available generation budget and dataset size requirements.
Together, these feedback loops improve both individual tasks and the dataset they form: sample-level checks guide candidate repair, while batch-level review guides future generation.Moving the training frontier The useful training distribution changes as the model improves.
AutoSynthData treats synthetic data generation as a search for tasks near the target model's capability boundary: difficult enough to expose weaknesses, but solvable enough for the teacher to provide reliable demonstrations.After post-training, we evaluate the updated model in the
Related
相關文章

OpenAI 通報逾百家第三方機構,自家智能體存在失控風險
作者:清源 責編:清源 評論: 10 月 2 日消息,當地時間 9 月 30 日晚,OpenAI 披露,旗下 AI 智能體或曾擅自嘗試繞過安全防護,或對百餘家機構的系統造成負面影響。OpenAI 此次披露的信息顯示,已知的 AI 智能體失控行為涉及範圍進一步擴大。

衝刺感恩節前掛牌,消息稱 Anthropic 尋求最早 11 月中旬上市
作者:清源 責編:清源 評論: 10 月 2 日消息,據彭博社今天(2 日)援引知情人士消息稱,在推遲原定計劃後,Anthropic 尋求最早於 11 月中旬上市。知情人士稱,Anthropic 最早可能在 11 月 9 日當週正式啟動 IPO 推介,並爭取在 11 月 26 日感恩節前掛牌交易。

加州檢察長向 OpenAI 發出傳票,調查 AI 網絡安全風險
作者:清源 責編:清源 評論: 10 月 2 日消息,據路透社今天(2 日)凌晨報道,加利福尼亞州總檢察長羅布 · 邦塔辦公室宣佈,邦塔本人已向 OpenAI 發出調查傳票,要求其就 AI 模型涉及的網絡安全事件和風險提供更多信息。邦塔上個月宣佈,加州司法部已就“Hugging Face 事件”正式展開調查。

懸在Muse頭上的劍
字母AI2026.10.01 16:16 · 來自北京全文4181字這麼可愛肯定是來騙我的吧!文 | 字母AI短短一週,Muse已經出現兩次“冒犯主人”的事情了。美國《Inc.》雜誌科技專欄作者傑森說,自己安裝Muse之後,和別人發信息聊天的時候,Muse居然突然彈窗:你們剛才聊的事情很適合做專欄哦。

Meta 否認 AI 智能體 Muse 未經許可讀取用戶私信:必須手動開多層權限才行
作者:遠洋 責編:遠洋 評論: 10 月 1 日消息,針對一名記者提出的指控,即 Meta 的 AI 智能體 Muse 在未經許可的情況下讀取用戶私人消息,Meta 予以反駁。注意到,此前《Inc.》雜誌專欄作家傑森 · 阿滕(Jason Aten)發佈報道詳細敘述了該問題;此後 Meta 負責公關事務的副總裁安迪 · 斯通(Andy Stone)出面回應,明確表示該公司並不認為自家產品在沒有取得用戶許可的前提下讀取過消息。

旗下智能體被指試圖入侵加拿大政府網站,OpenAI 回應稱正審查並已通報
作者:清源 責編:清源 評論: 10 月 1 日消息,據《華盛頓郵報》今天(1 日)上午報道,研究人員披露,AI 智能體曾在沒有收到指令的情況下試圖入侵加拿大政府網站。目前的類似事件不斷增加,許多與 OpenAI 有關的 AI 智能體都曾自行探測或入侵企業及政府的計算機系統。