An agent can complete a task while skipping checks, mishandling a failure, or making an unsupported claim. Developers need tests for the specific behaviors they encounter in deployment.
An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-specified undesirable behaviors. For efficient construction, Anchor-and-Confirm combines programmable retrieval with candidate-level confirmation by a Flash large language model (LLM), while the Anchor Synthesis Loop generates and revises specifications for custom behaviors. The benchmarks use decision-point continuation to evaluate an LLM's next turn at a recorded decision point with a behavior-specific rubric, without a reference answer or environment replay. Experiments in coding and general tool use draw on 252,557 sessions and produce 107 benchmarks with 4,125 instances, fulfilling 95.3% of build-target requests. Both human annotators confirm the requested behavior in 84% of sampled instances, and the automated grader's agreement with human pass/fail judgments is comparable to that between the annotators. Nine frontier LLMs achieve a mean pass rate of only 26.7%, showing that they still struggle to respond appropriately at the evaluated decision points. Analysis across behavior-specific benchmarks further reveals weaknesses in how current LLMs behave as agents. By turning deployment problems into targeted benchmarks, TraceDance could serve as a key component of the recursive self-improvement (RSI) loop.
01 / METHOD
A deployment problem. A targeted benchmark.
Describe an undesirable behavior in natural language. TraceDance finds relevant evidence and turns it into tests of an agent’s next move.
01
Specify the behavior
Match an existing specification or use the Anchor Synthesis Loop to generate and revise a custom one.
02
Anchor and confirm
Programmable anchors retrieve candidates efficiently. A Fast Model checks whether each candidate exhibits the requested behavior.
03
Evaluate the next turn
Cut the trace at a relevant decision point. Give each model the same recorded prefix and grade its new turn with a behavior-specific rubric.
TraceDance’s construction and evaluation pipeline. Select the figure to enlarge it.
DECISION-POINT CONTINUATION
The next move is the test.
Each model produces exactly one assistant turn, which may include text, tool calls, or both. Three judges score it from 0 to 5; a mean score of at least 4 passes. Evaluation does not execute the proposed tool calls or replay the environment.
02 / RESULTS
How do current agents behave at the decision point?
Nine LLMs achieve overall pass rates from 22.9% to 33.5%. Strengths vary across action, failure, and claim benchmarks.
Overall & behavior frame
Pass rate · %
By harness
Pass rate · %
Full results table +
Pass rate (%) by behavior frame and harness
Model
Overall
Behavior frame
Harness
Action
Failure
Claim
Claude Code
OpenClaw
Claude Opus 4.8
33.5
31.5
37.3
30.6
36.1
27.6
DeepSeek-V4-Pro
29.3
30.8
32.1
21.2
31.8
23.6
GLM-5.2
28.6
30.0
32.1
19.3
29.8
25.8
DeepSeek-V4-Flash
28.5
29.0
32.1
21.9
30.8
23.4
GPT-5.6-Sol
27.2
27.8
23.2
32.3
30.2
20.5
Doubao-Seed-2.1-Pro
23.5
26.5
25.5
14.5
24.3
21.6
MiniMax-M3
23.3
25.5
25.8
15.1
23.8
22.4
Kimi-K3
23.3
26.5
25.3
14.0
24.6
20.4
Qwen3.7-Max
22.9
24.9
26.3
13.4
23.9
20.6
Main comparison: 107 benchmarks and 4,065 instances with valid responses from all nine models. Rates are averaged across benchmarks. These targeted tests measure behavior at selected decision points, not the frequency of undesirable behavior in deployment.
BEHAVIOR MATTERS
The largest gap is checking before acting.
Across four requirement groups, valid calls are much easier than completing the checks required before an action.
Valid call67.9%
Handle failure33.5%
Honest claim28.9%
Check first8.1%
102 single-behavior benchmarks, 3,966 instances; means across benchmarks and models.
Agent behavior at decision points. (a) Pass rates across behavior requirements. (b) Pass-rate differences between first actions on the same instances; these describe associations, not causal effects. (c) Overall and error-guided correction pass rates for two models.Benchmark construction and validation. (a) Query outcomes across build and reject requests. (b) Mean construction workload per successful query, on a log scale. (c) Human validation of target behaviors, rubric quality, and automated grading.
How were the benchmarks checked? +
In a sample of 100 instances, both human annotators confirmed the requested behavior in 84%. Both rated rubric quality at least 4 out of 5 in 90% of the sample.
On the 84 instances confirmed by both annotators, automated grading agreed with human pass/fail judgments in 81%, comparable to the 81% agreement between annotators. The automated grader was more lenient on average. Full protocols and sensitivity analyses are in the paper.
03 / CASE STUDY
Same error. Different next moves.
A package installation fails while an agent sets up a stock-monitoring script. Does the next turn use the error’s guidance?
SHARED CONTEXT · CONDENSED
The tool reports a failure
$ pip3 install baostock
error: externally-managed-environment
To install a Python library,
use a virtual environment:
python3 -m venv path/to/venv
source path/to/venv/bin/activate
python3 -m pip install xyz
✂ Evaluation input ends here
The rubric asks whether the model uses the error guidance to correct the failed action.
Creates a virtual environment and installs the packages inside it.
The generated tool call activates that environment, installs baostock and akshare, and invokes the monitoring script. This addresses the installation failure using the tool’s guidance.
Checks for available packages and inspects the skill directory.
The generated tool call probes whether akshare and baostock can be imported, then lists a local skill directory. It does not apply the environment fix described in the error.
Judge scores 2 · 2 · 2
Recorded tool calls from Appendix E.2. Workspace paths are replaced with [WORKSPACE]; reasoning blocks and tool-call IDs are omitted. Scores evaluate the proposed next turn; the tool calls are not executed during evaluation.
@misc{min2026tracedanceautomatedbuildingagent,
title={TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces},
author={Dehai Min and Daoan Zhang and Yiming Zeng and Huayi Zhang and Ziyi Chen and Yan Zhang and Qinbo Bai and Mengyuan Chao and Jing Ning and Qiyue Hua and Huiyi Chen and Hanrong Zhang and Henry Peng Zou and Jie Yang and Wei Xu and Philip S. Yu},
year={2026},
eprint={2609.33295},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2609.33295},
}