Task success rate
Tasks passing agreed acceptance rules divided by all evaluated tasks. Define treatment of partial and canceled work in advance.
Connect task outcomes, model and tool calls, human involvement and resource use to identify quality and efficiency improvements.
X Worker digital workers and existing enterprise agents whose runtime data can be integrated.
Business criteria, runtime records and human review explain results together.
Agree on success and quality requirements with business teams before comparing time, human involvement and costs.
Tasks passing agreed acceptance rules divided by all evaluated tasks. Define treatment of partial and canceled work in advance.
Evaluate accuracy, completeness, supporting sources and end-to-end duration.
Track correction, review and supplementary effort, distinguishing required approval from failure recovery.
Total cost of the same evaluation set divided by successful tasks, including failures and retries. Add labor cost separately only when consistently estimated.
If no task succeeds, show that the metric cannot be calculated and explain why, rather than showing zero cost or infinite returns.

Build a baseline from explicit samples and rules, compare versions and monitor actual operation.
Select representative tasks and record difficulty, input versions, permitted sources and expected results.
Combine business rules, automated scores and human sampling, with each method's role and limitations explained.
Compare model or workflow versions under the same task and quality conditions.
Observe deviations in real tasks; review regressions and update evaluation sets.
Recommend changes to models, context, workflows and human collaboration, with a clear method for measuring improvement.
Link evidence to specific steps, failure causes or repeated consumption.
Adjust model capability, reduce redundant context or improve tool steps and review points.
Describe potential quality, time or cost improvements without assuming fixed percentages.
Validate with the original task set and new edge cases, recording changes and side effects.
Agree improvements with business and platform teams, update workflows or model policies, then evaluate the results.
Improve knowledge, skills, team roles and workflows using evidence.
Adjust routing within permitted models and quality thresholds.
Owners approve production changes, with versions and rollback paths retained.