Product /Agent Eval

Measure what agents accomplish before deciding what to optimize.

Connect task outcomes, model and tool calls, human involvement and resource use to identify quality and efficiency improvements.

Built for enterprise
01

What we evaluate

X Worker digital workers and existing enterprise agents whose runtime data can be integrated.

02

Decision evidence

Business criteria, runtime records and human review explain results together.

01 / Evaluation metrics

Agree on measures of successful delivery.

Agree on success and quality requirements with business teams before comparing time, human involvement and costs.

01

Task success rate

Tasks passing agreed acceptance rules divided by all evaluated tasks. Define treatment of partial and canceled work in advance.

02

Quality and time

Evaluate accuracy, completeness, supporting sources and end-to-end duration.

03

Human involvement

Track correction, review and supplementary effort, distinguishing required approval from failure recovery.

04

Cost per successful task

Total cost of the same evaluation set divided by successful tasks, including failures and retries. Add labor cost separately only when consistently estimated.

05

When calculation is unavailable

If no task succeeds, show that the metric cannot be calculated and explain why, rather than showing zero cost or infinite returns.

02 / Execution-path analysis

Find where tasks repeatedly consume resources.

Deliverables are evaluated against task objectives.
Concept illustration
03 / Evaluation methods

Combine automated scoring with business review.

Build a baseline from explicit samples and rules, compare versions and monitor actual operation.

01

Test sets

Select representative tasks and record difficulty, input versions, permitted sources and expected results.

02

Scoring

Combine business rules, automated scores and human sampling, with each method's role and limitations explained.

03

Baseline comparison

Compare model or workflow versions under the same task and quality conditions.

04

Post-launch monitoring

Observe deviations in real tasks; review regressions and update evaluation sets.

04 / Optimization reports

Turn evidence into testable recommendations.

Recommend changes to models, context, workflows and human collaboration, with a clear method for measuring improvement.

01

Findings

Link evidence to specific steps, failure causes or repeated consumption.

02

Recommend

Adjust model capability, reduce redundant context or improve tool steps and review points.

03

Expected impact

Describe potential quality, time or cost improvements without assuming fixed percentages.

04

Retesting

Validate with the original task set and new edge cases, recording changes and side effects.

05 / Feedback and next steps

Bring every change back to task outcomes.

Agree improvements with business and platform teams, update workflows or model policies, then evaluate the results.

01

X Worker

Improve knowledge, skills, team roles and workflows using evidence.

02

Route Engine

Adjust routing within permitted models and quality thresholds.

03

Enterprise review

Owners approve production changes, with versions and rollback paths retained.

Request an agent assessment