Skip to content
Agent Factory

Free tool

Agent Benchmark

15 operational metrics across 5 areas of running an agent. Type in the numbers you measure and it tells you, as you go, which are out of range and which article fixes each one.

Every threshold comes from a published article, and each one says which. 13 of 15 healthy thresholds are quoted from the library; the rest are ours, and are labelled as ours rather than borrowing authority from an article. Nothing is stored unless you ask for it.

/100

Awaiting input

0 of 15 metrics entered

Product health
0/3
Unit economics
0/3
Retrieval quality
0/3
Release safety
0/3
Commercial durability
0/3
Try an example:

Product health

Task completion rate

No data
%

Above 80% is healthy. Below 70%, the agent is not carrying the work.

Escalation rate

No data
%

Under 25% is healthy. Above 35%, people are still doing the job.

p95 latency

No data
s

Under 30 seconds is healthy. Past 60, users stop waiting.

Unit economics

Cost per completed task

No data
$

Under $0.50 per completed task is healthy. Above $2.00, pricing has to change.

Inference cost / revenue

No data
%

Below 20% of revenue is efficient. Above 40%, pricing or architecture needs fixing.

Gross margin incl. inference

No data
%

Above 70% after inference proves the unit economics. Below 50% does not.

Retrieval quality

Recall@k

No data
%

Above 80% is healthy. Below 70%, retrieval is the problem, not the prompt.

Faithfulness

No data
%

Above 90% is healthy. Below 80%, answers are drifting from the source.

Answer relevance

No data
%

Above 90% is healthy. Below 80%, answers are missing the question asked.

Release safety

Golden dataset pass rate

No data
%

95% is the deployment gate. Below 90%, nothing ships.

LLM-judge / human correlation

No data
r

Above 0.85 the judge can screen for you. Below 0.70 it cannot.

Override rate

No data
%

Under 10% earns more autonomy. Above 30%, the agent is miscalibrated for the job.

Commercial durability

Net revenue retention

No data
%

Above 110% grows on expansion alone. Below 90% is a churn problem sales cannot outrun.

Time to first win

No data
d

Inside 7 days retains three times better. Past 21, the momentum is gone.

CAC payback

No data
mo

Under 12 months works for SMB. Past 18, even enterprise economics break.

What to fix, in order

Nothing to rank yet. Enter any metric above — one is enough — and this becomes an ordered list of what to fix, worst first, each with the article that fixes it.

Or load an example to see a finished report.

Add this run to the dataset

Optional, and separate from using the tool — everything above was calculated in your browser and stored nowhere. Submitting is what makes percentiles possible, for you next month and for the next person.

What is stored
The metric values you entered, four coarse labels below, the threshold version, and the month — not the day or the time.
What is never stored
No company name, no email, no URL, no IP address, and no free-text box of any kind. There is no field to type one into.
What you get back
A place in the dataset the comparisons are drawn from. Reading your position against others, and your own run history, is a Pro view — submitting counts either way, and nobody sees a percentile until a cohort reaches 20.

Runs are grouped by a random key kept in this browser, never sent and never linked to you. Lose the browser storage and they become unreachable, by anyone, including us — that is the cost of not asking who you are. Submissions are deleted after 24 months, and you can delete yours at any time with the control below.

What does your agent do?
How big is the team behind it?
How far has it got?

A prototype and a production system should not be judged against the same numbers.

Industry

Optional. A rare industry in a thin month narrows the field more than it helps.

Enter at least one metric first.

How this works

Each metric has a healthy threshold and a lower bound. Meeting the healthy threshold scores 100, sitting exactly on the lower bound scores 50, and the far end of the scale scores 0, with a straight line between. That is what makes a cost gap and a latency gap comparable, and it is why the ranked list can order them against each other.

Blank is blank. A metric you do not enter is skipped entirely rather than counted as zero, so a partial answer gives you a real read on what you measured plus an honest count of how much that was. Nothing is imputed, defaulted, or extrapolated — the score never pretends to know more than you told it.

A value outside its expected range is scored as entered and flagged, not quietly corrected. If you type a percentage where a ratio belongs, you should see that rather than a plausible-looking score.

Comparisons need volume. A percentile is only shown once a category holds at least 20 submissions for that metric — below that it is noise, and it would start to identify the handful of people in it.