Task completion rate
No dataAbove 80% is healthy. Below 70%, the agent is not carrying the work.
Free tool
15 operational metrics across 5 areas of running an agent. Type in the numbers you measure and it tells you, as you go, which are out of range and which article fixes each one.
Every threshold comes from a published article, and each one says which. 13 of 15 healthy thresholds are quoted from the library; the rest are ours, and are labelled as ours rather than borrowing authority from an article. Nothing is stored unless you ask for it.
—/100
Awaiting input
0 of 15 metrics entered
Above 80% is healthy. Below 70%, the agent is not carrying the work.
Under 25% is healthy. Above 35%, people are still doing the job.
Under 30 seconds is healthy. Past 60, users stop waiting.
Under $0.50 per completed task is healthy. Above $2.00, pricing has to change.
Below 20% of revenue is efficient. Above 40%, pricing or architecture needs fixing.
Above 70% after inference proves the unit economics. Below 50% does not.
Above 80% is healthy. Below 70%, retrieval is the problem, not the prompt.
Above 90% is healthy. Below 80%, answers are drifting from the source.
Above 90% is healthy. Below 80%, answers are missing the question asked.
95% is the deployment gate. Below 90%, nothing ships.
Above 0.85 the judge can screen for you. Below 0.70 it cannot.
Under 10% earns more autonomy. Above 30%, the agent is miscalibrated for the job.
Above 110% grows on expansion alone. Below 90% is a churn problem sales cannot outrun.
Inside 7 days retains three times better. Past 21, the momentum is gone.
Under 12 months works for SMB. Past 18, even enterprise economics break.
Nothing to rank yet. Enter any metric above — one is enough — and this becomes an ordered list of what to fix, worst first, each with the article that fixes it.
Or load an example to see a finished report.
Optional, and separate from using the tool — everything above was calculated in your browser and stored nowhere. Submitting is what makes percentiles possible, for you next month and for the next person.
Runs are grouped by a random key kept in this browser, never sent and never linked to you. Lose the browser storage and they become unreachable, by anyone, including us — that is the cost of not asking who you are. Submissions are deleted after 24 months, and you can delete yours at any time with the control below.
Each metric has a healthy threshold and a lower bound. Meeting the healthy threshold scores 100, sitting exactly on the lower bound scores 50, and the far end of the scale scores 0, with a straight line between. That is what makes a cost gap and a latency gap comparable, and it is why the ranked list can order them against each other.
Blank is blank. A metric you do not enter is skipped entirely rather than counted as zero, so a partial answer gives you a real read on what you measured plus an honest count of how much that was. Nothing is imputed, defaulted, or extrapolated — the score never pretends to know more than you told it.
A value outside its expected range is scored as entered and flagged, not quietly corrected. If you type a percentage where a ratio belongs, you should see that rather than a plausible-looking score.
Comparisons need volume. A percentile is only shown once a category holds at least 20 submissions for that metric — below that it is noise, and it would start to identify the handful of people in it.