Compare the same tasks before and during the trial. Include checking and correction time, count only acceptable results, and record errors or unsupported questions separately.
Choose tasks before seeing the demo
Pick a recurring job such as obtaining an account balance, checking a sales total or preparing a follow-up list. Define the company, date range and result you expect. Include an ordinary question and a difficult one: similar customer names, a return or missing information.
Keep the task list stable during the comparison. Replacing failed tasks with easier ones makes the trial look better without telling you whether it helps the business.
Count the whole task, including review
In this fictional example, a staff member completes ten comparable tasks. Both methods start with the same prepared records, so additional data preparation time is zero in this example. Add any real preparation or upload time to your own trial. Previously each took 12 minutes, so the total was 120 minutes. During the trial, asking and receiving an answer takes three minutes per task; checking takes another four. Two tasks need five additional minutes of correction each.
| Work | Minutes |
|---|---|
| Previous method: 10 × 12 | 120 |
| Trial: 10 × (3 + 4) | 70 |
| Extra correction: 2 × 5 | 10 |
| Total trial time | 80 |
| Time difference | 40 saved |
The saving is 40 minutes across the ten tasks, about one-third of the previous time. It is not 90 minutes: that larger claim would ignore checking and corrections. These are example inputs, not results achieved by Dhandha customers.
Correctness is a separate condition
If a result remains wrong, do not count it as a completed task just because it arrived quickly. Record the error, whether the tool disclosed uncertainty and what the person did to recover. An honest refusal on a missing-data question may be the right behaviour, but it does not complete a task that the business needs answered.
Distinguish tasks answered correctly, appropriately refused, incorrectly answered and unfinished. This prevents a tool that refuses everything from appearing perfectly accurate and highly useful.
Then look at cost
Include setup effort, subscription or usage fees, record preparation and support. Time released is not automatically cash saved: the business may use it for sales, customer service or less overtime. State which outcome you actually observed.
Agree acceptable performance with the task owner. A casual message draft and an account balance do not have identical error consequences. For financial figures, compare with the matching source records before operational use.
Keep an evidence log
For each task, retain the question, inputs or snapshot date, answer, independent reference, review time and verdict. Use anonymised or approved sample data during evaluation. The NIST AI Risk Management Framework provides a broader reference for evaluating AI risks; this simple task log is an operating tool, not a NIST certification.
At the end of the trial, decide what to retain, fix or stop. “Fast, but needs a fresh export” is a more useful conclusion than “AI works.” It identifies the next change and keeps the owner in control of whether the result is worth the effort.
Questions you might ask
How should I measure time saved by a business assistant?
Compare the same completed tasks and include preparation, checking and correction time. Report errors and unfinished work separately.
Does an assistant with no wrong answers always pass a trial?
No. It may refuse useful tasks. Measure correct completion and appropriate refusals separately, using a fixed task list.