AI AutomationROIAuditabilityLocal Services

The DOGE Report's Lesson: Show Your AI Work

R
Reeve Team
5 min read

The GAO's findings expose a better standard for AI ROI: measurable savings, documented decisions, and results that survive an audit.


The savings claim is not the result

This week's GAO findings on DOGE's reported savings should matter to anyone buying AI for an operating business. The watchdog found that 96 percent of reported grant savings could not be verified. It also found that DOGE did not use its stated methodology for most contract savings and overstated savings from terminated leases. Ars Technica's report lays out the details.

The lesson is not about one government program. It is about how easily a large savings number can become detached from an observable business result.

Local service businesses are now hearing similar claims about AI. A system will save dozens of labor hours, recover missed revenue, eliminate administrative work, or produce a certain return in a few months. Some of those results may be real. The problem is that buyers often receive the conclusion without the evidence needed to test it.

If an AI system touches calls, scheduling, dispatch, quoting, or financial records, you should demand an explanation of how its savings were calculated. A dashboard showing a large number is not proof. It is the beginning of the investigation.

What most ROI calculations get wrong

The common mistake is treating activity as savings.

An AI receptionist may answer 500 calls. That does not mean it created 500 valuable opportunities. A scheduling system may create 200 appointments. That does not mean those appointments were incremental. A dispatch tool may recommend fewer manual interventions. That does not mean the business avoided a cost unless someone can show what would otherwise have happened.

There are four different categories hiding inside most automation claims:

  • Work completed by the system, such as calls answered or jobs created.
  • Work avoided, such as manual follow-up that no longer needed to happen.
  • Business value created, such as an additional booked job that would likely have been lost.
  • Costs introduced, such as review time, corrections, integration work, and exception handling.

Only the third category is revenue. Only the second category is a true efficiency gain. The first category is a useful operational metric, but it is not automatically worth money. The fourth category is where optimistic ROI models quietly lose credibility.

A serious evaluation keeps these categories separate.

Build an evidence chain for every claimed benefit

Before trusting an ROI estimate, ask the system to show its work. For each claimed benefit, you should be able to trace a path from an event to a financial or operational result.

For example, a credible missed-call recovery claim should answer:

  1. How many calls were previously missed during a defined baseline period?
  2. How many of those calls were legitimate service opportunities?
  3. How many were answered after automation was introduced?
  4. How many became qualified jobs?
  5. How many converted into completed, paid work?
  6. What gross margin did those jobs produce?
  7. What did the system cost, including supervision and exceptions?

That chain is more useful than saying the system recovered 40 percent more calls.

The same standard applies to dispatch. If automation claims to save dispatcher time, measure the time spent before and after implementation. Track reassigned jobs, delayed arrivals, manual corrections, and escalations. If the system reduces touches but increases mistakes, the business has not saved money. It has moved the cost somewhere less visible.

The calculation should also distinguish gross revenue from gross profit. A $10,000 increase in booked work is not a $10,000 return. Fuel, labor, disposal fees, subcontractor payments, refunds, and rework all affect the result.

Human judgment belongs in the measurement

A trustworthy automation system does not pretend that every decision is automatic. It records where confidence was low, where a rule was overridden, and where a person made the final call.

That information is valuable for two reasons. First, it prevents the business from counting uncertain or rejected outcomes as successful automation. Second, it shows where the workflow still depends on expertise.

Suppose an AI system routes a plumbing job to the wrong crew, and a dispatcher catches the error before the technician leaves. The final customer outcome may look fine. The automation did not create a successful dispatch, though. It created a recommendation that required human correction. Count it accurately, and the business can improve the routing rules. Count it as a success, and the reporting becomes fiction.

This is also why exception queues matter. You should know how many jobs, calls, quotes, or records required review, why they were escalated, and how long resolution took. A high exception rate may mean the system is poorly configured. A low exception rate may mean the system is working well, or it may mean exceptions are disappearing without being tracked.

The audit trail needs to preserve both outcomes and decisions.

Use a baseline, not a vendor benchmark

The strongest ROI measurement begins before implementation. Pick a representative period, usually 30 to 60 days, and record the current state.

For a local service business, that baseline might include:

  • Inbound calls by hour and day.
  • Missed calls and callback time.
  • Lead-to-booking conversion.
  • Time spent coordinating crews.
  • Rescheduled or unassigned jobs.
  • Quote turnaround time.
  • Manual data entry and reconciliation time.
  • Cancellations, refunds, and rework.

Then define the measurement window and keep the definitions stable. Do not compare a slow winter month with a peak summer month and call the difference automation savings. Do not change the meaning of a qualified lead halfway through the test. Do not count revenue that would have arrived through the old process.

A useful test also includes a control period. Keep one location, service line, or shift on the existing workflow when practical. If every part of the business changes at once, you cannot tell whether the result came from AI, seasonality, pricing, staffing, or a marketing campaign.

This is not academic measurement. It protects your budget and your credibility with the people who depend on your numbers.

Ask for an audit packet before you buy

You do not need a complicated data science program to evaluate an automation claim. You need a repeatable record.

Ask for these items:

  • The exact definition of each KPI.
  • The baseline period and comparison period.
  • The formula used to calculate savings or return.
  • A record of successful, failed, and human-corrected outcomes.
  • The cost of implementation, supervision, and exception handling.
  • Sample event logs that connect actions to results.
  • A process for correcting inaccurate records.
  • A monthly report that preserves historical values instead of rewriting them.

If the answer is only a polished ROI calculator, keep asking questions. A calculator can model assumptions. It cannot prove that the assumptions were true.

Our earlier post, Your AI Vendor's Next Outage Tests More Than Uptime, made the case that a reliability metric is incomplete without evidence of how a workflow behaves under stress. ROI deserves the same treatment. A savings percentage without event-level evidence is just another uptime number.

The practical standard: defensible improvement

The right question is not, How much does this AI promise to save?

Ask instead: Can we explain what changed, measure it against a credible baseline, identify the judgment calls, and reproduce the calculation six months from now?

That standard will reject some impressive-looking automation claims. Good. The purpose of measurement is not to make every project look successful. It is to help you find the projects that produce durable improvement.

Reeve is built around that operational view, connecting calls, jobs, dispatch decisions, quotes, and accounting records so the business can inspect what happened instead of relying on a headline number.

Before you approve an AI project, require a baseline and a monthly evidence report. If the system cannot show its work, do not trust its ROI.

Ready to streamline your operations?

See how Reeve handles calls, dispatch, and billing for local service businesses.

Related Articles