Peak hurricane season does not break average-path answering. It breaks exceptions: triage, dispatch packets, and books. Grade vendors on surge completion.
The peak window just opened
NOAA's Atlantic hurricane season has a long-standing climatological peak around September 10. The dangerous stretch is not a single landfall. It is the window that opens in the second half of August and runs through mid-September. This week, HVAC shops, plumbers, waste haulers, and restoration crews enter the period when emergency inbound volume, after-hours demand, and vendor contention spike at the same time.
That is the first live, unforgiving test of the 2026 AI reception and dispatch stacks you bought on a Tuesday.
Most seasonal coverage will tell you to staff overtime, check generators, and make sure someone picks up the phone. Pickup is not the failure mode. Your AI already answers. Vendors already demo a clean call: name, address, "we will send someone." Containment rate looks fine. Then a tropical system sits off the coast, 30 to 40 percent of inbound calls become messy emergencies, and the system that sounded competent at lunch last week starts producing work that cannot be dispatched, cannot be quoted honestly, and cannot be booked without creating a week of cleanup.
If you are still comparing vendors on how human the voice sounds or how many calls it contains, you are grading the wrong exam.
Tuesday AI fails the exception path
We already argued, in The $2B AI Bet Misses the Workflow, that a finished prompt is not a finished job. Storm week is that argument under load. The average path still works. A caller with a clear address, a standard service, and a flexible window gets a confirmation. The damage is in the exceptions, and exceptions become the median case when weather hits.
Here is what actually breaks.
Incomplete job packets. The caller is in a car. The street sign is down. They give a subdivision name, a cross street, or "the brick house past the church." A Tuesday agent can loop until it gets a geocodable address. A surge agent under queue pressure will close the call with a partial record. The crew then sits on a job that is not a job.
Priority inversion. A scheduled roll-off swap and a no-AC call on a 96-degree afternoon land three minutes apart. If your triage is first-in, first-out or "whoever sounded urgent," you will send a truck to a dumpster while a medically vulnerable customer waits. Storm season is a ranking problem. If the model cannot apply your priority rules (life safety, then occupied structures, then property, then routine), it will invert them at volume.
Double-booked vendors. Three offices, or one AI plus a night dispatcher plus a foreman texting crews, will assign the same plumber to two flood calls. The vendor accepts both because the system never showed a conflict. You do not learn this from a containment dashboard. You learn it when the second customer is still standing in water at 7 p.m.
Quotes without constraints. A generator install, a tarp, a panel repair, or a commercial pump-out gets a price on the call. The price ignores that the part is on allocation, the city has a temporary permit hold, or the dump is only taking storm debris on even hours. You just sold work you cannot execute. The callback is worse than a declined quote.
A ledger that surfaces after the weather clears. Jobs get done on paper, in texts, and in the field app. Invoices wait. Credits wait. Duplicate work orders sit next to never-dispatched ones. When the sky clears, your books are two weeks of exceptions. That is not an accounting inconvenience. It is unbilled revenue, disputed invoices, and a month of owner time reconciling what the AI "handled."
None of those failures show up in a "we answered 94 percent of calls" slide. Answering is the cheap part of the test.
What you should grade instead
Stop asking vendors for containment rate as the headline metric. Ask for surge SLAs and exception completion.
A useful scorecard for this window:
- Dispatchable packet rate. Of emergency calls, what share produce a record a crew can run without a callback: geocoded or verified address, access notes, hazard flags, customer contact, and a service code your board actually uses.
- Triage fidelity. Under a scripted mix of life-safety, property, and routine work, does the system rank jobs the way your board ranks them, or does it flatten everything into "soon"?
- Hold and escalate authority. When the request is ambiguous, over quote authority, or outside licensed scope, does the agent hold, transfer, or invent an answer? We already asked who actually controls the AI operation. Storm week makes that operational: who can freeze a quote, bump a job, or force a human handoff when the queue is ugly.
- Vendor conflict detection. Can the system see the same tech across phone, SMS, and the scheduling board, and refuse a second assignment inside the travel window?
- Constraint-aware quoting. Does a price attach only after parts, permit, dump, and access constraints are checked, or does it emit a number because the caller asked for one?
- Same-day ledger integrity. By close of business, can you see which jobs were completed, which were deferred, which were duplicates, and which have no invoice path? If that view only exists after the storm, you do not have an operating system. You have a call logger.
If a vendor cannot produce those numbers from a pilot, they are selling Tuesday.
Name the exception cases in the RFP
Make the vendor say what happens, in writing, for each of these:
- The caller cannot give a house number and is on a dying phone.
- Two emergencies target the last available electrician.
- A commercial customer requires a PO and the AI quotes anyway.
- The licensed tech for gas work is already en route; a second gas call arrives.
- The customer changes scope on site and the original quote is now wrong.
- The AI and a human dispatcher assign the same crew 90 seconds apart.
- Ten completed jobs have no invoice because the work order never closed.
If the answer is "it escalates," ask to whom, after how long, with what packet, and what the customer hears while they wait. Escalation that dumps a raw transcript on a busy owner is not a fallback. It is a delay with extra steps.
Run this bake-off before Labor Day
You do not have a month to swap a happy-path demo for a surge-ready workflow. You have days. Use them.
This week, take the last real storm or heat-emergency day you have recordings for. Pull 25 calls. Tag at least 10 as exceptions: incomplete address, priority conflict, parts or permit constraint, vendor collision, scope change. Play them into every vendor still in your bake-off. Score only four outcomes:
- Dispatchable packet, no callback needed
- Correct hold or escalation, with a usable packet
- Invented or incomplete close
- Silent drop or duplicate assignment
Anything in the last two buckets is a fail, even if the caller thanked the agent.
Ask these questions in the pilot, out loud, on a recorded session:
- Show me the job record produced from this incomplete-address call. Can a crew roll on it?
- Here are three calls in four minutes: no heat for an infant, a routine filter change, and a flooded basement. What is the rank order, and who got the next truck?
- What is the quote authority ceiling, and what happens one dollar over it?
- If I assign the same vendor from the board while the AI is still on the call, who wins, and how fast does the loser get told?
- At 6 p.m. on a 40-call day, can I list every open, completed, and unbilled job without opening email?
If they cannot run that script this week, they are not ready for the window that is already open.
Reeve is built around the packet, the dispatch, and the books, not the greeting. If you are comparing systems for this season, start with exception completion. The voice is the easy part.
Peak season will grade your stack whether you scheduled the exam or not. Run the messy calls now, while you can still change hold rules, authority limits, and escalation targets. After September 10, you will be reading the same checklist as a postmortem.