The Bookability Baseline: July 2026
What a homeowner with money and a deadline actually meets when she calls
The Front Door Integration Test, first wave. Next wave: October 2026. The rubric does not change between waves.
The report, 20 pagesThe two-page summaryThe walkthrough
Free, and no email required. The two-page summary is the same study on two printable pages: send it to whoever owns intake, and the seven listening checks cost them an hour.
What I did
In July 2026 I placed 55 mystery-shop calls to 48 residential service brands, HVAC, plumbing and roofing, owned by roughly 20 private-equity sponsors. Same method, same scoring rubric, every call transcribed and scored the same way. I called the number each brand’s marketing budget exists to make ring, and I presented as the caller every one of these businesses says it wants: a homeowner with a real problem, money to spend, and two or three companies on the list.
The phones mostly get answered, and the selling stops there. The layer between the answered call and the booked job does not exist as a system. What I found instead were scattered individual habits: one written recap, two human save attempts, in 55 calls. Administrative integration had shipped at most of these platforms. Customer-facing integration mostly had not, and the caller is the only person positioned to notice.
Where the calls stopped
- answered, with a person on the line49 of 55
The six that never answered were all at brands running paid ads that week.
- could name the price of a visit40 of 49
11 of these 40 read me the fee card but could only take a message.
- asked for an email address13 of 49
- tried to keep me when I hesitated8 of 49
Six of the eight were machines. Agents 6 of 13, people 2 of 33.
- sent anything in writing1 of 55
- offered a sister brandnone of 55
“We only handle hot water tanks or tankless water heaters, but I don’t have any plumbers on staff for anything else. So you’d have to call a plumber for that, unfortunately.”
The five findings
Nobody passed me to a sister brand
0 of 55In 55 calls, including 13 where I volunteered a second problem in the home unprompted, no human or AI agent made a live referral to a sibling brand in any phrasing. Three of the 13 got one-trip handling; the rest drew a second visit, a second fee, or a competitor's name.
One platform, four different front doors
4 brandsFour sibling brands under one sponsor, shopped inside 48 hours with the same scenario, held an identical diagnostic fee and diverged on everything after it. One texted a recap mid-call. One captured no contact detail at all.
The machines tried harder to keep me
6 of 13 against 2 of 33At the moment I deferred, scripted AI agents made a save attempt in 6 of 13 agent-run calls. Human answerers made one in 2 of 33. Three answered calls opened with an agent and passed to a person, so they sit in neither group. Fisher's exact p = 0.004; the direction holds, the multiple carries wide uncertainty.
The $8,500 caller got the service-call process
0 of 5Five calls presented a competitor's written $8,500 replacement quote, two with an offer to email the proposal. Zero platforms asked to see it, and soonest visits ran two to ten days for a buyer deciding that weekend.
A machine repeats its mistakes exactly
16 of 55 AI-frontedOne agent accepted an email address announced out loud as fake, and the appointment flow continued as designed. Another walked its own empty calendar into September without escalating to a human being.
On finding 01, the fair objection: some platforms keep their brands apart on purpose, for licensing or service-area reasons. If that is you, the zero is a decision rather than a gap. Either way the caller hangs up thinking the platform sells one trade, so somebody should choose which of the two is true.
The instrument · eight stages, one call
Answer
49 of 55Reached a live voice. Six dead ends, all at brands with live paid media that week.
Qualify
no scoreNot scored: the rubric has no clean denominator here, because what counts as a qualifying question varies by trade. Read qualitatively it was the strongest stage in the sample, though one weekend emergency still ended without a phone number taken.
Price
40 of 49Stated the visit fee, or free. Nine could not price a visit when asked.
Paid software ends here · workflow begins
Hold
8 of 49Tried to keep the caller at the moment they deferred. Six of the eight were scripted AI agents.
Capture
13 of 49Asked for an email. One address announced out loud as fake was accepted.
Recap
1 of 55A written recap received before hang-up. The best call in the sample, and the only one.
Follow-up
2 observedTwo promised callbacks observably arrived. The rest fell outside the study window and are unscored.
Route
0 of 55Live referrals to a sibling brand. Not one, in any phrasing.
The callback ledger
2 of 11 arrived
nine unresolved, outside the study window
Eleven brands promised to call me back. Two did, both same-day. Nine unobserved outcomes would make any published rate a guess wearing a percent sign, so the ledger stays a ledger and this report publishes no callback-kept rate, deliberately.
If your platform can produce its own callback-kept rate from the CRM, you can answer a question this study could not.
Every rate, with its interval
| Finding | Observed | Rate | 95% interval |
|---|---|---|---|
| Answered, with a person on the line | 49 of 55 | 89% | 78–96% |
| Could name the price of a visit | 40 of 49 | 82% | 68–91% |
| Tried to keep me when I hesitated | 8 of 49 | 16% | 7–30% |
| Save attempts, agents against people | 46% vs 6% | agents 19–75%, humans 1–20% | |
| Asked for an email address | 13 of 49 | 27% | 15–41% |
| Sent anything in writing | 1 of 55 | 1.8% | 0.0–9.7% |
| Offered a sister brand | 0 of 55 | 0% | upper bound 6.5% |
| Second problem handled in one trip | 3 of 13 | 23% | 5–54% |
| Asked to see a competitor's written quote | 0 of 5 | 0% | 0–52% |
Method · limitations first
You diligence companies for a living, so the weaknesses go first.
- It cannot tell you close rates.
- The shopper never books, so nothing in this study is a conversion claim. I measured whether each brand did the things that close callers: the ask, the hold, the capture, the recap, the follow-up.
- Per-brand reads are directional.
- Most brands got one call; a handful got two or three. One call proves nothing about one brand. Fifty-five calls with the same misses repeating across 48 brands and roughly 20 sponsors is a pattern.
- July is peak season, and that cuts both ways.
- A brand with a full board has a rational reason to let a deferring caller go, so some of what I scored is capacity rationing rather than broken intake. Slow soonest-visit dates belong in that column. The stages this study leans on do not: asking for an email, putting the visit in writing, naming the sibling brand that does the work. None of those consume a truck slot. Replacement intent sits outside the argument too, because an $8,500 install runs on install crews, and install revenue is the thesis the platform was bought on.
- The economics are modeled, never measured.
- No bookings were completed, so the economic section is a scenario model with disclosed assumptions. Confidence intervals are published on every rate this study leans on, including the two that read zero.
- How the calls were run.
- Calls went to each brand’s published consumer line between July 16 and 25, 2026, across dayparts including evenings and a weekend. A small cast of homeowner personas ran four standard scenarios, and every call was scored on the same eight stages regardless of scenario. Personas used shopper-controlled phone numbers and inboxes. No bookings were confirmed, no trucks rolled, and no free work was extracted. Brands and sponsors are anonymized throughout, and quotes are verbatim, cleaned only where the transcription engine mangled a word.
The full method detail and ethics notes are in the appendix. The waves run quarterly on the same rubric, so the zero in finding one has somewhere to move.
What to do Monday
Ten questions. Seven cost an hour of listening to your own calls.
Listen · one hour · five recorded calls per brand
- 01Does anyone ask for the caller's email address?13 of 49
- 02When the caller defers, is the next sentence a save attempt, or “no worries”?8 of 49 attempted
- 03Can the person answering state the price of a visit, and whether it credits against the work?9 of 49 could not price at all; another 11 calls reached someone who could only take a message
- 04When a second problem comes up mid-call, does it get one trip, a second fee, or “call someone else”?3 of 13 got one trip
- 05Does anything arrive in writing after the call?1 of 55
- 06When the caller holds a competitor's quote, does anyone ask to see it?0 of 5
- 07Call your own after-hours line tonight. Can the voice that answers quote the fee, book a job, or commit to a callback with a time on it?mostly no
Pull · three reports · formulas published in full
- 08Call-to-booked-job, by brand, on one definition across every brand. If the definitions differ by brand, that finding outranks the number.
- 09Speed to answer and true abandonment, by brand and daypart, with short abandons excluded before the rate means anything.
- 10Membership attach, where eligibility is written down per brand and existing members are out of the denominator.
If the seven listening checks come back clean at your platform, close the report and keep the hour. In this sample of 48 brands, none would have come back clean.
The front door is the integration test. Most of the platforms I called would not pass it today, and the person who finds that out is the customer you paid to reach.
I also run this as a two-week, fixed-fee read on your own numbers: the three pulls on your booking and phone data, a sample of your own recorded calls, and the same arithmetic done on your figures instead of mine. If you want it, reply with the name of your busiest brand and I will tell you inside a day whether your data can answer the seven questions.
Zaha Al-Hmoud
Independent consultant, Vancouver.
zaha@ztsadvisory.com
Cite as: ZTS Advisory, The Bookability Baseline, July 2026.