ZTS Advisory

Field study · 55 calls · 48 PE-owned residential service brands · about 20 sponsors · July 16–25, 2026 · nobody booked

The Bookability Baseline: July 2026

What a homeowner with money and a deadline actually meets when she calls

The Front Door Integration Test, first wave. Next wave: October 2026. The rubric does not change between waves.

The report, 20 pagesThe two-page summaryThe walkthrough

Free, and no email required. The two-page summary is the same study on two printable pages: send it to whoever owns intake, and the seven listening checks cost them an hour.

What I did

In July 2026 I placed 55 mystery-shop calls to 48 residential service brands, HVAC, plumbing and roofing, owned by roughly 20 private-equity sponsors. Same method, same scoring rubric, every call transcribed and scored the same way. I called the number each brand’s marketing budget exists to make ring, and I presented as the caller every one of these businesses says it wants: a homeowner with a real problem, money to spend, and two or three companies on the list.

The phones mostly get answered, and the selling stops there. The layer between the answered call and the booked job does not exist as a system. What I found instead were scattered individual habits: one written recap, two human save attempts, in 55 calls. Administrative integration had shipped at most of these platforms. Customer-facing integration mostly had not, and the caller is the only person positioned to notice.

Where the calls stopped

  • answered, with a person on the line49 of 55

    The six that never answered were all at brands running paid ads that week.

  • could name the price of a visit40 of 49

    11 of these 40 read me the fee card but could only take a message.

  • asked for an email address13 of 49
  • tried to keep me when I hesitated8 of 49

    Six of the eight were machines. Agents 6 of 13, people 2 of 33.

  • sent anything in writing1 of 55
  • offered a sister brandnone of 55
Each row draws its own denominator, so a row of 49 is shorter than a row of 55. The zero includes 13 calls where I raised a second problem in the house, unprompted. 16 of 55 calls were fronted by a scripted agent and 13 of those ran start to finish. Three opened with an agent then passed to a person, so they count in neither group.
“We only handle hot water tanks or tankless water heaters, but I don’t have any plumbers on staff for anything else. So you’d have to call a plumber for that, unfortunately.”
CSR at an HVAC brand whose platform sells plumbing under other flags. Verbatim, finding 01, July 2026.

The five findings

  1. Nobody passed me to a sister brand

    0 of 55

    In 55 calls, including 13 where I volunteered a second problem in the home unprompted, no human or AI agent made a live referral to a sibling brand in any phrasing. Three of the 13 got one-trip handling; the rest drew a second visit, a second fee, or a competitor's name.

  2. One platform, four different front doors

    4 brands

    Four sibling brands under one sponsor, shopped inside 48 hours with the same scenario, held an identical diagnostic fee and diverged on everything after it. One texted a recap mid-call. One captured no contact detail at all.

  3. The machines tried harder to keep me

    6 of 13 against 2 of 33

    At the moment I deferred, scripted AI agents made a save attempt in 6 of 13 agent-run calls. Human answerers made one in 2 of 33. Three answered calls opened with an agent and passed to a person, so they sit in neither group. Fisher's exact p = 0.004; the direction holds, the multiple carries wide uncertainty.

  4. The $8,500 caller got the service-call process

    0 of 5

    Five calls presented a competitor's written $8,500 replacement quote, two with an offer to email the proposal. Zero platforms asked to see it, and soonest visits ran two to ten days for a buyer deciding that weekend.

  5. A machine repeats its mistakes exactly

    16 of 55 AI-fronted

    One agent accepted an email address announced out loud as fake, and the appointment flow continued as designed. Another walked its own empty calendar into September without escalating to a human being.

On finding 01, the fair objection: some platforms keep their brands apart on purpose, for licensing or service-area reasons. If that is you, the zero is a decision rather than a gap. Either way the caller hangs up thinking the platform sells one trade, so somebody should choose which of the two is true.

The instrument · eight stages, one call

  1. Answer

    49 of 55

    Reached a live voice. Six dead ends, all at brands with live paid media that week.

  2. Qualify

    no score

    Not scored: the rubric has no clean denominator here, because what counts as a qualifying question varies by trade. Read qualitatively it was the strongest stage in the sample, though one weekend emergency still ended without a phone number taken.

  3. Price

    40 of 49

    Stated the visit fee, or free. Nine could not price a visit when asked.

  4. Paid software ends here · workflow begins

    Hold

    8 of 49

    Tried to keep the caller at the moment they deferred. Six of the eight were scripted AI agents.

  5. Capture

    13 of 49

    Asked for an email. One address announced out loud as fake was accepted.

  6. Recap

    1 of 55

    A written recap received before hang-up. The best call in the sample, and the only one.

  7. Follow-up

    2 observed

    Two promised callbacks observably arrived. The rest fell outside the study window and are unscored.

  8. Route

    0 of 55

    Live referrals to a sibling brand. Not one, in any phrasing.

Observed incidence, 55 calls placed as a customer to 48 PE-owned brands, July 2026. Denominators vary because six calls never reached a voice and some stages only trigger in certain scenarios. The first three stages run on software these platforms already bought. The last five are habits, scripts, and configuration.

The callback ledger

2 of 11 arrived

nine unresolved, outside the study window

Eleven brands promised to call me back. Two did, both same-day. Nine unobserved outcomes would make any published rate a guess wearing a percent sign, so the ledger stays a ledger and this report publishes no callback-kept rate, deliberately.

If your platform can produce its own callback-kept rate from the CRM, you can answer a question this study could not.

Every rate, with its interval

FindingObservedRate95% interval
Answered, with a person on the line49 of 5589%78–96%
Could name the price of a visit40 of 4982%68–91%
Tried to keep me when I hesitated8 of 4916%7–30%
Save attempts, agents against people46% vs 6%agents 19–75%, humans 1–20%
Asked for an email address13 of 4927%15–41%
Sent anything in writing1 of 551.8%0.0–9.7%
Offered a sister brand0 of 550%upper bound 6.5%
Second problem handled in one trip3 of 1323%5–54%
Asked to see a competitor's written quote0 of 50%0–52%
Exact binomial intervals, published on every rate this study leans on, including the two that read zero. A wide interval is the honest reading of a small denominator: five replacement-quote calls cannot carry a rate, and the interval is how you can tell. The eleven promised callbacks appear nowhere here, because nine of them fell outside the study window and no rate was computed.

Method · limitations first

You diligence companies for a living, so the weaknesses go first.

It cannot tell you close rates.
The shopper never books, so nothing in this study is a conversion claim. I measured whether each brand did the things that close callers: the ask, the hold, the capture, the recap, the follow-up.
Per-brand reads are directional.
Most brands got one call; a handful got two or three. One call proves nothing about one brand. Fifty-five calls with the same misses repeating across 48 brands and roughly 20 sponsors is a pattern.
July is peak season, and that cuts both ways.
A brand with a full board has a rational reason to let a deferring caller go, so some of what I scored is capacity rationing rather than broken intake. Slow soonest-visit dates belong in that column. The stages this study leans on do not: asking for an email, putting the visit in writing, naming the sibling brand that does the work. None of those consume a truck slot. Replacement intent sits outside the argument too, because an $8,500 install runs on install crews, and install revenue is the thesis the platform was bought on.
The economics are modeled, never measured.
No bookings were completed, so the economic section is a scenario model with disclosed assumptions. Confidence intervals are published on every rate this study leans on, including the two that read zero.
How the calls were run.
Calls went to each brand’s published consumer line between July 16 and 25, 2026, across dayparts including evenings and a weekend. A small cast of homeowner personas ran four standard scenarios, and every call was scored on the same eight stages regardless of scenario. Personas used shopper-controlled phone numbers and inboxes. No bookings were confirmed, no trucks rolled, and no free work was extracted. Brands and sponsors are anonymized throughout, and quotes are verbatim, cleaned only where the transcription engine mangled a word.

The full method detail and ethics notes are in the appendix. The waves run quarterly on the same rubric, so the zero in finding one has somewhere to move.

What to do Monday

Ten questions. Seven cost an hour of listening to your own calls.

Listen · one hour · five recorded calls per brand

  1. 01Does anyone ask for the caller's email address?13 of 49
  2. 02When the caller defers, is the next sentence a save attempt, or “no worries”?8 of 49 attempted
  3. 03Can the person answering state the price of a visit, and whether it credits against the work?9 of 49 could not price at all; another 11 calls reached someone who could only take a message
  4. 04When a second problem comes up mid-call, does it get one trip, a second fee, or “call someone else”?3 of 13 got one trip
  5. 05Does anything arrive in writing after the call?1 of 55
  6. 06When the caller holds a competitor's quote, does anyone ask to see it?0 of 5
  7. 07Call your own after-hours line tonight. Can the voice that answers quote the fee, book a job, or commit to a callback with a time on it?mostly no

Pull · three reports · formulas published in full

  1. 08Call-to-booked-job, by brand, on one definition across every brand. If the definitions differ by brand, that finding outranks the number.
  2. 09Speed to answer and true abandonment, by brand and daypart, with short abandons excluded before the rate means anything.
  3. 10Membership attach, where eligibility is written down per brand and existing members are out of the denominator.

If the seven listening checks come back clean at your platform, close the report and keep the hour. In this sample of 48 brands, none would have come back clean.

The front door is the integration test. Most of the platforms I called would not pass it today, and the person who finds that out is the customer you paid to reach.

I also run this as a two-week, fixed-fee read on your own numbers: the three pulls on your booking and phone data, a sample of your own recorded calls, and the same arithmetic done on your figures instead of mine. If you want it, reply with the name of your busiest brand and I will tell you inside a day whether your data can answer the seven questions.

Zaha Al-Hmoud

Independent consultant, Vancouver.
zaha@ztsadvisory.com

Cite as: ZTS Advisory, The Bookability Baseline, July 2026.