Field study // 55 mystery-shop calls // July 2026

The Front Door Integration Test

Where PE-backed home services platforms prove, or fail, the thesis they paid for

Calls
55
Brands
48
Sponsors
~20
Method
One rubric

29 pages · free · no email required. The board exhibit is the same study on two printable pages: send it to whoever owns intake, and the seven listening checks cost them an hour.

In July 2026 I placed 55 mystery-shop calls to 48 residential service brands, HVAC, plumbing and roofing, owned by roughly 20 private-equity sponsors. Same method, same scoring rubric, every call transcribed and scored the same way. I called the number each brand’s marketing budget exists to make ring, and I presented as the caller every one of these businesses says it wants: a homeowner with a real problem, money to spend, and two or three companies on the list.

The phones mostly get answered, and the selling stops there. The layer between the answered call and the booked job does not exist as a system. What I found instead were scattered individual habits: one written recap, two human save attempts, in 55 calls. Administrative integration had shipped at most of these platforms. Customer-facing integration mostly had not, and the caller is the only person positioned to notice.

Exhibit · The cliff, observed

From forty-nine answered calls to zero referrals, in four steps.

  1. 49/5589%

    answered by a live voice

    The phone system did its job.

    Paid software ends here

  2. 8/4916%

    tried to hold the lead

    At the moment I deferred. Six were machines.

  3. 1/551.8%

    put anything in writing

    A recap received before I hung up.

  4. 0/550%

    referred a sibling brand

    Including all 13 calls where I volunteered a second problem.

Observed July 2026. The first number runs on software the platform bought; the next three are workflow. Denominators vary because six calls never reached a voice, and some stages trigger only when the call creates their condition.

“We only handle hot water tanks or tankless water heaters, but I don’t have any plumbers on staff for anything else. So you’d have to call a plumber for that, unfortunately.”
CSR at an HVAC brand whose platform sells plumbing under other flagsVerbatim · Finding 01 · July 2026

The five findings

Each one reads standalone, with its denominator in view.

  1. The cross-sell thesis lacks front-door machinery

    In 55 calls, including 13 where I volunteered a second problem in the home unprompted, no human or AI agent made a live referral to a sibling brand in any phrasing. Three of the 13 got one-trip handling; the rest drew a second visit, a second fee, or a competitor's name.

    0/55Observed
  2. Same platform, four different front doors

    Four sibling brands under one sponsor, shopped inside 48 hours with the same scenario, held an identical $169 diagnostic fee and diverged on everything after it. One texted a recap mid-call. One captured no contact detail at all.

    n = 4 brandsObserved
  3. Scripted agents attempt the save more consistently

    At the moment I deferred, scripted AI agents made a save attempt in 6 of 13 agent-run calls. Human answerers made one in 2 of 33. Fisher's exact p = 0.004; the direction holds, the multiple carries wide uncertainty.

    6/13 vs 2/33Observed
  4. Replacement intent is not recognized at intake

    Five calls presented a competitor's written $8,500 replacement quote, two with an offer to email the proposal. Zero platforms asked to see it, and soonest visits ran two to ten days for a buyer deciding that weekend.

    0/5Observed
  5. AI scales configuration, including configuration errors

    One agent accepted an email address announced out loud as fake, and the appointment flow continued as designed. Another walked its own empty calendar into September without escalating to a human being.

    16/55 AI-frontedObserved

On finding 01, the fair objection: plenty of sponsors keep brands apart on purpose, for pricing segmentation or licensing lines across trades and states. Where the separation is deliberate, the zero is a strategy choice and this study prices what it costs. Where it isn’t, the zero is the integration gap showing up at the front door. You know which one you are running; the caller hears the same thing either way, which is that the platform sells one trade.

The instrument

Eight stages, one call. The selling stops at stage four.

Exhibit · The eight-stage intake rubric, observed Observed
  1. Answer

    49/55

    Reached a live voice. Six dead ends, all at brands with live paid media that week.

  2. Qualify

    no score

    Not scored: the rubric has no clean denominator here, because what counts as a qualifying question varies by trade. Read qualitatively it was the strongest stage in the sample, though one weekend emergency still ended without a phone number taken.

  3. Price

    40/49

    Stated the visit fee, or free. Nine could not price a visit when asked.

  4. Paid software ends here · workflow begins

    Hold

    8/49

    Tried to keep the caller at the moment they deferred. Six of the eight were scripted AI agents.

  5. Capture

    13/49

    Asked for an email. One address announced out loud as fake was accepted.

  6. Recap

    1/55

    A written recap received before hang-up. The best call in the sample, and the only one.

  7. Follow-up

    2 obs.

    Two promised callbacks observably arrived. The rest fell outside the study window and are unscored.

  8. Route

    0/55

    Live referrals to a sibling brand. Not one, in any phrasing.

Observed incidence, 55 mystery-shop calls to 48 PE-owned brands, July 2026. Denominators vary because six calls never reached a voice and some stages only trigger in certain scenarios. The first three stages run on software these platforms already bought. The last five are habits, scripts, and configuration.

What to do Monday

Ten questions. Seven cost an hour of listening to your own calls.

Listen one hour · five recorded calls per brand

  1. Does anyone ask for the caller's email address?This sample: 13 of 49
  2. When the caller defers, is the next sentence a save attempt, or “no worries”?This sample: 8 of 49 attempted
  3. Can the person answering state the price of a visit, and whether it credits against the work?This sample: 9 of 49 could not
  4. When a second problem comes up mid-call, does it get one trip, a second fee, or “call someone else”?This sample: 3 of 13 got one trip
  5. Does anything arrive in writing after the call?This sample: 1 of 55
  6. When the caller holds a competitor's quote, does anyone ask to see it?This sample: 0 of 5
  7. Call your own after-hours line tonight. Can the voice that answers quote the fee, book a job, or commit to a callback with a time on it?This sample: mostly no

Pull three reports · formulas published in full

  1. Call-to-booked-job, by brand, on one definition across every brand. If the definitions differ by brand, that finding outranks the number.
  2. Speed to answer and true abandonment, by brand and daypart, with short abandons excluded before the rate means anything.
  3. Membership attach, where eligibility is written down per brand and existing members are out of the denominator.

If the seven listening checks come back clean at your platform, close the report and keep the hour. In this sample of 48 brands, none would have come back clean.

ZTS AdvisoryField study 01July 2026

The front door is the integration test. Most platforms in this sample have not yet taken it, and the caller can tell.

ZTS runs revenue-capture work inside PE-backed residential service platforms, hands-on in ServiceTitan, HubSpot, and the CCaaS layer on top of them. This study points the same method outward.ztsadvisory.com