Field study // 55 mystery-shop calls // July 2026
The Front Door Integration Test
Where PE-backed home services platforms prove, or fail, the thesis they paid for
The deal model assumes shared demand, cross-sell between sibling brands, and one customer base served by many flags. The place a homeowner would first experience any of that is the phone. So I called it, 55 times, at 48 brands under roughly 20 sponsors, and scored every call against one eight-stage rubric.
29 pages · free · no email required
Exhibit · The cliff, observed
From forty-nine answered calls to zero referrals, in four steps.
Each stage carries its own denominator, because a stage only triggers when the call creates its condition. Read them as four separate measurements, never as one funnel.
-
49/55 89%
Answered by a live voice
The phone system did its job. Six calls dead-ended, all at brands running live paid media that week.
-
8/49 16%
Tried to hold the lead
A save attempt at the moment I deferred. Six of the eight were scripted AI agents.
-
1/55 1.8%
Put anything in writing
A recap received before hang-up. 36 of 49 answered calls never asked for an email address.
-
0/55 0%
Referred a sibling brand
Not once, in any phrasing. Three of the 13 volunteered second problems got one-trip handling; the rest drew a second visit, a second fee, or a competitor's name.
Stage results: answered by a live voice 49 of 55; tried to hold the lead 8 of 49; put anything in writing 1 of 55; referred a sibling brand 0 of 55.
Observed July 2026. The first number runs on software the platform already bought; the next three are workflow. Denominators vary because six calls never reached a voice, and some stages trigger only when the call creates their condition.
Exhibit 5 · Four sibling brands, one platform, 48 hours
The fee card shipped. The habits did not.
Four consumer HVAC brands under one sponsor, shopped inside 48 hours with the same scenario. The diagnostic fee was identical at all four. Everything after the fee was local habit.
| Brand R1 | Brand R2 | Brand R3 · best door | Brand R4 · worst door | |
|---|---|---|---|---|
| Diagnostic fee | $169, member discount offered | Same | Same | Same |
| Soonest visit | Tomorrow morning | Friday | Same evening | “Nothing after 4 p.m. for the foreseeable future” |
| Second problem | Bookable, second trade, second fee | “You’d have to call a plumber” | Bookable, separate appointment and fee | Declined, no referral |
| Membership | Scripted pitch with terms and savings math | Explained on request | Offered with terms | Never mentioned |
| Send it in writing? | No; waited while I found a pen | “No way to send you the information” | Texted a recap during the call | “Check the website” |
| Contact captured | Name, address, phone | Address only | Phone | Nothing |
“I don’t have any way to send you the information, but you can check it out on our website.”
Interpretation The fee card is a configuration artifact; the recap is a habit. Configuration shipped, habit did not, and the spread inside one platform on identical software measures how much standardization actually landed.
The framework
Operating partners do not manage eight stages. They manage five controls.
Every stage, finding, and fix in the study maps to exactly one of them. Three of the five were close to absent in this sample.
01 · Reach
Can the caller reach someone who can price, schedule, or commit to a next step? Answering is necessary and insufficient.
49/55 answered · but 11 calls reached a voice that could neither price nor book02 · Recognize
Does intake recognize what walked in: urgency, replacement intent, membership eligibility, a second service need?
The emptiest layer in the sample03 · Retain
Does the platform secure a next step before the call ends — a save at deferral, contact capture, a written recap, a follow-up with a time on it?
8/49 held · 13/49 captured · 1/55 recapped04 · Route
Does the caller reach the correct estimator, trade, sibling brand, or escalation path?
0/55 sibling referrals · and the agents' unescalated dead loops05 · Review
Does management measure call-to-booked-job on one definition across brands? Review is the control that makes the other four durable.
No platform could have produced its AI agent's booked-job rate on the human definitionFinding 04 · Recognize
The highest-value call in the trade, run through the $99 process.
-
01
Nobody asked to see the quote
Five designed calls presented a competitor’s written $8,500 full-system replacement quote, two with an explicit offer to email the proposal on the call. Zero platforms asked to see it. None asked what had been quoted, by whom, or when the decision would close.
0/5Observed -
02
The clock never changed
Soonest offered visits ran two to ten days, where a date existed at all, for a buyer who announced out loud that she was deciding that weekend. High-value intent should change the queue, the questions, and the clock. It changed none of the three.
2–10 daysObserved -
03
Fifteen to twenty times the ticket, same queue
A full replacement at $8,500 against service jobs of roughly $432, both quoted to me on these calls. The larger job travelled through the same queue, the same questions, and the same calendar as a $99 no-cool.
15–20×Observed
Observed With five trials, a true ask rate as high as half the time is still consistent with observing zero (exact 95% CI: 0% to 52%). Treat this as a hypothesis with five supporting trials and no contradicting ones — and the easiest one in the report to test in your own data.
“I was hoping to get an answer before Monday because I’m wanting to make up my mind this weekend about the purchase. So, I guess we’re out of luck here.”
Findings 03 and 05 · The configured layer
The machines sold better. They also failed identically, every time.
Sixteen of 55 calls were fronted by a scripted AI agent, 13 of them end to end. Both halves of what follows are configuration facts, which is why one compounds and the other scales.
-
01
Scripted agents attempt the save more consistently
At the moment the shopper deferred, agents made a live save attempt in 6 of 13 agent-run calls. Human answerers made a comparable attempt in 2 of 33. Fisher’s exact p = 0.004, odds ratio 13.3. The direction is well supported; the confidence intervals are wide on both sides (19–75% for agents, 1–20% for humans), so treat the roughly sevenfold multiple as direction with magnitude uncertainty.
6/13 vs 2/33Observed -
02
Nobody could tell me what their agent booked
Not one platform in the sample could state the AI agent’s call-to-booked-job rate on the human definition, on the observed evidence. The line item is approved and spent. The measurement that would tell anyone whether it worked was not attached to it.
0 platformsObserved -
03
Configuration errors execute identically on every call
An agent gated scheduling behind an email address, then accepted one announced out loud as fake. Another confirmed a town roughly 150 miles from the one I stated and moved on; a truck dispatched on that record burns a slot and a customer. A third looped through its own empty calendar, month after month into September, while a live replacement buyer waited, and never escalated to a human being.
16/55 AI-frontedObserved
“Would it help to just go ahead and get us on the schedule as a free backup comparison?”
The two human saves in the sample deserve naming, because they were the best selling observed anywhere in 55 calls. One rep held that evening’s slot with text-to-cancel and no commitment. Another waived the visit fee unprompted, and offered to book so I wouldn’t have to do the process all over again. Both moves are one sentence long. Neither appeared anywhere else.
Interpretation The behaviour that decides these calls is one sentence long, and the machines have it configured. Configuration compounds and coaching decays, so the move is to route better at-bats to the humans and instrument the agents like any other CSR — inside systems the platform already runs.
Exhibit 8 · Illustrative economic exposure
A scenario model with disclosed assumptions. Drag it until it breaks.
An earlier draft of this study circulated with $1.8M to $2.1M of install EBITDA “at risk.” That was the full expected pipeline riding on the observed handling, and full pipeline is the wrong loss estimate: mishandling does not zero a close rate. This replaces it.
$409K
EBITDA before implementation cost · $334K net at $75K implementation
Below 9,200 demand calls a year, this intervention does not clear its own implementation cost. That is a real result, and it is the honest reason not to run it at every platform.
| Scenario | Revenue lift | EBITDA before cost | Net of implementation |
|---|---|---|---|
| Low · 3 pts, 25% | $637K | $159K | $84K |
| Base · 7 pts, 27.5% | $1.49M | $409K | $334K |
| High · 12 pts, 30% | $2.55M | $765K | $690K |
The lift assumption already embeds recoverability: it prices only the incremental installs improved handling wins, never the full pipeline. The other four leaks run through the same formula with smaller tickets. Do not sum them — the populations overlap and a summed figure double-counts. Rank them, pilot the top one, and let measured lift replace modeled lift line by line.
Your ServiceTitan call reasons replace the 5 percent. Your pilot replaces the lift. Your finance team replaces the margin. Once those three land, the model stops being mine.
What to do Monday
Ten questions. Seven cost an hour of listening to your own calls.
Listen one hour · five recorded calls per brand
Does anyone ask for the caller’s email address?
This sample: 13 of 49When the caller defers, is the next sentence a save attempt, or “no worries”?
This sample: 8 of 49 attemptedCan the person answering state the price of a visit, and whether it credits against the work?
This sample: 9 of 49 could notWhen a second problem comes up mid-call, does it get one trip, a second fee, or “call someone else”?
This sample: 3 of 13 got one tripDoes anything arrive in writing after the call?
This sample: 1 of 55When the caller holds a competitor’s quote, does anyone ask to see it?
This sample: 0 of 5Call your own after-hours line tonight. Can the voice that answers quote the fee, book a job, or commit to a callback with a time on it?
This sample: mostly no
Pull three reports · formulas published in full
Call-to-booked-job, by brand, on one definition across every brand. If the definitions differ by brand, that finding outranks the number.
Speed to answer and true abandonment, by brand and daypart, with short abandons excluded before the rate means anything.
Membership attach, where eligibility is written down per brand and existing members are out of the denominator.
If the seven listening checks come back clean at your platform, close the report and keep the hour. In this sample of 48 brands, none would have come back clean.
The engine room
Limitations first.
- Observed
- Directly found in the 55-call field study. Comes with a numerator and a denominator, always.
- External benchmark
- Found in a cited external dataset or authoritative publication. Vendor research is flagged as vendor research.
- Interpretation
- A reasoned explanation of the evidence, with the competing interpretation stated next to it.
- Modeled
- Calculated using disclosed assumptions. Every input is observed, benchmarked, or flagged for replacement.
What it cannot prove
- Close rates. The shopper never books, so nothing here is a conversion claim.
- Any per-brand read. Most brands got one call; a handful got two or three. One call proves nothing about one brand.
- Industry representativeness. Brands were not drawn randomly from a defined population. This is a structured control-failure detection study.
- Causation. No claim that a missed behaviour caused a lost job.
How the calls were run
- Published consumer lines, US and Canada, July 16 to 25, 2026, across dayparts including evenings and a weekend.
- Four scenario families, a small cast of homeowner personas, shopper-controlled phone numbers and inboxes.
- Personas deferred at the close by design — without that, the hold, capture, recap and follow-up stages never trigger.
- No bookings confirmed, no trucks rolled, no free work extracted. Brands and sponsors anonymized. Quotes verbatim, cleaned only of transcription artifacts.
- Eleven calls ended with a promised callback; two arrived, both same-day, both observed. The rest fell outside the study window, so no callback-kept rate is published.
External benchmarks
- Cross-sell. McKinsey surveyed 200 executives who had led revenue-synergy programs on deals above $2B: most fell short of their aspiration, by an average of 23 percent. Separate McKinsey work puts cross-selling at roughly 20 percent of revenue-synergy value.
- Phone conversion. Invoca’s 2025 analysis of 60M+ calls found home services converting 46 percent of phone leads during the call, against a 37 percent cross-industry average. Vendor research, disclosed sample, and its definitions differ from booked jobs.
- Speed. The Oldroyd/InsideSales work and the 2011 HBR audit of 2,241 firms found response inside an hour multiplies qualification odds roughly sevenfold. Those measured web and B2B-heavy samples; the decay shape is the transferable lesson, not the multiplier.
- Governance. NIST’s AI Risk Management Framework and its Generative AI Profile (NIST-AI-600-1, July 2024), used as governance and never as a conversion benchmark.
Statistical uncertainty
- Live sibling referral, 0/55 — exact 95% CI 0.0% to 6.5%. Informative even at zero events: if the true rate were even 7 percent, seeing none in 55 tries would be unlikely.
- Written recap, 1/55 — 1.8%, CI 0.0% to 9.7%.
- Save attempt at deferral, 8/49 — 16%, CI 7% to 30%.
- Asked to see the competitor’s proposal, 0/5 — CI 0% to 52%. Suggestive only, and stated as a hypothesis rather than a rate.
- Agents versus humans, 6/13 against 2/33 — Fisher’s exact p = 0.004, odds ratio 13.3.
- Single-coder scoring is a limitation. A published supplement should add the scripts, distribution tables, scoring codebook, de-identified call-level data and second-coder agreement.
ZTS Advisory·Field study 01·July 2026
The front door is the integration test. Most platforms in this sample have not yet taken it, and the caller can tell.
ZTS runs revenue-capture work inside PE-backed residential service platforms, hands-on in ServiceTitan, HubSpot, and the CCaaS layer on top of them. This study points the same method outward. ztsadvisory.com