XAPP Labs · Longitudinal Study

How AI Surfaces Home Service Businesses: The Booking Gap

Ten months into the longitudinal study we introduced in October 2025 — discovery went native; booking didn’t move.

What do things look like now and what has changed since October 2025


The Big Picture

  • What we found: Our first results last October already showed AI assistants surfacing local home-service businesses on profile strength, not ad spend. A year on, that surfacing has gone fully native — mapped, rated, and increasingly wrapped in a full profile card — and the assistants now qualify the job like a concierge. What none of the five we tested can do is the last step: actually book the appointment.
  • Why it matters: Discovery is no longer the bottleneck; booking is. The homeowner arrives qualified and ready, and the conversation dead-ends at the calendar — leaving the most valuable step in the funnel unowned.
  • What’s next: We keep tracking the panel as native discovery evolves, and we watch for the first real booking integrations — the moment an assistant can close the loop instead of handing the homeowner back the phone.

Update

When we launched this study last fall, the question was still open: as homeowners move from search engines to AI assistants, would those assistants actually surface local home-service businesses — and if so, on what basis? Our first results, last October, answered it: yes, they surface them, and they do it on the strength of a business’s profile and proximity, not its ad budget. This round widened the panel to five assistants and three metros — and the picture both sharpened and moved.

The short version: the assistants have gotten good — genuinely good — at finding and vetting local businesses, and far richer at it than a year ago. They have made almost no progress at booking them. The open question in October was discovery; discovery is now largely solved, and the whole story has moved one step down the funnel — from can AI find you to can it book you.

What we tested this round

Following the framework from the original post, we ran homeowner scenarios across the three use cases we defined: direct service intent (“who should I call”), symptom-to-professional triage (“something is broken — what now”), and project planning (“I’m considering a big replacement”). This cycle sampled five of the six assistants in our panel — ChatGPT, Claude, Gemini, Perplexity, and Microsoft Copilot (Grok returns next round) — alongside Google Search and Google Maps as the search baseline for comparison.

We tested in three metros — Washington, D.C. (our original geography), plus Dallas–Fort Worth and Tampa — across windows, HVAC, and plumbing. We ran both cold prompts and hot, back-and-forth conversations, and we reproduced key results in memory-off sessions to make sure we were observing default behavior rather than personalization. These are qualitative observations, not a statistical panel; we report them as such.

Finding 1: Discovery is now native

A year ago, asking an assistant to “find a plumber near me” often produced a text list assembled from the open web. Today, ChatGPT answers a query like “my water heater is leaking in Washington, D.C. — who should I call?” by rendering its own interactive map of nearby, rated businesses — Magnolia Plumbing, Haynes Plumbing, and others — directly inside the conversation. Ask about whole-house window replacement near Dallas and you get a map of local installers with star ratings and “open now” status. Ask about a fifteen-year-old AC in Tampa and you get the same, tuned to HVAC.

Figure 1 — ChatGPT renders its own interactive map of rated local businesses (Dallas window installers) directly in the conversation.

This is a meaningful shift. The assistant is no longer just describing businesses; it is presenting a curated, ranked, mapped set of them as a first-class part of the answer, sourced from across the web and community discussion. For the “symptom triage” use case we flagged last year as AI’s natural strength, it now pairs that map with genuinely useful safety guidance — for the leaking water heater, it walked through shutting off the supply and gas-safety steps before listing 24/7 options.

Finding 1, continued: the map is becoming a business profile

Since our first post, ChatGPT’s native map grew a second half. Tap one of the pinned businesses and a knowledge-panel–style profile opens alongside the chat — logo, star rating, hours (“Open until 12:00 AM”), full address, phone, a short sourced description, and a categorized list of services, with Directions, Website, and Call actions. OpenAI began surfacing these local knowledge panels around the end of 2025; a year ago the assistant returned a list, then a map, and now it returns a profile.

In our Tampa HVAC session, selecting Total Air Solutions produced exactly that: a 4.9-star card citing 380-plus reviews, the Anderson Road address and an 888 number, an “around-the-clock emergency heating and cooling” description carrying inline source citations, and a Services list broken out into installation, replacement, repair, and maintenance.

Figure 1b — ChatGPT’s business knowledge panel for a Tampa HVAC company: structured fields (rating, hours, address, phone) beside semi-structured, cited attributes (description, services).

Two things are worth noting. First, the panel blends structured and semi-structured data: the rating, hours, address, and phone read like licensed place data, while the description and the service list are assembled from the open web and shown with citations — the assistant is composing a single canonical profile out of several sources. Second, that profile appears to lean heavily on the Google Business Profile: independent observers report changing a field in their GBP and seeing it reflected in ChatGPT shortly after. OpenAI hasn’t documented how the panel is built, so we report it as an observation rather than an architecture — but the direction is clear enough to matter. The assistant is no longer just pointing at businesses; it is presenting a profile it assembles about each one, and the inputs to that profile — your Google Business Profile, your website, your reviews — are the levers you actually control.

And notably, for all its richness, the panel still can’t book. Its actions are Directions, Website, and Call — the same handoff back to the phone or the website we describe below. A fuller profile, the same missing last step.

Finding 2: So is qualification

The assistants don’t stop at a list. They increasingly behave like a concierge: recommending which two or three companies to start with, explaining the trade-offs, and asking the qualifying questions a good intake agent would. In our HVAC and windows sessions, ChatGPT volunteered a plan (“get two to three in-home estimates”), then asked for location, project scope, budget tier, and preferred appointment times — unprompted.

Figure 2 — The same chat behaves like an intake agent, asking concierge-style qualifying questions before it recommends.

In other words, the two hardest-sounding parts of connecting a homeowner to a pro — finding the right businesses and understanding the job — are now handled competently by the assistant itself, at no cost to the homeowner.

Finding 3: And then everything stops at the booking

Here is where the year’s progress ends abruptly. In every scenario, across every assistant we tested and all three metros, the moment the homeowner asked to actually book, the assistant hit a wall — and said so plainly.

ChatGPT: “I can’t actually make appointments or pretend to be you by calling or submitting booking forms.” And, tellingly: “you’ll need to use a service or integration that supports booking.” Even for the leaking-water-heater emergency — the case where speed matters most — it declined, noting it could not submit a booking “without an appropriate connected service.”

Figure 3a — ChatGPT’s booking wall, in its own words.

Claude reached the same place from a different direction. It searched the web, listed companies with their own quote forms, and concluded: “I don’t have a direct booking integration for any of these.” When we let it search its connector registry for a booking tool, it came back with nothing suitable — the nearest match was a general task-marketplace it dismissed as wrong for the job. Its fallback was to offer to email the companies on the homeowner’s behalf and wait for a callback.

Figure 3b — Claude searches its connector registry for a booking tool and finds nothing suitable.

The pattern is consistent enough to state as a finding: today’s assistants can take a homeowner from problem to a short list of qualified, ranked local businesses — and then hand the homeowner back the phone. The last step, the one that actually produces a scheduled appointment, is missing everywhere we looked.

Widening the panel did not change this. Gemini narrowed to three companies and pointed to their websites; Perplexity listed candidates and told the homeowner to “call their estimate lines directly.” The one different note came from Microsoft Copilot, which offered to launch an autonomous “task” to research installers and schedule an appointment on the homeowner’s behalf — the first sign of an agentic attempt to close the gap. But it was a proposal, not a completed booking, and an agent that fills forms or places calls still runs into the same missing piece: nothing on the business’s side is built to accept the appointment. The gap is not that the assistants lack ambition to book; it is that there is no rail for them to book through.

Figure 3c — Microsoft Copilot proposes an autonomous “task” to research and schedule an appointment — the agentic attempt to close the gap.

Finding 4: The marketplaces are in the room, but not closing the gap

Since our first post, the large home-services marketplaces have arrived inside ChatGPT — Angi launched an app this spring, and Thumbtack partnered with OpenAI. But in our testing they did not change the outcome. Both operate as discovery-and-referral: they surface options and then send the homeowner out to their own site to continue. Neither completed a booking in the conversation, and neither surfaced automatically — even when we named one directly, the assistant used its own native results and treated the marketplace as a citation rather than handing off to it. In the app directory, the only home-services entries are those two marketplaces; searches for specific trades like “windows” return nothing. The booking step remains unowned.

Figure 4a — Angi in the ChatGPT app directory: install-only, not enabled by default. Figure 4b — a directory search for “windows” returns no results; no vertical-specific finder exists.

Finding 5: Why these businesses — and how it compares to search

A year ago we noted the assistants seemed to weight proximity heavily. That holds — but the deeper answer to “why these businesses” turns out to be: it depends on which retrieval layer the assistant reaches for, and the assistants switch between two very different ones.

There is a places layer — the Google Business Profile world of ratings, review counts, category, and proximity — and a web layer — organic results, ads, and “best of” listicles. Answering from the places layer returns well-rated local firms ranked much like Google Maps, indifferent to advertising. Answering from the web layer returns the brands that win organic search and buy ads. The two produce nearly non-overlapping lists — and, importantly, the same assistant will use either one depending on the query.

We saw both, across three metros:

  • Dallas / windows. ChatGPT answered from the places layer (Maverick Windows, Performance Windows Dallas, The Window Source of DFW). Claude, for the identical prompt, answered from the web layer (Brennan Enterprises, Renewal by Andersen, Window World, ZINTEX) — almost exactly Google’s paid-and-organic leaderboard. The only business both named was Maverick.
  • Tampa / HVAC. Here Claude switched to the places layer, rendering a Google-Maps-style map (The Comfort Authority, 3H AC, Kenny’s, REM) whose names line up with Google Maps’ own top results. ChatGPT also used the places layer but skimmed the very-highest-rated firms (Total Air Solutions, One Way Air at 5.0, Eskimo Bros), which don’t sit in Maps’ default top ranking.
  • Washington, D.C. / plumbing. Both ChatGPT and Claude used the places layer here, and they converged — Haynes Plumbing, District Capital Services, and Magnolia appeared in both, and in Google’s local pack.

Widening the panel reinforced that the split is about mechanism, not brand loyalty. For the same Dallas windows question, Gemini answered purely from the web layer (Renewal by Andersen, Window World, EcoView — the listicle-and-ad brands), while Perplexity showed both at once: a Places module of local shops beside a web-cited recommendation.

Two conclusions held across the panel. First, paid-search placement does not transfer to the places layer. The ad-only local players — Brennan and West Shore in Dallas, Alvarez and Cyprus Air in Tampa — never surfaced there. That is not the same as saying big brands can’t appear: Roto-Rooter showed up in D.C. on the strength of a genuine organic profile (nearly 3,000 reviews at 4.8 stars), and national advertisers like Renewal by Andersen and Window World reliably win the web layer — because it rewards the same SEO and content footprint their advertising rests on. What doesn’t carry over is buying your way in; the places layer runs on profile quality, and the web layer on content authority. Second, there is no single “AI ranking.” Visibility depends on which layer the assistant taps for a given phrasing, and a firm strong in one layer can be invisible in the other.

Our original observation about proximity was right but partial: proximity and ratings govern the places layer, while brand and content authority govern the web layer — and which one a homeowner encounters is, increasingly, the assistant’s choice, not the homeowner’s.

Figure 5a — Google Maps’ local pack for the Dallas windows query (the places layer). Figure 5b — Google Search ads and organic for the same query (the web layer). Figure 5c — Claude’s web-layer answer, naming the same ad/SEO brands Google Search surfaces and almost none of what the map shows.

What this means

The worry that first motivated this study — being invisible to AI — has receded, and our October results already showed it easing: the assistants actively surface local pros, with ratings and context, as the default answer to a home-service question. The new risk — and the new opportunity — sits one step later. The homeowner is now fully qualified, sitting in a chat, ready to book, and the conversation dead-ends because nothing connects the assistant to the business’s calendar.

That gap is not a small UX detail. It is the difference between a warm, ready-to-hire homeowner and a kept appointment. The assistants have told us, in their own words, exactly what is missing: a service that lets them complete the booking. Whoever supplies that — a bookable rail the assistants can call, backed by businesses whose agents can actually accept the appointment — will own the part of this funnel that the last year of progress left untouched.

What’s next

The next round completes the six-assistant panel (adding Grok), restores the full search baseline (Bing Search and Bing Maps alongside Google), adds repeat runs to measure how quickly the native-discovery experience is changing, tracks how the new business knowledge panels evolve — which fields they carry, which sources they cite, and how closely they mirror the Google Business Profile — and begins tracking the first booking integrations as they appear, because if the last year was about discovery becoming native, the next one is likely to be about who closes the loop. We’ll keep publishing what we see.

As in the original study, we publish high-level patterns here; the raw transcripts and full dataset remain proprietary and are available only on a limited basis to research partners.

Methodology note: qualitative, in-product testing on live consumer accounts across ChatGPT, Claude, Gemini, Perplexity, and Microsoft Copilot, July 2026, in Washington D.C., Dallas–Fort Worth, and Tampa, across windows, HVAC, and plumbing. Assistant behavior varies by account, model version, and platform experimentation, and will continue to shift; we reproduce key results in memory-off sessions and report observations as directional rather than statistical.