Independent test · July - August 2026 · 7 businesses · 4 widget platforms
15 questions each, drawn from the business's own pages · every problem re-checked in 3 separate sessions
Every business here is anonymised, and each was emailed its own results privately - no pitch, nothing to buy - before this was written. Assistant quotes are verbatim; they exist only in a chat window. Company page quotes are paraphrased, because a verbatim page quote is a search key that would identify the business in one click.
Read each business's public pages first, then build fifteen questions whose answers are already published there - prices, policies, hours, what is included. Ask as an ordinary customer, once each. Re-ask anything that looks wrong three times in separate fresh sessions: a one-off is noise, and if the answer varies then the inconsistency is itself the finding. About twenty-four messages per business, one day each. No lead forms, no signups.
Then the step that matters most: before calling anything wrong, open every link the assistant cites and re-check the live page. Four findings died there.
Asked: "What do I have to pay on day one to start a month-to-month membership?" 3 of 3
"For the month-to-month plan, you'll pay the first and last month upfront: $119 × 2 = $238 on day one. There are no hidden or annual fees."
$129.99 a month, with a matching first-and-last payment due at signup - $259.98.
The structure is right; the base price is out of date, so the total is short by $21.98 - and it shows its working, which makes the answer read as authoritative. Their FAQ page still lists membership prices that no longer match their checkout: every tier about $10 a month low, the pay-upfront option $50 low. The assistant is quoting the FAQ.
Asked: "How many visits a year does the residential maintenance plan include?" 3 of 3
"the Residential Preventative Maintenance Plan includes quarterly visits, which means 4 visits per year."
Residential is twice a year. Quarterly is the commercial programme, described a few paragraphs down the same page in near-identical language.
The two programmes differ by one word and the assistant collapsed them. It repeated the figure unprompted when asked about cost. A homeowner signs up expecting four visits, gets two, and finds out at the visit that never comes - holding the transcript.
Asked three times, on three different pages, in two wordings. Every time: 3 of 3
"I wasn't able to find specific opening hours in our knowledge base." - and then printed their phone number and street address.
In the header of every page of the site, in the same block as that phone number and address, a few inches above the chat window.
It has the block. It cannot find the hours inside it. Two phrasings, so it is not a matter of how you ask. "When are you open" is the most common question a service business gets, and every customer who asks is sent to the phone.
Asked: "Are taxes included in the nightly rate, or added on top?" 3 of 3
Tax of 11.5% is added on top; "the nightly rates shown on our website are before tax."
All taxes are already included, no extra charges. Their policies page says the opposite - and agrees with the assistant.
The purest case in the set: the assistant is right, according to one of the two pages. It picked the one the guest was not reading. A second finding on the same site - the FAQ says most properties have heated pools, and the assistant promises every guest that all of them do.
Asked: "How cold can it get before I can't use the battery? Not charging, just driving." 3 of 3
You can discharge down to -4°F (-20°C), and a temperature sensor shuts the battery down below that.
Discharge down to approximately -20°F - sixteen degrees colder than the assistant claims.
This looked like a unit-conversion slip. It probably is not: -20°C / -4°F is the standard published discharge floor for this battery chemistry, stated on the FAQ pages of at least two large competitors. The assistant is not reciting its client's page - it is answering from what it knows about the category, and overriding a page that claims something better.
This one cannot be fixed by tidying a website. The assistant understates its own client's product by sixteen degrees in the exact specification a cold-climate buyer shops on, using a number the company never published. It was also the most accurate assistant tested - fourteen of fifteen - which did not protect it where it mattered commercially.
A New England inn: fifteen of fifteen. Check-in times, the cancellation window and its per-room fee, pet weight limits and nightly charges, the smoking fine, the rollaway charge. Asked for a rate on a specific Saturday night, it said it had no access to real-time pricing and pointed at the booking page - the correct answer, and the one most of these assistants get wrong by guessing.
Its documentation agrees with itself. That is the whole difference, and it is what makes this a control rather than an anecdote.
| What looked wrong | What was actually true |
|---|---|
| A pet business quoting a fee for an extra animal that was not on the pricing page | Published - different page, different service |
| An inn offering a named package on no policy page, whose name other hotels also use | Published verbatim on the page the assistant's own citation linked to |
| A retailer saying you cannot buy directly from it, on what is plainly an online store | Correct - product pages list prices but have no cart, only a dealer locator |
| A gym quoting a day-pass price found nowhere on its site or checkout | Correct - that exact pass exists in the checkout. My search could not see it because the price sat behind markup |
The last is the instructive one. It survived the rule "open every link the assistant cites" - I did open it - and died only on a second, more careful read of the same page, hours before it would have been sent to the owner. Opening the page is not enough; you have to be sure your method can actually read it.
Seven businesses is seven businesses, and they were not randomly sampled - they were found by looking for small-business sites running chat widgets, which selects for a certain kind of operator. One person ran every test. No claim is made about how common this is, and any percentage extrapolated from it has gone past the evidence.
It is also not a criticism of the widget vendors. Four platforms are represented here, all behaved correctly, and all answered from the material they were given. The finding is about the material. What the test does show is a failure mode that is legible and repeatable - each one reproduced three times in fresh sessions, and in five of six cases traceable to a specific page the owner could fix this afternoon.
Take the five questions customers ask before buying - price, what is included, hours, cancellation, and the one specific to your trade - and ask your own assistant in a fresh incognito window. Check each answer against your own page.
When an answer is wrong, the useful question is not "why is the AI wrong." It is "which of my pages says that?" Five times out of six, one of them does.