What seven small-business chatbots got wrong

Independent test · July - August 2026 · 7 businesses · 4 widget platforms
15 questions each, drawn from the business's own pages · every problem re-checked in 3 separate sessions

7
businesses tested
6
had problems
13
separate findings
1
completely clean
Summary
  1. Six of seven assistants gave answers that contradict their own company's published pages. Thirteen findings in total.
  2. They are mostly not inventing things. Of the thirteen findings, only one was a figure with no traceable source anywhere.
  3. Five of the six failures trace back to the company's own documentation - a FAQ contradicting a policy page, a price list contradicting the live checkout. The assistant reads both, picks one, and states it with total confidence.
  4. One failure did not. That company's documentation is clean; the assistant answered from what it knows about the product category and overrode the client's page. This is the harder problem, because tidying your website does not fix it.
  5. The one clean business has documentation that agrees with itself throughout - fifteen of fifteen correct, including honestly refusing to quote a price it did not have. That is the control case.
  6. The costliest single answer: a gym's assistant tells prospects they owe $238 on day one. The checkout charges $259.98.
  7. Four findings died in verification and are listed at the end. One of them hours before it would have been sent to a business owner.

Every business here is anonymised, and each was emailed its own results privately - no pitch, nothing to buy - before this was written. Assistant quotes are verbatim; they exist only in a chat window. Company page quotes are paraphrased, because a verbatim page quote is a search key that would identify the business in one click.

Method

Fifteen questions the pages already answer

Read each business's public pages first, then build fifteen questions whose answers are already published there - prices, policies, hours, what is included. Ask as an ordinary customer, once each. Re-ask anything that looks wrong three times in separate fresh sessions: a one-off is noise, and if the answer varies then the inconsistency is itself the finding. About twenty-four messages per business, one day each. No lead forms, no signups.

Then the step that matters most: before calling anything wrong, open every link the assistant cites and re-check the live page. Four findings died there.

A US strength gym

Day one costs $259.98. It says $238.

Asked: "What do I have to pay on day one to start a month-to-month membership?" 3 of 3

The assistant said

"For the month-to-month plan, you'll pay the first and last month upfront: $119 × 2 = $238 on day one. There are no hidden or annual fees."

The checkout charges

$129.99 a month, with a matching first-and-last payment due at signup - $259.98.

The structure is right; the base price is out of date, so the total is short by $21.98 - and it shows its working, which makes the answer read as authoritative. Their FAQ page still lists membership prices that no longer match their checkout: every tier about $10 a month low, the pay-upfront option $50 low. The assistant is quoting the FAQ.

A US HVAC contractor

It sells the home plan as the commercial one

Asked: "How many visits a year does the residential maintenance plan include?" 3 of 3

The assistant said

"the Residential Preventative Maintenance Plan includes quarterly visits, which means 4 visits per year."

Their homepage says

Residential is twice a year. Quarterly is the commercial programme, described a few paragraphs down the same page in near-identical language.

The two programmes differ by one word and the assistant collapsed them. It repeated the figure unprompted when asked about cost. A homeowner signs up expecting four visits, gets two, and finds out at the visit that never comes - holding the transcript.

The same HVAC contractor

It cannot say when they are open

Asked three times, on three different pages, in two wordings. Every time: 3 of 3

The assistant said

"I wasn't able to find specific opening hours in our knowledge base." - and then printed their phone number and street address.

Where the hours are

In the header of every page of the site, in the same block as that phone number and address, a few inches above the chat window.

It has the block. It cannot find the hours inside it. Two phrasings, so it is not a matter of how you ask. "When are you open" is the most common question a service business gets, and every customer who asks is sent to the phone.

A US vacation-rental operator

Two of its own pages disagree about tax

Asked: "Are taxes included in the nightly rate, or added on top?" 3 of 3

The assistant said

Tax of 11.5% is added on top; "the nightly rates shown on our website are before tax."

Their FAQ says

All taxes are already included, no extra charges. Their policies page says the opposite - and agrees with the assistant.

The purest case in the set: the assistant is right, according to one of the two pages. It picked the one the guest was not reading. A second finding on the same site - the FAQ says most properties have heated pools, and the assistant promises every guest that all of them do.

A US battery brand · the exception

Clean documentation. Wrong anyway.

Asked: "How cold can it get before I can't use the battery? Not charging, just driving." 3 of 3

The assistant said

You can discharge down to -4°F (-20°C), and a temperature sensor shuts the battery down below that.

Their own FAQ says

Discharge down to approximately -20°F - sixteen degrees colder than the assistant claims.

This looked like a unit-conversion slip. It probably is not: -20°C / -4°F is the standard published discharge floor for this battery chemistry, stated on the FAQ pages of at least two large competitors. The assistant is not reciting its client's page - it is answering from what it knows about the category, and overriding a page that claims something better.

This one cannot be fixed by tidying a website. The assistant understates its own client's product by sixteen degrees in the exact specification a cold-climate buyer shops on, using a number the company never published. It was also the most accurate assistant tested - fourteen of fifteen - which did not protect it where it mattered commercially.

The control

The one that passed

A New England inn: fifteen of fifteen. Check-in times, the cancellation window and its per-room fee, pet weight limits and nightly charges, the smoking fine, the rollaway charge. Asked for a rate on a specific Saturday night, it said it had no access to real-time pricing and pointed at the booking page - the correct answer, and the one most of these assistants get wrong by guessing.

Its documentation agrees with itself. That is the whole difference, and it is what makes this a control rather than an anecdote.

Discarded evidence

Four findings that died in verification

What looked wrongWhat was actually true
A pet business quoting a fee for an extra animal that was not on the pricing pagePublished - different page, different service
An inn offering a named package on no policy page, whose name other hotels also usePublished verbatim on the page the assistant's own citation linked to
A retailer saying you cannot buy directly from it, on what is plainly an online storeCorrect - product pages list prices but have no cart, only a dealer locator
A gym quoting a day-pass price found nowhere on its site or checkoutCorrect - that exact pass exists in the checkout. My search could not see it because the price sat behind markup

The last is the instructive one. It survived the rule "open every link the assistant cites" - I did open it - and died only on a second, more careful read of the same page, hours before it would have been sent to the owner. Opening the page is not enough; you have to be sure your method can actually read it.

Limits

What this does not show

Seven businesses is seven businesses, and they were not randomly sampled - they were found by looking for small-business sites running chat widgets, which selects for a certain kind of operator. One person ran every test. No claim is made about how common this is, and any percentage extrapolated from it has gone past the evidence.

It is also not a criticism of the widget vendors. Four platforms are represented here, all behaved correctly, and all answered from the material they were given. The finding is about the material. What the test does show is a failure mode that is legible and repeatable - each one reproduced three times in fresh sessions, and in five of six cases traceable to a specific page the owner could fix this afternoon.

If you run one of these

The ten-minute version

Take the five questions customers ask before buying - price, what is included, hours, cancellation, and the one specific to your trade - and ask your own assistant in a fresh incognito window. Check each answer against your own page.

When an answer is wrong, the useful question is not "why is the AI wrong." It is "which of my pages says that?" Five times out of six, one of them does.