Field Study 001 · Local AI Search

One local search question. Eleven different answers.

We asked ChatGPT, Perplexity, Microsoft Copilot, and Claude to name the best landscaping service in Andover, Massachusetts. Only two businesses appeared on every list.

May 8 and August 2, 2026 · Andover, Massachusetts

4AI assistants
1identical query
12weeks between tests
0star-rating changes
Method

Five words, one ordinary buying decision

“Best landscaping service in Andover.” A homeowner types it, an engine answers, and somebody’s truck pulls into the driveway on Saturday.

We captured the full response from four assistants on May 8, 2026, then ran the same question again on August 2. We did not contact any of the businesses, buy anything, or steer the engines toward a preferred answer. Between the two runs, we changed nothing.

Andover proved to be a useful test market: a review economy measured in dozens rather than thousands, a few nearby towns in the mix, and seven or eight serious landscaping firms. Less than three months separated the snapshots. Plenty moved.

The evidence

The answer matrix

On August 2, four engines named eleven companies. T&B Landscape & Irrigation and Andover Landscape Design & Construction were the only two every engine surfaced.

BusinessCopilotChatGPTPerplexityClaude
T&B Landscape & Irrigation3126
Andover Landscape Design4211
Pelletier Landscaping235
Alvarez Landscaping14
Ground Care Landscaping54
C&C Landscaping Maintenance63
Downer Brothers Landscaping2
Jireh Landscaping & Irrigation4
Northeast Landscape Contractors3
DiCicco Landscape & Irrigation7
M J Colombo Landscaping5
Named first Position in responseClaude declined to crown one winner; the number reflects list order.
114 reviews · 4.8

Downer Brothers made one list out of four

It had the deepest review history any engine reported near this town, yet three engines left it out.

138 reviews · 4.9

Alvarez ranked first on Copilot and disappeared elsewhere

ChatGPT gave it a passing mention. Perplexity and Claude did not surface it at all.

Neither business had a bad month. Each had a retrieval problem on some engines and a strong result on another, at the same time.
01
Finding

The contest is over the shortlist, not the ranking

Copilot and Claude reported identical review data in August. Ground Care at 4.8 with 47 reviews. C&C at 4.9 with 30. Pelletier at 5.0 with 14. T&B at 4.7 with 74. Andover Landscape at 4.8 with 29. Digit for digit, the same underlying numbers.

Then Copilot led with Alvarez and never mentioned Downer Brothers or DiCicco. Claude led with Andover Landscape and never mentioned Alvarez. Two engines, one dataset, two different rosters.

The variable deciding who got recommended was not star rating or review volume. What differed was which businesses entered the running before scoring happened.

Most local search reporting still measures position. Position is the last step, and by then the interesting decision has already been made upstream.

A 4.9 does nothing for you inside a candidate set you never joined.
02
Finding

Rankings moved while the businesses stood still

Across the seven businesses tracked on both dates, not one star rating changed. Review counts moved by twelve in total.

Alvarez Landscaping135138
T&B Landscape7074
Ground Care4447
Jireh Landscaping3334
Andover Landscape2829
C&C Landscaping3030
Pelletier Landscaping1414

In May, ChatGPT opened with Downer Brothers as its top pick. By August, Downer Brothers had vanished and T&B held the top spot. Perplexity opened May with Jireh at 4.9. By August, Jireh had slid to fourth.

Copilot did not budge: Alvarez first in May, Alvarez first in August, with the same shortlist in a slightly shuffled order.

The engine leaning on structured listing data held steady. The two engines assembling answers from the open web reshuffled winners while the businesses underneath them did close to nothing.

A recommendation drop is not automatic evidence of a reputation problem. It can be a sourcing problem, and the repair work looks entirely different.

03
Finding

Same company, same week, opposite verdicts

T&B Landscape & Irrigation received two reviews of its standing in early August. They do not resemble each other.

Perplexity · August 2

Safest all-around choice

Perplexity put T&B ahead for lawn maintenance, irrigation, and snow removal. Nearly every claim, including the competitor table, cited one page on T&B’s own website.

Source diet: one self-published page
Claude · August 2

Sixth of seven

Claude put T&B last on rating at 4.7 and filed it under irrigation rather than general contracting. It treated the supply yard as a convenience, not a reason to hire them.

Source diet: aggregated review and profile data

Both answers hold up inside their own evidence. They were reading different documents, and that alone was enough to flip the verdict.

04
Finding

A vendor page became the reference source

T&B published a page answering the exact question a customer would type, in its own market, naming its own competitors. Perplexity found it and built an answer on top of it.

The page was already among Perplexity’s sources in May. By August, close to the whole answer rested on it. Yet the same page did nothing on Copilot, which ranked T&B third behind a business Perplexity never named.

Publish the definitive page and you move one kind of engine. The other kind may never see it.
05
Finding

ChatGPT changed what it accepts as proof

In May, ChatGPT cited company websites and nothing else. By August, the same query returned BBB accreditation, HomeAdvisor and Houzz ratings, a founding year, and regional design awards.

Its evidence diet shifted from vendor self-description to outside validation. Businesses without maintained third-party profiles did not make the cut, and Google star ratings did not rescue them.

06
Finding

The engines disagree about the facts too

Claude described Downer Brothers as holding the area’s highest review volume at 114. Copilot listed Alvarez at 138 on the same date. Both cannot be right unless the engines draw the map differently, and they do. Claude separated North Andover from Andover. Copilot did not.

Copilot also listed Alvarez with a 978 phone number in May and a 603 number in August. That could reflect a business change, a directory inconsistency, or something worse. We did not verify which.

What is clear: whatever sits in a listing record now flows directly into the sentence a homeowner reads before deciding who to call. Directory errors used to cause a misdial. They now propagate into recommendations.

What we take from it

Six changes to how local visibility gets measured

Review quality still matters. This study shows that it is one input among several, weighted differently by each engine.

01

Track presence before position

Absence is the loudest result in this dataset. Five of eleven businesses appeared on exactly one engine. A ranking report cannot show the three engines where you are missing entirely.

02

Audit each engine’s source diet

Copilot answered from listing data, ChatGPT from third-party validation, Perplexity from a published page, and Claude from aggregated profiles. Four pipelines produced four different answers.

03

Establish a repeated baseline

Two dates show movement, but not whether it came from a lasting change or run-to-run variance. Serious measurement needs a consistent sampling cadence.

04

Treat directory hygiene as answer hygiene

The Alvarez phone-number discrepancy flowed directly into a recommendation. Listing errors no longer stop at a directory. They become sentences customers act on.

05

Build the third-party footprint deliberately

BBB, Houzz, HomeAdvisor, and a regional award helped decide ChatGPT’s August shortlist. On that engine, those profiles were the basis for selection.

06

Publish the definitive page, with limits

T&B’s comparison page carried Perplexity almost single-handedly and did nothing for them elsewhere. It is worth doing, but it is not an across-the-board fix.

Method and limits

What this study cannot tell you

Two dates in one small town will only carry an argument so far.

  • One market, one query phrasing, and two dates. Other phrasings almost certainly produce other shortlists.
  • A single run per engine per date. We cannot separate twelve-week change from ordinary run-to-run variance.
  • No independent verification of review counts or ratings. We recorded what the engines reported.
  • No control for location, personalization, or account history on any engine.
  • Some engines included businesses outside Andover proper without saying so, inflating disagreement by an unknown degree.

What survives those limits: one ordinary question produced eleven businesses across four assistants, with half the matrix empty. The reasons for the divergence are legible, and they can be measured.

The ASK Index

AI is already answering questions about your brand.

Find out which engines surface you, what sources shape the answer, and where the gaps begin.