Downer Brothers made one list out of four
It had the deepest review history any engine reported near this town, yet three engines left it out.
We asked ChatGPT, Perplexity, Microsoft Copilot, and Claude to name the best landscaping service in Andover, Massachusetts. Only two businesses appeared on every list.
“Best landscaping service in Andover.” A homeowner types it, an engine answers, and somebody’s truck pulls into the driveway on Saturday.
We captured the full response from four assistants on May 8, 2026, then ran the same question again on August 2. We did not contact any of the businesses, buy anything, or steer the engines toward a preferred answer. Between the two runs, we changed nothing.
Andover proved to be a useful test market: a review economy measured in dozens rather than thousands, a few nearby towns in the mix, and seven or eight serious landscaping firms. Less than three months separated the snapshots. Plenty moved.
On August 2, four engines named eleven companies. T&B Landscape & Irrigation and Andover Landscape Design & Construction were the only two every engine surfaced.
| Business | Copilot | ChatGPT | Perplexity | Claude |
|---|---|---|---|---|
| T&B Landscape & Irrigation | 3 | 1 | 2 | 6 |
| Andover Landscape Design | 4 | 2 | 1 | 1 |
| Pelletier Landscaping | 2 | — | 3 | 5 |
| Alvarez Landscaping | 1 | 4 | — | — |
| Ground Care Landscaping | 5 | — | — | 4 |
| C&C Landscaping Maintenance | 6 | — | — | 3 |
| Downer Brothers Landscaping | — | — | — | 2 |
| Jireh Landscaping & Irrigation | — | — | 4 | — |
| Northeast Landscape Contractors | — | 3 | — | — |
| DiCicco Landscape & Irrigation | — | — | — | 7 |
| M J Colombo Landscaping | — | — | 5 | — |
It had the deepest review history any engine reported near this town, yet three engines left it out.
ChatGPT gave it a passing mention. Perplexity and Claude did not surface it at all.
Neither business had a bad month. Each had a retrieval problem on some engines and a strong result on another, at the same time.
Copilot and Claude reported identical review data in August. Ground Care at 4.8 with 47 reviews. C&C at 4.9 with 30. Pelletier at 5.0 with 14. T&B at 4.7 with 74. Andover Landscape at 4.8 with 29. Digit for digit, the same underlying numbers.
Then Copilot led with Alvarez and never mentioned Downer Brothers or DiCicco. Claude led with Andover Landscape and never mentioned Alvarez. Two engines, one dataset, two different rosters.
The variable deciding who got recommended was not star rating or review volume. What differed was which businesses entered the running before scoring happened.
Most local search reporting still measures position. Position is the last step, and by then the interesting decision has already been made upstream.
A 4.9 does nothing for you inside a candidate set you never joined.
Across the seven businesses tracked on both dates, not one star rating changed. Review counts moved by twelve in total.
In May, ChatGPT opened with Downer Brothers as its top pick. By August, Downer Brothers had vanished and T&B held the top spot. Perplexity opened May with Jireh at 4.9. By August, Jireh had slid to fourth.
Copilot did not budge: Alvarez first in May, Alvarez first in August, with the same shortlist in a slightly shuffled order.
The engine leaning on structured listing data held steady. The two engines assembling answers from the open web reshuffled winners while the businesses underneath them did close to nothing.
A recommendation drop is not automatic evidence of a reputation problem. It can be a sourcing problem, and the repair work looks entirely different.
T&B Landscape & Irrigation received two reviews of its standing in early August. They do not resemble each other.
Perplexity put T&B ahead for lawn maintenance, irrigation, and snow removal. Nearly every claim, including the competitor table, cited one page on T&B’s own website.
Source diet: one self-published pageClaude put T&B last on rating at 4.7 and filed it under irrigation rather than general contracting. It treated the supply yard as a convenience, not a reason to hire them.
Source diet: aggregated review and profile dataBoth answers hold up inside their own evidence. They were reading different documents, and that alone was enough to flip the verdict.
T&B published a page answering the exact question a customer would type, in its own market, naming its own competitors. Perplexity found it and built an answer on top of it.
The page was already among Perplexity’s sources in May. By August, close to the whole answer rested on it. Yet the same page did nothing on Copilot, which ranked T&B third behind a business Perplexity never named.
Publish the definitive page and you move one kind of engine. The other kind may never see it.
In May, ChatGPT cited company websites and nothing else. By August, the same query returned BBB accreditation, HomeAdvisor and Houzz ratings, a founding year, and regional design awards.
Its evidence diet shifted from vendor self-description to outside validation. Businesses without maintained third-party profiles did not make the cut, and Google star ratings did not rescue them.
Claude described Downer Brothers as holding the area’s highest review volume at 114. Copilot listed Alvarez at 138 on the same date. Both cannot be right unless the engines draw the map differently, and they do. Claude separated North Andover from Andover. Copilot did not.
Copilot also listed Alvarez with a 978 phone number in May and a 603 number in August. That could reflect a business change, a directory inconsistency, or something worse. We did not verify which.
What is clear: whatever sits in a listing record now flows directly into the sentence a homeowner reads before deciding who to call. Directory errors used to cause a misdial. They now propagate into recommendations.
Review quality still matters. This study shows that it is one input among several, weighted differently by each engine.
Absence is the loudest result in this dataset. Five of eleven businesses appeared on exactly one engine. A ranking report cannot show the three engines where you are missing entirely.
Copilot answered from listing data, ChatGPT from third-party validation, Perplexity from a published page, and Claude from aggregated profiles. Four pipelines produced four different answers.
Two dates show movement, but not whether it came from a lasting change or run-to-run variance. Serious measurement needs a consistent sampling cadence.
The Alvarez phone-number discrepancy flowed directly into a recommendation. Listing errors no longer stop at a directory. They become sentences customers act on.
BBB, Houzz, HomeAdvisor, and a regional award helped decide ChatGPT’s August shortlist. On that engine, those profiles were the basis for selection.
T&B’s comparison page carried Perplexity almost single-handedly and did nothing for them elsewhere. It is worth doing, but it is not an across-the-board fix.
Two dates in one small town will only carry an argument so far.
What survives those limits: one ordinary question produced eleven businesses across four assistants, with half the matrix empty. The reasons for the divergence are legible, and they can be measured.
Find out which engines surface you, what sources shape the answer, and where the gaps begin.