Blog/Local data
10 min read

Yelp API: The Four Ads Above the Results Had 5, 75, 80 and 199 Reviews

The top organic result had 2,156. The ads sitting above it had a fraction of that. Both arrive in the same response, in two separate arrays.

Yelp API: The Four Ads Above the Results Had 5, 75, 80 and 199 Reviews

Copy this line to your agent to build a local business list with ratings and reviews.

set up https://monid.ai/SKILL.md and use litescrape /yelp/search, then pull reviews for the top ten

We searched Austin for coffee on 2026-10-01, sorted by review count. The top organic result had 2,156 reviews. Sitting above the organic list, in their own array, were four paid placements with 5, 75, 80 and 199. The endpoint keeps them separate, which is the right call, and that separation is the only thing standing between a clean local lead list and one where a shop with five reviews outranks a local institution. This guide runs through Monid, the OpenRouter for agent tools.

What does the Yelp API return?

A ranked page of businesses, the ads that sit above them, and a key that unlocks the reviews.

The search call

find_loc takes a city or neighborhood, find_desc takes a name or search term, and cflt takes a Yelp category identifier if you would rather browse than search. sortby accepts recommended, rating or review_count, and start pages in offsets.

What a row carries

Everything you need for a lead list without a second call: title, rating, reviews, categories, price, the full address broken into fields, phone, gps_coordinates, a snippet, and the place_id.

The two arrays

{ "organic_results": [ /* 10 businesses */ ],
  "ads_results":     [ /* 4 paid placements */ ],
  "search_information": { "total_results": 240, "results_per_page": 10 } }

This separation is the most important thing in the response. Yelp's own page interleaves paid placements with organic results, and a scraper that reads the rendered page has to work out which is which. This endpoint has already done it.

The key that chains

Every organic row carries place_id, and the reviews endpoint takes exactly that value with no lookup step in between. Search then read, two calls, no glue.

📖 See also Google Maps Reviews Without a Places API Key

Why were the ads ranked above better businesses?

Because they paid, and because review count is not what put them there.

We asked for sortby: review_count, which the organic list honoured exactly:

organic  1  Mozart's Coffee Roasters   4.0   2,156 reviews
organic  2  Paperboy East              4.5   1,463
organic  3  Jo's Hot Coffee Good Food  4.0   1,003
         …
organic 10  Kick Butt Coffee           4.0     560

ads        Black Rock Coffee Bar       4.4       5
ads        Starbucks                   3.9      80
ads        Anderson's Coffee Company   4.6     199
ads        Capital One Café            4.4      75

The weakest ad has 5 reviews against the organic leader's 2,156. That is a 431 times gap in evidence, and on the rendered page those four appear first.

Why this matters more than it looks

Nobody sets out to rank a coffee shop below a bank's café lounge. It happens when a pipeline concatenates the two arrays:

// the bug
const all = [...res.ads_results, ...res.organic_results];

Now your "top four coffee shops in Austin" are four advertisers, one of which has five reviews, and every downstream score inherits it. The arrays exist separately so you can make a decision, and the decision has to be made explicitly.

When to keep them

Keep the ads when you are studying the market rather than the businesses. They tell you who is spending to acquire customers in a category, which is a buying signal of its own and a decent proxy for marketing budget. Just keep them in their own column, never merged into a ranking.

A sort honoured is not a filter applied

sortby: review_count ordered the organic list correctly and did nothing about relevance. Ranks 5 and 6 carry a Tacos category alongside the coffee one, and rank 8 is a Cocktail Bars listing. Yelp matched on overlapping categories, not on your intent, so a category check after the fetch is still your job.

What do the reviews actually contain?

More than the stars, including the part most people forget to ask for.

One call against the top result's place_id, asking for 20, returning 20:

{ "position": 1, "rating": 3,
  "date": "2026-09-08T…",
  "user": { … },
  "comment": { "text": "Beautiful location, mid pastries and coffee…" },
  "feedback": { … }, "reactions": { … },
  "owner_replies": [ … ] }

The fields worth planning around

rating and comment.text are what everyone comes for, and the pairing is the point: a 3-star review that says "beautiful location, mid pastries" is a different signal from a 3-star with no text.

owner_replies is the one people miss. Whether a business answers its reviews, and how fast, is a straightforward operational signal. For a lead list it separates the places that are actively managed from the ones running on autopilot.

date and local_date let you weight recency, which matters because a 4.0 built on 2,156 reviews over a decade means something different from a 4.0 built this year.

The paging shape

num accepts up to 49, so a 49-review sample is a single call. sortby and a rating array let you pull only the one and two star reviews, which is the efficient way to do complaint mining: you do not need the happy ones to find the pattern.

What one call costs you in coverage

Twenty reviews out of 2,156 is a sample, not a corpus. For a sentiment claim about one business, pull several pages. For a comparison across fifty businesses, twenty recent reviews each is usually enough to rank them, and that is the trade worth making consciously rather than by accident.

How do you build a local list?

For agents

Grab an API key at app.monid.ai, then paste this to your agent and hand it the key:

set up https://monid.ai/SKILL.md

It learns the whole discover, inspect, run workflow itself. More in the agent quickstart.

For humans

npm install -g @monid-ai/cli
monid keys add -k <your-api-key> -l main
monid run -p litescrape -e /yelp/search \
  --query '{"find_loc":"Austin, TX","find_desc":"coffee","sortby":"review_count"}'

Step 2, split the arrays before you do anything else

const businesses = res.organic_results;        // rank these
const advertisers = res.ads_results;           // keep, separately

Do this at the point of ingestion, not later. Once the two are merged, nothing downstream can tell them apart.

Step 3, page

results_per_page came back as 10 and total_results as 240, so a full sweep of that query is 24 calls. Use start as an offset.

Step 4, chain into reviews with the id you already have

monid run -p litescrape -e /yelp/reviews \
  --query '{"place_id":"-4ofMtrD7pSpZIX5pnDkig","num":49,"sortby":"relevance_desc"}'

Ask for 49 rather than the default if you are going to analyse the text. The call costs the same.

Step 5, filter on category after the fetch

The search matches overlapping categories, so a coffee query returned a cocktail bar and two places tagged for tacos. Check categories yourself against what you actually meant.

Give this to your agent

$Set up https://monid.ai/SKILL.md, and then use Monid to find coffee shops in Austin by review count, drop the paid placements, pull 49 reviews for each of the top ten, and tell me which ones reply to their reviews.

📖 See also Why One Google Maps Scraper Is Not Enough

Which endpoint should I use for which job?

EndpointWhat it doesInputReturnsLatency we sawBilling
litescrape/yelp/searchBusinesses in a placefind_loc, find_desc, sortby10 organic plus ads, with place_idAround 10 sPer call
litescrape/yelp/reviewsReviews for one businessplace_id, num, ratingUp to 49 reviews with owner repliesUnder 10 sPer call
clutch/get_company_reviewsB2B service reviewsA companyPaginated reviewsAround 18 sPer call
zillow/search_homes_for_saleWhat the area costslocationListings and countsUnder 6 sPer call

Every row was verified with monid inspect on 2026-10-01 and the latency column is what we actually waited. Billing shape is given rather than figures, because shape drives design and current numbers live on monid.ai/tools.

The first two rows are one pipeline: search gives you the place_id, reviews takes it, and there is no identifier resolution step between them. That is worth more than it sounds, because the usual shape of this job is "find the business, then work out what this source calls it", and here the answer arrives with the search result.

When should you not use this?

Four cases.

You need the whole category in a city. Our search reported 240 total results at 10 per page, so a full sweep is 24 calls before any reviews. That is fine for one market and expensive for fifty. Narrow by neighborhood or category first.

Yelp is not where your market lives. Coverage is strong for restaurants, bars and consumer services in US metros, and thin almost everywhere else. For B2B services a reviews source built for that is a better starting point, and for anything outside the US check a sample before you build on it.

You want one true rating. Ratings differ by platform because the populations differ, and a business that is 4.0 here can be 4.5 elsewhere. If the rating is going into a decision, read more than one source, which is the same argument as three sources disagreeing on a technographic count.

You are republishing the review text. The ratings and counts are facts about a business. The reviews are writing by individual people and carry both copyright and platform terms. Analysing them to produce a summary is ordinary; reposting them as your own content is not, and the fact that an API returned them does not change that.

And the disclosure: we resell these endpoints. The most useful thing here costs nothing, which is to keep ads_results out of your ranking, and that applies whichever source you end up buying.

Conclusion

One Yelp search on 2026-10-01 returned ten businesses and four ads, in two separate arrays. Sorted by review count, the organic leader had 2,156 reviews. The four paid placements above it had 5, 75, 80 and 199. The endpoint doing that separation for you is the single most valuable thing in the response, and the way to waste it is a spread operator that merges the arrays before anything ranks them.

The sort did what it was asked and nothing more. Ranks 5 and 6 came back tagged for tacos and rank 8 is a cocktail bar, because Yelp matched on overlapping categories rather than on what you meant by coffee. Filter on categories after the fetch.

The chain into reviews is clean: every organic row carries a place_id and the reviews endpoint takes exactly that, no lookup in between. Ask for 49 rather than the default 20, since the call costs the same, and read owner_replies while you are there. Whether a business answers its customers is free in the payload and absent from almost every lead list we have seen.

Free next step: run monid inspect -p litescrape -e /yelp/search, make one call, and print organic_results.length next to ads_results.length. Seeing the two arrives separately is what makes the rest obvious. Start at monid.ai.

FAQ

Should you keep the ads or throw them away?

Keep them, in their own column, and never let them into a ranking. They are genuinely useful data about a different question: who is paying to acquire customers in this category right now. That is a marketing-budget signal and a competitive one, and for a sales use case it may be more interesting than the organic list. What they cannot do is tell you who the significant businesses are, and our search is the clean illustration: the weakest paid placement had five reviews while the organic leader had 2,156. The failure mode is almost always a one-line merge at ingestion, after which nothing downstream can separate them again. Split at the point you read the response.

Why did a coffee search return a cocktail bar?

Because Yelp matches on a business's category list, and businesses carry several categories. Rank 8 in our search was tagged both as a cocktail bar and for breakfast and brunch, and two others in the top ten carry a tacos category next to the coffee one. From Yelp's side this is reasonable, since a place that serves coffee in the morning and cocktails at night is both. From your side it means the result set is wider than your query, and the fix is cheap: every row returns its categories, so check them against what you actually meant before the business reaches your list. Sorting does not help here, because a sort orders a set and does not decide what belongs in it.

How many reviews can you actually get?

The endpoint takes num up to 49 per call and a start offset for paging, so a business with 2,156 reviews is 44 calls to read in full. Almost nobody needs that. Two patterns cover most work. For ranking many businesses, pull one page of 49 for each, weighted toward recent, and compare. For understanding one business, use the rating filter to pull only the one and two star reviews: complaints cluster, and you will see the pattern in two pages without paying for the praise. The date and local_date fields let you weight recency, which matters because an average built over a decade describes a different business from one built this year.

What can you legally do with the review text?

The split worth holding in your head is between facts and writing. A business's name, address, phone, rating and review count are facts about a business, and using them in a lead list or an internal analysis is ordinary commercial activity. The review text is writing by an identifiable person, published on a platform with its own terms, and it carries copyright. Reading it to produce your own summary, a sentiment score or a list of recurring complaints is normal analysis. Republishing the reviews as content on your own site, or presenting them as reviews of you, is a different act and not one an API grants you. If the output is customer-facing rather than internal, get advice rather than taking a blog post's word for it.

Last updated October 2026.

yelp apiyelp scraperlocal business datayelp reviews apilocal lead list