Blog/Search & RAG
11 min read

Perplexity API: Sonar and the App Answered the Same Question Differently

Same six questions, same day. Sonar replied in about three seconds with 84 cited domains. The consumer app took up to 92 and named different vendors.

Perplexity API: Sonar and the App Answered the Same Question Differently

Copy this line to your agent to get a cited answer without holding a Perplexity account.

set up https://monid.ai/SKILL.md and get me a cited answer to this question from perplexity

"The Perplexity API" refers to two different products, and on 2026-09-28 we asked both of them the same six questions. The Sonar API answered in about three seconds and attached 84 distinct source domains. The consumer app, reached through a pay-per-call endpoint, took between 36 and 92 seconds and attached 55. On a question about tools similar to People Data Labs, the two surfaces returned lists with almost nothing in common, and the consumer app never named People Data Labs. If you are picking one for a product, the speed is the smaller difference. This guide runs through Monid, the OpenRouter for agent tools.

What do people mean by the Perplexity API?

Two things, and mixing them up produces numbers that describe nothing.

The Sonar API

Perplexity's developer product. You hold an account, you get a key, and you call /chat/completions with a model named sonar or one of its larger siblings. It is an OpenAI-shaped chat endpoint that runs retrieval before it generates, and it returns citations alongside the text. Billing is per token with a search fee, on your own Perplexity account.

The consumer app

What a person gets at perplexity.ai. Its own retrieval, its own routing between models, its own answer formatting. This is the surface people mean when they say "Perplexity recommended X", because it is the one humans use. It has no official API, and the pay-per-call endpoints that reach it are reading the product a person sees.

Why the distinction is not pedantic

Because the two produce different answers to the same question, which we measured rather than assumed. If you are tracking whether your brand gets named, the answer the Sonar API gives is not evidence about what a buyer sees, and a report that says "Perplexity" without saying which surface has skipped the only part that mattered.

The same trap exists across engines, which is the subject of our four-surface measurement from the same day.

📖 See also The Best Web Search API for AI Agents in 2026

How do Sonar and the consumer app differ?

Six questions, one day, both surfaces.

Sonar APIConsumer app via a pay-per-call endpoint
Latency2,551 to 3,467 ms36,238 to 92,310 ms
Answer length1,510 to 2,235 chars230 to 4,979 chars
Distinct cited domains, six answers8455
Vendors named, six answers15 distinct16 distinct
Account neededYours, at PerplexityNone beyond one key
ShapeOpenAI-style chat completionAnswer text plus a source list

The latency gap is the design constraint

A factor of more than twenty. Sonar is fast enough to sit in a request path a user is waiting on. The consumer-app route at up to 92 seconds per question is a batch job: a 60-question sweep is over an hour, so it belongs on a schedule with a queue, not behind a spinner.

The content gap is the bigger one

Two questions make it concrete.

"What tools are similar to People Data Labs?"
  Sonar        : people data labs, zoominfo, clay, apollo
  Consumer app : zoominfo, apollo, diffbot

"Is there a free alternative to Apify?"
  Sonar        : apify, scraperapi, firecrawl, octoparse, browse ai
  Consumer app : apify, in a 230-character answer

On the first, the consumer app did not name the company the question was about. On the second, Sonar produced five vendors and the consumer app produced one in a reply short enough to fit in a tweet. Across the six questions Sonar named 15 distinct vendors and the consumer app 16, and neither list contains the other.

Answer length is not a quality signal

The consumer app produced both the shortest answer in the set at 230 characters and the longest at 4,979. Sonar stayed between 1,510 and 2,235 on every question. If you are scoring answers, normalise for length before you compare, or you will conclude that the surface with more words said more about you.

Can you reach a sonar model pay per call?

We tested it directly, because it is the obvious thing to want.

There is a pay-per-call chat endpoint on the shelf that accepts any model id you pass. We sent four ids, one short question each:

Model id sentProvider statusChargedResult
sonar400noUnknown model
sonar-pro400noUnknown model
perplexity/sonar400noUnknown model
gpt-4o-mini200yes350-char answer in 4,956 ms

The refusal body is worth reading because it names its own fix:

{ "error": "Unknown model: sonar. Try one of: openai/gpt-6-astra,
   openai/gpt-6-sol, openai/gpt-6-luna, openai/gpt-5.6-sol.
   Full list: GET /v1/models" }

So the answer is no: that endpoint serves chat models and not Perplexity's. The pay-per-call route to Perplexity reaches the consumer app, not Sonar. If you need Sonar specifically, you need a Perplexity account and your own key.

Two smaller findings from the same four calls. The suggestion list in the error is a sample rather than an allowlist, because gpt-4o-mini worked and was not in it, so read the models endpoint rather than the error text. And all three refusals cost nothing while the one success was billed. The rule is about visibility rather than about the billing model: a provider that refuses with an HTTP error is free, and a 200 that hands you an empty record is not necessarily.

How do you call each one?

For agents

Grab an API key at app.monid.ai, then paste this to your agent and hand it the key:

set up https://monid.ai/SKILL.md

It learns the whole discover, inspect, run workflow itself. More in the agent quickstart.

For humans

npm install -g @monid-ai/cli
monid keys add -k <your-api-key> -l main

The consumer app, by the call

monid run -p cloro -e /perplexity/ask \
  -d '{"prompt":"what is the best web scraping api for ai agents","country":"US"}'

Returns { success, result: { text, sources } }. Set your client timeout above 120 seconds, because our slowest of six took 92 and a timeout of 30 would have recorded that question as an absence.

Sonar, with your own key

curl https://api.perplexity.ai/chat/completions \
  -H "Authorization: Bearer $PPLX_KEY" -H 'Content-Type: application/json' \
  -d '{"model":"sonar","messages":[{"role":"user","content":"what is the best web scraping api for ai agents"}]}'

Citations come back in citations, or in search_results depending on the model and version, so read both keys rather than one. A fresh account has a low request ceiling and a throttle returns a row with no answer text, which looks exactly like a genuine absence in your data. Retry on 429 before you record anything.

Give this to your agent

$Set up https://monid.ai/SKILL.md, and then use Monid to ask this question on both perplexity surfaces, then show me the two answers side by side with the vendors each one named and the domains each one cited.

📖 See also Answer Engine Optimization: We Asked Four Surfaces the Same Six Questions

Which endpoint for which job?

EndpointSurfaceInputReturnsLatency we measuredBilling
cloro/perplexity/askPerplexity consumer appprompt, countryAnswer text, sources36 to 92 sPer call
blockrun.ai/api/v1/exa/answerExa's answer modelqueryAnswer, citations3.7 to 5.1 sPer call
blockrun.ai/api/v1/chat/completionsChat models, no Perplexitymodel, messagesCompletion, no citations5 sPer call
mrscraper/google/ai-modeGoogle AI Modekeyword, countryText blocks, markdown6.5 to 11.4 sPer call
firecrawl/scrapeAny page, as texturlMarkdown plus metadataVariesTiered per call

Every row was verified with monid inspect on 2026-09-28 and the latency column is what we waited. Billing shape is given rather than figures, because shape drives design and current numbers live on monid.ai/tools.

Reading that table for a decision: row one is the only one that reads Perplexity as a person sees it, and it is slow enough to be a scheduled job. Row two is the fast option when you want a cited answer and do not care which engine wrote it. Row three is for when you want a model's opinion with no retrieval at all. Row four is the other surface a buyer actually uses, and pairing it with row one is the cheapest honest measurement of "what do the AI answers say about us".

When do you not need Perplexity at all?

Four cases, and the first is most of them.

You want documents, not an answer. If your next step is your own extraction or your own model, a generated answer is a lossy middle layer. Fetch the pages and keep the text, which is what a straight search endpoint is for.

You need the citations more than the prose. Some jobs only want the source list. A search endpoint returns ranked links faster and cheaper than any answer engine, and you skip paying for a paragraph you throw away.

You are building a SERP feature. Answer engines are not search engines. If you need positions, ranked results and the shape of the page, that is a SERP endpoint, and the two are not substitutes in either direction.

The question has a structured source. Company facts, pricing, profiles and filings live in endpoints that return typed fields. Asking a language model to summarise something you can look up gives you prose where you wanted a value.

And the disclosure: we resell one route to the consumer app and none to Sonar. If Sonar is what your product needs, go get a Perplexity key, and we are saying so because a pay-per-call endpoint that reads a consumer product is a different thing from a vendor's own API and should not be sold as the same.

Conclusion

"The Perplexity API" is two products. The Sonar API is a developer endpoint on your own account that answered our six questions in about three seconds each with 84 distinct source domains attached. The consumer app, reachable by the call, took between 36 and 92 seconds and attached 55. That speed gap decides your architecture, but the content gap decides your conclusions: on a question about tools similar to People Data Labs the two lists had little in common and the consumer app never named the company in the question.

You cannot buy Sonar by the call through a general chat endpoint. Three sonar model ids came back as HTTP 400 with the error naming its own fix, and none of the three was billed, while a chat model outside the error's own suggestion list worked fine. The pay-per-call route reaches the consumer app, which is the surface a buyer sees, and Sonar needs Perplexity's key.

The practical rule if you are measuring: pick the surface that matches the claim you want to make. If the claim is about what a human sees, read the consumer app and budget 90 seconds a question. If the claim is about your own product's answer quality, use Sonar and enjoy the three seconds. Never average the two, and never label either one simply "Perplexity".

Free next step: run monid inspect -p cloro -e /perplexity/ask, read the response shape, and set your timeout above 120 seconds before your first call. Start at monid.ai.

FAQ

Which surface should a citation-tracking tool read?

The consumer app, if the claim you want to make is about what buyers see, because that is the surface they use. Budget for it honestly: our slowest answer of six took 92 seconds, so a 60-question sweep is more than an hour of wall clock and belongs in a queue rather than a cron job you expect to finish in minutes. Use Sonar when you want a fast, repeatable read for engineering purposes, such as regression-testing your own documentation, and label the column with the surface name so nobody later reads one as the other. What you should not do is run whichever is convenient and report it as "Perplexity", which is the mistake that makes most AI visibility dashboards uncomparable with each other.

Why do the two disagree if they are the same company?

Because retrieval is most of the answer, and they do not share it. The consumer product runs its own search, its own routing between models, and its own answer formatting, and it optimises for a person reading a page. The Sonar API is built for developers who want a cited completion in a request path. Different documents pulled, different number of them, different prompt wrapping the result. In our six questions this showed up as different vendor lists rather than different phrasing of the same list, which is the version that matters: on one question Sonar named four vendors and the consumer app named three, with only two shared.

What do rate limits look like in your data?

Like an absence, which is the dangerous part. A fresh Perplexity account has a low request ceiling, and a throttled call returns a row with no answer text. If your pipeline records that row, you will later read it as the engine having said nothing about you on that question, on that day. Handle it at the fetch layer: retry on 429 with a backoff, and if the retries are exhausted write a failure marker rather than an empty answer. The same discipline applies to the pay-per-call route, where a client timeout below 120 seconds will silently manufacture the same false absence.

Is the answer stable enough to cache?

For minutes, yes. For weeks, no, and that is a feature rather than a bug given that the whole point is a live answer. Two consequences. First, cache aggressively inside a single run so you do not pay twice for the same question, and key the cache on the exact question string, because small rewordings genuinely change the answer. Second, do not compare an answer you cached last month with one you fetched today and call the difference a trend, unless the question string and the surface are identical and you recorded both. Store the raw text with a timestamp and a surface label, and you can always recompute a score later.

Last updated September 2026.

perplexity apiperplexity sonar apisonar modelai answer apiperplexity citations