TinyFish Made a Free Web Search API. One Query Is a Third.
We asked TinyFish Search one question five ways. It returned 27 pages, and the best single phrasing found 9 of them. Every call was free.

We asked TinyFish Search one question five ways. Not five questions: one question, phrased the way five different people would type it, all wanting the same number.
endpoint: tinyfish /search
intent: what is GitHub's REST API rate limit without a token
phrasing 1 8 results
phrasing 2 7 results
phrasing 3 9 results
phrasing 4 5 results
phrasing 5 9 results
----------
27 unique URLs across 13 domains
The best single phrasing returned 9 of those 27. Six of the ten pairs of phrasings had no URL in common at all.
Every one of those five calls was free, which is the part that matters. Not because free is cheap, but because free is what makes running all five the obvious move instead of an extravagance you have to justify. TinyFish prices search at zero on every plan, and this article is about what that decision does to the code you write, not to the invoice.
Fair disclosure before the numbers. This is Monid's blog, TinyFish is both a provider in our catalogue and a content partner, and their search endpoint is the one we measured. The endpoint is listed at zero per call, so we have no billing interest in how many times you call it, which is roughly the point of the article.
What does a free web search API actually cost?
Nothing, on this endpoint, and the interesting question is what that does to the design rather than to the invoice.
Free here is literal, not a trial tier
TinyFish Search is listed in our catalogue at zero per call, and their own pricing page says search is free on every plan and never draws from a wallet balance. We ran every measurement below against it and the balance did not move. There is no credit that runs out at the end of the month and no per-seat gate underneath.
That is a real position rather than a discount. Search is the cheap primitive at the front of a workflow; the money in web data sits in the steps after it, in rendering, extraction and multi-step automation. A vendor that gives away the front door is betting you will walk through it.
The cost moves to context, not to the bill
What a free query does still spend is the agent's attention. Twenty-seven results at roughly 150 characters of snippet each is about four thousand characters of material the model has to read and rank before it does anything. Free at the API is not free at the prompt, and an agent that fans out without deduplicating first pays for the fan-out in tokens.
So the honest framing is not "search is free, call it infinitely". It is that the constraint moved from your budget to your context window, and those two constraints reward completely different designs.
Does the same query return the same results twice?
On this endpoint, yes, exactly. That matters because it is the control for everything after it.
Three identical calls, three identical answers
We sent github rest api rate limit unauthenticated to TinyFish Search three times:
run 1 8 results
run 2 8 results
run 3 8 results
pairwise Jaccard 1.00 1.00 1.00
identical ordering true true true
same rank 1 yes
Same eight URLs, same order, same top result. Repeat a query and you get the same answer, which means a cache in front of it is safe and a difference between two runs is real signal rather than noise.
Not every search endpoint behaves this way
This is worth stating because it is not universal, and it is the one place where TinyFish beat a peer outright in our own testing. We ran the same experiment on a different search endpoint in an earlier piece and got a different top result on all three calls, with the result sets agreeing only 89 percent of the time. Same kind of product, same kind of query, completely different determinism.
If you are choosing between endpoints, that property is worth testing on your own queries before you build on either, because it decides whether "the answer changed" is information or weather. The vendor landscape for that choice is in our web search API comparison, and the taxonomy underneath it, the three different products that all call themselves a search API, is in what a SERP API actually is.
Why does rephrasing the same question return different pages?
Because the query string is doing far more work than it looks like it is doing, and a stable endpoint is stable per string, not per intent.
Five phrasings, twenty-seven pages
Here is the whole result:
ph_1 github rest api rate limit unauthenticated 8
ph_2 how many requests per hour to github api without a token 7
ph_3 github api anonymous request limits 9
ph_4 what is the github rate limit for requests with no auth 5
ph_5 github api limits without authentication 2026 9
38 rows, 27 unique URLs, 13 domains
best single phrasing covers 9 of 27 (33%)
six of the ten phrasing pairs share zero URLs
Read that last line twice. Two people asking the same question in their own words can receive result sets with nothing in common, from an endpoint that is perfectly repeatable when either of them asks twice.
Determinism and coverage are different properties
This is the part that gets conflated. The endpoint is deterministic: same input, same output, reliably. It is not complete: the input is a phrasing, and the phrasing selects a neighbourhood of the index.
An agent that writes one query and reads the results is not sampling the web. It is sampling one phrasing's neighbourhood, and our measurement puts that at about a third of what the same endpoint would have surfaced for the same question. The single query is not wrong. It is partial, and nothing in the response says so, which is the same shape of problem as a filter that silently returns the wrong slice.
Which is why the price is an architecture decision
When a query is metered, fan-out is a cost you defend in review, so teams write one careful query and move on. When a query is free, running five phrasings and taking the union is simply the correct implementation, and the argument never has to happen.
That is the real content of "search should be free for agents". Not a pricing preference: a claim that the right number of queries per intent is greater than one, and that metering pushes teams toward the wrong number. The metering question in general is worked through in pay per call versus a subscription; this is the case where the answer is neither.
Give this to your agent![]()
Set up https://monid.ai/SKILL.md, and then use Monid to search the live web for GitHub's unauthenticated REST API rate limit using five different phrasings of the question, then show me the union of URLs and which ones only one phrasing found.When can an agent skip fetching the page?
More than half the time here, and the exceptions are the dangerous part.
Sixteen of twenty-seven snippets carried the answer
The true answer is 60 requests per hour. Across the 27 unique results, 16 snippets contained that figure outright, in roughly 150 characters. For those, fetching the page would have bought nothing at all: the search response was the answer, and the fetch step is pure latency and tokens.
This is the underrated economics of a good snippet. A search endpoint that returns a useful 150 characters removes a whole page fetch and a whole extraction pass from the majority of lookups, which is a far larger saving than the price of the search call ever was.
The two rank-1 rows that were accurate and wrong
Now the part to design against. Rank 1 carried the correct figure on only three of the five phrasings. The two misses were not junk pages. Both were docs.github.com, and both quoted a real GitHub limit:
ph_2 rank 1 "no more than 80 content-generating requests per minute"
ph_4 rank 1 "a separate rate limiting bucket with a limit of 300
requests per minute"
Those are true sentences about GitHub. Neither answers the question that was asked. An agent that takes the top result and stops has a two-in-five chance here of returning a confident, well-sourced, wrong number, and no error will fire.
Fanning out fixes this almost for free: across five phrasings the correct figure appears sixteen times and the two distractors appear once each, so a simple majority over the union lands on 60 without any model reasoning at all. The general version of this, that agreeing sources are worth more than a ranked one, is the same argument as reconciling two vendors that answer the same question.
One parameter that did nothing, for the record
TinyFish Search takes an optional purpose field, described as an extra ranking signal. We sent the same query twice with opposite purposes, one operational and one historical, and got byte-identical rankings against the no-purpose control. That is one query and one negative observation, not a verdict, but it is a reminder to measure the knobs rather than trust the parameter list.
When is paying for search still the right call?
Three cases, and none of them are about the search step itself.
When you need the page, not the pointer
A search endpoint returns titles, URLs and snippets. When the answer is not in 150 characters, something has to read the page, and that read is the step with real cost in it. Some of those reads are cheap and some involve a rendered browser and a session, which is where the priced products in this category actually live. Picking between those is a separate decision, covered in which MCP server gives an agent live web data.
When the index is not the one you need
General web search finds public pages. It does not find a company's headcount as a field, a filing as structured rows, or a profile you can join on. Those come from endpoints built for the shape, not from a better query, and reaching for search when you needed a dataset is the most common version of this mistake. The practical demonstration of what a purpose-built endpoint returns instead is in giving an agent live web context.
When free comes with a limit you have not read
Free at the call is not free at any volume, and a generous allowance is still an allowance. Before a fan-out design goes to production, read the rate limit rather than the price, because the constraint that bites will be requests per minute rather than dollars. One balance across metered and unmetered endpoints is the reason the OpenRouter for agent tools exists as a shape: discovery and schema inspection cost nothing, so the comparison happens before the commitment. Current per-endpoint rates sit at monid.ai/tools, and the search-specific ones at the web search endpoints, rather than in a sentence here that will age.
Conclusion
Three numbers from one afternoon of free calls against TinyFish Search.
Jaccard 1.00 across three identical queries, so the endpoint is genuinely repeatable. Nine of twenty-seven, the share of the available pages that the best single phrasing found. And three of five, how often the top result actually answered the question that was asked, with the two misses being accurate sentences about something else.
Put together they say something narrower and more useful than "search should be free". They say the right number of queries per intent is not one, that a deterministic endpoint will not tell you when your single phrasing missed two thirds of the material, and that the only thing standing between most teams and fixing that was a meter. Take the meter away and the better design is also the lazier one, which is the rarest thing in engineering.
FAQ
If the search is free, what stops me calling it forever?
Rate limits rather than price, and they are the number to check before you design a fan-out. A free endpoint with a per-minute ceiling and a metered endpoint with none can support very different architectures, and the pricing page usually leads with the half that sounds better. Read both.
How many phrasings should an agent actually run?
Our five phrasings found 27 pages where the best single one found 9, so the second and third phrasings clearly earn their place. We did not test where the curve flattens, and it will differ by topic. The practical approach is to add phrasings until the union stops growing on your own queries, which costs nothing to measure on a free endpoint and is the only way to get an answer that applies to you.
Which web search API should I pick?
That depends on whether you need a Google mirror, an independent index, or a fetcher, and it is a different article. The 2026 comparison names a pick and the reasoning behind it.
Does the fan-out cost anything if the calls are free?
Yes, in context rather than in currency. Thirty-eight rows collapsed to 27 unique URLs in our run, so more than a quarter of what came back was duplicate material the model would have read twice. Deduplicate on URL before anything reaches the prompt and the fan-out is close to free in both senses. Skip that step and you have traded a small bill for a large context window.
Last updated September 2026.


