[{"data":1,"prerenderedAt":2083},["ShallowReactive",2],{"blog-guide-\u002Fblog\u002Fguides\u002Ffind-all-urls-on-a-domain":3,"blog-guide-related-\u002Fblog\u002Fguides\u002Ffind-all-urls-on-a-domain":1116},{"id":4,"title":5,"author":6,"body":7,"category":1096,"cover":1097,"description":1098,"draft":1099,"extension":1100,"image":1101,"launchCta":1101,"listingCover":1102,"meta":1103,"navigation":580,"ogImage":1101,"path":1104,"publishedAt":1105,"readTime":1106,"seo":1107,"stem":1108,"tags":1109,"toolCategory":1050,"updatedAt":1105,"__hash__":1115},"blogGuides\u002Fblog\u002Fguides\u002Ffind-all-urls-on-a-domain.md","Find All URLs on a Domain: Why the Sitemap Route Usually Stops","Jasper Li",{"type":8,"value":9,"toc":1065},"minimark",[10,15,62,74,79,82,87,90,94,102,108,119,123,126,140,144,147,151,157,259,281,285,288,292,307,310,314,336,340,343,347,356,364,372,376,433,437,446,460,465,499,534,540,550,554,559,564,611,622,627,631,636,669,675,683,687,691,694,698,701,705,712,716,719,723,728,732,752,756,902,911,918,922,925,934,940,955,958,962,968,971,982,996,1000,1007,1024,1030,1047,1055,1061],[11,12,14],"p",{"style":13},"font-size:18px !important;line-height:1.65 !important;margin:0 0 24px;color:inherit;","Copy this line to your agent to inspect a domain's sitemap structure.",[16,17,22],"pre",{"className":18,"code":19,"language":20,"meta":21,"style":21},"language-sh shiki shiki-themes material-theme-lighter material-theme material-theme-palenight","set up https:\u002F\u002Fmonid.ai\u002FSKILL.md and use api.strale.io \u002Fx402\u002Fsitemap-parse on a domain\n","sh","",[23,24,25],"code",{"__ignoreMap":21},[26,27,30,34,38,41,44,47,50,53,56,59],"span",{"class":28,"line":29},"line",1,[26,31,33],{"class":32},"s2Zo4","set",[26,35,37],{"class":36},"sfazB"," up",[26,39,40],{"class":36}," https:\u002F\u002Fmonid.ai\u002FSKILL.md",[26,42,43],{"class":36}," and",[26,45,46],{"class":36}," use",[26,48,49],{"class":36}," api.strale.io",[26,51,52],{"class":36}," \u002Fx402\u002Fsitemap-parse",[26,54,55],{"class":36}," on",[26,57,58],{"class":36}," a",[26,60,61],{"class":36}," domain\n",[11,63,64,65,73],{},"\"Parse the sitemap\" is the standard answer to this question, and on 2026-09-02 it produced a usable URL count for one of the four domains tested. The other three returned a sitemap index and no count at all. That is not a broken endpoint; it is what sitemaps are actually like once a site is bigger than a brochure, and it changes the plan. This guide runs through ",[66,67,72],"a",{"href":68,"rel":69},"https:\u002F\u002Fmonid.ai",[70,71],"noopener","noreferrer","Monid",", the OpenRouter for agent tools.",[75,76,78],"h2",{"id":77},"can-you-get-every-url-on-a-domain","Can you get every URL on a domain?",[11,80,81],{},"No, and starting from that makes the rest of the work sane.",[83,84,86],"h3",{"id":85},"there-is-no-authoritative-list","There is no authoritative list",[11,88,89],{},"A website is not a directory you can list. It is whatever a server returns for whatever paths exist, and only the operator knows the full set. Everything else is inference from three sources: what the site declares, what is linked, and what someone else already indexed.",[83,91,93],{"id":92},"the-three-routes-and-what-each-one-misses","The three routes, and what each one misses",[11,95,96,97,101],{},"A ",[98,99,100],"strong",{},"sitemap"," is a declaration. It contains what the site chose to publish, which is often a subset and occasionally a superset, since sitemaps go stale and list pages that now 404.",[11,103,96,104,107],{},[98,105,106],{},"crawl"," follows links from a starting point. It finds anything reachable by link and misses anything that is not: orphan pages, paginated tails behind a script, anything gated.",[11,109,96,110,113,114,118],{},[98,111,112],{},"search index"," returns what a search engine chose to keep. It is the most curated and the least complete, and ",[66,115,117],{"href":116},"\u002Fblog\u002Fguides\u002Fserp-api-for-ai-agents","what a SERP API actually returns"," covers how much of an index you can reach that way.",[83,120,122],{"id":121},"what-all-urls-usually-means-in-practice","What \"all URLs\" usually means in practice",[11,124,125],{},"Nearly always one of: every page that could rank, every page in a section, or every page that changed recently. Those are all answerable. \"Every URL that exists\" is not, and pursuing it is how a two-hour job becomes a week.",[127,128,129],"blockquote",{},[11,130,131,132,135,136],{},"📖 ",[98,133,134],{},"See also"," ",[66,137,139],{"href":138},"\u002Fblog\u002Fguides\u002Fserp-history-track-the-shape","SERP History: Track the Shape, Not the Number",[75,141,143],{"id":142},"why-did-three-of-four-sitemaps-return-no-urls","Why did three of four sitemaps return no URLs?",[11,145,146],{},"Because they are sitemap indexes, and an index does not contain URLs. It contains other sitemaps.",[83,148,150],{"id":149},"the-measurement","The measurement",[11,152,153,156],{},[23,154,155],{},"api.strale.io\u002Fx402\u002Fsitemap-parse"," against four domains on 2026-09-02:",[158,159,160,190],"table",{},[161,162,163],"thead",{},[164,165,166,170,175,180,185],"tr",{},[167,168,169],"th",{},"Domain",[167,171,172],{},[23,173,174],{},"type",[167,176,177],{},[23,178,179],{},"total_urls",[167,181,182],{},[23,183,184],{},"has_lastmod",[167,186,187],{},[23,188,189],{},"has_priority",[191,192,193,212,229,244],"tbody",{},[164,194,195,199,204,207,210],{},[196,197,198],"td",{},"scrapingbee.com",[196,200,201],{},[23,202,203],{},"urlset",[196,205,206],{},"1,113",[196,208,209],{},"true",[196,211,209],{},[164,213,214,217,222,225,227],{},[196,215,216],{},"apify.com",[196,218,219],{},[23,220,221],{},"sitemap_index",[196,223,224],{},"absent",[196,226,224],{},[196,228,224],{},[164,230,231,234,238,240,242],{},[196,232,233],{},"notion.com",[196,235,236],{},[23,237,221],{},[196,239,224],{},[196,241,224],{},[196,243,224],{},[164,245,246,249,253,255,257],{},[196,247,248],{},"monid.ai",[196,250,251],{},[23,252,221],{},[196,254,224],{},[196,256,224],{},[196,258,224],{},[11,260,261,262,265,266,269,270,265,273,276,277,280],{},"One flat sitemap, three indexes. For the indexes the response carries ",[23,263,264],{},"child_sitemaps"," and ",[23,267,268],{},"child_count"," instead: apify.com listed 13 children including ",[23,271,272],{},"pages.xml",[23,274,275],{},"actors1.xml"," through ",[23,278,279],{},"actors5.xml","; monid.ai listed 2.",[83,282,284],{"id":283},"why-indexes-are-the-normal-case","Why indexes are the normal case",[11,286,287],{},"The sitemap protocol caps a single file at 50,000 URLs, so any site past that must split, and most sites split long before it for their own convenience: one sitemap per content type, regenerated independently. Five numbered files of actor pages is a site telling you where its bulk is.",[83,289,291],{"id":290},"the-part-that-surprised-us","The part that surprised us",[11,293,294,295,298,299,302,303,306],{},"Passing a specific child sitemap URL does not work. Given ",[23,296,297],{},"https:\u002F\u002Fmonid.ai\u002Fpages.xml",", which a plain fetch confirms is a real ",[23,300,301],{},"\u003Curlset>",", the response came back describing ",[23,304,305],{},"https:\u002F\u002Fmonid.ai\u002Fsitemap.xml"," instead. The endpoint resolves the domain's root sitemap and ignores the path you hand it.",[11,308,309],{},"So this endpoint is a reconnaissance tool, not an enumerator. It tells you how a site organises itself and roughly where the volume sits. It will not walk the index for you, and on a site with an index it returns no URLs at all.",[83,311,313],{"id":312},"one-more-thing-worth-reading","One more thing worth reading",[11,315,316,317,265,320,323,324,327,328,330,331,335],{},"For scrapingbee.com, ",[23,318,319],{},"newest_lastmod",[23,321,322],{},"oldest_lastmod"," were both 2026-08-31. Every one of 1,113 URLs carried the same date. That is a build timestamp written at deploy, not a record of when anything changed, and it means ",[23,325,326],{},"lastmod"," on that site carries no information about content freshness. If you were planning to use ",[23,329,326],{}," to detect changes, check that the values actually differ before building on them. This is the same class of check as asserting on values rather than status codes, argued in ",[66,332,334],{"href":333},"\u002Fblog\u002Fguides\u002Famazon-asin-scraper-when-it-returns-nothing","the Amazon ASIN guide",".",[75,337,339],{"id":338},"how-do-you-actually-enumerate-a-site","How do you actually enumerate a site?",[11,341,342],{},"Three steps, and the order matters because the cheap one tells you whether the expensive one is needed.",[83,344,346],{"id":345},"for-agents","For agents",[11,348,349,350,355],{},"Grab an API key at ",[66,351,354],{"href":352,"rel":353},"https:\u002F\u002Fapp.monid.ai\u002F",[70,71],"app.monid.ai",", then paste this to your agent and hand it the key:",[16,357,362],{"className":358,"code":360,"language":361,"meta":21},[359],"language-text","set up https:\u002F\u002Fmonid.ai\u002FSKILL.md\n","text",[23,363,360],{"__ignoreMap":21},[11,365,366,367,335],{},"It learns the whole discover, inspect, run workflow itself. More in the ",[66,368,371],{"href":369,"rel":370},"https:\u002F\u002Fmonid.ai\u002Fdocs\u002Fguide\u002Fquickstart-skill",[70,71],"agent quickstart",[83,373,375],{"id":374},"for-humans","For humans",[16,377,381],{"className":378,"code":379,"language":380,"meta":21,"style":21},"language-bash shiki shiki-themes material-theme-lighter material-theme material-theme-palenight","npm install -g @monid-ai\u002Fcli\nmonid keys add -k \u003Cyour-api-key> -l main\n","bash",[23,382,383,398],{"__ignoreMap":21},[26,384,385,389,392,395],{"class":28,"line":29},[26,386,388],{"class":387},"sBMFI","npm",[26,390,391],{"class":36}," install",[26,393,394],{"class":36}," -g",[26,396,397],{"class":36}," @monid-ai\u002Fcli\n",[26,399,401,404,407,410,413,417,420,424,427,430],{"class":28,"line":400},2,[26,402,403],{"class":387},"monid",[26,405,406],{"class":36}," keys",[26,408,409],{"class":36}," add",[26,411,412],{"class":36}," -k",[26,414,416],{"class":415},"sMK4o"," \u003C",[26,418,419],{"class":36},"your-api-ke",[26,421,423],{"class":422},"sTEyZ","y",[26,425,426],{"class":415},">",[26,428,429],{"class":36}," -l",[26,431,432],{"class":36}," main\n",[83,434,436],{"id":435},"step-1-reconnoitre-with-the-sitemap","Step 1. Reconnoitre with the sitemap",[11,438,439,442,443,445],{},[98,440,441],{},"What it does."," Tells you the shape of the site in one call: flat or indexed, how many sections, and whether ",[23,444,326],{}," is real.",[11,447,448,135,451,456,457,335],{},[98,449,450],{},"The endpoints.",[66,452,454],{"href":453},"\u002Ftools\u002Fweb-extraction",[23,455,155],{},", billed per call. Takes ",[23,458,459],{},"url",[11,461,462],{},[98,463,464],{},"The call.",[16,466,468],{"className":378,"code":467,"language":380,"meta":21,"style":21},"monid run -p api.strale.io -e \u002Fx402\u002Fsitemap-parse --query '{\"url\": \"https:\u002F\u002Fexample.com\"}'\n",[23,469,470],{"__ignoreMap":21},[26,471,472,474,477,480,482,485,487,490,493,496],{"class":28,"line":29},[26,473,403],{"class":387},[26,475,476],{"class":36}," run",[26,478,479],{"class":36}," -p",[26,481,49],{"class":36},[26,483,484],{"class":36}," -e",[26,486,52],{"class":36},[26,488,489],{"class":36}," --query",[26,491,492],{"class":415}," '",[26,494,495],{"class":36},"{\"url\": \"https:\u002F\u002Fexample.com\"}",[26,497,498],{"class":415},"'\n",[11,500,501,504,505,508,509,511,512,511,514,511,517,511,519,511,521,511,523,526,527,508,530,265,532,335],{},[98,502,503],{},"What comes back."," Either ",[23,506,507],{},"type: \"urlset\""," with ",[23,510,179],{},", ",[23,513,184],{},[23,515,516],{},"has_changefreq",[23,518,189],{},[23,520,319],{},[23,522,322],{},[23,524,525],{},"top_path_segments"," and a sample; or ",[23,528,529],{},"type: \"sitemap_index\"",[23,531,264],{},[23,533,268],{},[11,535,536,537,539],{},"Read ",[23,538,525],{}," when you get it. It tells you how the site's URL space is divided, which is usually what someone actually wanted when they asked for every URL.",[11,541,542,545,546,335],{},[98,543,544],{},"What it costs."," Per call. Current figures at ",[66,547,549],{"href":548},"\u002Ftools","monid.ai\u002Ftools",[83,551,553],{"id":552},"step-2-fetch-the-child-sitemaps-yourself","Step 2. Fetch the child sitemaps yourself",[11,555,556,558],{},[98,557,441],{}," Fills the gap the endpoint leaves.",[11,560,561,563],{},[98,562,464],{}," No endpoint needed. Child sitemaps are plain XML over HTTP:",[16,565,569],{"className":566,"code":567,"language":568,"meta":21,"style":21},"language-python shiki shiki-themes material-theme-lighter material-theme material-theme-palenight","import re, urllib.request\n\ndef urls_in(sitemap_url):\n    xml = urllib.request.urlopen(sitemap_url).read().decode()\n    return re.findall(r\"\u003Cloc>(.*?)\u003C\u002Floc>\", xml)\n\nall_urls = [u for child in child_sitemaps for u in urls_in(child)]\n","python",[23,570,571,576,582,588,594,600,605],{"__ignoreMap":21},[26,572,573],{"class":28,"line":29},[26,574,575],{},"import re, urllib.request\n",[26,577,578],{"class":28,"line":400},[26,579,581],{"emptyLinePlaceholder":580},true,"\n",[26,583,585],{"class":28,"line":584},3,[26,586,587],{},"def urls_in(sitemap_url):\n",[26,589,591],{"class":28,"line":590},4,[26,592,593],{},"    xml = urllib.request.urlopen(sitemap_url).read().decode()\n",[26,595,597],{"class":28,"line":596},5,[26,598,599],{},"    return re.findall(r\"\u003Cloc>(.*?)\u003C\u002Floc>\", xml)\n",[26,601,603],{"class":28,"line":602},6,[26,604,581],{"emptyLinePlaceholder":580},[26,606,608],{"class":28,"line":607},7,[26,609,610],{},"all_urls = [u for child in child_sitemaps for u in urls_in(child)]\n",[11,612,613,614,617,618,335],{},"Regex on XML is normally a bad idea and ",[23,615,616],{},"\u003Cloc>"," is the exception that survives, because the element is flat and the protocol is rigid. Use a real parser if the sitemap carries image or video extensions you care about. The same durability argument applies to any selector you write, per ",[66,619,621],{"href":620},"\u002Fblog\u002Fguides\u002Fxpath-cheat-sheet-for-scraping","the XPath guide",[11,623,624,626],{},[98,625,544],{}," Nothing but your own bandwidth.",[83,628,630],{"id":629},"step-3-crawl-only-what-the-sitemap-does-not-cover","Step 3. Crawl only what the sitemap does not cover",[11,632,633,635],{},[98,634,441],{}," Finds linked pages the site never declared.",[11,637,638,135,640,645,646,651,652,511,654,511,657,511,660,511,663,265,666,335],{},[98,639,450],{},[66,641,642],{"href":453},[23,643,644],{},"context.dev\u002Fweb\u002Fcrawl"," returns crawled pages as Markdown; ",[66,647,648],{"href":453},[23,649,650],{},"mrscraper\u002Fscrape\u002Fmap"," discovers URL structure and takes ",[23,653,459],{},[23,655,656],{},"maxDepth",[23,658,659],{},"maxPages",[23,661,662],{},"limit",[23,664,665],{},"includePatterns",[23,667,668],{},"excludePatterns",[11,670,671,674],{},[98,672,673],{},"An honest note."," Both are asynchronous, and on 2026-09-02 the Monid CLI's polling path returned an HTML redirect rather than JSON on both 0.1.6 and 0.1.7, so neither is reported here as measured. The schemas are as described; the responses are not something this article verified, and saying so seemed better than describing output we did not see.",[11,676,677,682],{},[98,678,679,680,335],{},"The point of ",[23,681,668],{}," Faceted navigation generates effectively infinite URLs, and a crawl without exclusions on an ecommerce site will happily enumerate every colour-and-size combination until you stop paying. Set them before the first run, not after the bill.",[684,685],"skill-prompt",{"prompt":686},"check the sitemap structure for these 12 domains and tell me which ones are flat, which are indexed, and which have a useless lastmod",[75,688,690],{"id":689},"what-does-a-sitemap-tell-you-that-a-crawl-does-not","What does a sitemap tell you that a crawl does not?",[11,692,693],{},"Four things, and they are the reason to start here even though it does not enumerate.",[83,695,697],{"id":696},"intent","Intent",[11,699,700],{},"A sitemap is the site's own statement about which pages matter. A crawl cannot distinguish a flagship page from a tag archive; a sitemap that omits the tag archives has told you.",[83,702,704],{"id":703},"structure-cheaply","Structure, cheaply",[11,706,707,265,709,711],{},[23,708,264],{},[23,710,525],{}," describe how a site is organised for the price of one call. Deriving the same picture from a crawl costs thousands of requests.",[83,713,715],{"id":714},"orphans-by-subtraction","Orphans, by subtraction",[11,717,718],{},"Pages listed in a sitemap but not reachable by link are orphans, and they are usually a bug worth knowing about. You can only find them by comparing the two sources, which is a good reason to collect both rather than picking one.",[83,720,722],{"id":721},"freshness-claims-and-whether-to-believe-them","Freshness claims, and whether to believe them",[11,724,725,727],{},[23,726,326],{}," is the only cheap change signal available, when it is real. Today it was not on the one site where we could check it. That is worth knowing before designing an incremental pipeline around it.",[83,729,731],{"id":730},"the-metadata-that-tells-you-which-is-which","The metadata that tells you which is which",[11,733,734,735,738,739,742,743,747,748,751],{},"Both endpoints from this vendor return a ",[23,736,737],{},"_meta.provenance.source",". On the sitemap parser it reads ",[23,740,741],{},"http-fetch",". On ",[66,744,746],{"href":745},"\u002Fblog\u002Fguides\u002Fextract-pricing-and-free-trial-information","the pricing extractor"," it reads ",[23,749,750],{},"claude-haiku",", meaning a model produced that output. Same vendor, same envelope, and one field separating a deterministic parse from a generated one. That is the field to check before deciding how hard to verify a response.",[75,753,755],{"id":754},"which-endpoint-should-i-use-for-which-job","Which endpoint should I use for which job?",[158,757,758,780],{},[161,759,760],{},[164,761,762,765,768,771,774,777],{},[167,763,764],{},"Endpoint",[167,766,767],{},"What it does",[167,769,770],{},"Input",[167,772,773],{},"Output",[167,775,776],{},"Best for",[167,778,779],{},"Billing",[191,781,782,806,831,855,879],{},[164,783,784,790,793,797,800,803],{},[196,785,786],{},[66,787,788],{"href":453},[23,789,155],{},[196,791,792],{},"Sitemap reconnaissance",[196,794,795],{},[23,796,459],{},[196,798,799],{},"Type, counts or child list, path segments",[196,801,802],{},"Understanding a site in one call",[196,804,805],{},"Per call",[164,807,808,814,817,822,825,828],{},[196,809,810],{},[66,811,812],{"href":453},[23,813,650],{},[196,815,816],{},"Crawl for URL structure",[196,818,819,821],{},[23,820,459],{},", depth, limits, patterns",[196,823,824],{},"Discovered URLs",[196,826,827],{},"Enumerating linked pages",[196,829,830],{},"Per result",[164,832,833,839,842,847,850,853],{},[196,834,835],{},[66,836,837],{"href":453},[23,838,644],{},[196,840,841],{},"Crawl and convert",[196,843,844,846],{},[23,845,459],{},", limit",[196,848,849],{},"Pages as Markdown",[196,851,852],{},"You want content, not just URLs",[196,854,830],{},[164,856,857,865,868,871,874,877],{},[196,858,859],{},[66,860,862],{"href":861},"\u002Ftools\u002Fsearch",[23,863,864],{},"context.dev\u002Fweb\u002Fsearch",[196,866,867],{},"Search, optionally scrape",[196,869,870],{},"A query",[196,872,873],{},"Results with optional content",[196,875,876],{},"You want indexed pages only",[196,878,830],{},[164,880,881,888,891,894,897,900],{},[196,882,883],{},[66,884,885],{"href":861},[23,886,887],{},"ahrefs\u002Fsite-explorer\u002Fall-backlinks",[196,889,890],{},"Externally linked pages",[196,892,893],{},"A domain",[196,895,896],{},"Backlinks with target URLs",[196,898,899],{},"Finding pages others link to",[196,901,830],{},[11,903,904,905,908,909,335],{},"Every row was verified with ",[23,906,907],{},"monid inspect"," on 2026-09-02. The table gives billing shape rather than figures; shape drives design and current numbers live on ",[66,910,549],{"href":548},[11,912,913,914,335],{},"The last row is a genuinely different discovery source and it is the one people forget. Backlink data surfaces URLs that outsiders link to, which will include old pages a redesign orphaned and pages the current sitemap has dropped. If your goal is \"find everything that might still get traffic\", that set is not reachable from either a sitemap or a crawl. We used exactly that source to audit our own link graph, described in ",[66,915,917],{"href":916},"\u002Fblog\u002Freal-cost-of-scraping-youtube-yourself","the crawler build-versus-buy piece",[75,919,921],{"id":920},"when-should-you-not-enumerate-a-site","When should you not enumerate a site?",[11,923,924],{},"Three cases.",[11,926,927,930,931,933],{},[98,928,929],{},"You need a count, not a list."," If the question is \"how big is this site\", the sitemap total or a ",[23,932,268],{}," answers it for one call. A crawl to establish the same number is an expensive way to get one integer.",[11,935,936,939],{},[98,937,938],{},"The site is generated combinatorially."," Faceted search, calendars and parameterised filters produce unbounded URL spaces where \"all URLs\" is not a finite set. Decide which parameters matter and exclude the rest, or the crawl never ends.",[11,941,942,945,946,949,950,954],{},[98,943,944],{},"Someone else's site, at volume."," A thorough crawl is a load event for the operator. Rate-limit, respect ",[23,947,948],{},"robots.txt",", and prefer the sitemap for anything you can get from it. The politeness arithmetic is the same one in ",[66,951,953],{"href":952},"\u002Fblog\u002Fguides\u002Fdo-ai-agents-need-a-rotating-proxy","the rotating proxy guide",". The 502 we hit on a different endpoint yesterday was a reminder that origins fall over, and being the cause is avoidable.",[11,956,957],{},"And the disclosure: this is Monid's blog and we sell per-call access to all of the above. The honest summary of the paid sitemap endpoint is that it is reconnaissance rather than enumeration, it stops at the index on most real sites, and the follow-up step in this article is free HTTP you can do yourself.",[75,959,961],{"id":960},"conclusion","Conclusion",[11,963,964,965,967],{},"The sitemap route is the right first move and it is not the answer. On 2026-09-02, three of four domains returned a sitemap index rather than a URL list, and the endpoint reports the index without walking it. On the fourth, 1,113 URLs carried an identical ",[23,966,326],{},", which means the freshness field there is a deploy timestamp rather than a change signal.",[11,969,970],{},"So the working sequence is: one call to learn the shape, free HTTP to pull the child sitemaps, and a crawl only for what the site never declared, with exclusion patterns set before the first run. Reach for backlink data when the goal includes pages the site has forgotten about.",[11,972,973,974,976,977,265,979,981],{},"And check ",[23,975,737],{}," when a response offers it. ",[23,978,741],{},[23,980,750],{}," are very different promises about the same-looking JSON.",[11,983,984,985,988,989,992,993,335],{},"Free next step: ",[23,986,987],{},"curl https:\u002F\u002Fyourtarget.com\u002Frobots.txt"," and look for the ",[23,990,991],{},"Sitemap:"," lines. It costs nothing, it often names sitemaps the root file omits, and it takes about five seconds. Start at ",[66,994,248],{"href":68,"rel":995},[70,71],[75,997,999],{"id":998},"faq","FAQ",[1001,1002,1004],"faq-item",{"q":1003},"Why does a site: search show a different number every time?",[11,1005,1006],{},"Because it was never a count. The figure above search results is an estimate produced for display, it varies between data centres and refreshes, and it is not a promise about how many pages are indexed. It also caps out well before it shows you everything, so paging through the results is not a workaround. Treat it as a rough order of magnitude and get real numbers from the sitemap or a crawl.",[1001,1008,1010],{"q":1009},"Should you read robots.txt to find sitemaps?",[11,1011,1012,1013,1015,1016,1019,1020,1023],{},"Yes, and it is the cheapest step in the whole process. The ",[23,1014,991],{}," directive is how a site declares sitemaps that are not at the conventional location, and larger sites often list several there that ",[23,1017,1018],{},"\u002Fsitemap.xml"," does not reference. It costs one unauthenticated fetch and no API call. While you are there, read the ",[23,1021,1022],{},"Disallow"," rules, because they tell you both what the operator does not want crawled and, usefully, which path prefixes exist.",[1001,1025,1027],{"q":1026},"Which URLs will neither a sitemap nor a crawl find?",[11,1028,1029],{},"Anything behind a login, anything reachable only by form submission or a script-built link, orphan pages that nothing links to and no sitemap lists, and pages deliberately excluded from both. Old pages that still return content after a redesign are the common and costly case, since they can still hold links and traffic while being invisible to every discovery method except backlink data. If completeness matters, the operator's own server logs or CMS export are the only complete source, and if you have access to those you should use them instead of any of this.",[1001,1031,1033],{"q":1032},"How large does a site crawl actually get?",[11,1034,1035,1036,511,1038,265,1040,1042,1043,335],{},"Larger than the sitemap suggests, usually by a lot, because every faceted filter combination is a distinct URL. A catalogue with 5,000 products and four filter dimensions can expose millions of valid URLs, none of which are pages anyone wants. This is why ",[23,1037,659],{},[23,1039,662],{},[23,1041,668],{}," exist and why setting them is not optional: on per-result billing, an unbounded crawl is an unbounded bill. Cap the run, look at what came back, then widen deliberately, the same escalation logic as ",[66,1044,1046],{"href":1045},"\u002Fblog\u002Fguides\u002Fjob-scraping-software-per-board","job scraping per board",[1048,1049,1052],"tool-cta",{"category":1050,"title":1051},"web-extraction","Learn the shape before you crawl",[11,1053,1054],{},"One call tells you whether a site is flat or indexed, where its volume sits, and whether its freshness dates mean anything.",[11,1056,1057],{},[1058,1059,1060],"em",{},"Last updated September 2026.",[1062,1063,1064],"style",{},"html pre.shiki code .s2Zo4, html code.shiki .s2Zo4{--shiki-light:#6182B8;--shiki-default:#82AAFF;--shiki-dark:#82AAFF}html pre.shiki code .sfazB, html code.shiki .sfazB{--shiki-light:#91B859;--shiki-default:#C3E88D;--shiki-dark:#C3E88D}html .light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html.light .shiki span {color: var(--shiki-light);background: var(--shiki-light-bg);font-style: var(--shiki-light-font-style);font-weight: var(--shiki-light-font-weight);text-decoration: var(--shiki-light-text-decoration);}html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html pre.shiki code .sBMFI, html code.shiki .sBMFI{--shiki-light:#E2931D;--shiki-default:#FFCB6B;--shiki-dark:#FFCB6B}html pre.shiki code .sMK4o, html code.shiki .sMK4o{--shiki-light:#39ADB5;--shiki-default:#89DDFF;--shiki-dark:#89DDFF}html pre.shiki code .sTEyZ, html code.shiki .sTEyZ{--shiki-light:#90A4AE;--shiki-default:#EEFFFF;--shiki-dark:#BABED8}",{"title":21,"searchDepth":400,"depth":400,"links":1066},[1067,1072,1078,1085,1092,1093,1094,1095],{"id":77,"depth":400,"text":78,"children":1068},[1069,1070,1071],{"id":85,"depth":584,"text":86},{"id":92,"depth":584,"text":93},{"id":121,"depth":584,"text":122},{"id":142,"depth":400,"text":143,"children":1073},[1074,1075,1076,1077],{"id":149,"depth":584,"text":150},{"id":283,"depth":584,"text":284},{"id":290,"depth":584,"text":291},{"id":312,"depth":584,"text":313},{"id":338,"depth":400,"text":339,"children":1079},[1080,1081,1082,1083,1084],{"id":345,"depth":584,"text":346},{"id":374,"depth":584,"text":375},{"id":435,"depth":584,"text":436},{"id":552,"depth":584,"text":553},{"id":629,"depth":584,"text":630},{"id":689,"depth":400,"text":690,"children":1086},[1087,1088,1089,1090,1091],{"id":696,"depth":584,"text":697},{"id":703,"depth":584,"text":704},{"id":714,"depth":584,"text":715},{"id":721,"depth":584,"text":722},{"id":730,"depth":584,"text":731},{"id":754,"depth":400,"text":755},{"id":920,"depth":400,"text":921},{"id":960,"depth":400,"text":961},{"id":998,"depth":400,"text":999},"Search & RAG","\u002Fimg\u002Fblog\u002Ffind-all-urls-on-a-domain.png","Four domains, one sitemap endpoint. Three returned an index and no URL count at all. A sitemap gives you a map of a site, not a list of its pages.",false,"md",null,"\u002Fimg\u002Fblog\u002Ffind-all-urls-on-a-domain-card.png",{},"\u002Fblog\u002Fguides\u002Ffind-all-urls-on-a-domain","2026-09-02","11 min",{"title":5,"description":1098},"blog\u002Fguides\u002Ffind-all-urls-on-a-domain",[1110,1111,1112,1113,1114],"find all urls on a domain","sitemap parser","site crawl","url discovery","seo audit","Clki1TUxKeDKIFUE5ms6tswkSNOdzV3a8QBjEQemeu8",[1117,1131,1144,1158,1170,1182,1193,1195,1208,1220,1233,1245,1257,1269,1283,1295,1307,1317,1328,1340,1352,1362,1373,1384,1395,1405,1417,1427,1437,1446,1456,1468,1480,1493,1504,1514,1527,1538,1550,1561,1570,1580,1589,1598,1604,1613,1623,1634,1644,1655,1667,1674,1684,1692,1700,1710,1720,1731,1739,1748,1758,1769,1779,1789,1798,1808,1819,1828,1837,1848,1858,1866,1877,1887,1897,1905,1915,1925,1933,1942,1950,1960,1970,1982,1991,2001,2010,2019,2030,2039,2048,2056,2065,2074],{"path":1118,"title":1119,"description":1120,"publishedAt":1121,"author":6,"category":1122,"tags":1123,"readTime":1106,"cover":1129,"listingCover":1130,"image":1101},"\u002Fblog\u002Fguides\u002Fare-instagram-engagement-trackers-accurate","Are Instagram Follower and Engagement Trackers Accurate?","Two providers disagreed by 2 followers out of 104 million. The same 12 posts gave 0.096% or 0.334% engagement depending on one undocumented choice.","2026-09-03","Social data",[1124,1125,1126,1127,1128],"instagram engagement rate","follower count accuracy","influencer vetting","social data","instagram api","\u002Fimg\u002Fblog\u002Fare-instagram-engagement-trackers-accurate.png","\u002Fimg\u002Fblog\u002Fare-instagram-engagement-trackers-accurate-card.png",{"path":1132,"title":1133,"description":1134,"publishedAt":1121,"author":6,"category":1135,"tags":1136,"readTime":1106,"cover":1142,"listingCover":1143,"image":1101},"\u002Fblog\u002Fguides\u002Fbounce-rates-and-spam-traps-when-validating-email","Bounce Rates and Spam Traps When Validating an Email List","A verifier returned twelve fields for three addresses. None of them was spam_trap, and none could be. Here is what verification fixes and what it cannot.","Sales & enrichment",[1137,1138,1139,1140,1141],"spam traps","email bounce rate","email verification","deliverability","list hygiene","\u002Fimg\u002Fblog\u002Fbounce-rates-and-spam-traps-when-validating-email.png","\u002Fimg\u002Fblog\u002Fbounce-rates-and-spam-traps-when-validating-email-card.png",{"path":1145,"title":1146,"description":1147,"publishedAt":1121,"author":6,"category":1148,"tags":1149,"readTime":1155,"cover":1156,"listingCover":1157,"image":1101},"\u002Fblog\u002Fguides\u002Frate-limit-calls-to-a-third-party-api","How to Rate Limit Calls to a Third Party API","Six real failure responses collected over three days. Only one of them meant slow down, and telling them apart matters more than any backoff algorithm.","Product",[1150,1151,1152,1153,1154],"rate limiting","retry backoff","429 too many requests","api reliability","concurrency","12 min","\u002Fimg\u002Fblog\u002Frate-limit-calls-to-a-third-party-api.png","\u002Fimg\u002Fblog\u002Frate-limit-calls-to-a-third-party-api-card.png",{"path":1159,"title":1160,"description":1161,"publishedAt":1121,"author":6,"category":1135,"tags":1162,"readTime":1106,"cover":1168,"listingCover":1169,"image":1101},"\u002Fblog\u002Fguides\u002Fscrape-linkedin-without-getting-banned","How to Scrape LinkedIn Without Getting Your Account Banned","The ban risk comes from tools that drive your logged-in session. Three endpoints returned profile data from a URL alone, with no cookie involved.",[1163,1164,1165,1166,1167],"linkedin scraping","account ban","li_at cookie","sales navigator","linkedin api","\u002Fimg\u002Fblog\u002Fscrape-linkedin-without-getting-banned.png","\u002Fimg\u002Fblog\u002Fscrape-linkedin-without-getting-banned-card.png",{"path":1171,"title":1172,"description":1173,"publishedAt":1121,"author":6,"category":1148,"tags":1174,"readTime":1155,"cover":1180,"listingCover":1181,"image":1101},"\u002Fblog\u002Fguides\u002Fstop-depending-on-one-scraping-vendor","How to Stop Depending on One Scraping Vendor","Ten provider failures in three days, all while doing other work. None was a bad vendor. The defence is a second route you can reach without a rewrite.",[1175,1176,1177,1178,1179],"vendor lock-in","scraping reliability","actor removed","failover","data pipeline","\u002Fimg\u002Fblog\u002Fstop-depending-on-one-scraping-vendor.png","\u002Fimg\u002Fblog\u002Fstop-depending-on-one-scraping-vendor-card.png",{"path":745,"title":1183,"description":1184,"publishedAt":1105,"author":6,"category":1148,"tags":1185,"readTime":1155,"cover":1191,"listingCover":1192,"image":1101},"Extract Pricing and Free Trial Information From Websites","One endpoint returned four plans, numeric amounts and a free_trial_available flag. It also disclosed that a model read the page, which changes everything.",[1186,1187,1188,1189,1190],"pricing page extraction","competitor pricing","free trial data","saas pricing","llm extraction","\u002Fimg\u002Fblog\u002Fextract-pricing-and-free-trial-information.png","\u002Fimg\u002Fblog\u002Fextract-pricing-and-free-trial-information-card.png",{"path":1104,"title":5,"description":1098,"publishedAt":1105,"author":6,"category":1096,"tags":1194,"readTime":1106,"cover":1097,"listingCover":1102,"image":1101},[1110,1111,1112,1113,1114],{"path":1196,"title":1197,"description":1198,"publishedAt":1105,"author":6,"category":1199,"tags":1200,"readTime":1155,"cover":1206,"listingCover":1207,"image":1101},"\u002Fblog\u002Fguides\u002Fmls-real-estate-api-without-mls-access","MLS Real Estate API: What You Get When You Cannot Get MLS","MLS access needs a licence and a broker. Two public sources returned 40 listings and 48 agents with phones. What they carry, and what they do not.","Local data",[1201,1202,1203,1204,1205],"mls real estate api","idx feed","property data","real estate agents","listing data","\u002Fimg\u002Fblog\u002Fmls-real-estate-api-without-mls-access.png","\u002Fimg\u002Fblog\u002Fmls-real-estate-api-without-mls-access-card.png",{"path":1209,"title":1210,"description":1211,"publishedAt":1105,"author":6,"category":1148,"tags":1212,"readTime":1106,"cover":1218,"listingCover":1219,"image":1101},"\u002Fblog\u002Fguides\u002Fsimply-wall-st-graphql-api-fundamentals","Simply Wall St GraphQL API: What Returns Fundamentals Instead","There is no public Simply Wall St GraphQL API. An endpoint that does return fundamentals gave 26 line items across six periods, and named its own source.",[1213,1214,1215,1216,1217],"simply wall st api","stock fundamentals api","income statement api","financial data","graphql","\u002Fimg\u002Fblog\u002Fsimply-wall-st-graphql-api-fundamentals.png","\u002Fimg\u002Fblog\u002Fsimply-wall-st-graphql-api-fundamentals-card.png",{"path":333,"title":1221,"description":1222,"publishedAt":1223,"author":6,"category":1224,"tags":1225,"readTime":1106,"cover":1231,"listingCover":1232,"image":1101},"Amazon ASIN Scraper: What to Do When It Returns Nothing","A purpose-built ASIN endpoint returned 26 fields and zero values for two valid ASINs. A generic extractor read the same page fine. Assert on values.","2026-09-01","Ecommerce",[1226,1227,1228,1229,1230],"amazon asin scraper","asin lookup","ecommerce data","product data","silent failure","\u002Fimg\u002Fblog\u002Famazon-asin-scraper-when-it-returns-nothing.png","\u002Fimg\u002Fblog\u002Famazon-asin-scraper-when-it-returns-nothing-card.png",{"path":1234,"title":1235,"description":1236,"publishedAt":1223,"author":6,"category":1148,"tags":1237,"readTime":1155,"cover":1243,"listingCover":1244,"image":1101},"\u002Fblog\u002Fguides\u002Fdatacenter-proxies-vs-residential","Datacenter Proxies vs Residential: What Each One Actually Fixes","A real residential IP with a real Chrome user agent still got 403 from two sites. The IP type was not the constraint. What each type does fix.",[1238,1239,1240,1241,1242],"datacenter proxies vs residential","proxy types","rotating proxies","web scraping","geolocation","\u002Fimg\u002Fblog\u002Fdatacenter-proxies-vs-residential.png","\u002Fimg\u002Fblog\u002Fdatacenter-proxies-vs-residential-card.png",{"path":1246,"title":1247,"description":1248,"publishedAt":1223,"author":6,"category":1135,"tags":1249,"readTime":1106,"cover":1255,"listingCover":1256,"image":1101},"\u002Fblog\u002Fguides\u002Fextracting-leadership-contact-information","Extracting Leadership Contact Information Without Paying to Look","Hunter's multi-domain search returned 404,695 matching executives for free, names masked and titles visible. You pay only for the addresses you choose.",[1250,1251,1252,1253,1254],"leadership contact information","executive contact data","decision maker search","hunter io","b2b prospecting","\u002Fimg\u002Fblog\u002Fextracting-leadership-contact-information.png","\u002Fimg\u002Fblog\u002Fextracting-leadership-contact-information-card.png",{"path":1258,"title":1259,"description":1260,"publishedAt":1223,"author":6,"category":1148,"tags":1261,"readTime":1155,"cover":1267,"listingCover":1268,"image":1101},"\u002Fblog\u002Fguides\u002Fpuppeteer-extra-plugin-stealth-release-date","Puppeteer Extra Plugin Stealth: Check the Release Date First","The stealth plugin last shipped in March 2023. Puppeteer shipped last week. What that gap means, and what scrapy-playwright does differently.",[1262,1263,1264,1265,1266],"puppeteer-extra-plugin-stealth","scrapy-playwright","headless browser","bot detection","browser automation","\u002Fimg\u002Fblog\u002Fpuppeteer-extra-plugin-stealth-release-date.png","\u002Fimg\u002Fblog\u002Fpuppeteer-extra-plugin-stealth-release-date-card.png",{"path":1270,"title":1271,"description":1272,"publishedAt":1273,"author":6,"category":1096,"tags":1274,"readTime":1280,"cover":1281,"listingCover":1282,"image":1101},"\u002Fblog\u002Fguides\u002Fc-sharp-web-scraping-without-the-stack","C# Web Scraping: Keep HtmlAgilityPack, Drop the Fetch","The .NET parsing stack was never the problem. What breaks a C# scraper is the request leaving your server, and that half is not a language question.","2026-08-31",[1275,1276,1277,1278,1279],"c# web scraping","htmlagilitypack","dotnet","web extraction","playwright","10 min","\u002Fimg\u002Fblog\u002Fc-sharp-web-scraping-without-the-stack.png","\u002Fimg\u002Fblog\u002Fc-sharp-web-scraping-without-the-stack-card.png",{"path":1284,"title":1285,"description":1286,"publishedAt":1273,"author":6,"category":1135,"tags":1287,"readTime":1106,"cover":1293,"listingCover":1294,"image":1101},"\u002Fblog\u002Fguides\u002Flinkedin-post-scraper-what-comes-back","LinkedIn Post Scraper: What Actually Comes Back","Posts, reactions and comments are three billing decisions in one call. A live run shows the fields, and the two flags that quietly multiply the bill.",[1288,1289,1290,1291,1292],"linkedin post scraper","linkedin company scraper","social selling","harvestapi","linkedin data","\u002Fimg\u002Fblog\u002Flinkedin-post-scraper-what-comes-back.png","\u002Fimg\u002Fblog\u002Flinkedin-post-scraper-what-comes-back-card.png",{"path":1296,"title":1297,"description":1298,"publishedAt":1273,"author":6,"category":1096,"tags":1299,"readTime":1106,"cover":1305,"listingCover":1306,"image":1101},"\u002Fblog\u002Fguides\u002Fweb-scraping-vs-api-wrong-question","Web Scraping vs API: You Are Choosing a Maintainer","Both end in an HTTP request returning the same facts. What differs is who owns the parser when the page changes, and whether anyone promised you anything.",[1300,1301,1302,1303,1304],"web scraping vs api","data acquisition","buy vs build","scraping api","ai agents","\u002Fimg\u002Fblog\u002Fweb-scraping-vs-api-wrong-question.png","\u002Fimg\u002Fblog\u002Fweb-scraping-vs-api-wrong-question-card.png",{"path":620,"title":1308,"description":1309,"publishedAt":1273,"author":6,"category":1096,"tags":1310,"readTime":1106,"cover":1315,"listingCover":1316,"image":1101},"XPath Cheat Sheet: The Dozen Expressions That Survive","Most XPath references list what the spec allows. This lists what still matches after a redesign, ranked by how long each anchor survives.",[1311,1312,1313,1278,1314],"xpath cheat sheet","xpath","css selectors","scraping","\u002Fimg\u002Fblog\u002Fxpath-cheat-sheet-for-scraping.png","\u002Fimg\u002Fblog\u002Fxpath-cheat-sheet-for-scraping-card.png",{"path":1318,"title":1319,"description":1320,"publishedAt":1273,"author":6,"category":1199,"tags":1321,"readTime":1106,"cover":1326,"listingCover":1327,"image":1101},"\u002Fblog\u002Fguides\u002Fzillow-scraper-what-the-listing-carries","Zillow Scraper: What One Listing Record Carries","A live search returns zpid, unformatted price, beds and the detail URL. Five endpoints split by intent, and the sold one is the underrated half.",[1322,1323,1203,1324,1325],"zillow scraper","real estate api","local data","listings","\u002Fimg\u002Fblog\u002Fzillow-scraper-what-the-listing-carries.png","\u002Fimg\u002Fblog\u002Fzillow-scraper-what-the-listing-carries-card.png",{"path":1329,"title":1330,"description":1331,"publishedAt":1332,"author":6,"category":1333,"tags":1334,"readTime":1106,"cover":1338,"listingCover":1339,"image":1101},"\u002Fblog\u002Fguides\u002Fad-agent-auction-outside-the-account","Soku Runs the Campaign. The Auction Sits Outside It.","One small advertiser shares a single keyword with a bidder running 198,105 of them. An ad agent optimising inside its own account cannot see that.","2026-08-30","Ads intelligence",[1335,1336,1337,1304],"paid search","ad agents","competitor intelligence","\u002Fimg\u002Fblog\u002Fad-agent-auction-outside-the-account.png","\u002Fimg\u002Fblog\u002Fad-agent-auction-outside-the-account-card.png",{"path":1341,"title":1342,"description":1343,"publishedAt":1332,"author":6,"category":1096,"tags":1344,"readTime":1106,"cover":1350,"listingCover":1351,"image":1101},"\u002Fblog\u002Fguides\u002Fbing-serp-tracker-after-the-api","Bing SERP Tracker: What to Use After the API Retired","Microsoft retired the Search API, so a Bing position now comes from three places. One is free and first-party, and most articles do not mention it.",[1345,1346,1347,1348,1349],"bing serp tracker","bing rank tracking","serp api","bing webmaster tools","rank tracking","\u002Fimg\u002Fblog\u002Fbing-serp-tracker-after-the-api.png","\u002Fimg\u002Fblog\u002Fbing-serp-tracker-after-the-api-card.png",{"path":1353,"title":1354,"description":1355,"publishedAt":1332,"author":6,"category":1148,"tags":1356,"readTime":1106,"cover":1360,"listingCover":1361,"image":1101},"\u002Fblog\u002Fguides\u002Fbrowser-automation-without-an-api","Browser Automation When the Site Has No API","Check the premise first. The site owner has no API, but somebody may already have an endpoint for it, and a browser is the most expensive answer available.",[1266,1357,1358,1359,1304],"antidetect browser","nodriver","no api","\u002Fimg\u002Fblog\u002Fbrowser-automation-without-an-api.png","\u002Fimg\u002Fblog\u002Fbrowser-automation-without-an-api-card.png",{"path":1045,"title":1363,"description":1364,"publishedAt":1332,"author":6,"category":1135,"tags":1365,"readTime":1106,"cover":1371,"listingCover":1372,"image":1101},"Job Scraping Software: There Is No Single Job Feed","Postings are fragmented across boards by design, and no endpoint covers all of them. Pick per board, and the schema differences are the real work.",[1366,1367,1368,1369,1370],"job scraping software","indeed job scraper","jobspy","job postings api","hiring signals","\u002Fimg\u002Fblog\u002Fjob-scraping-software-per-board.png","\u002Fimg\u002Fblog\u002Fjob-scraping-software-per-board-card.png",{"path":1374,"title":1375,"description":1376,"publishedAt":1332,"author":6,"category":1135,"tags":1377,"readTime":1106,"cover":1382,"listingCover":1383,"image":1101},"\u002Fblog\u002Fguides\u002Fmid-form-enrichment-fewer-questions","TinyCommand Asks One Question. The API Fills Twenty-Six.","One domain went in. Twenty-six fields came back, including a 43-item tech stack. Every one of them is a question the form no longer has to ask.",[1378,1379,1380,1381],"lead enrichment","forms","progressive profiling","workflow automation","\u002Fimg\u002Fblog\u002Fmid-form-enrichment-fewer-questions.png","\u002Fimg\u002Fblog\u002Fmid-form-enrichment-fewer-questions-card.png",{"path":1385,"title":1386,"description":1387,"publishedAt":1332,"author":6,"category":1096,"tags":1388,"readTime":1280,"cover":1393,"listingCover":1394,"image":1101},"\u002Fblog\u002Fguides\u002Fnode-unblocker-what-it-does-not-do","Node Unblocker: What It Does, and What It Does Not","A proxy you host still leaves from your address. It changes the code path, not what refused you. Its own README lists the sites it cannot handle.",[1389,1390,1391,1392,1278],"node unblocker","web proxy","blocked scraper","nodejs","\u002Fimg\u002Fblog\u002Fnode-unblocker-what-it-does-not-do.png","\u002Fimg\u002Fblog\u002Fnode-unblocker-what-it-does-not-do-card.png",{"path":138,"title":1396,"description":1397,"publishedAt":1332,"author":6,"category":1096,"tags":1398,"readTime":1106,"cover":1403,"listingCover":1404,"image":1101},"SERP History: Track the Page's Shape, Not Just Your Rank","A flat position line hides the month an AI Overview appeared and took the clicks. We have 45 keywords on page one and zero clicks. Here is what to record.",[1399,1400,1401,1349,1402],"serp history","paa seo","ai overview","serp features","\u002Fimg\u002Fblog\u002Fserp-history-track-the-shape.png","\u002Fimg\u002Fblog\u002Fserp-history-track-the-shape-card.png",{"path":1406,"title":1407,"description":1408,"publishedAt":1409,"author":6,"category":1096,"tags":1410,"readTime":1106,"cover":1415,"listingCover":1416,"image":1101},"\u002Fblog\u002Fguides\u002Fbrave-search-api-what-you-can-keep","Brave Search API: Check What You Are Allowed to Keep","An independent index at a fair price, with one clause that decides whether it fits: result storage is granted per plan, and most builds need to store.","2026-08-27",[1411,1412,1413,1414,1349],"brave search api","search api","independent index","rag","\u002Fimg\u002Fblog\u002Fbrave-search-api-what-you-can-keep.png","\u002Fimg\u002Fblog\u002Fbrave-search-api-what-you-can-keep-card.png",{"path":1418,"title":1419,"description":1420,"publishedAt":1409,"author":6,"category":1096,"tags":1421,"readTime":1106,"cover":1425,"listingCover":1426,"image":1101},"\u002Fblog\u002Fguides\u002Fgoogle-search-url-parameters","Google Search URL Parameters: What Still Works in 2026","Google silently killed num in September 2025. Here is the current list, which ones are official, and which are community guesses you should not depend on.",[1422,1347,1349,1423,1424],"google search url parameters","udm","search operators","\u002Fimg\u002Fblog\u002Fgoogle-search-url-parameters.png","\u002Fimg\u002Fblog\u002Fgoogle-search-url-parameters-card.png",{"path":1428,"title":1429,"description":1430,"publishedAt":1409,"author":6,"category":1122,"tags":1431,"readTime":1106,"cover":1435,"listingCover":1436,"image":1101},"\u002Fblog\u002Fguides\u002Finstagram-api-which-one-in-2026","What Is the Best Instagram API in 2026?","Four products answer to the name: Meta's Graph API, the private mobile API, hosted actors, metered endpoints. One question decides which.",[1128,1432,1127,1433,1434],"instagram graph api","tikhub","apify","\u002Fimg\u002Fblog\u002Finstagram-api-which-one-in-2026.png","\u002Fimg\u002Fblog\u002Finstagram-api-which-one-in-2026-card.png",{"path":1438,"title":1439,"description":1440,"publishedAt":1409,"author":6,"category":1122,"tags":1441,"readTime":1106,"cover":1444,"listingCover":1445,"image":1101},"\u002Fblog\u002Fguides\u002Ftiktok-scraper-which-endpoint","TikTok Scraper: Which Endpoint for Which Question?","Profiles, videos, comments and search are four jobs with four cost curves. And the creator arrives nested inside every post, which changes the design.",[1442,1443,1127,1433,1434],"tiktok scraper","tiktok api","\u002Fimg\u002Fblog\u002Ftiktok-scraper-which-endpoint.png","\u002Fimg\u002Fblog\u002Ftiktok-scraper-which-endpoint-card.png",{"path":1447,"title":1448,"description":1449,"publishedAt":1409,"author":6,"category":1096,"tags":1450,"readTime":1106,"cover":1454,"listingCover":1455,"image":1101},"\u002Fblog\u002Fguides\u002Fweb-scraping-services-should-you-hire-it-out","Web Scraping Services: Should You Hire This Out?","A managed service is priced against the value of your data, not the cost of the requests. Three questions decide whether that is a bargain or a markup.",[1451,1452,1302,1453,1303],"web scraping services","managed scraping","data as a service","\u002Fimg\u002Fblog\u002Fweb-scraping-services-should-you-hire-it-out.png","\u002Fimg\u002Fblog\u002Fweb-scraping-services-should-you-hire-it-out-card.png",{"path":1457,"title":1458,"description":1459,"publishedAt":1460,"author":6,"category":1148,"tags":1461,"readTime":1106,"cover":1466,"listingCover":1467,"image":1101},"\u002Fblog\u002Fguides\u002Fbright-data-mcp-what-you-are-installing","Bright Data MCP: What Are You Actually Installing?","Sixty-nine tools arrive in one install, and every definition sits in the model's context before the agent does anything. When that trade is worth it.","2026-08-26",[1462,1463,1304,1464,1465],"bright data mcp","mcp server","web data","tool calling","\u002Fimg\u002Fblog\u002Fbright-data-mcp-what-you-are-installing.png","\u002Fimg\u002Fblog\u002Fbright-data-mcp-what-you-are-installing-card.png",{"path":1469,"title":1470,"description":1471,"publishedAt":1460,"author":6,"category":1135,"tags":1472,"readTime":1106,"cover":1478,"listingCover":1479,"image":1101},"\u002Fblog\u002Fguides\u002Fclay-alternatives-waterfall-enrichment","Clay Alternatives: You Are Buying the Waterfall","Clay's product is not the data, it is the fall-through logic between providers. Once they share a balance, that logic is about twenty lines of code.",[1473,1474,1475,1476,1477],"clay alternatives","waterfall enrichment","lead enrichment api","data enrichment","gtm","\u002Fimg\u002Fblog\u002Fclay-alternatives-waterfall-enrichment.png","\u002Fimg\u002Fblog\u002Fclay-alternatives-waterfall-enrichment-card.png",{"path":1481,"title":1482,"description":1483,"publishedAt":1460,"author":6,"category":1484,"tags":1485,"readTime":1280,"cover":1491,"listingCover":1492,"image":1101},"\u002Fblog\u002Fguides\u002Felevenlabs-alternatives-three-complaints","ElevenLabs Alternatives: Price, Voice, or Plumbing?","Three different complaints hide under one search. Only one of them needs a different provider, and the other two have cheaper fixes on the same one.","Generative media",[1486,1487,1488,1489,1490],"elevenlabs alternatives","text to speech","tts api","voice ai","generative media","\u002Fimg\u002Fblog\u002Felevenlabs-alternatives-three-complaints.png","\u002Fimg\u002Fblog\u002Felevenlabs-alternatives-three-complaints-card.png",{"path":1494,"title":1495,"description":1496,"publishedAt":1460,"author":6,"category":1096,"tags":1497,"readTime":1106,"cover":1502,"listingCover":1503,"image":1101},"\u002Fblog\u002Fguides\u002Frank-tracking-api-store-the-position","Rank Tracking API: The Hard Part Is Not the Fetch","Getting today's position is one cheap call. The product is yesterday's position, identical parameters, and knowing what moved. Here is the whole shape.",[1498,1499,1500,1501,1412],"rank tracking api","bing rank tracker","serp tracking","seo api","\u002Fimg\u002Fblog\u002Frank-tracking-api-store-the-position.png","\u002Fimg\u002Fblog\u002Frank-tracking-api-store-the-position-card.png",{"path":1505,"title":1506,"description":1507,"publishedAt":1460,"author":6,"category":1484,"tags":1508,"readTime":1106,"cover":1512,"listingCover":1513,"image":1101},"\u002Fblog\u002Fguides\u002Fstreaming-voice-agent-vs-batch-tts","Gradium Holds the Call. A Batch TTS API Renders the Rest.","Ten measured calls to a batch text to speech endpoint. The fastest never beat 3.05 seconds, and the model mattered more than four times the text.",[1509,1510,1511,1304],"text to speech api","voice agents","tts latency","\u002Fimg\u002Fblog\u002Fstreaming-voice-agent-vs-batch-tts.png","\u002Fimg\u002Fblog\u002Fstreaming-voice-agent-vs-batch-tts-card.png",{"path":1515,"title":1516,"description":1517,"publishedAt":1518,"author":6,"category":1135,"tags":1519,"readTime":1106,"cover":1525,"listingCover":1526,"image":1101},"\u002Fblog\u002Fguides\u002Fb2b-data-provider-dataset-vs-api","B2B Data Providers: Buy the Dataset or Call the API?","A dataset and an enrichment API are the same rows with a different contract. What changes is who carries the staleness and who pays for unread records.","2026-08-25",[1520,1521,1522,1523,1524],"b2b data","data provider","company enrichment","zoominfo","people data labs","\u002Fimg\u002Fblog\u002Fb2b-data-provider-dataset-vs-api.png","\u002Fimg\u002Fblog\u002Fb2b-data-provider-dataset-vs-api-card.png",{"path":1528,"title":1529,"description":1530,"publishedAt":1518,"author":6,"category":1096,"tags":1531,"readTime":1106,"cover":1536,"listingCover":1537,"image":1101},"\u002Fblog\u002Fguides\u002Fcrawl4ai-when-to-run-your-own-crawler","Crawl4AI: When Should You Run Your Own Crawler?","Crawl4AI does the parsing for free. The bill is the fleet you run around it. A decision rule for when to self-host a crawler and when to call one.",[1532,1533,1414,1534,1535],"crawl4ai","web crawler","markdown","open source","\u002Fimg\u002Fblog\u002Fcrawl4ai-when-to-run-your-own-crawler.png","\u002Fimg\u002Fblog\u002Fcrawl4ai-when-to-run-your-own-crawler-card.png",{"path":1539,"title":1540,"description":1541,"publishedAt":1518,"author":6,"category":1096,"tags":1542,"readTime":1155,"cover":1548,"listingCover":1549,"image":1101},"\u002Fblog\u002Fguides\u002Fis-web-scraping-legal-2026","Is Web Scraping Legal? What the Rulings Actually Say","Four US cases decided the modern law of scraping and none of them turned on scraping. They turned on login, publicness, copying and circumvention.",[1543,1544,1545,1546,1547],"web scraping legal","hiq linkedin","bright data","cfaa","dmca","\u002Fimg\u002Fblog\u002Fis-web-scraping-legal-2026.png","\u002Fimg\u002Fblog\u002Fis-web-scraping-legal-2026-card.png",{"path":1551,"title":1552,"description":1553,"publishedAt":1518,"author":6,"category":1096,"tags":1554,"readTime":1106,"cover":1559,"listingCover":1560,"image":1101},"\u002Fblog\u002Fguides\u002Fweb-scraping-news-articles","Web Scraping News Articles: Three Jobs, Three Endpoints","Scraping news is three jobs under one name. Find articles, read one article, watch one company. Each needs a different call and a different budget.",[1555,1556,1414,1557,1558],"news scraping","news api","media monitoring","web search","\u002Fimg\u002Fblog\u002Fweb-scraping-news-articles.png","\u002Fimg\u002Fblog\u002Fweb-scraping-news-articles-card.png",{"path":1562,"title":1563,"description":1564,"publishedAt":1518,"author":6,"category":1122,"tags":1565,"readTime":1106,"cover":1568,"listingCover":1569,"image":1101},"\u002Fblog\u002Fguides\u002Fyoutube-scraper-past-the-quota","YouTube Scraper: What the Data API Will Not Give You","The official API allows 100 searches a day and no money raises it. That ceiling, not price, is why YouTube scrapers exist. Here is what each one returns.",[1566,1567,1127,1434,1433],"youtube scraper","youtube data api","\u002Fimg\u002Fblog\u002Fyoutube-scraper-past-the-quota.png","\u002Fimg\u002Fblog\u002Fyoutube-scraper-past-the-quota-card.png",{"path":1571,"title":1572,"description":1573,"publishedAt":1574,"author":6,"category":1148,"tags":1575,"readTime":1155,"cover":1578,"listingCover":1579,"image":1101},"\u002Fblog\u002Fguides\u002Fagentic-browser-when-your-agent-needs-one","Agentic Browser: When Your Agent Actually Needs One","Agentic browser names two products: one you sit in front of and one your agent drives. Most jobs need neither. The test is whether the task carries state.","2026-08-21",[1576,1266,1304,1279,1577],"agentic browser","browserbase","\u002Fimg\u002Fblog\u002Fagentic-browser-when-your-agent-needs-one.png","\u002Fimg\u002Fblog\u002Fagentic-browser-when-your-agent-needs-one-card.png",{"path":952,"title":1581,"description":1582,"publishedAt":1574,"author":6,"category":1096,"tags":1583,"readTime":1155,"cover":1587,"listingCover":1588,"image":1101},"Do You Still Need a Rotating Proxy in 2026?","A rotating proxy is an input you buy by the gigabyte and hope works. Most teams wanted the outcome instead. Here is how to tell which one you need.",[1584,1585,1241,1304,1586],"rotating proxy","residential proxy","infrastructure","\u002Fimg\u002Fblog\u002Fdo-ai-agents-need-a-rotating-proxy.png","\u002Fimg\u002Fblog\u002Fdo-ai-agents-need-a-rotating-proxy-card.png",{"path":1590,"title":1591,"description":1592,"publishedAt":1574,"author":6,"category":1096,"tags":1593,"readTime":1155,"cover":1596,"listingCover":1597,"image":1101},"\u002Fblog\u002Fguides\u002Fpython-web-scraping-without-a-scraper","Web Scraping in Python Without Maintaining a Scraper","Python web scraping is two jobs wearing one name. Python is the best tool for one of them and the wrong place to solve the other. Here is the split.",[568,1241,1594,1595,1304],"beautifulsoup","selenium","\u002Fimg\u002Fblog\u002Fpython-web-scraping-without-a-scraper.png","\u002Fimg\u002Fblog\u002Fpython-web-scraping-without-a-scraper-card.png",{"path":116,"title":1599,"description":1600,"publishedAt":1574,"author":6,"category":1096,"tags":1601,"readTime":1155,"cover":1602,"listingCover":1603,"image":1101},"What Is a SERP API, and Which Kind Do You Need?","A SERP API is three products under one name: a Google mirror, an independent index, and a read-through fetcher. The difference decides which one is yours.",[1347,1412,1558,1304,1414],"\u002Fimg\u002Fblog\u002Fserp-api-for-ai-agents.png","\u002Fimg\u002Fblog\u002Fserp-api-for-ai-agents-card.png",{"path":1605,"title":1606,"description":1607,"publishedAt":1574,"author":6,"category":1096,"tags":1608,"readTime":1155,"cover":1611,"listingCover":1612,"image":1101},"\u002Fblog\u002Fguides\u002Fweb-scraping-tools-which-kind","Web Scraping Tools: Which Kind Do You Actually Need?","Web scraping tools come in four kinds and they differ on who operates them, not on features. Pick the operator first and the shortlist writes itself.",[1609,1610,1303,1304,1434],"web scraping tools","no code scraper","\u002Fimg\u002Fblog\u002Fweb-scraping-tools-which-kind.png","\u002Fimg\u002Fblog\u002Fweb-scraping-tools-which-kind-card.png",{"path":1614,"title":1615,"description":1616,"publishedAt":1617,"author":6,"category":1224,"tags":1618,"readTime":1155,"cover":1621,"listingCover":1622,"image":1101},"\u002Fblog\u002Fguides\u002Fautonomous-store-agent-outside-data","Runner AI Runs the Store. Monid Feeds It the Market.","One call returned 48 competitors: prices from $3.99 to $59.97 and an incumbent with 136,609 reviews. No store dashboard contains that, and agents need it.","2026-08-20",[1619,1304,1228,1620],"agentic commerce","autonomous business","\u002Fimg\u002Fblog\u002Fautonomous-store-agent-outside-data.png","\u002Fimg\u002Fblog\u002Fautonomous-store-agent-outside-data-card.png",{"path":1624,"title":1625,"description":1626,"publishedAt":1617,"author":6,"category":1096,"tags":1627,"readTime":1155,"cover":1632,"listingCover":1633,"image":1101},"\u002Fblog\u002Fguides\u002Fbing-search-api-retired-alternatives","Bing Search API Retired: What Actually Replaces It","Microsoft turned the Bing Search APIs off in August 2025 and the endpoints now return 410. What a replacement has to do, and which route fits which job.",[1628,1629,1630,1414,1631],"search","web","api","agents","\u002Fimg\u002Fblog\u002Fbing-search-api-retired-alternatives.png","\u002Fimg\u002Fblog\u002Fbing-search-api-retired-alternatives-card.png",{"path":1635,"title":1636,"description":1637,"publishedAt":1617,"author":6,"category":1096,"tags":1638,"readTime":1106,"cover":1642,"listingCover":1643,"image":1101},"\u002Fblog\u002Fguides\u002Fingest-mixed-documents-llm-embedding","Ingesting Mixed Documents for LLM Embedding: PDF to Markdown","PDFs, Office files and HTML all become one clean format before chunking. Where the pipeline actually breaks, and which endpoint handles which format.",[1414,1639,1640,1641,1534],"embeddings","pdf","ocr","\u002Fimg\u002Fblog\u002Fingest-mixed-documents-llm-embedding.png","\u002Fimg\u002Fblog\u002Fingest-mixed-documents-llm-embedding-card.png",{"path":1645,"title":1646,"description":1647,"publishedAt":1617,"author":6,"category":1122,"tags":1648,"readTime":1155,"cover":1653,"listingCover":1654,"image":1101},"\u002Fblog\u002Fguides\u002Finstagram-follower-engagement-api","Instagram Follower and Engagement Data: Which API in 2026?","Follower counts are accurate. What trackers build on top is inference. Which endpoints return which fields, and how to tell a real signal from a guess.",[1649,1650,1630,1651,1652],"instagram","social","creators","engagement","\u002Fimg\u002Fblog\u002Finstagram-follower-engagement-api.png","\u002Fimg\u002Fblog\u002Finstagram-follower-engagement-api-card.png",{"path":1656,"title":1657,"description":1658,"publishedAt":1617,"author":6,"category":1148,"tags":1659,"readTime":1155,"cover":1665,"listingCover":1666,"image":1101},"\u002Fblog\u002Fguides\u002Fllm-gateway-vs-mcp-gateway","LLM Gateway vs MCP Gateway: Four Families, One Word","An LLM gateway routes prompts to models. An MCP gateway routes tool calls to vendors. Four gateway families, and which problem each one solves.",[1660,1661,1662,1663,1664],"llm gateway","ai gateway","mcp","agent tools","routing","\u002Fimg\u002Fblog\u002Fllm-gateway-vs-mcp-gateway.png","\u002Fimg\u002Fblog\u002Fllm-gateway-vs-mcp-gateway-card.png",{"path":1668,"title":1669,"description":1670,"publishedAt":1617,"author":6,"category":1148,"tags":1671,"readTime":1155,"cover":1672,"listingCover":1673,"image":1101},"\u002Fblog\u002Fguides\u002Fmcp-vs-api-for-ai-agents","MCP vs API for AI Agents: Who Does the Wrapping?","MCP is not an alternative to APIs, it wraps them. The real question is who does the wrapping, and the answer decides how much code you write.",[1662,1630,1663,1304,1465],"\u002Fimg\u002Fblog\u002Fmcp-vs-api-for-ai-agents.png","\u002Fimg\u002Fblog\u002Fmcp-vs-api-for-ai-agents-card.png",{"path":1675,"title":1676,"description":1677,"publishedAt":1617,"author":6,"category":1148,"tags":1678,"readTime":1106,"cover":1682,"listingCover":1683,"image":1101},"\u002Fblog\u002Fguides\u002Fn8n-data-layer-scraping-enrichment","The Data Layer for n8n: One Key Instead of a Node Per Vendor","n8n workflows rarely break at the logic. They break at the data source. Build the scraping and enrichment layer so a dead provider is a config change.",[1679,1680,1681,1314,1631],"n8n","automation","enrichment","\u002Fimg\u002Fblog\u002Fn8n-data-layer-scraping-enrichment.png","\u002Fimg\u002Fblog\u002Fn8n-data-layer-scraping-enrichment-card.png",{"path":1685,"title":1686,"description":1687,"publishedAt":1617,"author":6,"category":1148,"tags":1688,"readTime":1155,"cover":1690,"listingCover":1691,"image":1101},"\u002Fblog\u002Fguides\u002Fopenrouter-alternatives-after-stripe","OpenRouter Alternatives in 2026: Models, and Then Tools","Stripe agreed to buy OpenRouter for over $7 billion. The real alternatives on the model side, and the routing layer nobody is selling yet.",[1689,1660,1663,1662,1664],"openrouter","\u002Fimg\u002Fblog\u002Fopenrouter-alternatives-after-stripe.png","\u002Fimg\u002Fblog\u002Fopenrouter-alternatives-after-stripe-card.png",{"path":1693,"title":1694,"description":1695,"publishedAt":1617,"author":6,"category":1148,"tags":1696,"readTime":1155,"cover":1698,"listingCover":1699,"image":1101},"\u002Fblog\u002Fguides\u002Fopenrouter-mcp-server-tool-layer","Is OpenRouter an MCP Server? What It Actually Exposes","Yes, OpenRouter has an MCP server, and it exposes OpenRouter. Here is what it does, what it does not, and how to give the same agent real tools.",[1689,1662,1663,1697,1465],"claude code","\u002Fimg\u002Fblog\u002Fopenrouter-mcp-server-tool-layer.png","\u002Fimg\u002Fblog\u002Fopenrouter-mcp-server-tool-layer-card.png",{"path":1701,"title":1702,"description":1703,"publishedAt":1617,"author":6,"category":1135,"tags":1704,"readTime":1155,"cover":1708,"listingCover":1709,"image":1101},"\u002Fblog\u002Fguides\u002Fprompt-instead-of-filters-people-search","A Wrong Enum Returns Zero: Ploid and Prompt-Shaped APIs","A wrong enum returned 0 rows. A wrong range format returned 160,884. Neither raised an error. Why prompt-shaped APIs like Ploid exist, and what they cost.",[1705,1706,1707,1304],"people search api","agentic api","prospecting","\u002Fimg\u002Fblog\u002Fprompt-instead-of-filters-people-search.png","\u002Fimg\u002Fblog\u002Fprompt-instead-of-filters-people-search-card.png",{"path":1711,"title":1712,"description":1713,"publishedAt":1617,"author":6,"category":1135,"tags":1714,"readTime":1155,"cover":1718,"listingCover":1719,"image":1101},"\u002Fblog\u002Fguides\u002Fproxycurl-shutdown-linkedin-data-alternatives","Proxycurl Shut Down: Where LinkedIn Data Goes Now","Proxycurl closed in July 2025 to settle with LinkedIn. What a replacement has to return, which endpoints cover which half, and the risk to read first.",[1715,1681,1630,1716,1717],"linkedin","sales","leads","\u002Fimg\u002Fblog\u002Fproxycurl-shutdown-linkedin-data-alternatives.png","\u002Fimg\u002Fblog\u002Fproxycurl-shutdown-linkedin-data-alternatives-card.png",{"path":1721,"title":1722,"description":1723,"publishedAt":1617,"author":6,"category":1122,"tags":1724,"readTime":1155,"cover":1729,"listingCover":1730,"image":1101},"\u002Fblog\u002Fguides\u002Ftiktok-data-behind-an-ai-video-ad","Oumomo Generates the Ad. TikTok Data Decides Which One.","An AI video tool makes whatever you ask. One TikTok search call returned twenty videos with a fifteenfold spread in likes, and that spread is the brief.",[1725,1726,1727,1728],"tiktok data api","ai video ads","tiktok shop","creative research","\u002Fimg\u002Fblog\u002Ftiktok-data-behind-an-ai-video-ad.png","\u002Fimg\u002Fblog\u002Ftiktok-data-behind-an-ai-video-ad-card.png",{"path":1732,"title":1733,"description":1734,"publishedAt":1617,"author":6,"category":1122,"tags":1735,"readTime":1155,"cover":1737,"listingCover":1738,"image":1101},"\u002Fblog\u002Fguides\u002Ftiktok-scraper-catalogue-audit","Oumomo Remakes the Winner. A TikTok Scraper Tells You Which.","Fourteen videos, twenty one days, and one was 65 percent of the account. You cannot see that in the clips. One call puts the catalogue in a table.",[1442,1736,1727,1726],"content audit","\u002Fimg\u002Fblog\u002Ftiktok-scraper-catalogue-audit.png","\u002Fimg\u002Fblog\u002Ftiktok-scraper-catalogue-audit-card.png",{"path":1740,"title":1741,"description":1742,"publishedAt":1617,"author":6,"category":1148,"tags":1743,"readTime":1745,"cover":1746,"listingCover":1747,"image":1101},"\u002Fblog\u002Fguides\u002Fwhat-is-an-mcp-gateway","What Is an MCP Gateway? Two Products, One Name","An MCP gateway is one endpoint that fronts many tools. Two very different products share the name, and the difference decides which one you need.",[1662,1744,1663,1304,1664],"mcp gateway","13 min","\u002Fimg\u002Fblog\u002Fwhat-is-an-mcp-gateway.png","\u002Fimg\u002Fblog\u002Fwhat-is-an-mcp-gateway-card.png",{"path":1749,"title":1750,"description":1751,"publishedAt":1617,"author":6,"category":1122,"tags":1752,"readTime":1155,"cover":1756,"listingCover":1757,"image":1101},"\u002Fblog\u002Fguides\u002Fx-intent-search-social-selling","Monid Finds the Tweet. Volumn Tells You Who Sent It.","A tightened X search returned 19 of 20 keyword-relevant tweets. Only 2 were someone actually asking. Keyword match is not intent, and that gap is the work.",[1753,1754,1290,1755],"twitter scraper","x api","lead generation","\u002Fimg\u002Fblog\u002Fx-intent-search-social-selling.png","\u002Fimg\u002Fblog\u002Fx-intent-search-social-selling-card.png",{"path":1759,"title":1760,"description":1761,"publishedAt":1762,"author":6,"category":1096,"tags":1763,"readTime":1106,"cover":1767,"listingCover":1768,"image":1101},"\u002Fblog\u002Fguides\u002Fapify-alternatives","Apify Alternatives: You Probably Want a Different Way to Buy It","Every AI answer to this question names Apify. We ran an Apify actor without an Apify account to show what the real alternative is.","2026-08-19",[1764,1241,1765,1766],"apify alternatives","actors","pay per call","\u002Fimg\u002Fblog\u002Fapify-alternatives.png","\u002Fimg\u002Fblog\u002Fapify-alternatives-card.png",{"path":1770,"title":1771,"description":1772,"publishedAt":1762,"author":6,"category":1224,"tags":1773,"readTime":1106,"cover":1777,"listingCover":1778,"image":1101},"\u002Fblog\u002Fguides\u002Fconnect-claude-to-amazon-search-data","Connect Claude to Amazon Search Data: OpenWeb Ninja and Monid","Amazon has no keyword API for sellers. Here is what a search endpoint actually returns, how to hand it to an agent, and which route fits which job.",[1774,1775,1776,1304],"amazon search api","amazon keyword data","ecommerce","\u002Fimg\u002Fblog\u002Fconnect-claude-to-amazon-search-data.png","\u002Fimg\u002Fblog\u002Fconnect-claude-to-amazon-search-data-card.png",{"path":1780,"title":1781,"description":1782,"publishedAt":1762,"author":6,"category":1135,"tags":1783,"readTime":1106,"cover":1787,"listingCover":1788,"image":1101},"\u002Fblog\u002Fguides\u002Fenrich-a-list-from-email-addresses","Enriching a List When All You Have Is Email Addresses","We enriched four real email addresses and one returned a person. The misses billed nothing. What that changes about how you plan a batch.",[1784,1524,1785,1786],"email enrichment","zoominfo alternatives","lead data","\u002Fimg\u002Fblog\u002Fenrich-a-list-from-email-addresses.png","\u002Fimg\u002Fblog\u002Fenrich-a-list-from-email-addresses-card.png",{"path":1790,"title":1791,"description":1792,"publishedAt":1762,"author":6,"category":1096,"tags":1793,"readTime":1106,"cover":1796,"listingCover":1797,"image":1101},"\u002Fblog\u002Fguides\u002Ffree-api-extract-page-content-rag","A Free API to Extract Page Content for RAG: Read This First","We scraped a Wikipedia page and got 74,552 characters starting with the nav menu. The official API gave 689 clean ones. When each is right.",[1414,1278,1794,1795],"chunking","llm context","\u002Fimg\u002Fblog\u002Ffree-api-extract-page-content-rag.png","\u002Fimg\u002Fblog\u002Ffree-api-extract-page-content-rag-card.png",{"path":1799,"title":1800,"description":1801,"publishedAt":1762,"author":6,"category":1199,"tags":1802,"readTime":1106,"cover":1806,"listingCover":1807,"image":1101},"\u002Fblog\u002Fguides\u002Fgoogle-maps-scraper-alternatives","Google Maps Scraper Alternatives: Building a Local Lead List","Two Google Maps actors, same catalogue, different jobs. One returned nothing on a plausible query. What the fields tell you about a business, measured.",[1803,1804,1434,1805],"google maps scraper","local leads","lead list","\u002Fimg\u002Fblog\u002Fgoogle-maps-scraper-alternatives.png","\u002Fimg\u002Fblog\u002Fgoogle-maps-scraper-alternatives-card.png",{"path":1809,"title":1810,"description":1811,"publishedAt":1762,"author":6,"category":1333,"tags":1812,"readTime":1106,"cover":1817,"listingCover":1818,"image":1101},"\u002Fblog\u002Fguides\u002Fmeta-ad-library-longest-running-ads","Mining the Meta Ad Library with Wireflow: Ninety Days Means It Works","Run duration is the profitability signal competitors publish by accident. How to read it, when it lies, and how to turn a long runner into creative.",[1813,1814,1815,1816],"meta ad library","competitor ads","ad creative","facebook ads","\u002Fimg\u002Fblog\u002Fmeta-ad-library-longest-running-ads.png","\u002Fimg\u002Fblog\u002Fmeta-ad-library-longest-running-ads-card.png",{"path":1820,"title":1821,"description":1822,"publishedAt":1823,"author":6,"category":1148,"tags":1824,"readTime":1106,"cover":1826,"listingCover":1827,"image":1101},"\u002Fblog\u002Fguides\u002Fai-agent-needs-two-integrations","A Model Is Half an Agent: AIHubMix for Models, Monid for Tools","A model gateway gets your agent talking. It still cannot look anything up. How to wire both halves, with AIHubMix on models and Monid on tools.","2026-08-18",[1304,1825,1663,1662],"model gateway","\u002Fimg\u002Fblog\u002Fai-agent-needs-two-integrations.png","\u002Fimg\u002Fblog\u002Fai-agent-needs-two-integrations-card.png",{"path":1829,"title":1830,"description":1831,"publishedAt":1823,"author":6,"category":1148,"tags":1832,"readTime":1106,"cover":1835,"listingCover":1836,"image":1101},"\u002Fblog\u002Fguides\u002Fapi-marketplace-for-ai-agents","The Best API Marketplace for AI Agents in 2026","Marketplaces were built for developers who integrate once. An agent chooses at run time. One ordinary task touched five providers across four steps.",[1833,1304,1662,1834],"api marketplace","tool use","\u002Fimg\u002Fblog\u002Fapi-marketplace-for-ai-agents.png","\u002Fimg\u002Fblog\u002Fapi-marketplace-for-ai-agents-card.png",{"path":1838,"title":1839,"description":1840,"publishedAt":1823,"author":6,"category":1135,"tags":1841,"readTime":1106,"cover":1846,"listingCover":1847,"image":1101},"\u002Fblog\u002Fguides\u002Fbusiness-entity-search-api","Business Entity Search by API: What Exists and What Does Not","State registries are the most searched company lookup and the least available by API. We ran the US endpoint on a private company and it returned nothing.",[1842,1843,1844,1845],"business entity search","company registry","sec edgar","company data","\u002Fimg\u002Fblog\u002Fbusiness-entity-search-api.png","\u002Fimg\u002Fblog\u002Fbusiness-entity-search-api-card.png",{"path":1849,"title":1850,"description":1851,"publishedAt":1823,"author":6,"category":1096,"tags":1852,"readTime":1106,"cover":1856,"listingCover":1857,"image":1101},"\u002Fblog\u002Fguides\u002Fscraper-blocked-what-gets-through","Your Scraper Is Blocked: What Actually Gets Through in 2026","A raw request to a Cloudflare-protected page returns 403. We ran the same URL through a managed endpoint and read what came back, field by field.",[1241,1853,1854,1855],"cloudflare","proxies","blocked","\u002Fimg\u002Fblog\u002Fscraper-blocked-what-gets-through.png","\u002Fimg\u002Fblog\u002Fscraper-blocked-what-gets-through-card.png",{"path":1859,"title":1860,"description":1861,"publishedAt":1823,"author":6,"category":1096,"tags":1862,"readTime":1106,"cover":1864,"listingCover":1865,"image":1101},"\u002Fblog\u002Fguides\u002Fweb-scraping-api-for-ai-agents","The Best Web Scraping API for AI Agents in 2026","An agent cannot pick a scraper from a list it has never seen. What changes when the catalogue is discoverable at run time, measured across three phrasings.",[1863,1304,1662,1680],"web scraping api","\u002Fimg\u002Fblog\u002Fweb-scraping-api-for-ai-agents.png","\u002Fimg\u002Fblog\u002Fweb-scraping-api-for-ai-agents-card.png",{"path":1867,"title":1868,"description":1869,"publishedAt":1870,"author":6,"category":1135,"tags":1871,"readTime":1106,"cover":1875,"listingCover":1876,"image":1101},"\u002Fblog\u002Fguides\u002Fapi-to-find-a-company-website","The Best API to Find a Company's Website From Its Name","Name to homepage is a resolution problem, not a search one. Which endpoint does it, why fuzzy matches are a feature, and how to pick the right row.","2026-08-17",[1872,1873,1874,1681],"company website api","company lookup","domain resolution","\u002Fimg\u002Fblog\u002Fapi-to-find-a-company-website.png","\u002Fimg\u002Fblog\u002Fapi-to-find-a-company-website-card.png",{"path":1878,"title":1879,"description":1880,"publishedAt":1870,"author":6,"category":1096,"tags":1881,"readTime":1106,"cover":1885,"listingCover":1886,"image":1101},"\u002Fblog\u002Fguides\u002Fcompany-news-api-press-releases","PR Newswire API: How to Read Releases, Not Send Them","The newswires sell distribution, not access. Reading a company's releases and its coverage is a different endpoint, and one field separates the two.",[1882,1883,1557,1884],"company news api","press release api","signals","\u002Fimg\u002Fblog\u002Fcompany-news-api-press-releases.png","\u002Fimg\u002Fblog\u002Fcompany-news-api-press-releases-card.png",{"path":1888,"title":1889,"description":1890,"publishedAt":1870,"author":6,"category":1135,"tags":1891,"readTime":1106,"cover":1895,"listingCover":1896,"image":1101},"\u002Fblog\u002Fguides\u002Fcrunchbase-api-alternatives-funding-data","Crunchbase API Alternatives: Getting Funding Data by the Call","Crunchbase and PitchBook sell seats. If you only need funding fields on a domain, an enrichment call carries them. What it returns, and what it gets wrong.",[1892,1893,1894,1522],"crunchbase api","pitchbook api","funding data","\u002Fimg\u002Fblog\u002Fcrunchbase-api-alternatives-funding-data.png","\u002Fimg\u002Fblog\u002Fcrunchbase-api-alternatives-funding-data-card.png",{"path":1898,"title":1899,"description":1900,"publishedAt":1870,"author":6,"category":1135,"tags":1901,"readTime":1106,"cover":1903,"listingCover":1904,"image":1101},"\u002Fblog\u002Fguides\u002Fjob-postings-api-hiring-signals","Job Postings API: Turning Hiring Into a Buying Signal","A job ad says what a company is building before its website does. Which endpoint returns postings, what the record contains, and the fields nobody reads.",[1369,1370,1902,1477],"jobs data","\u002Fimg\u002Fblog\u002Fjob-postings-api-hiring-signals.png","\u002Fimg\u002Fblog\u002Fjob-postings-api-hiring-signals-card.png",{"path":1906,"title":1907,"description":1908,"publishedAt":1870,"author":6,"category":1135,"tags":1909,"readTime":1106,"cover":1913,"listingCover":1914,"image":1101},"\u002Fblog\u002Fguides\u002Ftechnographic-data-platforms-vs-per-call","Technographic Data Platforms vs One API Call: What You Give Up","We ran a stack detection on a major site and half the fields came back empty. What per-call detection sees, and when a platform earns its price.",[1910,1911,1912,1477],"technographic data","tech stack detection","b2b targeting","\u002Fimg\u002Fblog\u002Ftechnographic-data-platforms-vs-per-call.png","\u002Fimg\u002Fblog\u002Ftechnographic-data-platforms-vs-per-call-card.png",{"path":1916,"title":1917,"description":1918,"publishedAt":1919,"author":6,"category":1122,"tags":1920,"readTime":1106,"cover":1923,"listingCover":1924,"image":1101},"\u002Fblog\u002Fguides\u002Fbest-social-media-scraping-api-2026","What Is the Best API for Social Media Scraping in 2026?","No single best one, because the platforms are not one problem. Which endpoint covers which network, and where a per-call bill differs from per-result.","2026-08-14",[1921,1127,1922,1649],"social media scraping api","tiktok","\u002Fimg\u002Fblog\u002Fbest-social-media-scraping-api-2026.png","\u002Fimg\u002Fblog\u002Fbest-social-media-scraping-api-2026-card.png",{"path":1926,"title":1927,"description":1928,"publishedAt":1919,"author":6,"category":1096,"tags":1929,"readTime":1106,"cover":1931,"listingCover":1932,"image":1101},"\u002Fblog\u002Fguides\u002Fmcp-server-live-web-data-agents","Which MCP Server Gives an AI Agent Live Web Data?","Most MCP servers wrap one vendor. The question is whether your agent needs a scraper or a catalogue it can search at runtime, and how to tell which.",[1463,1304,1464,1930],"model context protocol","\u002Fimg\u002Fblog\u002Fmcp-server-live-web-data-agents.png","\u002Fimg\u002Fblog\u002Fmcp-server-live-web-data-agents-card.png",{"path":1934,"title":1935,"description":1936,"publishedAt":1919,"author":6,"category":1148,"tags":1937,"readTime":1280,"cover":1940,"listingCover":1941,"image":1101},"\u002Fblog\u002Fguides\u002Fpay-per-call-data-api-vs-subscription","Which Data API Lets You Pay Per Call Instead of a Subscription?","Metered beats a subscription when your usage is bursty and loses when it is steady. The shapes, the crossover, and the measurement that decides it.",[1766,1938,1833,1939],"data api pricing","metered billing","\u002Fimg\u002Fblog\u002Fpay-per-call-data-api-vs-subscription.png","\u002Fimg\u002Fblog\u002Fpay-per-call-data-api-vs-subscription-card.png",{"path":1943,"title":1944,"description":1945,"publishedAt":1919,"author":6,"category":1135,"tags":1946,"readTime":1106,"cover":1948,"listingCover":1949,"image":1101},"\u002Fblog\u002Fguides\u002Fpeople-data-labs-apollo-zoominfo-alternatives","People Data Labs, Apollo, ZoomInfo: Which Should You Actually Buy?","Four providers, one honest split: search versus enrich versus contract. What each is genuinely best at, and why the price gap is not what it looks like.",[1524,1947,1523,1520],"apollo alternatives","\u002Fimg\u002Fblog\u002Fpeople-data-labs-apollo-zoominfo-alternatives.png","\u002Fimg\u002Fblog\u002Fpeople-data-labs-apollo-zoominfo-alternatives-card.png",{"path":1951,"title":1952,"description":1953,"publishedAt":1919,"author":6,"category":1122,"tags":1954,"readTime":1280,"cover":1958,"listingCover":1959,"image":1101},"\u002Fblog\u002Fguides\u002Freddit-scraping-api-alternatives","Is There an Alternative to Apify for Scraping Reddit?","Yes, several. The more useful question is why a Reddit scraper returns posts that miss your keywords, because that is a search problem, not a scraper one.",[1955,1956,1957,1764],"reddit api","reddit scraper","social listening","\u002Fimg\u002Fblog\u002Freddit-scraping-api-alternatives.png","\u002Fimg\u002Fblog\u002Freddit-scraping-api-alternatives-card.png",{"path":1961,"title":1962,"description":1963,"publishedAt":1964,"author":6,"category":1135,"tags":1965,"readTime":1106,"cover":1968,"listingCover":1969,"image":1101},"\u002Fblog\u002Fguides\u002Fbest-linkedin-scraper-api-2026","What Is the Best LinkedIn Scraper API in 2026?","There is no single best LinkedIn scraper API. There are three jobs, and the honest answer is which endpoint fits which job, and what each one gets wrong.","2026-08-12",[1966,1715,1716,1967],"linkedin scraper api","data","\u002Fimg\u002Fblog\u002Fbest-linkedin-scraper-api-2026.png","\u002Fimg\u002Fblog\u002Fbest-linkedin-scraper-api-2026-card.png",{"path":1971,"title":1972,"description":1973,"publishedAt":1974,"author":6,"category":1135,"tags":1975,"readTime":1106,"cover":1980,"listingCover":1981,"image":1101},"\u002Fblog\u002Fguides\u002Fapollo-scraper","Apollo Scraper: Export Apollo.io Leads Safely by API in 2026","An Apollo scraper extension can get your account flagged. How to pull the same Apollo.io people and company data by metered API, with search free.","2026-08-10",[1976,1977,1978,1979],"apollo scraper","apollo.io","lead-generation","b2b-data","\u002Fimg\u002Fblog\u002Fapollo-scraper.png","\u002Fimg\u002Fblog\u002Fapollo-scraper-card.png",{"path":1983,"title":1984,"description":1985,"publishedAt":1974,"author":6,"category":1122,"tags":1986,"readTime":1106,"cover":1989,"listingCover":1990,"image":1101},"\u002Fblog\u002Fguides\u002Freddit-scraper","Reddit Scraper: How to Get Reddit Data After the API Lockdown","Reddit closed self-service API access in 2025. The honest split between the licensed Data API and a public-page scraper you run metered per call.",[1956,1987,1988,1650],"reddit data api","web-scraping","\u002Fimg\u002Fblog\u002Freddit-scraper.png","\u002Fimg\u002Fblog\u002Freddit-scraper-card.png",{"path":1992,"title":1993,"description":1994,"publishedAt":1974,"author":6,"category":1135,"tags":1995,"readTime":1280,"cover":1999,"listingCover":2000,"image":1101},"\u002Fblog\u002Fguides\u002Ftechnographic-data-api","Technographic Data API: Detect Any Website's Tech Stack in One Call","Skip the enterprise technographic platform for one domain. Detect a site's CMS, framework, analytics, CDN and payments by metered API call.",[1996,1997,1998,1979],"technographic data api","tech stack","company-enrichment","\u002Fimg\u002Fblog\u002Ftechnographic-data-api.png","\u002Fimg\u002Fblog\u002Ftechnographic-data-api-card.png",{"path":2002,"title":2003,"description":2004,"publishedAt":2005,"author":6,"category":1122,"tags":2006,"readTime":1106,"cover":2008,"listingCover":2009,"image":1101},"\u002Fblog\u002Fguides\u002Fapify-instagram-scraper","The Apify Instagram Scraper: Which Actor to Use in 2026","Apify has five Instagram scrapers, not one. Which actor fits which job, how to run them without getting blocked, and what two measured runs actually cost.","2026-08-04",[2007,1649,1988,1631],"apify instagram scraper","\u002Fimg\u002Fblog\u002Fapify-instagram-scraper.png","\u002Fimg\u002Fblog\u002Fapify-instagram-scraper-card.png",{"path":2011,"title":2012,"description":2013,"publishedAt":2014,"author":6,"category":1135,"tags":2015,"readTime":1155,"cover":2017,"listingCover":2018,"image":1101},"\u002Fblog\u002Fguides\u002Fbest-linkedin-outreach-tools-2026","The Best LinkedIn Outreach Tools in 2026","LinkedIn outreach tools split into two layers in 2026: the data layer that finds and verifies people, and the sending layer that sequences the messages.","2026-07-29",[2016,1715,1716,1967],"linkedin outreach","\u002Fimg\u002Fblog\u002Fbest-linkedin-outreach-tools-2026.png","\u002Fimg\u002Fblog\u002Fbest-linkedin-outreach-tools-2026-card.png",{"path":2020,"title":2021,"description":2022,"publishedAt":2023,"author":6,"category":1224,"tags":2024,"readTime":2027,"cover":2028,"listingCover":2029,"image":1101},"\u002Fblog\u002Fguides\u002Famazon-pa-api-alternatives","Amazon's PA-API Retires in 2026: How to Move to Monid","PA-API retired in May 2026. Which endpoints actually replace GetItems and SearchItems, tested on the day of writing, including one returning empty prices.","2026-07-21",[2025,2026,1967,1776],"amazon product advertising api","pa-api alternative","9 min","\u002Fimg\u002Fblog\u002Famazon-pa-api-alternatives.png","\u002Fimg\u002Fblog\u002Famazon-pa-api-alternatives-card.png",{"path":2031,"title":2032,"description":2033,"publishedAt":2034,"author":6,"category":1224,"tags":2035,"readTime":1745,"cover":2037,"listingCover":2038,"image":1101},"\u002Fblog\u002Fguides\u002Fbest-amazon-reviews-api-2026","The Best Amazon Reviews API in 2026 (We Tested Them)","There is no official Amazon reviews API. Here are the real ways to get review text, what all of them costs, and the pagination limit nobody mentions.","2026-07-09",[2036,1967,1631,1776],"amazon reviews api","\u002Fimg\u002Fblog\u002Fbest-amazon-reviews-api-2026.png","\u002Fimg\u002Fblog\u002Fbest-amazon-reviews-api-2026-card.png",{"path":2040,"title":2041,"description":2042,"publishedAt":2034,"author":6,"category":1135,"tags":2043,"readTime":1155,"cover":2046,"listingCover":2047,"image":1101},"\u002Fblog\u002Fguides\u002Fbest-email-verification-api-2026","Best Email Verification API in 2026: ZeroBounce, NeverBounce, Kickbox, or One Call on Monid?","ZeroBounce, NeverBounce, Kickbox, Bouncer, a DIY SMTP check and Strale compared: credit packs versus one metered per-call endpoint.",[2044,1967,1631,2045],"email verification api","email","\u002Fimg\u002Fblog\u002Fbest-email-verification-api-2026.png","\u002Fimg\u002Fblog\u002Fbest-email-verification-api-2026-card.png",{"path":2049,"title":2050,"description":21,"publishedAt":2051,"author":1101,"category":1148,"tags":2052,"readTime":1101,"cover":2055,"listingCover":1101,"image":1101},"\u002Fblog\u002Fakta-pro-is-now-available-on-monid","Akta Pro Is Now Available On Monid","2026-07-07",[1631,2053,2054,1967],"partner-tools","private-markets","\u002Fimg\u002Fblog\u002Fakta-pro-is-now-available-on-monid-v2.png",{"path":2057,"title":2058,"description":2059,"publishedAt":2060,"author":1101,"category":1148,"tags":2061,"readTime":1101,"cover":2064,"listingCover":1101,"image":1101},"\u002Fblog\u002Fyour-claude-code-can-now-make-phone-calls","Your Claude Code can now make phone calls","Saperly is now available on Monid. Your agent can now make phone calls for you.","2026-07-05",[1631,2053,2062,2063],"voice","phone","\u002Fimg\u002Fblog\u002Fyour-claude-code-can-now-make-phone-calls.png",{"path":2066,"title":2067,"description":2068,"publishedAt":2069,"author":1101,"category":1148,"tags":2070,"readTime":1101,"cover":2073,"listingCover":1101,"image":1101},"\u002Fblog\u002Fintroducing-suzanne-chatgpt-for-3d-models","Introducing\nClaude for 3D models","Suzanne is now available on Monid. Turn any idea into a production-ready 3D model in one prompt.","2026-06-25",[2071,1631,2072],"3d","creative-tools","\u002Fimg\u002Fblog\u002Fintroducing-suzanne-chatgpt-for-3d-models.png",{"path":2075,"title":2076,"description":2077,"publishedAt":2078,"author":1101,"category":1148,"tags":2079,"readTime":1101,"cover":2082,"listingCover":1101,"image":1101},"\u002Fblog\u002Fminimax-is-now-available-on-monid","MiniMax is now available on Monid","Create images and music with MiniMax models through Monid.","2026-06-24",[1631,2072,2080,2081],"image-generation","music-generation","\u002Fimg\u002Fblog\u002Fminimax-is-now-available-on-monid.png",1788552965760]