docs
Finder

How Finder finds websites

What a scan searches for, what it leaves out and why, how the two numbers on a website row are made, how podcasts are found, and when the next scan runs.

A scan is Finder searching for people who already write about products like yours. This page covers the website half: blogs, review sites and newsletters with a site. YouTube and podcasts have their own searches and show up in the same list. Podcasts are near the end.

What a scan searches for

You gave Finder two lists when you set up your brand: your competitors and your keywords. A scan turns them into Google searches, in your brand's country and language:

  • Each keyword, as you wrote it.
  • For each competitor, five searches: "Name review", "Name alternatives", "Name vs", "Name affiliate program" and "Name coupon", each with the word your keywords share most on the end. Let's say your keywords are about affiliate software and one competitor is called Impact. The search is "Impact review affiliate", not "Impact review", so it finds reviews of the platform and not of a power tool.

Then Finder asks who already links to each competitor. See Sites that link to your competitors.

A scan makes at most 40 website searches. Keywords go first. Let's say you listed five competitors and ten keywords: Finder makes 35 searches, keeps the top 20 results from each, treats every site that ranks as a candidate, and keeps the pages that ranked as its evidence.

Your first scan starts as soon as you've saved your competitors and keywords, so you can write your offer while it runs. Website, YouTube and podcast searches run at the same time. None of them waits. A finished source can add its results while another source is still working.

A scan has a hard five-minute boundary, measured from the moment Finder accepts it. Provider work stops after 270 seconds. The final 30 seconds are reserved for saving the result and marking the scan done or failed. A slow source cannot keep the scan open or write results after its timeout. If one source finishes and another times out, Finder keeps the finished results and shows a warning for the source that stopped. A website source that runs short of time stops searching 20 seconds early and keeps the pages it already paid for, so you get rows instead of nothing.

One site, one row

A site is identified by its domain. www.example.com, blog.example.com and example.com/reviews are all example.com, because one company with one inbox owns them.

There's one exception. Some hosts rent a subdomain to anyone, like Substack or GitHub Pages. There, jane.substack.com and joe.substack.com are two different people, so each gets its own row.

When next week's scan finds the same site again, its row is updated with the new pages. You never get a duplicate. Anything you did to the row stays: saved, hidden, notes, the stage it's in. A page that hasn't ranked for 30 days drops off the row.

What gets left out

Most sites a scan finds aren't people you'd email. So Finder sorts them before you see the list, and it tells you what it did with each one, because a tool that quietly drops results is a tool you can't check.

Never listed at all:

  • Your own site, and the competitors you named.
  • Anyone who's already a partner in your AffiliateRail program, and anyone in a list you uploaded.
  • Big platforms that rank for everything: social networks, app stores, software marketplaces, national news sites.

Listed, but hidden. These are written to your account with the reason. You can see them under Hidden and bring any of them back:

ReasonWhat it means
Sells its own productThe site belongs to a software company. Their "best alternatives" post is marketing, and the person behind it is on a rival's payroll
Off topicThe site matched one of your words by accident. A competitor called Tolt shares its name with a horse gait
Wrong audienceOnly for brands that sell affiliate software. "13 affiliate programs to join" is written for affiliates, and you're looking for founders. So is a review that names your competitor but promises the reader money, like "Everflow review: how to make money as an affiliate"
Content farmA channel with thousands of videos that almost nobody watches
DirectoryA site that lists every tool on the same template, like a software marketplace. There's no person to pitch
AgencyA company that sells services, like "we run your affiliate program for you". It isn't a creator with buyers reading
Not your buyersThe AI reads the titles and judges who they're for. A "how to create affiliate links on Impact" tutorial, a Shopify store's setup video when you sell SaaS software, or a one-off how-to about a single setting in an app. None of them talk to people who'd buy what you sell. The AI has to say so twice, asked separately, before a row is hidden
PlatformA big platform's own site or channel, like a payments company's developer channel

Under each hidden row you'll see which reason it was and what Finder saw, in a few words. A row you hid yourself says so.

A later scan can hide a row too. When Finder's rules get better, the next scan checks the rows it finds again and hides any that now fail, with the reason. It never hides a row you saved, a row you hid yourself, or a row you brought back.

Why hide vendors? We ran three test scans on real brands before building this, and 30 of the 40 top-ranked websites sold software in the same category. They rank because they write comparison pages about each other. Hide them, and the reviewers you want move up.

Finder makes the vendor call in two ways. First it reads the evidence: a site that compares itself to a rival ("Acme vs Rival" on acme.com) is a vendor. Then an AI model reads the domain, the page titles and the short description Google shows under each one, with your competitors' names and domains beside them, and says who owns the site. YouTube channels get the same check: a company's own channel is hidden the same way. The AI only sorts. It never writes a number, a rate or a claim about a person. If a call is wrong, bring the row back from Hidden and it stays back.

Three rules don't need the AI at all, so they still work on a day it doesn't answer:

  • A competitor's own site. impact.com is hidden when you named Impact, even if you never gave its domain. A site that only shares a word, like impactplus.com, isn't.
  • The platform your keyword names. If one of your keywords is "stripe integration", stripe.com is the platform, not a reviewer. It's hidden as a platform.
  • Glossary pages. A site where only /glossary/ pages rank is a software company's own marketing.

If the AI fails to answer for a batch of sites, Finder asks it once more before it gives up.

Coupon and deals sites are not hidden. They stay in the list with a low score.

The two numbers on a website row

NumberWhere it comes fromBands
Estimated monthly visitsA third-party estimate of visits from Google search. It's an estimate, and it misses visits from email and socialLow: under 20,000. Medium: 20,000 to 99,999. High: 100,000 or more
Relevant ranking pagesPages on the site that ranked for your searches in this scan. Counted, not estimatedLow: under 10. Medium: 10 to 24. High: 25 or more

If the estimate isn't available, the row says monthly visits not known. Finder doesn't guess. It never shows a 0 for a number nobody measured, because a 0 reads like a count of zero visits, and you'd be right to call that false. The drawer names the source and the date of the estimate.

A website row has no country. Why? The search runs in your country, so all it can tell us is yours. So country filters and the country boost apply to YouTube only. Save a website, though, and its drawer shows which countries its search traffic comes from, as What a save reads explains.

Each page on the row has a type: review, comparison, alternatives, listicle, guide or landing page. A listicle is a ranked list like "7 best screen recorders". A landing page is a company's own product or pricing page.

Finder also spots two finer kinds of page. An affiliate program page is someone writing about one competitor's affiliate program, like "Tolt affiliate program review", for readers who might join that program or already have. So it counts like a review. A brand feature profiles or interviews one company without judging its product, like "How Tolt grew to $1M". It shows as a guide and earns no review points, because a friendly profile isn't a verdict on the product. Fair enough.

You'll sometimes see a chip like "30% recurring" or "affiliate link". That's what the page itself says about what its author already promotes. The chip comes from the search result, so the scan never loads the page to get it, and it doesn't infer anything.

Some rows say Already carries tracking links to Tolt, or just "already carries an affiliate tracking link". That line means Finder saw a real tracking link in what the person published. It's the strongest thing a row can tell you, because it isn't a guess about whether they'd promote software. They already do, and they already get paid for it. You don't have to explain how any of this works to them.

A tracking link is a normal web link with an id on the end, so the software knows who to pay. https://tolt.io/?ref=jane is one. The ref=jane part is Jane's id. Finder looks for three things: an id like ?ref=, ?via= or ?aff=; a redirect path like /go/ or /recommends/; and links through an affiliate network, which is a company that sits in the middle and handles the paying.

Finder names a competitor only when it can prove which one. If the link goes to your competitor's own site, it says so. If it goes through a network, the network hides who's on the other end, so the row says a link exists and names nobody. A plain link to a competitor's homepage is not a tracking link, and Finder doesn't count it.

Where this comes from:

Row typeWhere Finder reads it
YouTubeThe video descriptions and the channel's About text, which Finder already has
WebsiteThe search result's title and description, then their own pages once you save them

The scan itself costs nothing extra and adds no waiting. It never visits anyone's page. A search result is a title and a couple of sentences, though, and it rarely has links in it. So when you save a website, Finder reads up to four of their pages: the three that earned the score and their home page. It looks at every link on them. It uses the same rules as above. The row then says when it saw the link: "seen on their page 23 Sep 2026".

That page check only runs for sites you saved, so it never costs you anything you didn't pick. The next section says exactly what it reads, and what it leaves alone.

You'll see the link on the row, with the id taken off. Finder never stores or shows somebody's affiliate id.

A row that carries a link to one of your competitors gets up to 10 extra relevance points. A link Finder can't tie to a competitor gets 5. Nothing loses points for having no links, and plenty of good affiliates have none.

What a save reads

A scan only reads search results. Saving a website is what lets Finder look closer, and only at that one site, never the rest of the list.

Their pages. Up to four: the three pages that earned the score, then their home page.

  • Before each page, Finder reads the site's robots.txt. That's the file where a site tells bots which pages to leave alone. If it says no, Finder doesn't load the page, and the drawer says how many it skipped.
  • If robots.txt can't be reached because the site is down or slow, Finder skips the whole site. No file at all means no rules.
  • Finder introduces itself as AffiliateRailFinder/1.0 (+https://affiliaterail.com), so a site owner can see who called.
  • It reads public pages only. It gives up on a slow page after 6 seconds and reads at most the first 512 KB.
  • It never looks for or keeps an email address. Finding an address is what an email credit does.

What the pages say. From the same pages, with the page each fact came from:

  • Their social profiles on YouTube, X, LinkedIn, Instagram and TikTok. A profile counts only if their home page links to it, or two of their pages do. A review links to other people's channels all the time, so one mention isn't enough.
  • Whether they run a newsletter on Substack, beehiiv, Kit or Mailchimp. A signup form counts on any page. A plain link counts only on their home page.
  • Tracking links, as in Already carries tracking links.

Where their search traffic comes from. DataForSEO, the search data company behind the scan, estimates how many visits a site gets from Google in each country. The drawer shows the top five as shares, labelled "Search traffic by country (DataForSEO estimate)". A country under 5% of the traffic or 200 visits is counted and left out. Nobody counted these visits. Read them as a rough split.

  • Big sites can have more countries than DataForSEO returns in one answer. When the list comes back full, it may be missing the biggest markets, so Finder shows no split at all rather than a wrong one.
  • Most small sites are covered, and small sites are the ones worth emailing.

What they rank for. Up to five searches they show up for in your country, with where they rank and how many people search it a month. A search with under 100 searches a month is counted and left out.

How big they are, from real visits. Google's Chrome UX Report (CrUX) ranks sites by real visits from Chrome users: search, email, social, the lot. It puts each site in a bucket, like "top 50,000". The drawer shows the bucket for your country if the site is in it, or worldwide if not, credited "Chrome UX Report (CrUX) by Google, CC BY 4.0". Most small sites aren't in any list. That's fine. Finder never calls a site "unranked" or "small" because it's missing.

If a website has no traffic estimate, its CrUX bucket fills in for it in the score's audience fit: the top 10,000 counts as High, the top 100,000 as Medium and the top 1,000,000 as Low. A real traffic estimate always wins.

What it costs, and when it runs. The page reading and CrUX are free. The traffic split and the keywords cost us about 3 cents a site, so a trial gets them on its first 50 saved websites. Every saved site still gets its pages read. Finder asks DataForSEO again once a month for a saved site, so the numbers stay current.

Some rows say Links to your competitors where the others say "Weekly scan". These didn't come from Google. Finder asked a backlink index: a huge list of which pages link to which sites, built by a company that reads the web all day. DataForSEO runs the one Finder uses.

For each competitor, Finder asks the index who links to them, and keeps a site only if the link is one you could check yourself:

  • A tracking link to the competitor's own site, like ?via=, ?fpr= or ?ref=, or a redirect path like /go/.
  • A link through the competitor's own page on a network. PartnerStack gives each brand an address like rewardful.partnerlinks.io, and Impact does the same on sjv.io and pxf.io. Finder only asks about an address like that once the scan has already seen it.
  • A plain link from a real piece of writing, like a review or a guide. A plain link from a footer, a tag page or a list of links doesn't count.

Two things never count. One is a link to another business's programme that happens to run on a competitor's platform, like a signup page on stasher.tapfiliate.com. That makes the site Stasher's partner, not a Tapfiliate affiliate. The other is a page on a customer's corner of a shared host, like acme.freshdesk.com. It belongs to Acme, not to Freshdesk.

A row found this way says why, in one line. Let's say a blog has a ?via= link to Rewardful on its roundup page. The row says "Links to Rewardful with a ?via= tracking link, from /best-affiliate-software". Open it. The link's there. It's the same check as Already carries tracking links, and the id is taken off the same way.

A site that already has a tracking link to a competitor is the strongest row Finder can show you, because they've already signed up for a programme like yours and they're getting paid from it right now. Same checks, though. A vendor, a directory or a site on the wrong side of the deal is hidden the same way as any other row.

What it can't see:

  • A site that hides the link behind its own redirect, like theirblog.com/go/tolt. The index sees a link to their own site, so it doesn't know where it goes.
  • A link a page adds with code after it loads.
  • A brand's own custom address, like try.brand.com.

When it runs. On your first scan, every weekly scan, and your own competitor searches. It doesn't run on topic or lookalike searches, which aren't about your competitors. It runs after the Google searches and shares their budget. If the scan runs out of money, this part stops first, and the Google results are kept.

What it costs. About 3 cents a competitor, charged before each question. It never spends more than 40 cents in one scan. On the free trial, the Google searches come first, so most scans can ask about four or five competitors.

Podcasts

Some rows are podcasts. They come from Podcast Index, a free public list of nearly every podcast in the world, and from each show's RSS feed: the public file a show publishes to list its episodes, which is what every podcast app reads to show you new ones. Everything on a podcast row comes from that feed.

How a show gets on the list.

  • Finder searches episode titles and descriptions for each competitor's name. A one-word name gets your category's word on the end, the same as on Google. So "Impact" becomes "Impact affiliate".
  • It searches show titles for your keywords.
  • It reads the feeds those searches point to, at most 40 of them.
  • A show makes the list only if one of its own episodes names a competitor, or matches a keyword in its title. That episode is the row's evidence. Open it and check.

A search hit alone never counts. A show about "rewarding habits" isn't about Rewardful, and an episode on "the impact of AI" isn't about Impact. The name has to mean your competitor, in the episode itself.

What's left out. A competitor's own show, a platform's own show, a show with no new episode in a year, and a show in another language. These are counted on the scan and never written, so they can't take a place in your top 50.

Links in the show notes. Many shows put links under each episode. A tracking link to a competitor there, like ?via= or ?fpr=, is the same proof as on a website: they already earn from a programme like yours. The row says Already carries tracking links, "in their show notes". A link to someone else's programme doesn't count as naming a competitor.

What a podcast row doesn't have: an audience number. Only the host sees a show's downloads. So there's nothing honest to show, and the row says Audience fit: not measured. Its 30 points follow relevance and recent activity instead, which means a show earns them only by staying on topic and putting out episodes often. Recent activity counts episodes a month over the last ten. Weekly is High. Fortnightly is Medium.

How many. Up to 40. Podcasts never take a place from YouTube or websites, which keep their shares of the scan exactly as before; podcasts fill whatever room those two leave, best scores first.

The contact. Apple asks every show for an owner address, and the feed publishes it. When you reveal a podcast's email, Finder reads it then, never before. The address comes from podcast_rss_owner. If it belongs to a hosting company rather than the show, Finder reads the show's own website instead. It costs the same one credit as any reveal.

What it costs us. Nothing. Podcast Index is free, and so are the feeds.

What the weekly scan re-does

The weekly scan runs your whole search again, so it can tell you three things a single scan can't.

  • A traffic trend. Each weekly scan adds that month's estimate to the row instead of replacing it. It also asks for the last twelve months for the top 25 rows on your list, so the drawer can draw a trend straight away. That lookup costs about 15 cents a weekly scan. We quote the higher of DataForSEO's two published prices for it, and your first scan never makes it. A trend needs two months. Until then the drawer says "No trend yet".
  • Where they rank now. Each page keeps every search that found it and where it ranked, so the row's "Ranks 3rd for 'affiliate software'" is this week's place, not a guess.
  • Who stopped showing up. A missed scan is a weekly scan that didn't find a row you can see. After two in a row, the row says Not seen in the last 2 scans and its recent activity drops to 3 of 20. Find them again and it resets. Only the weekly scan counts: a search you start is narrow on purpose, so not finding a row says nothing. A source that failed that week doesn't count either.

When the next scan runs

Finder scans your brand again about a week after your last full scan, on a weekday that stays the same for your brand. The first weekly scan lands 7 to 13 days after your first one, at 04:30 UTC; every one after that is 7 days on. Brands are spread across the week so one morning never runs more scans than YouTube allows in a day. The header on Discover counts down to the real date. Rows found by the latest scan carry a NEW chip.

If your trial has ended, the weekly scan doesn't run. Everything you've found stays where it is.

Where a YouTube channel ranks

Finder remembers where YouTube placed each video for each search, from 1st to 50th. A channel scores on the same steps as a website: 1st to 3rd, 4th to 10th, 11th to 20th. The row says "3rd in YouTube for 'Impact review affiliate'". Below 20th it says "Found in YouTube for" and no number, because a 34th place isn't worth quoting. In the drawer, each video shows its place beside its views.

The score on a channel is Finder's own, worked out from YouTube's numbers. It isn't a YouTube rating. So the drawer shows YouTube's own counts beside it: subscribers, total views and videos.

When YouTube's daily searches run out

YouTube caps how many searches an app can make in a day, and every Finder account shares that cap. On a busy day your scan can find it used up. When that happens:

  • The website half of your scan runs as normal.
  • The YouTube half waits. The banner says YouTube results arrive tomorrow morning.
  • At 04:30 UTC the next day, the morning scan runs your YouTube search before anyone's weekly scan. You get the same number of searches your first scan was owed.

If only a few searches are left, Finder runs the most useful ones that fit, and waits only when fewer than 10 are left.

When a scan can't run

A scan that can't start, or fails part-way, tells you why in one sentence:

  • Nothing to search for. Add at least one competitor or keyword.
  • Plan can't scan. Your trial ended, or your plan lapsed. Pick a plan to scan again.
  • Research budget used. Finder caps what it spends on outside data for each trial. If a scan reaches the cap it stops cleanly, keeps what it found, and says so.
  • Provider calls are off. We switched them off on our side. Nothing was charged. Try again later.

A search that cost you a search credit and then failed gives the credit back on its own, and the message says so.

What we keep

Website rows, their scores and their evidence stay for as long as your account does. Normalized traffic facts and the site's about text stay with the row too. The raw data a provider sent us is deleted after 90 days, but removing that raw payload does not change a later score.

Podcast rows work the same way. The feed's raw data goes after 90 days. From Podcast Index itself, Finder keeps two things, the feed's address and the show's id in the index, and never a copy of its listing.

YouTube rows follow YouTube's rules, which are stricter. YouTube lets an app keep its data for 30 days unless it refreshes it. A weekly scan refreshes every channel it finds again. A channel not found for 30 days loses everything YouTube gave us: its name, country, about text, subscribers, average views, upload cadence, and each video's title, date and views. Its score goes too. We built it from those numbers. What stays is the channel's link and anything you did to the row, like saving it or adding a note. The row says YouTube's details were removed. Find the channel again in a later scan and it all comes back.