The question that separates a real alternative from a skin
Every search engine needs an index — its own stored, searchable copy of the web, assembled by a crawler (a program that fetches pages and follows the links out of them). Building one means running a crawler continuously against hundreds of millions or hundreds of billions of URLs, storing the text, and maintaining the machinery that ranks it. It is the single most expensive thing a search engine does, and it is the only part a company cannot fake.
Most of the search engines you have heard of do not do it. They buy finished results from someone who does, put their own interface, privacy policy and filters around them, and sell advertising against the page. That is a legitimate product — it is precisely what Startpage sells, openly — but it is not an alternative index, and treating it as one is the mistake nearly every “best Google alternatives” list makes.
This site records the answer for every engine it covers, as one of four states:
- Own index — the engine crawls the open web itself and ranks its own copy of it.
- Hybrid index — the engine runs a real crawler, but blends its own results with a partner's feed.
- Resold results — no crawler at all; the engine re-serves another company's results.
- Not a web index — the product answers from curated or computed data rather than a crawl, as Wolfram Alpha does.
None of these is a quality grade. Resold results can be excellent results — they are usually Google's or Microsoft's. But the consequence is structural and it is absolute: an engine with no index of its own cannot differ from its supplier on coverage. If Bing has never crawled a page, no Bing reseller can show it to you, no matter how good its ranking, how private its policy or how clean its design. What a reseller can change is the interface, the privacy handling, the advertising, and what it chooses to filter out. That is the whole of the difference, and it is worth knowing before you switch.
The independent web indexes, named
There are roughly a dozen organisations on earth running a general web crawl at scale. This is the list.
- Google — Googlebot, the reference index for the industry, and the largest by a wide margin.
- Microsoft Bing — bingbot. The second index the Western web runs on, and the wholesale supplier behind a large share of the “alternative” market.
- Yandex — YandexBot. A full independent index, strongest in Russian.
- Baidu — Baiduspider, with its own webmaster platform. Dominant in mainland China, and filtered in line with Chinese law.
- Brave Search — Bravebot. Built on the Tailcat stack Brave acquired from Cliqz in March 2021; began as a hybrid at about 87% own results in 2021 and announced the removal of its last Bing API calls on 27 April 2023, adding its own image and video indexes that August.
- Mojeek — MojeekBot, run by a small British company that has never resold anyone's results at any point in its history. Its own About page gave 9 billion pages for 2025.
- Marginalia Search — crawler and robots.txt token
search.marginalia.nu, run by one engineer in Sweden. Around 300 million documents as of 2024, deliberately ranked to favour text-heavy non-commercial pages. - Seznam — SeznamBot. Czech, and arguably the only national-scale independent index operated by an EU-headquartered company. It licensed its fulltext engine until 2005, then built its own.
- Naver — the Yeti crawler and the Naver Search Advisor webmaster platform. It does crawl the open web, though its results page is dominated by Naver's own blogs, cafes and Knowledge iN content.
- Cốc Cốc — a documented crawler fleet (
coccocbotand siblings) that, by Vietnam's own help pages, deliberately limits itself to Vietnamese content. A genuine index, but a national one. - EUSP / Staan — the European index built by European Search Perspective SAS, the 50:50 Paris joint venture that Ecosia and Qwant own together. It went live serving real traffic on 6–7 August 2025, starting in France. It is not a consumer search engine; it is an index its two parents draw on and sell access to.
Two entries need an asterisk. Sogou crawls the Chinese web itself but has served English-language results from Bing since a 2016 Microsoft partnership, and its most distinctive verticals are licensed corpora from Tencent and Zhihu — a blend, not a pure own index. Perplexity is a newer and stranger case: it spent roughly its first two years as a Bing reseller with a summary layer on top, and by September 2025 its own engineering write-up described an exabyte-scale crawl tracking over 200 billion unique URLs. That is a real index, built by an AI company rather than a search company, and the residual third-party mix is not disclosed.
That is the complete set. Every other search engine in this directory — and every entry on every listicle — is downstream of one of the names above.
Who resells whom
Naming suppliers is the part most write-ups skip, so here it is plainly. Bing supplies Yahoo, which has not run its own crawler since 2010 and has served Bing results under the Search Alliance signed on 29 July 2009 — including behind Yahoo Scout, the AI answer engine launched in January 2026, which Yahoo describes as Bing's index with Yahoo's ranking and experience over it. Bing also supplies Swisscows, which states its “exclusive partnership with Bing” on its own blog, and Petal Search, Huawei's product, whose standalone web engine was quietly discontinued on 25 June 2023. Bing supplies the web links behind DuckDuckGo and, historically, the bulk of Ecosia's and Qwant's results.
Google supplies Startpage, which relays your query anonymously and returns what comes back — Google only from 7 July 2009 until March 2023, and Google and Bing since. It supplies Kiddle, which is a Google Programmable Search instance with SafeSearch and a hand-curated allow list on top. And by the account of its own founder in 2012, and of independent testing by Larry Sanger in 2022, it supplies Gibiru — which markets itself as an uncensored independent engine and has never published a crawler, a user agent or an index size in seventeen years of operating.
Yandex supplies Rambler, which shut down its own search technology in 2011 and has run on Yandex's results ever since.
And one company supplies a surprising number of the rest. System1 owns InfoSpace, and with it Dogpile, WebCrawler and MetaCrawler; it supplies search to Excite, whose privacy policy states that “Excite search services are provided through a third party, System1 LLC”; and it is the majority owner of Startpage. Several distinct-looking brands on a “Google alternatives” list are the same corporate entity buying the same feeds. WebCrawler is the sharpest illustration of how far a name can drift from a product: the site named after the web's first full-text crawler stopped crawling in 2001. HotBot is worse — the domain was sold in 2016 and now operates as an AI chat aggregator with no relationship to the engine of the 1990s, which was never running its own technology anyway; it was Inktomi's.
Two entries deserve to be separated from the rest, because “resold” does not mean the same thing for them. SearXNG has no crawler and no index, but it is free software you install, it queries a set of engines you choose, and it publishes every line of its source under the AGPL. That is transparent, non-commercial, user-configured aggregation, and it is a different proposition from buying a feed and putting a logo on it. Microsoft Copilot also has no index of its own and grounds its answers on Bing search queries — but Bing is its own parent's index, so it is not reselling a rival's work so much as writing prose over the house feed.
One more, flagged rather than settled: Presearch has claimed since September 2025 to have moved from metasearch to its own index. No crawler name, user agent, robots.txt token, crawl IP range or index size has been published. Until one is, the documented mechanism is still metasearch — and its own privacy policy still logs which provider answered each query.
The hybrids, and why they are the hardest to describe honestly
The middle category is where careless writing does the most damage, because both simple summaries are wrong.
DuckDuckGo runs two documented crawlers, DuckDuckBot and DuckAssistBot, and maintains its own indexes — but those power the Instant Answers, not the ranked links. Its own help page says it sources “more traditional links and images” largely from Bing. When a Bing API outage hit in 2024, DuckDuckGo stopped returning web results, alongside other Bing-fed engines. So “DuckDuckGo has its own index” is misleading and “DuckDuckGo is a Bing skin” is also wrong; the accurate line is that Bing supplies the links and DuckDuckGo supplies everything around them. Its own sourcing page is the best primary citation on the question: duckduckgo.com/duckduckgo-help-pages/results/sources.
Qwant is the cautionary tale. It marketed itself as an independent European engine for years; in early 2019 an investigation found it had neither a crawler nor an indexer and was wholly reliant on Bing. It has since built one. In September 2023 Qwant published its own figures — around 20 billion pages indexed, over a billion pages crawled a day — alongside a self-measured 51% overlap with Bing's results, and it still says it uses Bing to supplement queries where its own relevance is weak and for images. Since August 2025 it also serves results from Staan.
Ecosia is the inverse: it has no crawler of its own but half-owns one. Its help centre, updated in August 2026, names three providers — Microsoft Bing, Google, and EUSP. The share coming from the index it co-owns has not been published; only rollout targets have, and targets are not results.
Kagi runs Kagibot into its own Teclis web index and TinyGem news index, and its documentation says the bulk of a results page comes from anonymised requests to third-party commercial indexes, Brave's among them. ChatGPT Search runs OAI-SearchBot and also buys web results; OpenAI's own help page says it “sometimes partners with other search providers.”
What the index decides, and what it does not
Sorting engines this way is only useful if you know what follows from it. The index determines coverage (whether a page is in the system at all), freshness (how quickly a new page can appear), and recall (whether an obscure page can be found by an obscure query). None of these can be improved by a reseller, because none of them lives in the layer a reseller controls.
What a reseller genuinely can change: whether your query and IP address are linked to an identity; whether you are profiled and personalised; what the page looks like and how many ads are on it; what gets filtered out, boosted or removed; and whether an AI summary is placed above the links. These are not small things. Startpage's product is exactly this trade — mainstream result quality without the mainstream engine watching you — and for a reader whose objection to Google is surveillance rather than coverage, it answers the question well.
But two consequences follow that most readers never hear. First, a reseller is not a second opinion. Checking a Google result against Startpage, or a Bing result against Yahoo or DuckDuckGo, tells you almost nothing: you are asking the same index twice. If you want to know whether a page exists outside the duopoly, you have to ask an engine that crawled the web itself — Mojeek and Marginalia are the two most useful for exactly this, precisely because they will show you things Google has quietly stopped showing.
Second, switching to a reseller does not reduce the concentration of search. Every Startpage query is ultimately a Google or Bing query, and Startpage pays for it. The data flow changes; the money flow and the market structure do not. Anyone choosing an engine on sovereignty or diversity grounds rather than privacy grounds needs a different shortlist — and it is a short one.
Why there are so few, and why the number is not fixed
The usual explanation is that only a company with Google's money can index the web. That is not quite right, and Marginalia is the proof: one engineer in Sweden, running costs he puts at roughly $200 a month, a few hundred million documents, a documented crawler with a published IP range and a real ranking function. The floor for an independent index is remarkably low. It is the ceiling that is expensive — comprehensiveness, freshness at news velocity, and every non-English language at once.
What is genuinely getting harder is access to content, and this is the part worth understanding. In 2024 Reddit blocked non-Google crawlers while Google retained access under a paid deal, cutting Mojeek and every other small crawler off from an enormous body of text. Consolidation in search is now maintained by content licensing at least as much as by technology, and a licensing barrier is one a clever engineer cannot route around.
The supply side moved too. Microsoft raised Bing API pricing sharply in 2023 and then retired all public Bing Search APIs on 11 August 2025, pointing customers at an LLM-grounding product in Azure instead. Longstanding direct syndication contracts — DuckDuckGo's, for instance — continued, but the open route to reselling Bing narrowed considerably. Anyone assessing an engine that “uses Bing” after that date should ask under what contract.
Two things push the other way. Ecosia and Qwant's Staan index is the first serious attempt at a new general web index in a decade, and it exists largely because Bing pricing made renting one untenable. And in the US antitrust remedies of 2 September 2025, Judge Mehta ordered Google to share search index and click-and-query data with qualified competitors and to license its results and search text ads on five-year terms, initially capped at 40% of queries — calling that data “the raw material that Google uses to improve search.” Those remedies took effect on 3 February 2026 and are under appeal by both sides in the D.C. Circuit, with no decision as of August 2026. Nobody should predict the outcome. But it is the first time a court has treated index access as the thing that matters, which is the argument this page has been making throughout.
How to check an engine for yourself
Claims about independence are marketing until they are documented. Four checks separate the two, and none requires technical skill.
- Look for crawler documentation. An operator that genuinely crawls publishes a named user agent, a robots.txt token, usually an IP list, and a contact address, because it needs site owners to cooperate. Mojeek, Marginalia, Seznam, Brave, Google, Bing, Yandex, Baidu, Naver and Cốc Cốc all publish this. Gibiru never has. Presearch has not, a year after announcing an index.
- Read the help pages, not the homepage. Ecosia, DuckDuckGo, Startpage and Swisscows all state their sources in their own support documentation — the marketing page rarely does. Ecosia's is unusually candid: support.ecosia.org/article/579-search-results-providers.
- Search for something rare and compare. Take an unusual exact phrase and run it on the engine, on Google and on Bing. A reseller returns the supplier's set, in something close to the supplier's order. Mojeek or Marginalia will return a visibly different set, often including pages the big two omit entirely.
- Watch what breaks. Supplier outages are the most honest disclosure in this industry. The 2024 Bing incident that silently emptied several engines' results pages told users more in an afternoon than years of about pages had.
And one rule that saves the most trouble: do not infer an index from an announcement. A press release is not a crawler. Qwant claimed independence for six years before it had one.
Frequently asked questions
Which search engines have their own index?
Google, Microsoft Bing, Yandex, Baidu, Brave Search, Mojeek, Marginalia, Seznam, Naver and Cốc Cốc all crawl the open web themselves and rank their own copy of it, as does the Staan index that Ecosia and Qwant jointly own through their European Search Perspective venture. Perplexity has built one since 2025. That is close to the complete list worldwide. Everything else you are likely to encounter buys its results from one of them.
Which search engines use Bing's results?
Yahoo has served Bing's web results since the 2009 Search Alliance and has not crawled the web since 2010. Swisscows states an exclusive Bing partnership. DuckDuckGo sources its ranked web links from Bing while running its own crawlers for other purposes. Huawei's Petal Search used Bing before its standalone web engine was discontinued in June 2023, and Microsoft Copilot grounds its answers on Bing queries. Ecosia and Qwant both still blend Bing into their results.
Does Startpage use Google?
Yes, and since March 2023 also Bing. Startpage has no crawler and no index; it relays your query to those engines from its own servers without your IP address or cookies and returns their results. It was Google-only from 7 July 2009 until the Microsoft partnership announced on 14 March 2023 added Bing web results and Bing ads. Startpage does not publish which provider answered a given query, so there is no way for a user to tell.
Is DuckDuckGo its own search engine or does it use Bing?
Both statements are partly right, which is why it is classified as a hybrid. DuckDuckGo operates its own crawlers, DuckDuckBot and DuckAssistBot, and its own indexes behind the Instant Answer boxes. But its own help page says the traditional links and images are largely sourced from Bing, and a Bing outage in 2024 took its web results down. It is a second interface onto Bing's index, not a second index.
Why does it matter whether a search engine has its own index?
Because coverage cannot be changed by a reseller. If the supplier never crawled a page, no engine buying that supplier's feed can show it to you, however good its ranking or privacy policy. A reseller can change the interface, the tracking, the advertising and the filtering, which are real differences. It cannot give you results the supplier does not have, and it cannot serve as an independent check on the supplier.
Are resold results worse than an engine's own results?
Usually the opposite, on breadth. Resold results are normally Google's or Bing's, which means better long-tail recall and freshness than any small independent index can match. “Resold” is a description of where results come from, not a judgement of quality. What it does mean is that the engine adds nothing to the diversity of search, and that its supply depends entirely on a contract it does not control.
How can I tell if a search engine really crawls the web?
Look for published crawler documentation: a named user agent, a robots.txt token, an IP range site owners can verify against, and a contact address. Every genuine crawler publishes these because it needs webmasters to cooperate. Then compare results for a rare exact phrase against Google and Bing. An engine that returns the same set in roughly the same order is not running an index of its own, whatever its homepage says.
Why are there so few independent search indexes?
Less because of raw cost than people assume — Marginalia runs a few hundred million documents on roughly $200 a month — and more because of scale and access. Comprehensiveness, news-speed freshness and full multilingual coverage are genuinely expensive. And large content platforms increasingly block small crawlers while licensing access to Google, as Reddit did in 2024, which is a barrier no amount of engineering can solve.
Sources
- duckduckgo.com/duckduckgo-help-pages/results/sources
- support.ecosia.org/article/579-search-results-providers
- brave.com/blog/search-independence/
- betterweb.qwant.com/en/2023/09/18/web-indexing-where-is-qwants-independence/
- mojeek.com/bot.html
- about.marginalia-search.com/article/crawler/
- learn.microsoft.com/en-us/lifecycle/announcements/bing-search-api-retirement
- blog.ecosia.org/launching-our-european-search-index/