This is the full directory: fifty-eight search engines, live and dead, each carrying a single label that says where its results come from. That label is the organising principle of the whole site, so it is worth explaining before the cards below make sense.
Why the own-index question organises everything
A search engine has three parts that can be owned separately. A crawler fetches pages and follows links. An index is the engine's own stored, searchable copy of what was fetched. A ranking system chooses what to show. Crawling and indexing the open web is a permanent, heavy, largely invisible cost — servers, bandwidth, storage, and a crawl that must keep running because an index that stops being refreshed rots within weeks. Ranking, branding, privacy engineering and AI answer layers are comparatively cheap and are what users see.
The predictable result is that far more companies operate search products than operate search indexes. Buying a results feed from Google or Microsoft, or querying other engines at the moment you press enter, produces a working search engine without any of the infrastructure. This is a legitimate business and several engines here do it deliberately and say so. But it means the practical answer to "is this a real alternative?" is not found in the design or the privacy policy. It is found in whether there is an index underneath, and whose.
The five states
- Own index — the engine runs its own crawler and ranks its own stored copy of the web. Google, Bing, Yandex, Baidu, Naver, Seznam, Cốc Cốc, Brave Search, Mojeek and Marginalia are here. So, on the strength of its published crawl and index architecture, is Perplexity.
- Hybrid index — a real crawler and a real index of its own, blended with somebody else's feed. DuckDuckGo crawls and maintains indexes but sources its ranked web links largely from Bing. Qwant crawls at scale and still supplements with Bing where its own relevance is weak. Kagi runs its own crawler and two small indexes and calls out to commercial indexes for general coverage. Hybrid is the most common honest position for an ambitious independent, and it is also the state most often described inaccurately in both directions.
- Resold results — no crawler, no index; results come from somebody else. This covers Yahoo (Bing since the migration completed in the early 2010s), Startpage (Google and Bing, relayed anonymously), Swisscows (an exclusive Bing partnership, stated on its own blog), Microsoft Copilot (grounded on Bing), Dogpile, MetaCrawler, WebCrawler, SearXNG and others. It is not a criticism. An anonymising relay in front of Google is a coherent product that cannot exist any other way.
- Not a web index — the engine answers questions but does not crawl the web at all. Wolfram Alpha computes answers from curated data and stated methods rather than retrieving pages. Jayde indexes only business sites submitted and reviewed by hand. Neither is a web search engine, and forcing them into the other four states would misdescribe both.
- No longer available — you cannot use it as a web search engine today. That covers outright shutdowns such as AltaVista, and companies that withdrew from consumer search while thriving elsewhere, such as You.com and Northern Light. Each profile says which.
The labels are dated, not permanent
An engine's state changes, sometimes sharply. Brave Search launched in 2021 taking roughly an eighth of its results from a third-party API and announced the removal of the last of those calls in April 2023 — it was a hybrid and became independent. Qwant marketed itself as an independent European engine for years, and a 2019 investigation found it had neither its own crawler nor its own indexer at that time; it has since built both, and remains a hybrid. Perplexity began by summarising results retrieved through a third-party API and now publishes the engineering of a large crawler and index of its own. Yahoo ran a genuine crawler from 2004 and gave it up. Read every label with its date attached.
The ten categories
Categories describe what an engine is for, not how good it is, and they cut across the index question — there are resold engines and independent ones in almost every group.
- Major global search engines — the general-purpose engines with worldwide reach and no regional specialisation.
- Privacy-focused search engines — engines whose stated proposition is what they do not collect. Their underlying indexes vary enormously.
- AI-native search engines — products whose primary output is a written answer with citations rather than a list of links.
- Regional and international search engines — engines with a dominant or substantial share in one country, usually built on a language advantage, a portal business, or state and trade policy.
- Meta-search engines — engines that hold no index and query others at the moment you search.
- Vertical and specialist search engines — search over one domain of knowledge or one curated collection rather than the open web.
- Kids and safe search engines — filtered front ends intended for children, almost always built over a mainstream engine's results with an allow-list on top.
- People and public-record search — search over identity and record data rather than web pages, with a quite different set of legal and ethical questions.
- Legacy engines still online — 1990s brands that still resolve, still have a search box, and no longer have any search technology behind them.
- Search engines you can no longer use — the twenty-three that stopped, from ALIWEB and Archie to Neeva, Gigablast and Ask.com.
Two categories are worth reading against each other. The legacy group and the defunct group look similar from the outside and are structurally opposite: a legacy engine's domain is a live commercial property with somebody else's results behind it, while a defunct engine's domain is usually parked, redirected, or repurposed into something unrelated. Confusing the two is why people still describe Lycos as a search engine and AltaVista as merely unpopular.