What a meta search engine actually is
A meta search engine owns no index. It has no crawler of its own, stores no copy of the web, and computes no ranking from scratch. What it owns is a dispatcher: it takes your query, sends it to several other search engines at once, waits for their answers, strips out the duplicates, and shows you one merged list.
That is the whole architecture, and it has not changed since 1995. Everything interesting about metasearch happens in three places — which engines get asked, how the answers are merged into a single ranking, and whether you are told where each result came from. Engines that answer those three questions openly are useful research tools. Engines that answer none of them are advertising properties with a search box.
The distinction that matters most is that metasearch is not a bigger search engine. It is a thinner one. Each upstream engine returns only its top slice, so a merged metasearch page usually contains far fewer total results than any one of its sources would have given you directly. Whether that trade is worth making depends entirely on whether the sources genuinely differ, which is the question the rest of this page turns on.
Why it made sense when no engine had the whole web
In the mid-1990s no single search engine came close to indexing the web. Each crawler found a different, partly overlapping slice, and the slices were small. Querying six engines and merging the answers was not a gimmick under those conditions — it was the only practical way to see a reasonable fraction of what existed.
The founding academic work states the case directly. When Erik Selberg and Oren Etzioni launched MetaCrawler at the University of Washington on 7 July 1995, it queried six services — Galaxy, InfoSeek, Lycos, Open Text, WebCrawler and Yahoo — because no one of them was sufficient. Their system also did something almost nobody does now: it fetched each returned URL to check that the page still existed before showing it to you. Dead links were a serious problem in 1995, and verification was a real feature. The whole thing was 3,985 lines of C++ and handled 50,878 completed queries between 7 July and 30 September 1995.
The best-known measurement of engine disagreement came later and from an interested party. In 2005, research commissioned by Dogpile and carried out with researchers at Queensland University of Technology and Penn State reported that only about 1.1% of first-page results overlapped across Google, Yahoo! and Ask Jeeves simultaneously, and 3.2% across any two. The study is real and constantly cited — but it was funded by a metasearch engine to justify metasearch, and should always be quoted with that attached. It shows that the top results differed; it does not show that metasearch coverage was broader.
The founding generation, 1994 to 2001
WebCrawler is the odd one out and belongs here for the opposite reason. Launched on 21 April 1994 by Brian Pinkerton at the University of Washington with an index of roughly 4,000 sites, it was the first web search engine to offer full-text search of the pages it indexed — earlier tools indexed filenames, titles or human-written descriptions. It crawled its own web until 2001, when its new owner switched it to Excite's database. The site named after the first web crawler has not crawled anything since; today it is a metasearch front end over feeds it does not disclose.
MetaCrawler (7 July 1995) is the genuine origin point. It was built as a research instrument, not a business, and its 1995 conference paper established the pattern every later metasearch engine follows: parallel dispatch, collation, de-duplication. Despite the name, it has never had a crawler.
SavvySearch came out of Colorado State University in 1995, created by graduate student Daniel Dreilinger with Professor Adele Howe. Its distinguishing idea was that it learned which engines to ask: it kept statistics on which back ends produced useful results for which kinds of query and routed accordingly, while modelling network load so it degraded gracefully instead of timing out. That work was published in AI Magazine in June 1997, which makes it one of the few 1990s engines whose "intelligent" label was earned in peer review rather than in marketing. CNET absorbed it into Search.com in October 1999 for a reported US$22 million. Note a source conflict worth printing: Search Engine Watch has published both March 1995 and May 1995 as its launch month, so "1995" is as precise as the record allows.
ProFusion, built at the University of Kansas by Susan Gauch with Guijun Wang and Mario Gomez in 1995–96, solved the same routing problem differently. It classified each query into one of thirteen categories derived from Usenet newsgroup hierarchies, picked the three engines calibrated as best for that category, merged results using each engine's own relevance score discounted by its measured reliability, and removed duplicates three ways — identical URLs, normalised variants, and n-gram similarity across URL paths. Its authors reported it beating MetaCrawler and SavvySearch, on twelve queries judged by themselves; repeat that only with the sample size attached. Intelliseek bought it in 2000 (March per Information Today, April per Search Engine Watch) and repositioned it as a gateway to the "Invisible Web" — the patent, imagery and archive databases crawlers could not reach.
Mamma.com launched in Montreal in July 1996, out of a 1995 Carleton University master's thesis by Herman Tumurcuoglu, and marketed itself for years as "the Mother of All Search Engines." Dogpile followed in November 1996, created by Aaron Flin, and won J.D. Power awards for residential online search in 2006 and 2007.
Two more deserve rescuing from obscurity. Inference Find queried Yahoo, AltaVista, Lycos, WebCrawler, Infoseek and Excite and then did something better than ranking them: it clustered results into groups, including by site type — commercial, educational, non-profit. It also let you set the timeout yourself, from one to thirty seconds, exposing the coverage-versus-speed trade-off directly to the user. It carried no advertising on its search pages as a matter of principle, and that is what killed it: spun out of Inference Corporation in February 2000, it ran out of capital and closed in March 2001, reportedly still serving around 35,000 visitors a day won entirely by word of mouth. Highway 61 ranked by agreement — a page returned by more of the underlying engines placed higher — reported how many hits each engine contributed, and labelled its timeout control "Your patience level." Nobody knows who ran it: no founder, company, launch date or shutdown date survives in any source, and the domain now resolves to a blank page.
Who owns the survivors: the InfoSpace lineage
Anyone comparing Dogpile, MetaCrawler and WebCrawler as competing options is comparing three skins on one company's infrastructure. The consolidation happened fast and has never been undone.
- January 1997 — MetaCrawler is sold to Go2Net, Inc. (A common error to avoid: Excite acquired NetBot, MetaCrawler's commercial spinout, in October 1997 — after the MetaCrawler asset had already gone to Go2Net. Two transactions, routinely merged in secondary sources.)
- August 1999 — Go2Net acquires Dogpile, putting the two best-known metasearch brands under one roof.
- July 2000 — InfoSpace, Inc. acquires Go2Net in a stock deal reported at around $4 billion, at the peak of the dot-com market.
- 2001 — InfoSpace picks up WebCrawler out of the Excite@Home bankruptcy and switches it to Excite's database.
- June 2012 — InfoSpace, Inc. renames itself Blucora; the search properties become a segment rather than the company's identity.
- July 2016 — Blucora sells the InfoSpace search business, brands included, to OpenMail LLC for $45 million cash. That price, for Dogpile plus MetaCrawler plus WebCrawler, is the clearest available measure of what the category was worth by then.
- 2017 — OpenMail renames itself System1; the company lists on the NYSE via a SPAC merger in January 2022.
What that means in practice: these are advertising properties. Dogpile's own About page claims Google and Yahoo! as sources and its app listing names Bing, but the copy is undated legacy text, and Yahoo! has itself been Bing-powered since 2010, so three named engines may describe two upstreams. MetaCrawler's equivalent page still names Truveo, a video service shut down years ago, and contains an unreplaced internal template string. Neither homepage discloses its sources. You cannot tell whose results you are reading or how heavily paid listings are mixed in, which is exactly the transparency that made 1995 metasearch worth using.
This site's own domain has a small stake in that history: from around 1999 to the mid-2000s searchengines.net was a webmaster tool that generated "add this search engine to your site" HTML snippets, one page per engine, and the pages that attracted the most links from other sites were the WebCrawler, Dogpile and MetaCrawler ones.
The modern form: SearXNG and Presearch
The idea did not die; it moved to open source. SearXNG is metasearch software you install rather than a website someone runs for you. It began as searx, started by Adam Tauber in October 2013, and was forked into SearXNG in mid-2021; the original searx repository was archived on 7 September 2023, so anyone still running plain searx is running abandonware. SearXNG is AGPL-licensed, sells nothing, has no company behind it, and its documentation describes aggregating results from up to 269 search services. Which of those are enabled, how they are weighted, and whether anything is logged are all set by whoever runs the instance.
That configurability is the point and the catch. There is no "the SearXNG" — there are self-hosted copies and a few dozen volunteer-run public instances. Self-hosted, the privacy guarantee is a property of your own machine rather than a promise from a company, which is categorically stronger. On a public instance, you have swapped trusting Google for trusting an anonymous volunteer with root access to the server your queries pass through, and nothing in the software prevents them logging everything. The trade-off runs both ways: a single-user instance makes every query attributable to your server's IP address at the upstream engine, so self-hosting gives you the best protection from operators and the worst anonymity in a crowd.
SearXNG also demonstrates the structural weakness of all metasearch better than any historical example. In June 2023 the project documented Google actively blocking instances — fresh installs blocked after roughly five consecutive searches, and the number of tracked public instances with a working Google engine collapsing from 66 of 91 to 25 within days. Rate-limit and error discussions continue. A metasearch engine depends on the continued tolerance of companies with a commercial interest in refusing it. That is not a bug awaiting a patch.
Presearch is the crypto-incentivised variant: founded in 2017 by Colin Pape in Canada, run from San Diego under different ownership since 2023–24, it pays users in a token for searching and routes queries through roughly 40,000 volunteer-run "nodes." Its own node documentation describes those nodes sending queries "to a range of search engines and APIs" — federation, which is metasearch — with crawling listed as a deferred future capability. Since September 2025 the company has claimed to be building its own index, and in December 2025 said the transition was complete. As of August 2026 that claim is unverified in the way that matters: no crawler name, no user-agent string, no published IP range, no index-size figure and no site-owner opt-out has appeared, and Presearch's live privacy policy still records which provider served each search.
Why it matters much less now
Metasearch was a strategy for a world in which nobody was good. That world ended. Two organisations now run the general web indexes that almost everything else consumes — Google and Microsoft — and a handful of others crawl independently at meaningful scale, including Yandex, Baidu, Brave and Mojeek. When one index is dramatically better than the alternatives and a second is close behind, merging six partial views stops adding information and starts adding noise, because you are blending rankings computed by systems that disagree about what relevance means.
Three consequences follow, and they are the honest case against most metasearch products:
- It cannot rank better than its sources. There is no ranking insight in a blend. Merged lists frequently rank worse than the best single source, because the good ranking is diluted by weaker ones.
- It offers no privacy by itself. Your query is, by definition, transmitted to third parties. Metasearch reduces your exposure only if the operator proxies the request and strips identifiers — which SearXNG does by design and which the System1 properties do not claim to do at all.
- It is hostage to its suppliers. Feeds can be repriced, rate-limited or withdrawn, and the results page just quietly gets thinner. The user is not told.
Two uses remain defensible. The first is deliberate result diversity: if you want to see what several engines say without visiting each, a metasearch page still does that, and it is genuinely interesting to watch how much they disagree. The second is control — SearXNG lets a technically comfortable person decide which engines answer, in what proportion, on hardware they own, which nothing else on offer does.
What is not defensible is treating a legacy metasearch brand as an "alternative" to Google. Dogpile, MetaCrawler and WebCrawler are one advertising company's properties, serving undisclosed feeds bought from the same duopoly you were trying to leave, with paid listings mixed in at a ratio nobody publishes. And Mamma is the sharpest warning of all: the site is alive, commercially active and still uses search language in its own copy, but as of August 2026 it has no search box and returns no results. A search brand outliving its search engine is now the normal end state for this category, not the exception.
Frequently asked questions
What is a meta search engine?
A meta search engine is a service that has no index of its own and instead sends your query to several other search engines at once, merges their answers, removes duplicates and shows you a single list. MetaCrawler, launched in July 1995, is the founding example. The defining fact is the absence: no crawler, no stored copy of the web, no ranking computed from scratch.
Do meta search engines give you more results than Google?
No, and this is the most common misunderstanding. Each upstream engine returns only its top slice of results, so a merged metasearch page normally contains far fewer total results than any single source would have returned directly. What metasearch gives you is a wider spread of sources across a shallower set of results, not greater coverage of the web.
Is Dogpile still working in 2026?
Yes. Dogpile still loads and returns results as of August 2026, and it has run since November 1996. But it is owned by System1, an advertising company, along with its sister sites MetaCrawler and WebCrawler. It publishes no current list of its result sources on its homepage, and its About page copy is undated legacy text, so you cannot tell whose results you are reading.
Are Dogpile, MetaCrawler and WebCrawler competitors?
No. They have shared an owner since 1999–2000, when Go2Net acquired Dogpile to sit alongside MetaCrawler and InfoSpace then acquired Go2Net. All three are System1 properties today, running on the same infrastructure with different branding. Comparing them as independent options is comparing three skins on one company's search product.
Is a meta search engine more private?
Not inherently. A metasearch engine forwards your query to third-party engines by definition, so the query still reaches them. Privacy only improves if the operator proxies the request and strips your IP address, cookies and fingerprint — which SearXNG does by design. The commercially operated legacy brands make no such claim, and one of them is owned by an advertising company.
What happened to Mamma.com?
The brand survived; the search engine did not. Mamma launched in Montreal in July 1996 and was a well-known metasearch engine for years. As of 19 August 2026 mamma.com is a live affiliate content site with no search box and no results, operated by a company called Trillion. Listing it as a working search engine, as many articles still do, is wrong.
Is SearXNG a search engine?
SearXNG is metasearch software, not a service. There is no company, no canonical website and no central operator — you either self-host a copy or use one of the public instances volunteers run. It aggregates results from other engines and holds no index. Two public instances can return materially different results and offer materially different privacy, because every setting is instance-level.
Why did most 1990s meta search engines disappear?
Because the problem they solved disappeared. Metasearch was valuable when every engine indexed a small, different slice of the web. Once one index became dramatically better than the rest, merging weaker sources stopped helping. They also owned nothing defensible: no crawler, no index, no ranking, and total dependence on suppliers who could cut them off or compete with them directly.
Sources
- homes.cs.washington.edu/~etzioni/papers/Overview.html
- ojs.aaai.org/aimagazine/index.php/aimagazine/article/view/1290
- jucs.org/jucs_2_9/profusion_intelligent_fusion_from/Gauch_S.html
- en.wikipedia.org/wiki/MetaCrawler
- en.wikipedia.org/wiki/Dogpile
- en.wikipedia.org/wiki/InfoSpace
- docs.searxng.org/
- github.com/searxng/searxng/issues/2515