SearchEngines.Net logo — an independent reference on search enginesSearchEngines.NetWho runs which index

How search engines work

How big are search engine indexes?

Almost every index-size figure you have read is a company describing itself, with no method attached and no way to check it.

The honest answer: nobody outside the company knows

This is a question with a large amount of confident-looking information behind it and almost no verifiable information behind it. Search engines are not audited on index size. No regulator requires the disclosure, no accounting standard defines the unit, and no external party has a mechanism to check a figure once it is published. Every number in circulation is a company describing itself.

That does not make the numbers worthless — a dated, self-published figure from an operator with a track record is real evidence about that operator. It makes them a different kind of evidence than they look like. "Nine billion pages" reads like a measurement. It is a statement.

What follows is what the major and independent engines have actually claimed, what those claims count, why the units are not comparable, and why the one launch that led with index size as its headline is now the standard cautionary tale in the field.

Three different quantities, all reported as "index size"

The single largest source of confusion is that "how big is the index" conflates at least three measurements that differ by orders of magnitude.

  • URLs seen. Every distinct address a crawler has ever encountered in a link, a sitemap or a redirect. This is the biggest number and the least meaningful, because it includes addresses never fetched, addresses that returned errors, and infinite auto-generated URL spaces.
  • Documents crawled and stored. Pages actually fetched and retained. Smaller, and closer to what people mean.
  • Documents retrievable by a query. What a searcher can actually reach. Smaller again, because engines apply quality filters, deduplicate, and cap how deep results go.

Google's July 2008 post is the clearest illustration on record. It reported 1 trillion unique URLs — and then immediately explained that this was a URL count, not a page count: "many pages have multiple URLs with exactly the same content or URLs that are auto-generated copies of each other." The same post gave the historical series: 26 million pages in the first index in 1998, and one billion by 2000. Those early figures are page counts; the trillion is not. They are not points on the same curve, and they get plotted on the same curve constantly.

Google has since given occasional round figures rather than a series. Wikipedia's article on Google Search records a 2012 statement of over 30 trillion web pages alongside 100 billion queries a month — again with no published method, and again a discovery figure rather than a retrievable-document figure. The same article notes the common assumption that Google indexes only a small fraction of what exists, the remainder sitting behind logins, paywalls and forms in what is loosely called the deep web. Every one of these is a company statement with a date on it, and that is all it is.

Cuil, and what happens when you lead with the number

The definitive case is Cuil, launched on 28 July 2008 by former Google search engineers and covered at the time as a "Google killer". Its headline claim was an index of roughly 120 billion web pages — the precise figure circulated was 121,617,892,992 — raised to 127 billion by February 2009.

Cuil launched three days after Google's trillion-URLs post, into a press cycle that read the two numbers as directly comparable. They were not. Cuil was claiming indexed pages; Google had reported URLs seen. Neither company published a methodology that would make the comparison mean anything, and the widely repeated framing that Cuil was "three times bigger than Google" came from journalists comparing against older Google disclosures rather than from any like-for-like measurement.

What sank Cuil was not that the index did not exist. It plainly did: Cuil built and ran its own crawler, Twiceler, and its own inverted index, and never resold anyone else's results — a genuinely rare achievement. What sank it was that leading with index size guaranteed that every reviewer's first act would be to run test queries, and the queries returned thin, stale or missing results. Hands-on reviews at Technologizer and Fortune reached that conclusion within forty-eight hours. Traffic-panel figures cited by Wikipedia show a collapse from roughly 0.2% of worldwide internet users in launch week to 0.005% by 13 October 2008. The service shut down on 17 September 2010.

The lesson generalised: an index-size number without a retrieval-quality number beside it is not a claim about a product. Cuil supplied only the former, and it was tested on the latter.

The engines that do publish figures, with dates

Several independent engines publish index milestones, and because they are small and specific they are the most useful numbers available.

Mojeek, the British engine that has crawled and ranked its own index since the mid-2000s, publishes a milestone series: one billion pages in 2015; 2.3 billion in May 2019; 4 billion in June 2021; 6 billion in October 2022; 8 billion in 2024; and 9 billion per its own About page in 2025. No 2026 figure has been published, and extrapolating one would be inventing it.

Marginalia Search, run by a single engineer, reports roughly 300 million documents as of 2024 occupying about a terabyte — with the notable honesty that the FAQ page carrying the figure flags itself as outdated, and the current site says only "hundreds of millions of documents". Marginalia treats the gap with Google as a design position rather than a shortfall: "an index with a million documents that are all of high quality is better than an index with a billion documents where only a fraction of them are interesting."

Qwant has self-reported an index of around 20 billion pages, dated September 2023.

Gigablast shows how the numbers fall apart under scrutiny even for a single engine in a single year. Its architecture was described as designed to scale to 200 billion pages — a ceiling, not a count. For 2015, Gigablast's own press release of 1 July said "its own searchable index of over a billion pages", while Wikipedia records "over 12 billion" for the same year. Those figures differ by an order of magnitude and both are attributed to 2015. This site prints the conflict rather than picking a side, because there is no basis for picking one.

Why an index cannot be independently audited

The reason nobody can check these numbers is not secrecy alone. It is that the only interface an outsider has to an index is the query box, and a query box cannot enumerate a set.

  • Result counts are estimates. The number an engine prints above the results is a projection from sampled statistics, not a count. It changes between refreshes of the same query and is not offered as precise.
  • Result depth is capped. Engines stop serving results after a few hundred entries regardless of the estimate, so you cannot page to the end of anything and count.
  • Sampling methods rest on assumptions. The academic approach — issue random queries, observe overlap between engines, infer relative sizes — dates back to work by Bharat and Broder in the late 1990s. It estimates relative coverage under assumptions about query and document distributions that were shaky then and are far shakier now that engines rewrite queries, personalise, and answer without listing.
  • The unit is undefined. Is a PDF one document or several? Is a paginated article one or twenty? Do stored-but-filtered pages count? Every engine answers differently and none publishes its answer.
  • The target moves hourly. An index is continuously crawled, refreshed and pruned. Any figure is a snapshot of a quantity that changed while it was being measured.

Add to that the absence of any incentive to be audited. There is no commercial upside to letting a third party verify that your index is smaller than you said, and no legal instrument compelling it. So the field runs on self-report, and the correct posture toward every figure on this page is to cite it with its date and its source and never to average two of them together.

Size is not the question that decides anything

Index size is a poor proxy for the thing people actually want to know, which is whether the engine can find what they are looking for. Three factors matter more.

Retrieval quality. Cuil's 120 billion pages did not produce good answers. Marginalia's few hundred million produce excellent ones within their domain — non-commercial, text-heavy, independently published pages — and nothing at all outside it. Neither outcome is predicted by the page count.

Coverage of what you specifically need. Access to particular corpora is now negotiated rather than crawled. In 2024, Reddit blocked non-Google crawlers, cutting Mojeek and other independent engines off from Reddit content while Google retained access under a paid arrangement. That is a coverage gap no amount of crawling can close, and it does not show up in an index-size figure at all.

Freshness. A large index refreshed slowly is worse for news than a small index refreshed quickly. Marginalia's refresh cycle runs to weeks, which is fine for its purpose and useless for anything current.

There is one figure in this whole subject that is genuinely load-bearing, and it is a cost rather than a size: Marginalia's operator has reported running the project on the order of a couple of hundred dollars a month. Independent web indexing is not impossible. It is not even especially expensive at the scale of hundreds of millions of documents. What is expensive is the last few orders of magnitude — and that, not any published page count, is the real shape of the search market.

Primary sources worth reading in the original: Google's 2008 post on the trillion-URL count and Mojeek's own index milestone announcements.

Frequently asked questions

How many pages does Google have in its index?

Google has never published an audited figure. It reported 26 million pages in 1998, one billion in 2000, and one trillion unique URLs seen in July 2008 — a URL count, not a page count. Later statements have described tens of trillions of pages and hundreds of billions in the index. All are company statements with dates, published without a stated method.

What was Cuil's 120 billion page claim?

Cuil launched on 28 July 2008 claiming an index of roughly 120 billion pages, precisely 121,617,892,992, raised to 127 billion by February 2009. The index was real and independently crawled, but the claim invited reviewers to test retrieval quality, which failed. Traffic collapsed within months and the service shut down on 17 September 2010.

What is the difference between URLs seen and pages indexed?

URLs seen counts every distinct address a crawler has encountered, including ones never fetched, ones returning errors, and auto-generated infinite URL spaces. Pages indexed counts documents actually stored and retrievable. Google made the distinction explicitly in 2008, noting many pages have multiple URLs with identical content. The two quantities differ by orders of magnitude and are routinely compared as if identical.

How big is Mojeek's index?

Mojeek publishes dated milestones: one billion pages in 2015, 2.3 billion in May 2019, four billion in June 2021, six billion in October 2022, eight billion in 2024, and nine billion per its About page in 2025. No 2026 figure has been published. Mojeek crawls the web itself and has never resold another engine's results.

Can anyone independently verify a search index size?

No. The only public interface to an index is the query box, result counts are estimates rather than counts, and engines cap how deep results go, so an index cannot be enumerated. Academic sampling methods estimate relative coverage under assumptions that modern query rewriting and personalisation undermine. There is also no unit definition and no obligation to disclose.

Does a bigger index mean better search results?

Not reliably. Cuil's claimed 120 billion pages produced poor answers; Marginalia's few hundred million produce good ones within a narrow domain. Retrieval quality, freshness and coverage of the specific corpora you need matter more. Reddit's 2024 block on non-Google crawlers created a coverage gap for independent engines that no index-size figure reflects.

Why do published index sizes conflict with each other?

Because the unit is undefined and the sources are self-reported. Gigablast is the clearest example: its own July 2015 press release said "over a billion pages" while Wikipedia records "over 12 billion" for the same year, and its 200 billion figure was an architectural ceiling rather than a count. Conflicting figures should be printed as conflicts, not averaged.

Sources

Top