SearchEngines.Net logo — an independent reference on search enginesSearchEngines.NetWho runs which index

The shift to AI search

AI assistant accuracy compared, as measured in October 2025

A point-in-time result. The assistants were tested on their free consumer tiers in October 2025, and the models behind them have changed since.

What was measured, and what it found

The largest published comparison of AI assistant accuracy found that 45 percent of AI assistant responses to news questions contained at least one significant issue, and that 81 percent had some issue. The single largest category of problem was sourcing, present in 31 percent of responses.

The study is News Integrity in AI Assistants, published in October 2025 by the European Broadcasting Union and the BBC. It matters more than the many smaller accuracy tests in circulation for one reason: it put the same questions to several assistants under the same conditions and had the answers graded by the same protocol, which makes it a genuine comparison rather than a collection of separate impressions.

It is also a snapshot, and the snapshot has a date on it. The assistants were tested on their free consumer tiers in October 2025. Every major assistant in the test has since shipped new model versions. Nothing on this page should be read as a claim about how these products behave now; it is a record of what was measured, when, and by whom.

Some vocabulary is needed before the numbers are readable. A significant issue, in the study's protocol, is an error serious enough to materially mislead a reader about the subject of the question — a wrong fact, a wrong attribution, a claim presented as current that is out of date. The wider some issue category includes lesser problems that a journalist would nonetheless correct. A sourcing issue means a problem in how the answer attributed its material: a missing attribution, a source that does not support the claim it is attached to, or a reference that cannot be located.

Two things follow from the headline figures that are easy to miss. The first is that the 81 percent and the 45 percent describe the same responses at different thresholds, so the gap between them is the share of answers with problems that a careful reader might tolerate. The second is that sourcing being the largest category is a finding about the mechanism by which these systems fail, not a detail. The answers were more often wrong about where information came from than about anything else, which is precisely the aspect a reader has been trained to treat as the reassuring part of the interface.

The per-assistant breakdown, as of October 2025

The comparison by assistant exists only in the report itself. Press coverage reported the 45 percent headline and generally stopped there, which left the most useful part of the study unread.

Share of responses containing at least one significant issue, as measured in October 2025, on free consumer tiers:

  • Google Gemini — 76 percent. The worst performer in the test by a wide margin, at roughly two and a half times the rate of the best.
  • Microsoft Copilot — 37 percent.
  • ChatGPT — 36 percent.
  • Perplexity — 30 percent. The best performer in the test.

The spread is the interesting quantity. Three of the four assistants cluster between 30 and 37 percent, a range narrow enough that the ordering among them should not be over-read; differences of a few points in a graded sample of this kind are not a durable ranking. Gemini's 76 percent sits outside that cluster entirely, and a gap of that size is not explained by grading noise.

The label on this section is not a formality. This ranking is a point-in-time result and has no standing as a current description of these products. Three things could each move it on their own: the underlying models have been revised since October 2025, the free consumer tiers tested are not the paid tiers most heavily marketed, and the retrieval systems that supply these assistants with source material are updated continuously and independently of the models. An assistant that placed last in one month's test has no guaranteed relationship to the same assistant a year later.

It is also worth being precise about what was ranked. These are error rates on news questions — current events, attribution, who said what and when — which is the category where accuracy is easiest to verify against a known record and where being out of date is itself a failure. It is not a general measure of capability, and nothing in the study supports extending the ordering to coding, summarization, translation or any other task.

How the study was built

The design is the reason this study carries more weight than the accuracy comparisons that circulate on a monthly basis, and it is worth setting out in full.

2,709 core questions were put to the assistants. The responses were evaluated by professional journalists from 22 public service media organizations across 18 countries and 14 languages. The evaluators were assessing answers about news in their own markets, in their own languages, against records they had independent means of checking.

Three properties of that design do real work. Because the questions were the same across assistants, the comparison is internally valid: a difference between two assistants is a difference in their answers rather than a difference in what they were asked. Because the graders were journalists rather than crowdworkers or an automated scoring model, the judgment about whether an attribution is wrong was made by people whose profession is checking attributions. And because the sample spans 14 languages, it escapes the English-only limitation that affects nearly every other published audit of these systems — a real constraint, since retrieval quality and the availability of trustworthy source material differ enormously between languages.

The limits are equally concrete. The test covered free consumer tiers only, so it does not describe the paid products or the enterprise configurations, which may use different models and different retrieval settings. It was conducted in October 2025 against model versions that have since been superseded. And it measured news questions specifically, a domain chosen because it is verifiable, not because it is representative of everything people ask an assistant.

There is one further property of this kind of research that applies to all of it and is rarely acknowledged. An assistant's answer to the same question is not stable: these systems are generative, so repeated asking produces varied wording and sometimes varied substance, and a graded sample is therefore an estimate of a distribution rather than a reading of a fixed value. That is an argument for reading the large gap between Gemini and the rest as meaningful and the small gaps within the cluster as provisional, which is the same conclusion the spread itself suggests.

The volume these error rates sit on

An error rate is only as consequential as the number of answers it applies to, and until recently there was no credible public figure for how many questions these assistants handle. There is now, and it comes from a court.

The memorandum opinion on remedies in United States of America v. Google LLC, No. 1:20-cv-03010-APM, ECF No. 1436, issued by the US District Court for the District of Columbia in September 2025, records daily query volumes for AI products as of March 2025:

  • ChatGPT — roughly 1.2 billion queries per day.
  • Meta AI — above 200 million per day.
  • Google Gemini app — approximately 140 million per day.
  • Grok — 75 million per day.
  • DeepSeek — 50 million per day.
  • Perplexity — 30 million per day.

These numbers exist in public only because they were produced under seal by OpenAI, Google and others in the course of litigation and entered the record through the remedies opinion. No commercial data provider publishes comparable per-product figures, and the estimates that circulate from vendors and analysts are modeled rather than observed.

Three cautions belong with them. The figures describe March 2025 and these volumes move faster than almost any other number in this field. The version of the opinion most readily available is hosted on a news organization's mirror of the court's filing rather than on a court server. And the remedies judgment is under appeal.

Placing the two sets of figures beside each other requires care, because they do not line up cleanly. The query volumes are for whole products across all uses; the error rates are for news questions on free consumer tiers in a different month. Multiplying one by the other would produce a number with no meaning. What the volumes establish is scale rather than arithmetic: the products being graded are being asked questions in quantities comparable to a mid-sized search engine, and at that volume a measured error rate is a description of a large number of individual answers rather than a laboratory curiosity.

Sourcing was the largest single failure

Sourcing problems appeared in 31 percent of the responses graded by the European Broadcasting Union and BBC study, making them the largest category of issue found. A separate body of research explains part of why.

An audit published in 2026 by researchers at UC Berkeley and Cornell University, LLM Hallucinations in the Wild: Large-Scale Evidence from Non-Existent Citations, examined 111 million references across 2.5 million papers on arXiv, bioRxiv, SSRN and PubMed Central. It conservatively estimated 146,932 hallucinated, non-existent citations in 2025 alone — references to works that do not exist, in documents that were actually published. The same audit found that fabricated citations disproportionately credit prominent and male scholars, which is a second-order result worth recording because it suggests the fabrication follows the statistical shape of real citation practice rather than being random.

Two limitations have to travel with that figure and are the reason it appears here as context rather than as a finding about search. It is an arXiv preprint and has not completed peer review. And it measures citation hallucination in academic writing, not in consumer AI search answers. Those settings are related but not the same: an academic manuscript passes through drafting, co-authors and review, while an assistant's answer is generated and displayed in a second. The direction of that difference is not obvious — review should catch fabrications, but nobody checks an assistant's citations at all — so the academic number is not a proxy for a consumer one in either direction.

What the two studies together establish is narrower than the headlines they produced, and more useful. A reference appearing beneath an AI-generated answer can fail in at least three distinct ways: it can be a real source that does not support the claim attached to it, a real source that has been misread, or a work that does not exist. The first is the most common in ordinary use and the hardest to notice, because the link resolves, the page loads, and only reading it reveals the mismatch. The third is the easiest to detect and the most alarming when found.

This is a description of what has been measured about the failure mode. What a reader does about it is outside the scope of this page.

Citations raise trust without raising accuracy

The design assumption behind every AI answer product is that showing sources makes an answer trustworthy. That assumption has been tested directly, and the result is not what the design implies.

A preregistered randomized experiment by researchers at MIT Sloan and the MIT Initiative on the Digital Economy, circulated in April 2025 as an arXiv preprint, spanned approximately 12,000 search queries and 80,000 results across seven countries. It found that adding reference links and citations significantly increased human trust in generative AI search results even when those links and citations were incorrect or hallucinated.

Preregistered means the researchers published their hypotheses and analysis plan before collecting data, which rules out the possibility that the finding was selected after the fact from among many tested. That matters here, because the result is the kind that would be easy to discover by accident and hard to believe on one showing.

The implication is precise and should not be inflated. A citation displayed beside an AI-generated answer is a trust signal — it changes how confident a reader feels — and the experiment found no corresponding improvement in the reader's ability to distinguish a correct answer from an incorrect one. The presence of a reference and the correctness of the claim are, behaviorally, separate things. A wrong answer with three sources under it reads as more reliable than a right answer with none.

Read alongside the European Broadcasting Union and BBC finding that sourcing was the largest single category of failure at 31 percent, the two results describe a closed loop. The element of the interface that most raises reader confidence is the element the systems get wrong most often, and the confidence it produces does not depend on whether it is correct.

The usual caveats apply and are not small. This is an April 2025 preprint that has not completed peer review, and it is older than most of the products it describes, though the finding concerns human behavior rather than model capability and human behavior of this kind does not usually reverse in a year. It measures perceived trust and discrimination between correct and incorrect answers; it does not measure whether readers subsequently acted on what they read.

What these measurements support, and what they do not

Set out together, the four studies support a short list of claims and no more than that.

  • Supported: AI assistants make significant errors on news questions at a substantial rate. 45 percent of responses carried at least one significant issue in the October 2025 test, and 81 percent carried some issue.
  • Supported: the assistants differed from one another at that moment. Gemini at 76 percent sat well outside a cluster of Copilot at 37, ChatGPT at 36 and Perplexity at 30.
  • Supported: sourcing is the most common failure. 31 percent of responses had a sourcing problem, and a separate audit of academic writing found non-existent references being published in the tens of thousands per year.
  • Supported: citations increase reader trust independently of whether they are correct. A preregistered randomized experiment across roughly 12,000 queries found exactly that.
  • Not supported: any statement about which assistant is most accurate today. The comparison describes free consumer tiers in October 2025 and the models have changed.
  • Not supported: any extension beyond news questions. News was chosen because it is verifiable. Nothing here measures coding, summarization, translation or reasoning.
  • Not supported: a count of wrong answers in the wild. The court-verified query volumes and the graded error rates come from different months, different populations of question and different scopes, and multiplying them produces a number that means nothing.

The gap between the first list and the second is where most writing on this subject goes wrong. A graded sample on one topic, in one month, on one product tier, is a real and valuable measurement, and it is not a league table with ongoing validity. Both halves of that sentence are load-bearing.

When each figure on this page was measured

Every figure here is dated, because a figure about AI capability without a date is not usable.

  • Accuracy rates and the per-assistant breakdown — tested October 2025, free consumer tiers, published by the European Broadcasting Union and the BBC in News Integrity in AI Assistants, 2025.
  • Daily query volumes — as of March 2025, recorded in the remedies memorandum opinion in United States of America v. Google LLC, issued September 2025. The judgment is under appeal, and the readily available copy is a news organization's mirror of the court filing.
  • Hallucinated citations in academic writing — 2025 publication year, audited in a 2026 arXiv preprint from UC Berkeley and Cornell University. Preprint status, and the setting is academic writing rather than consumer AI answers.
  • The trust experiment — April 2025 arXiv preprint from MIT Sloan and the MIT Initiative on the Digital Economy. Preprint status.

This page was checked against the source documents in September 2026. None of the four measurements has a published successor covering the same ground on the current generation of models, which is itself worth knowing: the reason a two-year-old audit is still the best comparison available is that nobody has repeated it at that scale.

Frequently asked questions

Which AI assistant is the most accurate?

No current answer to that exists. In the October 2025 test by the European Broadcasting Union and the BBC, measured on free consumer tiers, Perplexity had the lowest rate of significant issues at 30 percent, followed by ChatGPT at 36 percent, Microsoft Copilot at 37 percent and Google Gemini at 76 percent. That ordering describes news questions on those tiers in that month. All four products have shipped new model versions since, so the result is a point-in-time snapshot rather than a standing ranking.

How often do AI assistants get things wrong?

In the largest published comparison, 45 percent of AI assistant responses to news questions contained at least one significant issue and 81 percent contained some issue. That study, News Integrity in AI Assistants from the European Broadcasting Union and the BBC, graded 2,709 core questions using professional journalists from 22 public service media organizations across 18 countries and 14 languages, in October 2025, on free consumer tiers. The rate applies to news questions specifically, which is a domain chosen because it can be verified against a known record.

What kind of mistakes do AI assistants make most often?

Sourcing errors, present in 31 percent of graded responses and the largest single category in the European Broadcasting Union and BBC study. A sourcing problem means the answer mishandled attribution: a missing source, a source that does not support the claim attached to it, or a reference that cannot be located at all. That is the failure mode a reader is least equipped to spot, because checking it means opening the link and reading the page rather than glancing at the fact that a link exists.

Can AI assistants invent sources that do not exist?

The largest real-world measurement of that is an audit of academic writing rather than of search answers. Researchers at UC Berkeley and Cornell University examined 111 million references across 2.5 million papers on arXiv, bioRxiv, SSRN and PubMed Central and conservatively estimated 146,932 hallucinated, non-existent citations in 2025 alone. It is an arXiv preprint, and it measures published academic manuscripts, not consumer AI answers, so it establishes that the failure mode is real and common at scale without giving a rate for what an assistant shows a reader.

Do citations under an AI answer mean it is accurate?

They do not, and there is a direct experiment on the point. A preregistered randomized study from MIT Sloan and the MIT Initiative on the Digital Economy, spanning roughly 12,000 search queries and 80,000 results across seven countries, found that adding reference links and citations significantly increased users' trust in generative AI search results even when those links and citations were incorrect or hallucinated. Citations function as a trust signal rather than an accuracy signal. The study is an April 2025 arXiv preprint.

How many questions do AI assistants handle?

As of March 2025, according to figures produced under seal and entered into the public record through the remedies memorandum opinion in United States of America v. Google LLC: ChatGPT roughly 1.2 billion queries per day, Meta AI above 200 million, the Google Gemini app approximately 140 million, Grok 75 million, DeepSeek 50 million and Perplexity 30 million. No commercial data provider publishes comparable figures. The volumes move quickly, the data is from March 2025, and the remedies judgment is under appeal.

Is the October 2025 accuracy ranking still valid?

It should not be treated as current. The test covered free consumer tiers in October 2025, and the models behind all four assistants have been revised since, as have the retrieval systems that supply them with source material. The ranking is a record of what was measured at a particular moment. No successor study covering the same ground at the same scale has been published, which is why the October 2025 figures remain the best available comparison despite their age.

Why is this study more useful than other AI accuracy tests?

Because it is a controlled comparison rather than a set of separate impressions. The same 2,709 core questions went to each assistant, the answers were graded by the same protocol, and the graders were professional journalists checking news in their own languages and markets across 18 countries. Most published accuracy tests use one language, a small question set, and an automated or informal grading method, which makes their results incomparable with one another. Its limits are that it covers news questions, free consumer tiers, and one month.

Sources

Top