A people search site is not a search engine
The name is the first thing to correct. A web search engine crawls pages that publishers chose to put online, keeps a copy so it can answer queries quickly, and sends you back to the source. It answers a question with links to documents. Take the source page down and the entry drops out of the index at the next crawl, because the index is a pointer, not the thing itself.
A people search site does something categorically different. It assembles a durable record about a person — names, ages, current and former addresses, phone numbers, email addresses, likely relatives, likely associates, property and court references — from many separate sources, stores that record, and sells access to it. The dossier is the product. It does not update when the underlying source changes, because it was copied, not linked. The industry term for an operator like this is a data broker, and that is the accurate description regardless of what the site calls itself.
Three practical consequences follow from that distinction, and they explain most of the frustration people have with these sites:
- Removing a source does not remove the record. Deleting an old profile, closing an account, or getting a page taken down does not touch a copy that was made years ago and sold on.
- De-indexing is not deletion. Persuading a search engine to stop showing a page does not delete the broker's underlying record, which remains searchable on the broker's own site and licensable to its customers.
- Opting out of one site does nothing about the others. Brokers buy from each other and from the same upstream suppliers. A record removed from one operator's front end can be re-derived from the same source data by that operator's next refresh, and by everyone else's.
Where the data actually comes from
There is no single database of people. What these sites sell is a merge of several very different kinds of source, and the differences matter because they have wildly different accuracy and different legal status.
Public records. The largest and most legitimate layer. In the United States a great deal of government-held record is public by design: property deeds and tax assessment rolls, court dockets and filings, bankruptcy records, marriage and divorce records in many states, business registrations, professional licences, and — in some states, under some conditions — voter registration. These are genuinely public, and aggregating them is generally lawful. Aggregation is nonetheless the point: an address on a county deed record is public in the sense that a determined person could go and look it up; a searchable national index of every deed cross-referenced with every phone number is a different object with different consequences, and that is the change these companies actually sell.
Commercial and marketing data. Name-and-address files, phone listings, purchase histories, warranty and loyalty registrations, survey responses, subscription lists, and modelled attributes bought from marketing data suppliers. Much of this is self-reported, some of it is inferred by statistical model, and none of it was collected for the purpose of identifying a specific person accurately. It is the least reliable layer and it is invisible to the person it describes.
Scraped web and social content. Profile pages, usernames, photographs, employer and school listings, and anything else publicly reachable. Scraping is what people intuitively assume these sites are doing, and it is usually the smallest of the three layers rather than the largest.
Other brokers. Data flows between operators, so an error introduced once tends to propagate across the industry and to reappear after a deletion, because the record is being re-imported from a supplier who never saw the deletion request.
Why the results are so often wrong
Almost everyone who looks themselves up on one of these sites finds errors: addresses they never lived at, an age that is off by several years, relatives who are strangers, a middle initial that belongs to someone else. This is not sloppiness at the margins. It is a direct consequence of how the records are built.
There is no universal identifier. Records arrive from dozens of sources with no shared key. The broker must decide whether the Jennifer Miller on a 2011 utility record is the same Jennifer Miller on a 2019 property deed, using fuzzy matching on names, dates of birth, partial addresses and phone numbers. Every such decision is a probability. At national scale, with common names, a system tuned to merge aggressively will fuse several people into one profile, and a system tuned conservatively will split one person into several. Both failure modes are visible on every one of these sites.
Old data is never retired. Address history accumulates. There is no signal in a county record saying "this person moved out," so addresses stay attached indefinitely, and a profile can list a dozen addresses of which two are current and one belongs to somebody else entirely.
Relatives and associates are inferences, not facts. The "possible relatives" and "known associates" lists that make these profiles feel authoritative are usually derived from shared address history and surname similarity. That reliably produces former roommates, previous tenants of the same apartment, landlords, and people who happened to receive mail at the same address. It is presented in the interface with the same confidence as a court record.
Nobody in the transaction is the person described. The buyer cannot verify the profile, the subject is not consulted, and the operator's incentive is to produce a dossier that looks complete rather than one that is correct. A blank field sells nothing; a speculative one does. Accuracy has no natural feedback loop here, which is precisely why the law that does exist in this area is built around accuracy obligations.
There is also a business-model effect on top of the data problem. The common pattern is a free search that returns a teaser — a name, a rough location, and a count of records "found" — followed by a paywall or a subscription, often one that renews automatically. Teasers are written to imply that something significant is waiting behind the payment. The gap between the implication and the contents is where a lot of consumer complaint in this sector has come from.
What US law says, in general terms
The United States has no comprehensive federal privacy statute. What exists instead is a patchwork of sector-specific federal law plus a growing and inconsistent set of state laws, and people-search operators sit deliberately in the gaps between them.
The most important federal statute in this area is the Fair Credit Reporting Act. The FCRA regulates consumer reports — information assembled and used for decisions about credit, insurance, employment, or tenancy — and imposes real obligations on the companies that supply them: reasonable procedures to ensure accuracy, a duty to investigate disputes, notice to the subject when a report is used against them. This is why essentially every people-search site carries a prominent disclaimer stating that it is not a consumer reporting agency and that its data must not be used to make employment, tenant, credit or insurance decisions. That disclaimer is the entire legal architecture of the business. Complying with the FCRA is expensive; disclaiming its scope is free. The Federal Trade Commission has brought enforcement actions in this sector against operators that marketed their reports for FCRA-covered purposes while claiming to be outside the Act, and the question of whether inaccurate information alone is a concrete enough injury to sue over reached the US Supreme Court in Spokeo, Inc. v. Robins in 2016.
Other federal statutes carve out specific record types rather than the industry as a whole. The Driver's Privacy Protection Act restricts the disclosure and resale of motor-vehicle record data. Financial privacy rules constrain what may be derived from banking relationships. Neither touches the bulk of what these sites hold.
At state level three things have been happening, and if you want to understand your position this is where to look:
- Registration. Some states now require data brokers to register with the state and disclose that they exist. Vermont enacted the first such law and California followed with its own registry. The practical value is that a public list of registered brokers is the closest thing to an inventory of an industry that otherwise has no membership roll.
- Comprehensive privacy statutes. California's law came first and a substantial and growing number of states have since passed their own. The rights they grant are broadly similar in shape — to know what is held about you, to have it corrected or deleted, and to opt out of its sale or sharing — but they vary in who they cover, what thresholds apply, and whether an individual can enforce them personally or must rely on a state regulator.
- Centralised deletion. California has legislated for a single state-run mechanism through which a resident can make one deletion request that registered brokers must check and honour, rather than filing hundreds of individual requests. This is the most consequential structural idea in the area, and it is being phased in; anyone relying on it should check the current position with the state agency responsible, because the timetable and the obligations have moved.
Separately, several states have enacted protections for specific occupational groups whose home addresses create physical danger — New Jersey's Daniel's Law, passed after the murder of a federal judge's son at the family home, allows judges, prosecutors and law-enforcement officers to require removal of home addresses and unpublished phone numbers, and other states have adopted comparable schemes.
Outside the US the framing is completely different. Under the EU and UK data protection regimes, compiling and selling personal data requires a lawful basis and individuals hold a right to erasure by default, which is why the largest US people-search operations generally do not offer their consumer products in those markets.
What opting out does and does not achieve
Most operators publish an opt-out or suppression process, some because state law requires it and some voluntarily. The mechanics vary but the pattern is consistent: you locate the specific profile, submit a removal request through the operator's own form, and confirm it — often by email, sometimes by a code sent to a phone number, occasionally by supplying identity documents. Being asked to hand over more personal data in order to remove personal data is a genuine and unresolved tension in this process.
What that achieves is narrower than people expect, and the honest summary is worth stating plainly:
- It suppresses a display, not a source. The underlying public record — the deed, the court filing, the voter roll entry — is unaffected. It remains public and remains collectable.
- It is per-operator. There are many operators, several of them running multiple consumer brands over the same database, so a removal that appears to have worked may only have covered one front end.
- It decays. Records commonly reappear after a refresh cycle, because the operator re-imports from a supplier that has no record of the suppression. Treat opt-out as maintenance rather than a one-off.
- It does not affect search engines directly. If a broker page has been indexed, the index entry generally clears only after the page itself is gone or has been de-indexed, which is a separate request to a separate company. Major engines do operate policies for removing certain categories of personal information, such as contact details published to harass someone, but those policies are narrow and are not a general right to be forgotten in the US.
The broader point for anyone trying to understand this category: these sites are not a specialised corner of web search and should not be evaluated as if they were. They are a distinct industry that borrowed the search interface because a search box is a familiar and reassuring way to present a dossier. Judging them by the standards you would apply to a search engine — coverage, freshness, ranking quality — misses what they are. The relevant standards are the ones you would apply to any company that holds a file on you: who are they, what is in it, how did they get it, is it correct, and what can you make them do about it.
Frequently asked questions
Are people search sites search engines?
No. A search engine crawls published web pages and returns links to them; the index is a pointer to a source that still exists. A people search site compiles and stores a persistent record about an individual from public records, commercial data and scraped content, then sells access to it. The record is the product, and it survives changes to the sources it came from.
Where do people search engines get their information?
Three main layers. Public records — property deeds, tax rolls, court filings, business registrations, licences, and in some states voter registration. Commercial marketing data — address files, phone listings, purchase and survey data, and modelled attributes. Scraped public web and social profiles. Operators also buy data from one another, which is why an error can reappear after being removed.
Why is the information on people search sites wrong?
Because records from many sources are merged with no shared identifier, using probabilistic matching on names, dates of birth and addresses. Aggressive matching fuses different people with similar names; cautious matching splits one person across several profiles. Old addresses are never retired, and "relatives" and "associates" are usually inferences from shared address history, which produces former roommates and previous tenants.
What is a data broker?
A data broker is a company that collects personal information about people who are not its customers, compiles it into records, and sells or licenses access to those records. People search sites are the consumer-facing end of that industry. The defining feature is that the subject of the data has no relationship with the company and usually does not know it holds a file on them.
Can employers legally use people search sites for background checks?
Almost every such site explicitly forbids it in its terms, because using data for employment, tenancy, credit or insurance decisions brings the report under the US Fair Credit Reporting Act and its accuracy, dispute and notice obligations. Operators disclaim consumer reporting agency status specifically to stay outside that regime. The FTC has taken enforcement action where operators marketed reports for those purposes anyway.
How do I get removed from people search sites?
Each operator runs its own opt-out process, typically requiring you to find the specific profile and confirm a removal request. It is per-operator, not industry-wide, and it suppresses a display rather than deleting the underlying public record. Records frequently reappear after a data refresh, so removal is ongoing maintenance rather than a permanent fix.
Does removing a result from Google delete the underlying record?
No. De-indexing removes a page from a search engine's results; it does not touch the data broker's record, which remains on the broker's own site and available to its customers. The two are separate requests to separate companies. Major engines do have narrow policies for removing certain personal information, but those are not a general deletion right in the United States.