Web evidence in litigation
Abstract layered wave illustration representing Impressions and Exposure Analysis

How far it carriesReproducible, assumptions statedIt works, and the assumptions it rests on have to be argued.

Impressions and Exposure Analysis

Short answer
The record shows delivery events, on each operator's own definition, not readers
Where it comes from
Search Console, server access logs, platform counters, third-party estimates
Who holds it
Google, the platform, the site's own host, and the estimate vendor
What process returns
Search Console is owner-only; no operator publishes a per-viewer log for production
What will not work
No available metric counts readers, and no standard converts one into a count
Applies to
Damages and the scope of the injury, not whether publication occurred

An impression, a view and a reader are three different things, and each operator defines its own

Three words that are not synonyms

This page is written for counsel who has been handed a number — impressions, views, estimated visits — and needs to know what it measures. The most useful thing an examiner does in this area is keep three concepts apart, because everyone in the room conflates them, including the client.

  • An impression is a delivery event. The platform placed a link or a piece of content somewhere a user could have seen it. It does not require that the content entered anyone's field of view, still less that anyone read it.
  • A view is a platform-defined event. The definition varies by operator and changes over time. It is not a common unit, and adding view counts across platforms sums quantities that are not the same quantity.
  • A reader is a person who read the statement and understood it. No platform metric measures this.

That third line is the whole subject. How publication and its extent must be proved in a given forum is a question for counsel; the forensic point does not change from matter to matter — extent goes to damages and to the scope of the injury, and no available metric measures the thing the element is actually about.

What Google says an impression is, and whose site it describes

Before the definition, the constraint that decides the shape of most matters: Search Console is visible only to a verified owner of the property it reports on. It describes search traffic to the plaintiff's site and says nothing about the defendant's page. The same asymmetry runs through the stack — access logs belong to whoever hosts the page in dispute, platform insights to the account holder — which makes exposure a discovery problem before it is an analysis problem, and one on a short clock, because none of these records is preserved by default.

Google does publish its definitions, which makes them quotable. From the Search Console help documentation, read 15 August 2026:

"An impression means that a user has seen (or potentially seen) a link to your site in Search, Discover, or News."

The parenthetical is Google's own hedge, in Google's own help text, and it is the sentence to reach for whenever an impression count is described as eyeballs. An impression is counted when the item appears in the current results page whether or not the user scrolls, except inside expandable elements such as carousels.

The documented limitations matter more than the definitions, and Google states them plainly in its data documentation: most reports "only cover a representative sample of URLs"; the Performance report "doesn't show all data" and "might not track some queries that are made a very small number of times or those that contain personal or sensitive information"; tables "can show a maximum of 1,000 rows."

Read against a defamation fact pattern the consequence is large. Branded queries in these matters are low volume and frequently contain a personal name, which makes them precisely the queries most likely to be withheld as rare or as personal. The absence of a query from Search Console is not evidence that nobody searched it, and the sum of a filtered table will systematically undercount the unfiltered total.

The retention window, and the figure that is not published

A 16-month retention figure for Search Console performance data circulates universally in the search industry. The Google pages read in this research on 15 August 2026 — the data documentation page and the Performance report page — state no retention period at all. So the honest sentence for a report is that Google does not publish a retention figure on the pages that document the data, and any number quoted came from somewhere else.

Nothing operational depends on the exact number. What is established is that the window is rolling, finite and measured in months rather than years, so if the statement is a year old and nobody has exported Search Console data, part of the exposure record has already aged out and no process recovers it.

The instruction that follows is dull and it is the one that changes outcomes: export early, export monthly, and store each export with its hash and a collection log. That is a chain-of-custody act as much as an analytics act.

What a view count counts, per operator

The operators that publish a definition publish a narrow one. X states, on its view counts page, read 15 August 2026:

"Anyone who is logged into X who views a post counts as a view, regardless of where they see the post… or whether or not they follow the author." "Multiple views may be counted if you view a post more than once, but not all views are unique." "If you're the author, looking at your own post also counts as a view."

Parsed, an X view count is a non-unique impression count of logged-in users, inclusive of the author's own views and exclusive of embedded posts — simultaneously the best publicly available exposure figure for an X post and a figure that cannot support a person count without an assumption the operator does not license.

YouTube's published position is different and equally useful: metrics are "algorithmically confirmed," and to verify accuracy YouTube "may temporarily slow down, freeze, or change your metric count," discarding low-quality playbacks. A YouTube view count is a moderated figure the operator reserves the right to revise retroactively. If one is going to be an exhibit, capture it repeatedly with timestamps and expect movement.

OperatorPublished definition of the counterWhat it does not tell you
Google SearchImpression: a link "seen (or potentially seen)"Whether anyone looked or read
XView: any logged-in viewer, non-unique, author includedHow many distinct people
YouTubeView: algorithmically confirmed, playbacks discardedA stable figure; it may be revised
Facebook, InstagramNone located, read 15 August 2026Anything, until read on the day
TikTok, LinkedInNone located, read 15 August 2026Anything, until read on the day

A remembered definition is worth nothing in a report: read the operator's page on the day the report is written, and cite it with that date.

First-party logs, and the referrer gap

Where the plaintiff controls a server, the raw access logs are the only exposure data in this method that is first-party, complete and under the party's own control. Requests carrying a Referer header that names the page in dispute show that at least some readers traveled from it — the closest available thing to a direct measurement of causal traffic.

Four things degrade it, and a report that does not name them is incomplete:

  • Referrer data is systematically incomplete. Modern defaults such as strict-origin-when-cross-origin strip the path and leave only the origin, search engines pass no query, and a "direct" bucket absorbs an unmeasurable share.
  • The client IP may not be the client. Apache's documentation warns that where a proxy sits between user and server, the logged address is the proxy's. Most high-traffic sites sit behind a CDN, and a log that does not record X-Forwarded-For behind one holds no client addresses at all.
  • Two fields are claims, not observations. Apache describes Referer and User-Agent as what the client reports, and the sender sets both.
  • Automated traffic is not marginal. Cloudflare's Radar year in review, published 15 December 2025, reported that AI crawlers generated roughly 20% of verified bot traffic, that other AI bots accounted for 4.2% of HTML request traffic and Googlebot for 4.5%. An unfiltered hit count overstates human reach by an unknown amount, and filtering on user agent alone is the weakest filter available.

Client-side analytics has a different defect. Google's documentation on data thresholds states that rows are withheld — withheld, not zeroed — to prevent anyone inferring the identity of individual users from small counts, and the low-traffic page is exactly the page thresholding suppresses. Logs and analytics will therefore disagree, and the gap is structural.

Third-party estimates and the one measured comparison

When nobody has the defendant's logs, the fallback is a traffic estimator. These tools model traffic to properties nobody in the room controls, so their error rate is not an academic question.

The best-documented comparison in this research is Omniconvert's study, first published 22 November 2017 and updated 5 June 2026: read-only analytics access to more than 4,000 ecommerce sites, with an estimator's figures cross-referenced for 1,787 of them. The estimator reported about 94% more sessions than the sites themselves tracked, the overstatement was systematic rather than random, and accuracy was worst for sites under 10,000 monthly sessions.

Three caveats belong with that figure in any report: it is a vendor's study rather than a peer-reviewed one, the sample is ecommerce only, and the update date does not identify which findings were re-measured. Cite it as one measured comparison against ground truth rather than as an established error rate — and note that this research found no first-party accuracy statement from any major estimator vendor, and no peer-reviewed validation of one.

The size problem is what decides matters. A defamatory page on a niche complaint site is a low-traffic page, which is the case the study reports the tool handles worst. Offering an estimate for it as a measure of exposure is offering the tool's least reliable output in the situation it handles least well. Say so first.

What an exposure analysis can defensibly say

In descending order of confidence, and the ordering is the deliverable:

  1. "This page was indexed and returned for these queries, at these positions, on these dates." Directly observable, reproducible from dated result-page captures.
  2. "Traffic to the plaintiff's own property from this referrer was N requests over this period." First-party, complete to the limits of referrer policy.
  3. "The platform reports N views, on its published definition, which is [quote]." Re-observable rather than reproducible, and not a person count.
  4. "Search Console reports N impressions for these queries over this period, subject to Google's published anonymization and row limits." First-party, and systematically under-inclusive for this query type.
  5. "A third-party tool estimates N visits, from a methodology the vendor does not fully disclose, which one 1,787-site comparison found overstated sessions by roughly 94%."

Anything below line five — a "reach" figure, a "potential audience," an impressions-times-click-rate-times-population model — is a model rather than a measurement. Models are sometimes worth building, and their assumptions belong to the analyst rather than to the data. That distinction has to be visible on the face of the report.

Where no standard exists

Three absences belong in this subject, stated as absences rather than papered over. First, there is no accepted methodology, no standards-body publication and no validated model for converting page-level traffic figures into a count of people who read a statement and understood it to refer to a particular plaintiff.

Second, this research located no controlling or widely cited US opinion setting a standard for proving online audience size in a defamation case; targeted case-law searching returned secondary commentary rather than holdings. That is an absence in the research, not a statement about what any court would do.

Third, the strictest audited definition of "someone saw this" in commercial use is weaker than the phrase suggests. The Media Rating Council's Viewable Ad Impression Measurement Guidelines, version 2.0 final, dated 18 August 2015, set the threshold at 50% of the advertisement's pixels in an in-focus browser tab for one continuous second, and two continuous seconds for video. The same document requires a separate "Viewable Status Undetermined" bucket and acknowledges unexplained inconsistencies between vendors applying it. That is the industry's own best case: a pixel test rather than a person test. Everything a consumer platform publishes is weaker.

Frequently Asked Questions

Can an impression count show how many people read a defamatory post?

No. Google's own help text defines an impression as a link a user has seen, or potentially seen, in Search, Discover or News, and that hedge is the operator's own wording rather than an interpretation of it. An impression is a delivery event: the platform put something where a user could have encountered it. It does not establish that anyone looked, scrolled, read or understood. Impression data is genuinely useful for showing that content was being served for particular queries over a particular period, which is a different proposition and a defensible one.

Does Google publish how long Search Console keeps performance data?

Not on the pages read in this research on 15 August 2026. A 16-month figure circulates universally in the search industry, and neither the data documentation page nor the Performance report page states a retention period. So the accurate sentence is that the operator does not publish a figure, and any number quoted came from somewhere else. The operationally important part survives the uncertainty: the window is rolling, finite and measured in months, so exports should start immediately and repeat monthly, each stored with a hash and a collection log.

Can view counts from different platforms be added together?

No, and doing it is one of the easier things to take apart. Each operator defines a view by its own published rule, and the rules are not the same quantity. X counts any logged-in viewer, non-unique, including the author's own views, and excludes embedded posts. YouTube algorithmically confirms views and reserves the right to slow, freeze or change the count and to discard low-quality playbacks. Several major operators publish no current definition at all. Summing those figures produces a number with no definition behind it.

Is a third-party traffic estimate good enough to show exposure?

It is the weakest line in the analysis and it should be labeled as such. The one substantial comparison against ground truth in this research covered 1,787 sites and found the estimator reported about 94% more sessions than the sites themselves tracked, systematically rather than randomly, with the worst accuracy on sites under 10,000 monthly sessions. A defamatory page on a niche complaint site is precisely that case. No first-party accuracy statement from any major estimator vendor was located, and no peer-reviewed validation of any of them.

What is the strongest exposure evidence available in most matters?

Ranking observation and first-party server logs. That a page was indexed and returned for named queries at named positions on named dates is directly observable and reproducible from dated captures. Requests to the plaintiff's own property carrying a referrer that names the page in dispute are first-party and complete to the limits of referrer policy. Both are also perishable: logs rotate in days to weeks and nothing preserves them by default, so a preservation step aimed specifically at access logs, CDN logs and analytics exports belongs in the first week.

Does the number of readers decide whether publication occurred?

As the research reports the rule, the publication element is satisfied by communication to one person other than the plaintiff, and audience size goes to damages and to the scope of the injury rather than to whether publication happened. Whether and how extent must be proved in a given forum is a question for counsel. The forensic point is the one that does not vary: the record almost never identifies a reader, so an analysis that promises a reader count is promising something no available data source measures.

Why is exposure analysis marked reproducible only with stated assumptions?

Because the observations reproduce and the conversions do not. Another analyst can re-run a ranking capture, re-read a platform counter, re-export Search Console data for a verified property and re-parse a log file, and reach the same table. The moment the table is converted into a statement about how many people saw or read something, an assumption enters that no operator publishes and no standard validates. Keeping the two halves visibly separate — observation, then stated assumption — is what makes the work usable rather than impressive.
Keep reading

The guides run the sequence

A page here covers one element, or one method. A guide covers the order the work happens in — what has to be collected before it changes, and which analysis is worth paying for at all.

Top