What reach evidence is for, and what it is not
This page is written for counsel litigating an internet defamation matter, on either side, and it addresses one question: what the technical record can show about how far a statement traveled. Whether any element is satisfied is a question for counsel. What follows is what the records show.
Reach evidence and the fact of a communication are different subjects, and separating them is the first thing I do on an engagement of this kind. Access logs, platform counters, analytics accounts and modeled estimates describe how widely something was delivered. None describes whether a particular person received the statement and understood it.
What I can produce is bounded: every record that bears on reach, what its own operator says it counts, what it demonstrably does not count, and how much of it still exists. What I cannot produce is a number of readers. No record in this subject holds one.
The standard that does not exist
The research behind this page went looking for the case law on proving online audience size and did not find any. Targeted searches for US opinions treating hits, page views or impressions as proof of the extent of an online publication returned secondary commentary and no opinion setting a standard. I state that as a finding rather than a gap, because it changes how the subject has to be approached.
One caution travels with it: a search returning nothing describes what a careful search found, and is not proof that nothing exists. If you have an opinion on point I would rather read it than defend the negative.
No benchmark declares one method of measuring audience sufficient and another inadequate, and that cuts both ways — an opposing examiner has no benchmark either. Absent a standard, the defensible position is to describe each figure by its operator's published definition and refuse to convert it into a quantity the record does not contain.
What one line in a server access log establishes
The most probative artifact of reach is the web server's access log, held almost always by the defendant or a third-party host rather than by the plaintiff — a discovery target, not something an examiner turns up alone. The de facto formats are Apache's Common and Combined Log Formats, documented in Log Files, Apache HTTP Server 2.4 (read 15 August 2026). Combined is %h %l %u %t "%r" %>s %b "%{Referer}i" "%{User-agent}i", and the fields are where cross-examination happens.
%h, the client address. Apache warns that this "is not necessarily the address of the machine at which the user is sitting" and that where a proxy intervenes it is the proxy's address. Most high-traffic sites sit behind a content delivery network, so the origin log holds the network's address unlessX-Forwarded-Forlogging was configured — a log without it has no client addresses at all.%t, the arrival time, comes from the server's clock in its configured zone, and whether that clock was disciplined to a time source is a question for the custodian.%r,%>sand%b— request line, status code, response size — are server-generated and the most reliable fields in the line.%land%uare a hyphen on a public page.RefererandUser-Agent, which Apache describes as what the client reports — trivially settable and routinely spoofed. A user-agent string is a claim, not an observation.
One line supports a finding that a request with these characteristics was received and a response of this status and size was sent — not that a person was present, that the page rendered, or that two entries from one address are one person. A status 200 with a full byte count is equally consistent with a crawler, a prefetch and someone who closed the tab in half a second. Automated requests are in there too, and not marginally: on Cloudflare's Radar figures for 2025 (published 15 December 2025, read 15 August 2026), "other AI bots accounted for 4.2% of HTML request traffic" and "Googlebot alone accounted for 4.5%." A raw line count is an upper bound on human reach and nothing more. I do not offer a global bot-share percentage; the numbers in circulation trace to vendor reports rather than to a measurement anyone can check.
There is no standard retention period for these logs, no regulation setting one for an ordinary publisher, and nothing that preserves them by default.
An impression, a view and a reader are three quantities
Three words are used as synonyms and are not. An impression is a delivery event: the operator placed content where a user could have seen it, with no requirement that it entered a field of view. A view is an operator-defined event whose definition differs by platform and changes over time, which makes adding view counts across platforms indefensible. A reader is a person who took the statement in and understood it — the quantity that matters, and the one nothing here measures.
The strictest definition in commercial use comes from the advertising industry's own accreditation body, which makes it a useful ceiling rather than a floor. The Media Rating Council's Viewable Ad Impression Measurement Guidelines, Version 2.0 Final, 18 August 2015 (read 15 August 2026) requires, for a display advertisement:
"Greater than or equal to 50% of the pixels in the advertisement were on an in-focus browser tab on the viewable space"
"The time the pixel requirement is met was greater than or equal to one continuous second, post ad render."
For video the threshold is two continuous seconds. The same guidelines require measurers to report impressions with "Viewable Status Undetermined" as a separate category, and acknowledge "unexplained inconsistencies in viewable impression reporting among measurers." So the audited best case in commercial measurement is half the pixels on screen for one second, with the standards body noting that its own accredited measurers disagree about applying it. A consumer platform's public counter is looser than that by a margin the platform does not quantify.
What the operators say their own counters count
Where a counter becomes an exhibit, the operator's published definition belongs beside it. From X's help page on view counts (read 15 August 2026):
"Anyone who is logged into X who views a post counts as a view, regardless of where they see the post… or whether or not they follow the author."
"Multiple views may be counted if you view a post more than once, but not all views are unique."
"If you're the author, looking at your own post also counts as a view."
Read together, an X view count is a non-unique count of logged-in users that includes the author's own views and excludes embedded copies. It is at once the best publicly available exposure figure for a post on that service and, on the operator's own account, not a count of people. X also publishes that its reporting "is finalized within 24-48 hours of when impressions are served."
YouTube's engagement metrics page (read 15 August 2026) states that "to verify that metrics are accurate, YouTube may temporarily slow down, freeze, or change your metric count, and discard low-quality playbacks," giving as examples "streaming the same video across several windows and tabs." A view count observed on a date is provisional and retroactively adjustable by an undisclosed process. Capture it repeatedly with timestamps and expect the number to move.
The plaintiff's own instruments, and what they withhold
Search Console reports Google's own record of impressions, clicks and position for a property whose ownership has been verified, which decides more matters than any methodological question: it describes search performance for the plaintiff's site, not for the defendant's page. If the content sits on a complaint site, that page's data belongs to the complaint site.
Google's definition of the unit is the sentence to keep at hand. An impression, per Google's documentation (read 15 August 2026), "means that a user has seen (or potentially seen) a link to your site in Search, Discover, or News." The parenthetical is Google's own hedge. Google publishes the limits too: reports "only cover a representative sample of URLs," and the performance report may not track "queries that are made a very small number of times or those that contain personal or sensitive information" — which in a defamation matter describes the queries at issue exactly. A query missing from the report has not been shown to be a query nobody made. Because the performance window rolls forward and is measured in months, an export taken today may be the only surviving version of last year's figures. Take them early and hash them.
Analytics has the same problem in a different shape. Google's documentation on data thresholds (read 15 August 2026) states that thresholds are applied "to prevent anyone viewing a report or exploration from inferring the identity or sensitive information of individual users." Rows are withheld rather than zeroed when user counts are small — and a single defamatory URL on a small site is precisely a low-count page. Log-derived request counts and analytics-derived session counts will disagree; produce both and explain that the gap is structural rather than inventing an adjustment.
Estimates, and the word the vendor uses in its own filing
Where nobody has the defendant's logs, the substitute is Similarweb, Semrush, Ahrefs or a comparable tool. These are estimates built from panels, purchased clickstream data and modeling, and the vendors say so where saying otherwise carries consequences. Similarweb's Form 20-F for the fiscal year ended 31 December 2023 (read 15 August 2026) discusses in its risk factors the possibility that its data is not "current, sufficiently accurate, comprehensive or reliable," and that "information collected in future periods is not comparable with information collected in prior periods." The company's own term for the output is "estimated insights." A methodology change at the vendor can move a series without anything changing on the internet.
The best measured comparison located in the research is a vendor study rather than a peer-reviewed one, and it belongs in a report as one comparison and never as an established error rate: Omniconvert cross-referenced Similarweb figures against read-only Google Analytics access for 1,787 ecommerce sites and reported about 94 percent more sessions than the sites themselves tracked, with accuracy worst for sites under 10,000 monthly sessions and the overreporting systematic rather than random. A defamatory page on a niche complaint site is exactly that low-traffic case.
What an exposure opinion can defensibly say
A report in this area is a list of statements ranked by how little has to be assumed to make them. At the top sit direct observations: this URL was indexed and returned at these positions for these queries on these dates, captured with the query string, location parameter, device, session state and time zone recorded. Next sit first-party counts from records the plaintiff controls — requests carrying a given referrer over a stated period, complete only to the limits of referrer policy.
Below those sit operator-reported figures, re-observable but never person counts, quoted with the operator's own definition attached. Below those sit Google's reports for a verified property: first-party, and structurally under-inclusive for exactly the queries a defamation matter turns on. At the bottom sit vendor estimates, models the vendor declines to describe in full.
Anything under that last rung — a reach figure, a potential audience, an impressions-times-click-rate-times-population calculation — is a construction rather than an observation, and its assumptions belong to whoever built it. It should not be called a measurement.
Every record on that ladder is held by someone else and expires on a schedule nobody publishes. The distance between what could have been shown and what can be shown is usually created in the first sixty days.
Frequently Asked Questions
Can an expert tell the jury how many people saw the post?
No, and I would treat an expert who offers to as a risk. Every available record counts something other than people: a server log counts requests, a platform counter counts events on the operator's own definition, and a third-party tool reports a model. What an examiner can do is produce each figure, quote the operator's published definition of it, state what filtering was applied, and describe the gap between the figure and a reader. A range of delivery events with its assumptions on the face of the report is defensible. A headcount is not.Is there a legal standard for proving online audience size?
The research behind this page could not locate a US opinion setting one. Targeted searches for decisions treating hits, page views or impressions as proof of the extent of an online publication returned only secondary commentary. I report that as a finding, with the caveat that a search returning nothing is not proof that nothing exists. The practical consequence is symmetrical: there is no benchmark validating one measurement approach over another, and an opposing examiner works in the same empty space. That makes the disclosure of method, rather than the size of the number, the thing worth arguing about.Does a platform's view count prove how many people read the post?
It does not, and the operators say so themselves. X publishes that anyone logged in who views a post counts as a view, that multiple views may be counted for the same person, and that the author's own views count. YouTube publishes that it may slow, freeze or change a metric count and discard low-quality playbacks such as the same video streamed across several windows. So a counter is a non-unique, operator-defined, retroactively adjustable number. It is worth capturing, repeatedly and with timestamps, and it is worth quoting the definition beside it.Are Similarweb or Semrush estimates good enough to establish reach?
They are estimates and should be labeled as such. Similarweb's own annual filing calls the output estimated insights and warns that figures from different periods may not be comparable after a methodology change. One vendor study that cross-referenced Similarweb against Google Analytics for 1,787 ecommerce sites reported roughly 94 percent more sessions than the sites tracked, with the worst accuracy on sites under 10,000 monthly sessions. A defamatory page on a niche complaint site is that low-traffic case. The estimate can still be used; presenting it as a measurement is what fails.What does a server access log actually establish?
That a request with certain characteristics was received and a response of a certain status and size was sent. It does not establish that a person was present, that the page rendered, that anything was read, or that two entries from one address are the same person. Behind a content delivery network the logged address may be the network's rather than the client's unless forwarded-header logging was configured. The referrer and user-agent fields are supplied by the client and are routinely spoofed. The log is still the strongest reach artifact available, which says something about the rest.Which reach records disappear first if nothing is done?
Web server and CDN access logs, because no rule or default preserves them and rotation is automatic. Then analytics event data, which expires on a retention clock the account owner configured. Then platform-side figures, which are re-observable only while the content remains up. Search Console performance data sits on a rolling window measured in months. None of that is an opinion about what should be demanded or from whom; it is the order in which the technical record decays, and it is why the distance between what could have been shown and what can be shown is usually set in the first sixty days.Published