Who this is for, and what sort of inventory this is
This page is written for an attorney litigating an internet defamation case, on either side, who needs to know what the technical record can carry before deciding what to plead and what to demand. It is not written for the subject of an online attack and there is nothing here about removal.
I am not an attorney, and whether an element is satisfied or a claim is worth bringing are judgments I do not make. What I do is sort a case by record, around the elements rather than the platform, because platforms change and the elements do not.
Three states run through everything below.
- The record establishes it. Primary records settle the point, and a second examiner working from the same records reaches the same result.
- Record plus testimony. The records support the point and cannot close it. What closes it is discovery, a custodian, or an ordinary witness.
- Not a technical question. It turns on law or on the fact-finder. No expert settles it, and one who offers to hands the other side its cross-examination.
Publication is satisfied by one reader, and the record rarely names a reader
Publication as an element means communication of the statement to someone other than the person it concerns. The consequence is constantly missed: one reader satisfies it, and almost nothing in the record identifies a reader. Server logs record requests. Platforms record delivery events and their own counters. Neither records a person who read a sentence.
Everything an expert produces about reach — access logs, analytics exports, platform counters, third-party estimates — therefore goes to the scope of the injury rather than to whether publication occurred. A report leading with an audience figure as proof of publication is answering a question nobody asked.
What the record carries is presence and address: a response containing this text was served from this URL, with this status code and byte count, at this time on the server's clock. Worth having, and not the element.
Record plus testimony. The records place content at an address on a date; who saw it comes from elsewhere or from nowhere.
Falsity: the record fixes the timeline and then stops
Three classes of record get confused in falsity disputes, and keeping them apart is most of the value an examiner adds.
- Records of what was published — captures, archives, platform exports. Content at a time.
- Records of the underlying fact — registries, dockets, licensing boards, filings, employment files. The state of the world at a time, and usually not technical at all: a paralegal pulls them and a custodian authenticates them.
- Records of what changed — version history, edit logs, diffs. Sequence.
An expert contributes to classes one and three. If the statement is that a clinic lost its license in a given year, a capture shows the sentence was published on a date; the licensing register shows the fact. Only one of those exhibits is mine.
Archives carry limits that cut both ways. Absence of a capture proves nothing, because crawls are sparse and the Internet Archive publishes no coverage guarantee. A replayed page is reassembled at view time from files captured on possibly different dates, so an exhibit captioned as the page on a date overstates unless it names the capture timestamp.
Record plus testimony for what was said and when. Not a technical question for whether it was true.
Identification: linkage is observable, understanding is not
Whether a reasonable reader would understand the content to refer to the plaintiff is the element, and it belongs to the fact-finder. What is machine-observable is narrower and useful.
The strongest signal is the publisher's own structured data. Where the markup declares its subject — an about or mentions property, a sameAs pointing at the plaintiff's own profiles, a Person entity with an identifier — the publisher has stated in machine-readable form who the page is about. That is the defendant's own declaration, sitting in the HTML, capturable and hashable, and it is the most underused artifact in this area.
Below it: anchor text on inbound links, which is a third party's contemporaneous statement of what the content is about; on-page context such as employer, license number or a phrase lifted from the plaintiff's own site; and reader statements in place — comments, replies and quote-posts naming the plaintiff. That last category is the most direct and the most fragile. Capturing the comment thread rather than the article is the difference between having it and not.
Exact-string searching misses misspellings, initials, former names, handles and deliberate obfuscation with zero-width or look-alike characters. Normalizing text before searching is a real step, and its absence is a real criticism.
The record establishes linkage. Not a technical question: what a reader understood.
Attribution: four records and one inference
Attribution is a chain of five inferences, and the opinion is worth what its weakest link is worth.
- Content to account — the post is associated with an account or a session.
- Account to registration records — an email, a phone number, a payment instrument, a creation time.
- Account to address and time — the platform logged an address from which the account acted.
- Address and time to subscriber — the carrier maps that address at that instant to a billing account.
- Subscriber to person — somebody in that household, office or building did it.
Links one through four are records. Link five is not a record and never has been. It is an inference from circumstance, and it is where these cases are decided — usually with device examination on a machine obtained in discovery, an admission, or ordinary witness testimony, rather than with network evidence.
Every identifying field a platform holds is self-asserted except the payment instrument, the one with an out-of-band verification step behind it — so attribution succeeds far more often against accounts that bought something than against accounts that bought nothing. The transaction, not the login, is where these chains close. And a return of a recent login address is not a login history; the address used at the moment of the post rarely arrives without a specific, timestamped request.
Record plus testimony at best, and frequently not even that.
Republication: a diff is a fact, republication is a conclusion
This is the most productive intersection in the subject, because the legal test is asked about facts only the record answers.
The leading American application of the single publication rule to the web is Firth v. State of New York, decided in 2002, in which the court held the rule applies to material on a public website, stated that republication retriggering the limitations period occurs upon a separate aggregate publication from the original on a different occasion, and said the mere addition of unrelated information to a website cannot be equated with repetition of defamatory matter in a separately published edition. Whether the rule governs in a forum, and whether a change amounts to republication, are for counsel; I report the decision as the research reports it.
What I produce is the change record. A version-by-version diff shows what text changed and when. Redirect chains, canonical tag changes, sitemap modification dates and captures at both addresses show that content moved. Newsletter archives, a publisher's own timeline, feed re-entry and a modified-date property in the markup show it was pushed out again. Each is a dated fact; whether any is republication is an argument, and not mine.
The record establishes the change and its date. Not a technical question: what the change means.
Damages: every counter is a proxy, and one standard does not exist
The finding that matters most here is an absence. The research behind this site went looking for a controlling or widely cited US opinion setting a standard for proving how many people saw something online, and could not locate one; repeated case-law searches returned secondary commentary and nothing else. I state that rather than filling the space.
What can be said precisely is that impressions, views and readers are three quantities. An impression is a delivery event — content placed where someone could have seen it. A view is a platform-defined event whose definition varies by operator and changes without notice, which is why adding view counts across platforms is not a defensible operation. A reader is a person who read the statement and understood it, and no metric measures that.
Every method that reaches a number runs through an instrument whose owner reserves the right to change it, and through a counterfactual nobody observed. A before-and-after comparison assumes the instrument held still across the break, that the pre-period covers seasonality, that pricing and media spend did not move, and that the conversion rate of retained traffic applies to lost traffic. Each is an assumption. And where a defamatory article, a regulatory complaint and a bad quarter land together, no method separates them.
Record plus testimony for measurable first-party series. Not a technical question for what any of it is worth.
What will not work
Four things I am asked for regularly and decline, because each fails on its own terms.
- A folder of screenshots assembled months later. A screenshot shows that a screen once looked a certain way. It does not establish the URL unless the address bar is legible in the same frame, does not establish the date, and does not distinguish itself from a page edited before capture.
- A third-party traffic estimate for a small complaint page. The one large comparison against ground truth I can point to covered 1,787 sites and found one estimator reported roughly 94 percent more sessions than the sites tracked, worst below ten thousand monthly sessions.
- A retention figure for an operator that publishes none. Any such number quoted to you for Google, Reddit or Yelp originated somewhere other than the operator.
- Writing-style attribution of a two-sentence review. No published error rate exists at that length.
Where these cases are actually decided
Almost every strong record here sits in someone else's hands and expires on a schedule the holder sets. What is published about preservation, read on 15 August 2026, is thinner than the commentary around it suggests:
| Operator or statute | What is published |
|---|---|
| Meta | 90 days for account records in official criminal investigations, pending formal legal process |
| X | 90 days, a temporary snapshot, pending valid legal process; IP logs a very brief period, no figure |
| Accepts preservation requests; no window published | |
| No preservation window and no log-retention figure published | |
| Yelp | No retention window published; retention stated as long as reasonably necessary |
18 U.S.C. 2703(f) | 90 days, extended by 90 on renewed request, on the request of a governmental entity |
Two things follow. The statutory preservation letter at 18 U.S.C. 2703(f) runs to a governmental entity, so a private plaintiff has instead a request the provider may honor, and a court order. And Meta's published guidelines condition retention for law enforcement purposes on a valid preservation request arriving ahead of the user's own deletion — which, where a post has come down, is the difference between a record and no record.
Every figure in that table has a date attached because operators revise these pages without notice.
Frequently Asked Questions
Which parts of an internet defamation case can a technical expert actually settle?
Presence, address, sequence and integrity. I can show that a response containing specific text was served from a specific URL at a specific time, that a page changed in a specific way on a specific date, that content appears at other addresses, and that files collected have not been altered since collection. I can show machine-readable linkage between content and a named person. What I cannot settle is whether the statement was true, whether a reader understood it to refer to the plaintiff, or how many people read it.Does audience evidence prove publication?
No, and the two are frequently conflated. Publication as an element is communication to one person other than the subject, and one reader satisfies it. Audience evidence goes to the scope of the injury instead. Logs, analytics and platform counters record requests and delivery events, not readers, so an expert offering an impression figure as proof that publication occurred is answering a different question than the one the element asks. The distinction matters because it changes what you need to demand in discovery.Is there a recognized standard for proving how many people saw an online statement?
The research behind this site looked for a controlling or widely cited US opinion setting such a standard and could not locate one. That absence is the finding, and I state it rather than substituting a plausible method. What exists instead is a set of proxies with published limitations: platform counters defined by their operators, first-party search performance data that is sampled and thresholded, server logs that include automated traffic, and third-party estimates the vendors themselves describe as modeled. None of them counts readers.Why does an IP address not identify the poster?
Because it identifies a connection, and only at an instant. A household router puts an entire household behind one public address, and mobile and many broadband carriers put many unrelated subscribers behind one address at the same time using carrier-grade network address translation. Behind a content delivery network, an origin server log may record the network's address rather than any client at all. Even a clean carrier return names a billing account, and the step from a billing account to a person at a keyboard is an inference, not a record.What should be preserved first in an internet defamation matter?
The records that expire without anyone touching them. On the plaintiff's own side that means web server access logs, content delivery logs, analytics exports, search performance exports and content management system revisions, all of which rotate or are pruned by routine maintenance. On the other side it means the specific record types named in a preservation demand rather than a general instruction. Meta publishes that it does not retain data for law enforcement purposes absent a preservation request received before the user deleted the content.Which elements are not technical questions at all?
Whether the statement was false, whether a reasonable reader would understand it to refer to the plaintiff, whether a change to a page amounts to republication, and what any measured decline is worth. Records can fix a timeline precisely enough for the factual record to be laid against it, show machine-readable linkage, produce a dated change record, and produce first-party series. Every one of those stops short of the element itself. An expert who crosses that line is offering the other side an easy hearing.How current is the platform information on this page?
The preservation and retention positions summarized here were read from the operators' own published pages on 15 August 2026, and the table says which operators publish a figure and which publish nothing. Two publish 90 days. Three publish no window at all. I date these statements because operators revise the pages without notice, and because a figure quoted from memory or from a blog is the kind of detail that gets an opinion picked apart later.Published