Web evidence in litigation
Proving an Internet Defamation Case

What a Defamation Expert Can and Cannot Say

How far the records carry, which assumptions have to be said out loud, and what an expert who overreaches sounds like

Who this is for, and what an opinion in this field consists of

This page is written for counsel evaluating whether an internet defamation expert has anything useful to offer in a matter, and what the resulting opinion would assert. It is about substance — the reach of the records and the reach of the methods — rather than procedure, and it does not tell a defamation victim what to do about content.

An opinion in this field is a set of bounded statements about records. Each has a source, a date, a method another examiner can rerun, and a limit. The work is not in producing a conclusion. It is in knowing where each statement stops, and writing so that the stopping point is visible rather than buried. A finding has to hold up when someone attacks it, and the findings that hold up are the ones specified before they were run.

I am not an attorney. Whether an element is met and whether a claim is viable are not questions I answer, in either direction, for either side.

The sentence a collection actually supports

Traditional custody assumes a physical object that is seized, sealed and stored. Web evidence has none of those properties: the original is on a server I do not control, it may differ between requesters, and it can change or vanish without notice. So the claim has to be stated narrowly, and stated narrowly it is provable:

On this date, at this time, from this network, using this software, a request to this URL returned these bytes, and those bytes have not changed since, as shown by these hashes.

That is defensible. This is what was on the internet is not, and no product makes it so. Several vendors sell captures delivered with a signed certificate, a hash manifest and sometimes a trusted timestamp token. What those services genuinely add is a third-party operator who can be questioned, a uniform documented process, and a clock the party does not control. What they do not add is any special status, and marketing that calls a capture product court-ready evidence is selling authenticity and implying the rest.

One point follows from the sentence above and is worth stating plainly: a hash proves integrity forward from the moment of hashing, not backward. It says the file has not changed since it was hashed. It says nothing about whether the file faithfully represents what the server sent. Conflating those two claims is the most common overstatement in this discipline, and it is why the interval between capture and hashing matters — a hash computed a week later covers that week.

What the records carry without help

A short list, and it is shorter than most reports imply. Working from primary records, and with the collection done properly, I can say:

  • That content existed at an address at a time. Served bytes, response headers, status code, and a hash computed at collection.
  • That content changed, and how. A version-by-version comparison, with the source of each version identified.
  • That the same or near-identical text appears at other addresses. Normalized text, hashed and compared for exact matches; shingled and scored for near matches.
  • That separate sites share an operator, sometimes. Account-scoped identifiers — analytics measurement IDs, ad publisher IDs, affiliate tags — are issued to a registered party, which makes their reuse across unconnected sites one of the few open signals leading to an account holder rather than a cluster.
  • That a page was returned for a query at a position on a date. Reproducible only if the capture recorded the exact query, the location parameter, the device and viewport, the session state, whether personalization was suppressed, the timestamp with its zone, and the full page.
  • That traffic reached the plaintiff's own property from a named referrer. First-party logs, complete to the limits of referrer policy.

Everything else in a typical report is one or more inferential steps beyond that list, and each step needs to be named.

What has to be stated as an assumption, out loud

Some of the most useful work in this field rests on choices that are not facts. The choice is legitimate. Hiding it is not.

Similarity thresholds. Near-duplicate detection requires a shingle size and a similarity cutoff, and I could find no standards body or forensic publication setting either. The number has to be justified from the data in the matter rather than cited to an authority, and an opinion that does not disclose it is not reproducible.

Instrument continuity. Any before-and-after comparison assumes the measuring instrument held still across the break. Analytics migrations, consent banners, tag changes, bot-filter changes and browser privacy releases all produce step changes in measured traffic with no change in actual traffic. Continuity has to be established affirmatively, not assumed.

Control series. Counterfactual methods need controls the event did not touch, and in these matters that requirement fails in predictable directions: competitors as controls rise because of the event and inflate the estimate; the plaintiff's other lines fall with it and shrink the estimate toward zero; category demand is contaminated if the event was newsworthy. There is no statistical repair. The control set has to be argued substantively, with a placebo run on a date when nothing happened and a sensitivity check across alternative controls.

Filtering. Any count from server requests includes automated traffic, and filtering by user-agent alone is the weakest form because that field is supplied by the client. What was filtered, and what could not be, belongs in the opinion.

Methods that reach an exclusion but not an identification

Writing-style analysis is the clearest case. On long documents with a small set of known candidates it is a real research field with real results. On a sixty-word review it is being asked to do something no published error rate covers, and the best large-scale figure in the literature comes from work on thousands of words per candidate. Text drafted or edited by a language model compounds it: there is no accepted method, no standard and no error rate for authorship analysis of machine-assisted text, and an examiner cannot currently rule out that a questioned post was machine-drafted.

The defensible position, and it is mine: style comparison generates leads and supports an exclusion far more comfortably than an identification. Where writing similarity is used at all it should be described concretely — a shared misspelling, a shared idiom, an identical factual error — rather than dressed in a similarity score the reader cannot check. Identical errors are worth more than identical phrasing, because a phrase can be copied on purpose and an error usually is not noticed.

The same logic governs automated account classification. One published evaluation of the most widely used bot-detection tool, resampling to an assumed fifteen percent bot population at a common threshold, reported precision of 0.59 — about four in ten accounts classed as bots were human — with recall of 0.2, and scores that crossed the threshold repeatedly within a three-month window. That behavior can support a screening step whose output is then examined by hand. It cannot support an assertion that a particular account is inauthentic, and the report should say which was done.

What a platform counter means, in the operator's own words

The most common evidentiary error in this area is treating a public counter as a count of people. The operators say otherwise on their own help pages, and quoting them is more persuasive than arguing about them.

X publishes that anyone logged in who views a post counts as a view regardless of where they see it, that multiple views may be counted so not all views are unique, and that an author looking at their own post counts as a view; embedded posts do not add to the count. Read together, an X view count is a non-unique impression count of logged-in users, inclusive of the author's own views and exclusive of embeds. It is the best public exposure figure for a post there, and it cannot carry a person count.

YouTube publishes that it may temporarily slow down, freeze or change a metric count, that it discards low-quality playbacks, and that it is constantly confirming and adjusting engagement events. A count captured on one day may not match the same counter later, which makes an exhibit built on one an exhibit about a moment.

Google's definition of a search impression is that a user has seen, or potentially seen, a link — the hedge is in the operator's help text. And the strictest audited definition of viewability in commercial use requires half an advertisement's pixels in an in-focus tab for one continuous second, two for video, with a separate reported bucket for impressions whose viewable status could not be determined. That is the industry's best case, and it does not claim to identify a person or a reading.

Numbers I will not produce

Refusals are part of the opinion, and stating them at the start is cheaper for everyone than finding them at the end.

  • A count of people who read a statement. No metric measures it and no validated method converts traffic estimates into it.
  • A percentage adjustment for analytics undercounting. Blocking rates are site-specific and no reliable published figure exists; an expert applying a generic uplift invented the uplift.
  • A preservation window nobody has published. On the pages I read on 15 August 2026, three of the operators most often at issue set out no such period at all.
  • A probability attached to a timing cluster. Nobody has published a base rate for coincidental posting overlap, so the denominator does not exist and the number would not be a probability in any technical sense.
  • A location precise enough to name a person. The dominant commercial provider puts its own accuracy at about 80 percent for a state or region and about 66 percent for a city within 50 kilometers, and says its data is never precise enough to identify or locate a specific household, individual or street address.
  • A total count of copies. No method enumerates every copy; copies inside apps, private groups, newsletters and messaging are invisible from outside. Every count is a floor, and the correct phrasing is that at least this many copies were located by these named methods.

What overreach sounds like

A short field guide, useful in both directions — for reading an opposing report and for reading mine.

  • I traced the poster. Device identifiers, cookie IDs, advertising IDs, internal linkage graphs and address histories are platform-internal. Ask which record supports the trace, and where it came from.
  • These accounts are the same person. An outside analyst can establish that a pattern exists and describe how unlikely it looks under an independence assumption. Common control is an identity claim, and outside a platform production it is stacked on top of the pattern claim.
  • The content was copied 47 times. A floor presented as a total.
  • The reviews are similar, therefore coordinated. The usual defect is that only the negative reviews were collected. If the positive reviews are equally similar — and platform prompts, industry vocabulary and shared experiences make that common — the finding evaporates.
  • Brand search demand fell. The usual instrument reports relative, sampled and scaled figures, excludes searches made by very few people and queries containing apostrophes, and warns that its data sometimes reflects statistical noise rather than search interest at low volume. A business name is a low-volume query.
  • The page as it appeared on this date. An archived page is a reconstruction assembled from files captured on potentially different dates, and the archive says so.

Concluding that the data does not support an estimate

The most valuable and least supplied output in this market is a conclusion that the records cannot carry the number being asked for. It is a legitimate opinion, frequently the correct one, and far more robust than a figure built on four stacked assumptions.

It is also useful earlier than any figure would be, because it changes what gets pleaded and what gets demanded. If a reputational event is one of four inside a quarter, no method separates them. If the only exposure data for a page on someone else's site is a third-party estimate, and the page is small enough that the estimator performs worst there, the estimate is the tool's weakest output applied to the case it handles least well.

Where the records do carry something, the strongest form of the analysis is the narrowest: first-party impressions and clicks for the queries at issue, over a long pre-period, with instrument continuity documented and no revenue extrapolation attached. The extrapolation is where these opinions fail.

Where my opinion stops and counsel's argument starts

The boundary is not a disclaimer; it is the reason the rest gets believed. A diff shows what changed on what date; whether the change is republication is an argument. Structured data shows that a publisher declared a page to be about a named person; whether a reasonable reader made that connection is for the fact-finder. A control-adjusted series produces a range; whether anything is recoverable is law. A capture with headers and hashes shows these bytes were served and have not changed; whether the statement was false is proved with the register, the docket or the filing, which is documentary evidence rather than a technical exhibit.

The practical version, for a first phone call: I will tell you what the records can carry and what they cannot, in that order, and I would rather tell you the second part before you have spent anything on the first.

Frequently Asked Questions

What can an internet defamation expert say about a web page capture?

That on a stated date and time, from a stated network, using named software, a request to a specific URL returned specific bytes, and that those bytes have not changed since, as shown by hashes computed at collection. That claim is provable and narrow. The broader claim that a capture shows what was on the internet is not provable, because a server can return different content to different requesters, and no capture product changes that. Hashes prove integrity forward from hashing, not backward to the server.

Can an expert identify an anonymous author from writing style?

Not on short platform text, in my view. The strongest published work in the field uses thousands of words per candidate author, and there is no published error rate for a sixty-word review. Text drafted or edited by a language model makes it worse, because no accepted method or error rate exists for authorship analysis of machine-assisted writing. Style comparison generates leads and supports an exclusion more comfortably than an identification, and where it is used it should point to concrete shared features rather than an unverifiable score.

Is a platform view count evidence of how many people saw a post?

No, and the operators say so. X publishes that a view is counted for any logged-in user who sees the post, that multiple views by the same person may be counted, and that the author's own views count. YouTube publishes that it may slow, freeze or change a metric count and discard low-quality playbacks. Google defines a search impression as a link a user has seen or potentially seen. These are delivery and interaction counters with operator-specific definitions, and adding them across platforms is not a defensible operation.

Should a bot-detection score appear in an expert report?

Only as a screening step that was then checked by hand, and labeled as such. A published evaluation of the most widely used tool reported precision of 0.59 at a common threshold under an assumed fifteen percent bot population, meaning roughly four in ten flagged accounts were human, with recall of 0.2 and scores that crossed the threshold repeatedly over three months. Those numbers cannot support a claim that a specific account is inauthentic. They can support a decision about which accounts to examine individually.

What makes a similarity finding reproducible?

Naming the method and its parameters, and disclosing every choice that was not dictated by the data. For near-duplicate text that means the normalization applied, the shingle size and the similarity cutoff. I could locate no standards body or forensic publication setting an accepted threshold, so the number has to be justified from the material in the matter rather than cited to an authority. The deliverable should be the dated dataset and the code that produced the table, so that another examiner can rerun it.

Is concluding that damages cannot be estimated a real expert opinion?

Yes, and it is often the correct one. Where several reputational events fall inside the same quarter, no method separates their effects, and saying so is more defensible than a figure resting on stacked assumptions about instrument continuity, control series, conversion rates and duration. It is also more useful early, because it changes what gets demanded in discovery. Where the records do carry something, the narrowest version — first-party query data over a long pre-period, with no revenue extrapolation — is the version that holds.

How should an opposing expert report be read?

Look for the joins rather than the conclusions. Which record supports each step, what was assumed rather than measured, whether a count is presented as a total when it can only be a floor, whether only the material supporting the theory was collected, and whether thresholds and parameters are disclosed. In coordinated-account work the most common defect is sampling only the negative reviews and never checking whether the positive ones are equally similar. In damages work it is the extrapolation attached at the end.
Keep reading

The pages behind this guide

Every element and every method named here has its own page, with what legal process returns quoted, the date it was read, and the row that names what the record will not establish.

Top