Web evidence in litigation
How far it carries

How Far These Methods Carry

Written for counsel: six forensic methods sorted by how far each carries, and the retention finding — three major operators publish no window at all.

Overview

This is written for counsel deciding what to demand, what to collect, and whether the technical work in a matter is worth paying for. The six methods here are sorted on one axis: how far each carries before it needs an assumption, and where one of them stops being accepted at all.

A method that reproduces is a different asset from a method that models. Both can be worth having, and the mistake counsel most often inherits is a report presenting the second kind in the voice of the first. Three tiers run through these pages.

  • Reproducible. Another analyst working from the same primary records reaches the same result. Any dispute is about what the result means, not about the result.
  • Reproducible with stated assumptions. The method works, and the assumptions have to be said out loud: the scope, the definitions, and what the operator published on the day the request went out.
  • Contested. The method is disputed, or no accepted standard exists for it. One of the six sits here, and it is the one this market sells hardest.

The two that reproduce

Authenticating a web page and documenting chain of custody are the only methods here where a second examiner working from the same records lands in the same place, because both are process rather than inference. A collection that recorded the served response and its headers, not only the rendered view, and that was hashed at the moment of capture, supports a narrow claim in which every component is documented.

Narrow is the operative word. A flawless custody record shows a set of files is unchanged since collection. It says nothing about whether the page was true, who wrote it, or how many read it. Products marketed as producing court-admissible evidence are selling about a quarter of the problem, and there is no US standard governing chain of custody for web evidence for them to comply with. SWGDE and NIST publish guidance; guidance is not a rule.

What does not work is the screenshot on its own. SWGDE's guidance on acquiring online content ranks acquisition methods and puts screenshots last, calling them the least forensically sound approach and instructing that they be used alongside better methods, never instead of them. A screenshot holds what one browser displayed to one examiner; what the server sent sits outside it.

What process returns, and the three operators that publish nothing

Whether a production is reproducible turns on something no analyst controls: what the operator holds, and what it published on the date process was served. State that and the method is sound; leave it out and the production is a stack of records of unknown completeness.

That position is thinner than most practitioners expect: two of the operators that come up most here publish a preservation window, and three publish none at all. Every page below was read on 15 August 2026.

Operator or authorityPublished preservation windowWhat is published about retention
Meta90 days, criminal investigations, pending formal processSays it does not retain for law enforcement purposes absent a preservation request received before the user deleted the content
X90 days, a “temporary snapshot,” pending valid processSays IP logs may be stored only “a very brief period”; publishes no figure
GoogleNone publishedAccepts preservation requests; no period on the transparency FAQ
RedditNone publishedPublishes a notice-to-user policy and no retention period
YelpNone publishedRetains “as long as reasonably necessary”; residual backup copies
18 U.S.C. § 2703(f)90 days, extended by 90 on renewed requestOperates on the request of a governmental entity

The only numbers in that table are the statutory 90 and the two operators mirroring it, so anyone who has quoted you a retention period for Google, Reddit or Yelp is quoting something other than the operator. The figures circulating for those three do not trace to any current first-party page found anywhere in this research, and I decline to print them: a plausible number that cannot be sourced is what gets an expert taken apart.

Section 2703(f) itself operates on the request of a governmental entity. A private litigant's preservation letter is not a § 2703(f) request and does not carry its statutory force, whatever an operator may do as a matter of policy. What duty an opponent or an operator is under is a question for counsel; the technical point is that the letter and the statute are not the same instrument.

Where the assumptions live

On metadata, the useful record is usually not in the copy that was posted. Platforms process uploads and processing discards fields, and no major operator publishes a policy describing what it strips, so the honest answer for a given service is a dated experiment on a stated upload path. The stronger records sit elsewhere: the native file held by a custodian, document properties, and the operator's account records, which hold an upload timestamp the file never carried. EXIF is a file's statement about itself, and one command rewrites any field.

On exposure, an impression, a view and a reader are three quantities rather than three words for one. A search impression, in Google's published wording, covers a link the user has “potentially seen.” X counts views without deduplicating them, the author's own included. Every counter here is an operator's definition of its own event, revisable by the operator. An exposure analysis can say what was indexed and returned for which queries on which dates, and what a first-party log recorded. Converting any of that into a count of people who read a sentence and understood it to concern the plaintiff is a model, and its assumptions belong to whoever built it. No US opinion setting a standard for proving online audience size turned up anywhere in the research behind this site.

The one that is contested

Coordinated account analysis carries the contested verdict, and it is sold most confidently in this market. From outside a platform an analyst can observe join dates, posting times, wording overlap, username schemes and handle reuse. What that supports is a pattern. Common control, the claim that one hand operated the accounts, rests on registration data, address logs and payment instruments that exist only platform-side.

The reasons are documented rather than rhetorical. No operator publishes its own detection method, so there is nothing to replicate. Automated scoring is measurably unreliable: a peer-reviewed evaluation of the most widely used bot-detection tool reported precision of 0.59 at a plausible threshold, meaning roughly four in ten flagged accounts were human, with recall of 0.2, meaning most actual bots went unflagged. And the commonest defect in real reports is methodological: the analyst collects the negative reviews, finds them similar to one another, and never samples the positive reviews to see whether they are similar too. Every pattern consistent with coordination is also consistent with a prompt template, a shared industry vocabulary, or two people on one office network.

None of that makes the analysis worthless; it makes it a screening step whose output is a list of what to demand from the operator. That is the pattern across all six methods: the ones that hold up were specified before they were run.

The entries

All 6 entries


Keep reading

Where these fit together

A page here answers one question about one element or one method. The guides run them in sequence, which is where the order starts to matter: several of these steps cannot be taken out of turn.

Top