How I Work: From the First Call to a Delivered Opinion
The sequence an internet defamation engagement actually runs in, and where each stage can go wrong
Who this page is for, and what it covers
This page is for an attorney deciding whether to retain a technical expert in an internet defamation matter. It describes the engagement in the order it runs, and it is about the substance of the work rather than court procedure.
One structural fact goes first: what is provable in year two is decided in the first week. Preservation windows, where published at all, are short. Server logs rotate in days. Pages get edited and taken down. Nothing later recovers a record that was gone before anyone asked for it.
The conflict check comes before the conversation
Before I hear the facts, I need the names: the parties, the sites or platforms involved, opposing counsel, and any related entities. That is the conflict check, and it comes first for a practical reason as much as an ethical one. If I have worked for the other side, or on the same domain, or for a company connected to a party, the engagement is over before it starts — and if I heard your theory first, that is your problem as much as mine. So a first message is light on facts and complete on names, and the substantive call happens once the check clears.
Scope, written down before anything is collected
The scoping call fixes four things, written down before I touch a keyboard:
- The question. Not “look at this website” but something a record can answer — was this page published on this date; is this exhibit what it purports to be; do these accounts share registration records.
- The universe. Which URLs, which accounts, which date range, which platforms.
- The viewing conditions. Logged in or logged out, which browser, whether content differs by viewer. That sounds like a detail and is frequently the whole dispute.
- What would make the answer no. Deciding in advance what result defeats the theory is what keeps an analysis from becoming a search for confirmation.
Scope decided after collection is scope decided by what was found, and that is the first thing an opponent looks for.
What records exist, and who holds them
Before collection, records get inventoried by holder, because who holds a record determines both how you reach it and how long you have: what your client already has (its own site, server logs, search console data, email), what is publicly visible right now (the live page, the account, the replies), and what only process reaches (platform account records, registration addresses, session histories, registrar and hosting records, payment instruments). The first is what ordinary housekeeping destroys. The second belongs to whoever can delete it, so collect it today. The third resolves attribution and cannot be self-collected.
18 U.S.C. § 2703(f) requires a provider to preserve records for 90 days, extendable by a further 90 on renewed request — but it runs to a request by a governmental entity, and a private litigant's letter is not a § 2703(f) request.
As of 15 August 2026, Meta and X publish a 90-day preservation window on their law-enforcement pages. Google, Reddit and Yelp publish no preservation window and no log-retention figure on their legal-request pages. Any number you have been quoted for those three came from somewhere other than the operator. Meta also publishes that it does not retain data for legal purposes where a user deleted the content before a valid preservation request arrived — which is why a preservation letter is a first-week task.
Collection, and hashing at the moment of capture
Collection is where most exhibits are won or lost. There is no US standard specific to chain of custody for web evidence and no accreditation for it, so a product sold as “compliant” with one is describing something that does not exist. What exists is defensible practice, drawn from SWGDE's digital-evidence guidance and NIST's four-phase structure.
- Strongest method first. A programmatic interface where one exists, full browser-based capture second, screenshots as a supplement and never alone. A screenshot establishes that a screen once looked a certain way — not the URL, not the date, not that a server sent it.
- Capture the record, not the picture. Full-page rendering, served HTML, post-script DOM, response headers, and where possible a
WARCfile — the archival format standardized as ISO 28500, which stores a page together with the requests and responses that produced it, so the archive holds the interaction and not a picture of it. - Hash at the moment of collection. SHA-256, every file and the manifest, written into the log at capture time. A gap between capture and hashing is the most common real defect: a hash computed a week later proves integrity from that week, not from collection. Not MD5 and not SHA-1, because a practical SHA-1 collision was demonstrated publicly in 2017.
- Log the environment — time with timezone and time source, machine, operating system, browser, tool and versions, network and address, proxy state, any account logged in — repeat the acquisition where content is dynamic, and keep the failures: a log showing a URL returned a 404 is evidence, and one that omits the attempt is a gap someone else will find.
All of it exists so one narrow claim can be defended: on this date, from this network, using this software, a request to this URL returned these bytes, and those bytes have not changed since. “This is what was on the internet” is a different claim, and no product makes it.
The analysis, and what each method assumes
With a defensible collection in hand, the analysis falls into four groups.
Authenticity and provenance
Whether an exhibit is what it purports to be: captures against response headers, archived copies against each other, current state against historical records. Where a page changed, the change is often the finding.
Attribution
Registration history, infrastructure overlap, timing patterns and posting behavior generate leads and quantify patterns. They rarely name a person, because the records that resolve attribution sit with platforms, registrars and access providers. An address identifies a connection, not a person at a keyboard.
Exposure
Impressions, views and readers are three different things with three definitions, each written by the platform reporting it, and none of them counts readers. No accepted standard exists for converting any of it into a reader count, and third-party traffic estimators carry documented, large error — worst on the small sites most defamation targets run.
Coordination
Whether a group of accounts behaves as a group: registration timing, posting synchronization, content reuse, network structure. An outside analyst can observe behavior and cannot see device identifiers, addresses or a platform's internal linkage assessments.
A method whose threshold was chosen after seeing the result is not a method, so the assumptions get stated as assumptions.
What gets delivered
The deliverable is a written opinion containing four things:
- The findings, each tied to the records it rests on, in language distinguishing what a record shows from what it is merely consistent with.
- The collection — exhibit set, hash manifest and collection log, so another analyst working from the same primary records gets the same result.
- The assumptions, including the ones I would rather not draw attention to.
- What could not be established, and what record would have established it — the part attorneys tell me is most useful, because it converts into what to ask for in discovery.
Well before anything is drafted, you get a short written statement of what the records support.
What happens when the records do not support the theory
Sometimes the page cannot be tied to the defendant. Sometimes the content came down before anyone preserved it and the archived copies do not show what counsel hoped. Sometimes the exposure data supports a fraction of the audience the complaint assumes. Sometimes the question is not technical and no expert can settle it.
When that happens I say so in writing, as early as I know it and before an opinion is drafted, and the engagement can end there. An expert who finds a way to support the theory anyway is worth less than nothing: the weakness does not disappear, it gets found by someone else at a worse moment.
What it costs, structurally
I work against a retainer, billed for time, and I do not quote a figure on a website because it would be meaningless without the scope. What drives the number is worth knowing, since three of these are within your control:
- Volume. Ten URLs and two accounts is a different engagement from four hundred URLs across six platforms.
- Timing. Collecting live content is cheap. Reconstructing content that is gone costs several times more and produces a weaker result. This is the largest single lever, and it is set by how early the call happens.
- Whether productions arrive, and in what shape. Platform productions vary wildly in format and completeness, and reconciling a messy one is real work.
- Whether testimony is needed. If the matter reaches that stage, it is scheduled and billed separately.