What attribution is being asked to do
This page is written for counsel weighing whether technical work can put a name to an anonymous or pseudonymous post in a defamation matter. It is about what the records support and where they stop, not about what the person targeted can do about the content.
The sentence to keep in mind: an IP address identifies a connection, not a person. Everything here is an attempt to close the distance between those two things, and the honest description of the state of the art is that records close most of it and never the last part.
That is not a reason to skip the technical work, only to be exact about what it is for. In most matters that resolve, the records narrow the field to an account, a connection and a window of time, and the name arrives from discovery, from a device produced in litigation, from an admission, or from a witness. Framing the engagement that way saves more than any method choice made later.
The chain of five, and the link that is not a record
Attribution is not one inference but five, and an opinion is only as strong as the weakest.
- Content to account. The post is associated with an account or a session.
- Account to registration records. The account carries an email address, a phone number, a payment instrument and a creation timestamp.
- Account to address and time. The platform logged an address from which the account acted, at a time.
- Address and time to subscriber. The ISP maps that address at that instant to a billing account.
- Subscriber to person. Someone in that household, office or building did it.
Links one through four are records held by identifiable custodians. Link five is not a record and never has been. It is an inference from circumstance, and it is where matters are won and lost — on device forensics, admissions, or testimony, rather than on anything from the network.
The chain tells counsel which custodian holds which link, which links are expiring, and which link a proposed subpoena would advance. A subpoena that returns a registration email for a disposable address has bought link two and moved nothing else.
What an address does not establish
Four problems sit between a logged address and a subscriber, each a separate line of cross-examination.
Proxies and address sharing. Apache's documentation warns that the address in an access log "is not necessarily the address of the machine at which the user is sitting" — behind a proxy or a delivery network it is the intermediary's. A home router puts a household behind one public address; carriers put many unrelated subscribers behind one, using Carrier-Grade NAT. The IETF allocated a block for it: RFC 6598 (April 2012, read 15 August 2026) states that 100.64.0.0/10 was allocated "to accommodate the needs of Carrier-Grade NAT (CGN) devices." For a post over a mobile connection, an address and a timestamp may not resolve to one subscriber unless the carrier also recorded the source port range assigned at that instant — a question to put to the carrier, since I have not seen one publish an answer.
Dynamic assignment. Residential addresses are leased and reassigned, so the lookup is "who held it at this instant," which makes the timestamp and its zone load-bearing. A platform log in UTC and a request in local time with no offset return the wrong window.
Geolocation is not identification, and the vendor says so. MaxMind publishes its own accuracy figures (read 15 August 2026): 99.8% at country level, "around an 80% accuracy at the state/region level, and a 66% accuracy for cities (within a 50km radius)." And, in the vendor's own words:
"GeoIP geolocation data is never precise enough to identify or locate a specific household, individual, or street address"
A VPN, MaxMind adds, geolocates to the VPN server, not the user. A one-in-three error rate at city level is not a footnote.
When the address logs run out
Most attribution attempts fail on the clock, not the method. Where an operator publishes nothing, the honest answer is that nothing is published — not a figure from a blog.
| Custodian | What is published about the records that matter to attribution |
|---|---|
| Comcast | "IP address logs are retained for 180 days"; preserved material held 90 days, extendable by 90 |
| X | IP logs "may only be stored for a very brief period of time" — no figure; a 90-day snapshot on valid process |
| Meta | 90 days pending formal legal process; nothing retained for that purpose if the user deleted first |
| Accepts preservation requests; no window and no log-retention figure published | |
| No window and no retention figure on the pages located; a notice-to-user policy is published | |
| Yelp | No window; retention stated only as "as long as reasonably necessary," plus residual backups |
All six read 15 August 2026. Only one custodian there publishes a figure for the record attribution needs — the address log — and it is the ISP, not a platform. Anyone quoting a retention period for Google, Reddit or Yelp is estimating.
The statute behind the 90 days is 18 U.S.C. § 2703(f) (read 15 August 2026), and its first qualifier is what surprises people. The duty to "take all necessary steps to preserve records and other evidence" arises "upon the request of a governmental entity," and what is preserved is held "for a period of 90 days," extendable once on renewed request by that entity. A private litigant has no § 2703(f). Whether a civil party can obtain an order compelling preservation is for counsel. The technical finding is that the default is deletion, the clock is short, and the renewal is the requester's calendar entry, not the platform's.
What a production about an account contains
What a service holds about an account is narrower than most assume. The basic subscriber records at 18 U.S.C. § 2703(c)(2) are a fair proxy: name; address; session times and durations; length and types of service; subscriber number "including any temporarily assigned network address"; and means and source of payment. Meta's own guidelines put the same tier in operational language, ending with the phrase that matters here: "a recent login/logout IP address(es), if available."
Read the qualifier: a recent login or logout address, if available. That is not a login history, and the most valuable field — the address used at the moment the disputed post was made — is the one least likely to appear without a specific, timestamped request.
The more actionable point: every identifying field on a consumer account is self-asserted except the payment instrument. A name is typed in, an email address is free and disposable, a phone number can be a voice-over-IP line obtained in minutes. Payment is the one field with an out-of-band verification step behind it, which is why attribution succeeds far more often against accounts that bought something — a subscription, a promoted post, a domain, a review package — than against accounts that never transacted. Look for the transaction, not the login.
The dataset the plaintiff already owns
One attribution dataset needs nobody's permission to preserve, and it is the one most often lost. Anonymous posters frequently visit the target's own website — before posting, and repeatedly afterward to watch. Those requests sit in the target's access logs, with referrer strings, user-agent strings, timestamps and addresses attached. Preserving them the day the content is discovered is the highest-value, lowest-cost attribution step in this subject, and almost never taken, because logs rotate away while everyone is looking at the post.
Two caveats belong beside it. The correlation runs the wrong way — a visit from an address is not authorship. And the browser-fingerprint research that makes this dataset interesting is old — Peter Eckersley's How Unique Is Your Web Browser? (PETS 2010) reported that 83.6% of 470,161 browsers had an instantaneously unique fingerprint, but plugin enumeration is gone and browsers have shipped countermeasures since, so that is not a current rate. The structural point survives: fingerprinting is a collection technique, not an analysis technique, so it yields evidence only where somebody was running the collection on a page the poster visited — here, the target's own site and essentially nowhere else.
Writing style, timing, and what neither carries
Two techniques get asked for most, and both exclude more comfortably than they identify.
Stylometry. The best published internet-scale accuracy figures come from very large candidate sets and long writing samples — the right author ranked first roughly 20% of the time out of 100,000 candidates, on thousands of words per candidate (Narayanan and colleagues, 2012). A two-sentence review is not that. There is no published error rate for authorship attribution on texts that short, so an examiner attributing a brief anonymous post by writing style is offering an opinion with no known error rate on a sample smaller than anything in the literature. Stylometry generates leads and supports an exclusion far more comfortably than an identification. A newer problem has no settled answer at all: text produced or edited by a language model does not carry its nominal author's stylistic signal, and there is no accepted method and no error rate for attributing machine-assisted text.
Timing correlation. No base rate for coincidental timing overlap has been published. If six accounts all post between nine and eleven in the evening, the pattern is consistent with one person and equally consistent with six people in one time zone who have jobs. Nobody has published the denominator, so a probability attached to a timing cluster is not a probability in any technical sense. Platform timestamps are also not uniform: some are absolute and in UTC, some render as relative text, some are normalized to the viewer's zone. A report that does not state, for every timestamp, where it came from and in what zone, is not reproducible.
The calendar is the binding constraint
The procedural sequence, not the method, usually decides whether attribution is possible. In federal court discovery generally cannot begin until the parties have conferred, and they cannot confer without a defendant: Fed. R. Civ. P. 26(d)(1) (read 15 August 2026) provides that "a party may not seek discovery from any source before the parties have conferred as required by Rule 26(f)," except as "authorized by these rules, by stipulation, or by court order."
Against an unnamed defendant the case runs through that last clause: a motion filed, briefed and decided before any subpoena issues. Laid against a published 180-day retention window for one ISP's address logs, and preservation snapshots measured in 90-day increments, the calendar does more work than any forensic choice. What standard governs the motion, and how it meets the tests courts apply before unmasking an anonymous speaker, are legal questions I do not answer.
Two features of that landscape, as the research found them. There is no federal statute governing civil unmasking of anonymous online speakers, no uniform test across the states, and no Supreme Court decision setting one; the standards come from state and lower federal case law and are not consistent. And at least one legislature has codified a test — Va. Code § 8.01-407.1 (read 15 August 2026) requires a showing "that other reasonable efforts to identify the anonymous communicator have proven fruitless."
That element answers a question attorneys ask me directly: what is a technical investigator for against an unnamed defendant, if the records will not produce the name? Where a test reads that way, part of the answer is documentary. The investigation produces the record that cheaper routes were tried and did not work — a deliverable with a defined shape, and a different engagement from one sold on the promise of a name.
Frequently Asked Questions
Can an expert identify who wrote an anonymous post?
Rarely from technical records alone. Attribution runs through five inferences — content to account, account to registration records, account to address and time, address and time to subscriber, and subscriber to person — and only the first four are records. The last is an inference from circumstance. In matters that resolve, the records narrow the field to an account, a connection and a window, and the name comes from discovery, from examination of a device produced in the case, from an admission, or from a witness. An examiner who promises the name at the outset is promising the one link that has never been a record.What does an IP address prove about the person who posted?
It identifies a connection at a moment, not a person. A household router puts everyone behind one address, and Carrier-Grade NAT puts many unrelated subscribers behind one address — the IETF allocated a dedicated block for it in RFC 6598. Behind a content delivery network the logged address may be the network's rather than the client's. Residential addresses are leased and reassigned, so the lookup depends on a precise timestamp with a stated time zone. And geolocation is not identification: MaxMind, whose data is widely used, publishes roughly 66% accuracy at city level within a 50-kilometer radius.How long do platforms and ISPs keep the records that matter?
Less time than the procedural calendar usually allows, and most operators publish nothing. Meta and X each publish a 90-day preservation window pending formal legal process. X will say only that IP logs may be kept for a very brief period, and publishes no figure for it. Google, Reddit and Yelp published no preservation or log-retention window on the pages located when this was researched on 15 August 2026. Comcast's law enforcement guide, updated December 2024, states that IP address logs are retained for 180 days. Any other number you have been quoted came from somewhere other than the operator.Can a civil plaintiff send a statutory preservation letter?
No. 18 U.S.C. § 2703(f) requires a provider to preserve records, in the statute's own words, upon the request of a governmental entity, for 90 days extendable by a further 90 on renewed request by that entity. A private litigant has no equivalent under that section. What a private party has is a request the provider may honor or decline as a matter of its own policy, and whatever a court orders. Whether and how to seek such an order is a question for counsel. The technical finding underneath is that the default is deletion, the default clock is short, and calendaring the expiry is the requester's job.Is writing-style analysis reliable enough to name an author?
Not at the length of text a defamation matter usually involves. The strongest published internet-scale results come from long samples and report the right author ranked first roughly 20% of the time out of 100,000 candidates. For a two-sentence review there is no published error rate at all, and for text produced or edited by a language model there is no accepted method and no error rate either. My position is that stylometry generates investigative leads and supports an exclusion far more comfortably than an identification, and that an opinion naming an author from a short post offers a conclusion with no measured error behavior behind it.What is a technical investigator for if the records will not produce a name?
Three things. Narrowing the field, by tying accounts to registration artifacts, connections and time windows so that a subpoena asks for something specific rather than something broad. Preserving what is expiring, starting with the one dataset nobody's permission is needed for — the target's own server logs, which frequently record the poster's own visits. And producing the record of what was tried: at least one legislature has codified an unmasking test requiring a showing that other reasonable efforts to identify the speaker proved fruitless, which makes a documented investigation a deliverable in its own right.Which account field is most likely to lead somewhere?
The payment instrument. Every other identifying field on a consumer account is self-asserted: a name is typed in, an email address is free and disposable, and a phone number can be a voice-over-IP line obtained in minutes. Payment is the one field with an out-of-band verification step behind it, which is why attribution succeeds far more often against accounts that bought something — a subscription, a promoted post, a domain registration, a review package — than against accounts that never transacted. The practical instruction is to look for the transaction rather than the login, and to name payment records specifically wherever records are sought.Published