Two claims, and only one is provable from outside
This page is written for counsel weighing whether to commission an analysis of a cluster of accounts. "Coordinated account analysis" covers two claims worth separating before any technique is discussed.
- Claim A, weaker and usually supportable: these accounts behaved in a way unlikely to have arisen independently. A statistical claim about a pattern.
- Claim B, stronger and usually not supportable from outside: these accounts are one person, or were directed by the defendant. An identity claim, and outside a platform production it is an inference stacked on Claim A.
The most common failure in this work is presenting Claim A's evidence and asserting Claim B's conclusion.
The verdict is structural rather than a comment on any analyst. The platforms publish takedown reporting on coordinated inauthentic behavior; they do not publish the detection method behind it. An outside analyst sees registration dates, timing and content similarity, with no published error rate and no accepted standard for weighing them — no standards body publication, no validation study, no agreed threshold at which a pattern becomes a finding. A report that does not say so is claiming a footing the field does not have.
What an outside analyst can actually observe
Working only from public data, the observable signals are these:
- Account creation timing. Where the platform displays a join date, a cluster of accounts created inside a narrow window and used only to post about one subject is the most probative public signal available.
- Posting timing. Same-hour or same-minute posts, identical intervals, activity confined to a working-day window consistent with one time zone, or a burst against a background of nothing.
- Content similarity. Shared phrases, shared misspellings, identical formatting habits, identical factual errors. Shared errors are worth more than shared phrasing, because a phrase can be copied deliberately and an error usually is not noticed.
- Naming conventions and profile completeness. Usernames drawn from one scheme; no avatar, no bio, no prior activity, review or karma counts near zero.
- History overlap. On review platforms, two accounts that reviewed the same three obscure businesses is a strong signal, and the coincidence rate is computable.
- Cross-platform reuse. The same username, bio text or avatar image elsewhere — avatar reuse, comparable by perceptual hash, is underweighted and sometimes decisive.
All of it comes from data anyone can re-collect, so the deliverable is a dated dataset and the code that produced the table.
What the platforms publish about their own detection
The operators are more open about outcomes than about method. Meta's Inauthentic Behavior standard, read 15 August 2026, prohibits "using a connected network of inauthentic Meta assets to deliver substantial quantities of fake engagement in ways designed to look authentic," and states that enforcement applies "agnostic of content, political or otherwise." Its account integrity policy, change log dated 28 May 2026, describes acting against accounts "assessed to have common ownership and content as previously removed accounts." Meta makes those assessments; it does not publish them for a given account, and nothing in the policy gives an outside analyst the signals behind one.
Its threat reporting shows the same shape. The report published in December 2025 documents per-network takedowns with account counts — one scam network of "over 6,400 Facebook accounts and Pages" — and states that operations are identified using "behavioral and technical signals." The counts are published; the signals are named but not specified. Detection is behavior-based and network-scale, which is why a complaint about three accounts rarely engages it. This research found no example of that program applied to a private reputation attack.
Yelp is the most transparent of the consumer platforms, because it names a signal an outside analyst cannot replicate. Its recommendation software is "completely automated, no Yelp employee can manually override the software," and among the stated reasons a review may not be recommended are reviews "originating from the same IP address." Yelp's consumer alerts include a Suspicious Review Activity alert triggered by "large numbers of reviews coming from a single IP address, or reviews from users who may be connected to a group that coordinates incentivized fake reviews," and Yelp says it will "provide the evidence we find whenever possible."
| Operator | What is published | Reproducible by an outside analyst? |
|---|---|---|
| Meta | Policy criteria, takedown counts, "behavioral and technical signals" | No — the signals behind an assessment are unpublished |
| Yelp | Automated software; IP clustering named; alert triggers | No — IP addresses are platform-side |
| Legal-request and public-content policies only | Nothing published to reproduce | |
| X | A transparency landing page; no figures located | Nothing published to reproduce |
No current manipulation-enforcement figures were located for Reddit or X, so no numbers for either appear anywhere in this work.
What no outside analysis can see
The signals that would settle the question are all platform-side:
- IP addresses. Not visible on any consumer platform. The single most probative linking signal, available only through legal process.
- Device and browser fingerprints, cookie identifiers, advertising identifiers. Platform-internal, without exception.
- Registration email addresses and phone numbers. Platform-internal, and a shared registration email is often the strongest single link.
- The platform's own linkage determinations. These exist — the policy language quoted above describes making them — and they are not published.
- Payment instruments. Where accounts bought promotion or advertising, the card is frequently the cleanest link in the file, and it is production-only.
Stated plainly: an outside analyst can establish that a pattern exists and quantify how unlikely it is under an independence assumption, and cannot establish common control. An expert who claims to have traced an anonymous poster from public signals alone should be asked precisely which record that rests on.
Every pattern consistent with coordination is consistent with something else
This is the section most reports leave out, and the one that decides whether an analysis is worth commissioning. Every signal above has innocent generators.
- Template effects. Platforms prompt reviewers with the same questions, industries share vocabulary, and people describing the same experience describe it in the same words. Similarity has an ordinary source.
- Shared infrastructure is not shared authorship. Two people on one office network, VPN exit node or carrier-grade NAT share an address, and even with the IP — which an outside analyst will not have — the inference is not automatic.
- A real event produces a real burst. A news story or a genuinely bad month generates a cluster of new accounts posting about one subject in a narrow window — the headline signal, arriving for the ordinary reason.
- Empty profiles are the default in some places. Where throwaway accounts are normal, no avatar and no history describe the median user.
- Copying is not coordination. Independent posters copy phrasing from a post they read, reproducing wording with nobody directing anybody.
- The visible record is already filtered. Platform countermeasures remove and hide content continuously, so what an analyst collects is what survived moderation.
None of that makes the work pointless. It makes the deliverable a quantified pattern with its alternative explanations named and, where possible, tested — which is not the same thing as an attribution.
Why a bot score is not a finding
It is tempting to run a public bot-detection tool over a set of accounts and report the score. The measured error rates argue against it, and they are published, which makes them easy to find. Rauchfleisch and Kaiser, "The False positive problem of automatic bot detection in social science research," PLOS ONE 15(10): e0241045, published 22 October 2020, evaluated the most widely used tool of its kind. On resampled data assuming a 15% bot population at a 0.76 threshold, precision was 0.59 — 41% of the accounts classified as bots were human — while recall was 0.2, meaning roughly 80% of actual bots were missed. Scores were also unstable across a three-month window. The authors conclude that most studies using the tool "will unknowingly count a high number of human users as bots and vice versa."
A tool with a 41% false-positive rate at a plausible threshold cannot support an assertion that a specific account is inauthentic. It can support a screening step whose output is then examined by hand, and the report should say which was done.
Stylometry sits in the same place. Authorship attribution by writing style has real results on long documents by known candidate authors. Applied to a 60-word review it is much weaker, it is defeated by an author who varies style deliberately, and this research found no peer-reviewed error rate for short platform-mediated texts. It belongs in a report as one described signal — shared misspellings and idioms, quoted side by side — never as a standalone attribution.
How this analysis fails, and what protects it
Three defects account for most of the bad reports in this area, and all three are avoidable at no cost.
Base-rate neglect. "Both accounts posted on a Tuesday afternoon" is not a signal. Every similarity claim needs a comparison to how often that similarity occurs by chance in the relevant population; without one, the analysis has produced a coincidence and called it a pattern.
Selection on the dependent variable. The analyst collects the negative reviews, finds them similar, and concludes coordination — without sampling the positive reviews to see whether they are equally similar. That is the most common defect in real reports and the first thing a competent opponent tests.
Confirmation bias in a paid engagement. The client believes the reviews are fake, and the engagement exists because of that belief. An analysis whose parameters were chosen after the analyst saw the data is not a measurement.
Contested does not mean unusable: the conclusion is arguable while the underlying work still has to be verifiable. Five habits carry that difference — write the protocol, comparison set and thresholds down before collecting; name the metric and its parameters, because "cosine similarity over TF-IDF character 3-grams, threshold 0.85" is reproducible and "very similar" is not; ship the dated dataset with the code that produced the table; collect more than once so change is documented rather than argued; and preserve the negatives. Reproducing a table is not the same as accepting the inference drawn from it, and a finding here still has to survive challenge — what these habits buy is that the argument happens about the inference rather than about whether the numbers can be checked at all.
Where the record has to come from instead
If a matter needs common control rather than a pattern, the platform holds the records that supply it. The registration record is the highest-value single item: creation timestamp, registration IP, email address, phone number and the user-agent string at signup are what tie separate accounts together, and a subpoena reaches the basic subscriber records that include what the statute calls temporarily assigned network addresses.
Even that degrades against documented tradecraft. In an enforcement matter announced on 21 October 2019, the Federal Trade Commission described company managers posting reviews "using fake accounts created to hide their identity," with instructions to create multiple accounts with false identities and to use VPNs to conceal location — which defeats IP clustering.
The conduct does at least have defined federal vocabulary now. The Commission's Rule on the Use of Consumer Reviews and Testimonials, published 22 August 2024, addresses fake reviews including misrepresenting "that the reviewer or testimonialist exists" (§ 465.2), compensation conditioned on reviews expressing a particular sentiment — which reaches paid negative reviews (§ 465.4) — and review suppression through "an unfounded or groundless legal threat" (§ 465.7), which runs against the party demanding removal. What any of that means for a given matter is a question for counsel; the forensic point is that an analysis framed to those categories describes conduct rather than falsity.
Frequently Asked Questions
Can an expert prove that a group of fake reviews came from one person?
Not from public data. An outside analyst can show that a set of accounts behaved in a way unlikely to have arisen independently, and can quantify how unlikely under a stated independence assumption. Establishing that the accounts are one person, or were directed by a particular party, requires records only the platform holds: registration IP addresses, registration email addresses and phone numbers, device identifiers, and payment instruments. Presenting the pattern evidence and asserting the identity conclusion is the standard failure in this area, and it is the one an opponent looks for first.What can an outside analysis show without a subpoena?
Account creation dates where the platform displays them, posting timestamps and intervals, content similarity including shared misspellings and shared factual errors, username schemes, profile completeness, overlapping review histories, and cross-platform reuse of usernames, bio text and avatar images. All of it is re-collectable, so the deliverable should be a dated dataset plus the code that produced the table. What it supports is a quantified pattern with its alternative explanations named — useful for deciding where to aim process, and not an attribution.Do the platforms publish how they detect coordinated accounts?
They publish outcomes, not methods. Meta publishes policy criteria and per-network takedown counts and says operations are identified using behavioral and technical signals, without specifying them, and it does not publish its assessment for any given account. Yelp goes furthest by naming IP clustering among the reasons a review may not be recommended — a signal no outside analyst can replicate. Current manipulation-enforcement figures for Reddit and X were not found in this research, so no numbers for either should be quoted. There is no published error rate for any of it.What innocent explanations produce the same pattern?
Several, and they have to be addressed in the report rather than in response to a question. Platforms prompt reviewers with identical questions and industries share vocabulary, so similarity has ordinary sources. People on one office network, VPN or carrier-grade NAT share an address. A real news event produces a genuine burst of new accounts posting about one subject. On some platforms an empty profile is the median user. Independent posters copy phrasing they read elsewhere. And moderation has already filtered what is visible, so the collected sample is not the population.Is a bot-detection score worth putting in a report?
As a screening step, described as one. The published evaluation of the most widely used tool — Rauchfleisch and Kaiser, PLOS ONE, 22 October 2020 — found precision of 0.59 at a plausible threshold, meaning 41% of accounts classified as bots were human, and recall of 0.2, meaning roughly 80% of bots were missed. Scores also moved across a three-month window for a substantial share of accounts. A tool with that false-positive rate cannot support an assertion that a particular account is inauthentic; it can prioritize accounts for manual examination.Why is this method marked contested?
Because no accepted standard exists for it. There is no standards-body publication governing how coordination signals should be weighed, no validation study establishing an error rate, and no agreed threshold at which a pattern becomes a finding. The platforms that do this at scale publish their takedowns and not their detection method, so there is nothing external to calibrate against. The underlying collection and computation can still be fully reproducible; what is contested is the inference drawn from the table, and a report should say so in those terms.What records actually link two accounts to one operator?
Registration records first: the creation timestamp, the registration IP address, the email address and phone number used, and the user-agent string captured at signup. Login and session history extends the picture where retention allows, and payment instruments are often the cleanest link where an account ever bought promotion or advertising. All of these are platform-side and reachable only through process. Nothing in the public record substitutes for them, which is why an outside analysis is most useful when it is used to decide where process should be aimed.Published