What survives an upload, and what does not
This page is written for counsel deciding whether a metadata examination is worth commissioning in an internet defamation matter, and what it will produce if it is. Two findings sit at the front of the subject, and they are separate findings. Almost nothing survives an upload to a major consumer platform. And almost no major platform publishes a policy telling you what it strips.
The consequence is immediate. If the item in dispute is a photograph or a scanned document that a defendant posted, the copy you download from the platform will normally carry no camera make or model, no GPS coordinates, no author field and no creation timestamp from the originating device. The posted copy is the one copy in the matter engineered to carry the least. Running a metadata tool over it is a reasonable first step and almost always a dead end, and I would rather say that before an engagement than after one.
That makes this a production question before it is an analysis question. The metadata worth having sits on the author's device, in a native file somebody still holds, or in the platform's own server-side records — and each of those needs process or a cooperating custodian rather than a closer look at what is public.
Where the record actually lives
Three separate places hold what people mean when they say metadata, and they fail in different ways.
- The native file. The original photograph or document, in its original format, held by whoever made or sent it. This is the only copy carrying the full field set, and it is normally reachable only through a custodian or a device image.
- The device. File system timestamps, browser artifacts, application databases and the deleted-file remnants around them — in practice a forensically imaged device obtained in discovery, not a machine anyone can sit down at.
- The platform. The server-side upload timestamp, the registration record, and the login and session history. The platform recorded when the file arrived even after it discarded what the file said about itself.
The most useful instruction I can give at the start of a matter is to ask for files in native format with hashes, rather than as printouts or re-downloads. A Word document printed to PDF discards its core properties, an image re-saved by a viewer is re-encoded, and a photograph forwarded in image mode arrives with almost nothing left. Each step is ordinary, none is sinister, and all three destroy the record before an examiner sees the file.
The one measured study, and what it measured
The best-sourced measurement in this research is Nishchal Soni, "Forensic Value of Exif Data: An Analytical Evaluation of Metadata Integrity across Image Transfer Methods," published in Perspectives in Legal and Forensic Sciences on 18 June 2025. The findings, as the author reports them, are unusually clean.
Transfers by USB and as an email attachment retained 100% of EXIF fields, with file hashes unchanged. Transfers in "document mode" through three messaging applications also retained 100%, because the application treats the file as a file rather than as an image to compress. Chat and image modes across the platforms tested retained approximately 16.67% of fields — effectively resolution and nothing else. The paper's own summary is that compression-based transfer modes "effectively remove metadata, changing the file integrity."
The implication worth carrying into a matter is that the transfer mode matters more than the platform. The same file, sent by the same person on the same day, survives intact as a document and arrives empty as a photo — which is why a witness who still has the original email attachment is often worth more than the platform copy.
| Transfer path | Measured EXIF retention (Soni, 2025) | Operator policy on stripping |
|---|---|---|
| USB copy or email attachment | 100% of fields; hashes unchanged | Not applicable |
| Messaging apps, document mode | 100% of fields | None located |
| Messaging and chat apps, image mode | ~16.67% of fields | None located |
| Instagram, Facebook Messenger, Snapchat (image mode) | ~16.67% of fields | None located |
| Facebook, X, Reddit, LinkedIn, Google Photos, Flickr | Not measured in that study | None located as of 15 August 2026 |
No operator publishes a stripping policy, so the answer is an experiment
Searching for an operator-published statement of metadata handling returned help-forum threads and vendor blog posts, and no policy page, for Facebook, Instagram, X, Reddit, LinkedIn, Google Photos and Flickr, read 15 August 2026. That is a genuine absence rather than a search that needs repeating. Any confident sentence of the form "that platform strips EXIF" traces back to somebody's test, not to the operator, and a report that states it as operator policy is stating something the operator has not said.
So I test it. The defensible way to describe a platform's behavior is to upload a file with known EXIF to that platform on a stated date, download the file back, and diff the metadata — then document that as an experiment with its own date, its own tool versions and its own hashes, because the behavior changes without notice and nobody announces the change. That experiment is reproducible: another examiner can run the same upload and compare. What it cannot do is speak to how the platform behaved eighteen months ago, when the file in dispute was posted.
That gap is the assumption this method rests on, and it is why the verdict on this page is what it is. The technique reproduces. The inference from today's behavior back to the behavior on the posting date is an assumption, it has to be stated in the report as an assumption, and it may be argued.
EXIF is evidence of what a file says about itself
The most common overstatement in this area is treating EXIF as an observation of what
happened. It is a set of fields inside a file, and a single exiftool command
rewrites any of them, including DateTimeOriginal and the GPS coordinates.
Nothing binds a consumer photograph to its metadata cryptographically.
Four failure modes come up repeatedly:
- The clock belongs to the device. A camera's date and time are
user-set and drift. Timezone handling is inconsistent across manufacturers:
DateTimeOriginalhistorically carried no timezone field, and the newerOffsetTimeOriginaltag is unevenly populated. A stated EXIF time without a timezone is ambiguous by up to a day, which in a timeline dispute is the whole argument. - A screenshot carries the wrong EXIF. A screenshot of a web page carries the capturing device's metadata, not the page's. Reading provenance out of it is a category error, and it appears in real reports.
- Re-encoding destroys and re-creates. Any save through an editor may
strip original fields and write new ones naming the editor as
Software. - Absence has three explanations. Missing camera EXIF is consistent with platform stripping, with ordinary editing, and with a file that was generated rather than photographed. Those are very different stories and the file alone does not choose between them.
What turns a field into a finding is corroboration from a source the author did not control: file system timestamps on a device image, a second copy from a different custodian, or the platform's server-side upload time. On its own, an EXIF field is a lead.
Document metadata is usually the stronger record
Where the statement began life as a letter, a complaint, a "report" or a dossier that was later posted, the document's own properties are often the strongest authorship evidence in the matter — and they are frequently the place where a party's own production reveals more than intended.
Office Open XML files (.docx, .xlsx, .pptx) are
ZIP containers. docProps/core.xml and docProps/app.xml carry
the creator, the last-modified-by name, creation and modification times, the revision
number, total editing time and the template used. PDFs carry an Info dictionary or XMP
with Producer, Creator, CreationDate and ModDate, and where a PDF was saved
incrementally, earlier revisions may be recoverable from the file body itself.
The failure modes are the same in shape as EXIF's and different in detail. Fields are editable. A last-modified-by value names an account, not a human being. Shared templates propagate a prior author's name into unrelated documents, which is how a document that nobody in the dispute wrote can appear to name a party. Many organizations run automated metadata scrubbers on outbound files, so absence again proves little. And a total editing time describes a session rather than a person: a document left open overnight reports hours of work nobody did.
Browser artifacts, and the epochs that produce wrong timelines
On a device obtained through discovery and imaged properly, the browser is usually the
richest single source. History databases — History for Chrome and Edge,
places.sqlite for Firefox, History.db for Safari — carry visit
timestamps and transition types that distinguish a typed URL from a link click. Cache may
hold a rendered copy of a page as it appeared at a specific visit, and it is occasionally
the only surviving copy of deleted content. Cookies and local storage show which accounts
were signed in on the machine. Downloads records, form autofill, profile data and
session-restore files fill in the rest.
It fails in four ways worth naming. Private browsing leaves far less. History is user-deletable and routinely deleted. Timestamps are stored in several different epochs — WebKit microseconds since 1601, Unix seconds, PRTime microseconds — and mis-conversion is a real and repeated source of timelines that are wrong by years rather than minutes. And an artifact establishes that a browser profile on a machine did something, which is one inferential step away from a person doing it.
On tool validation, the honest position is narrow. NIST runs the Computer Forensics Tool Testing program, which publishes specifications and test reports, and its published corpus is strongest on disk imaging, write blockers and mobile extraction. Whether it has published a test specification covering browser-artifact parsing or online content acquisition specifically is not something this research confirmed, so I do not claim that coverage for the tools used in this part of the work.
What the platform holds that the file does not
The account-side records usually beat the device-side records: a third party with no stake in the outcome holds them, and they are internally consistent across accounts. A subpoena reaches the basic subscriber records at 18 U.S.C. § 2703(c)(2), which expressly include "temporarily assigned network addresses" — IP addresses. In its law enforcement guidelines, read 15 August 2026, Meta describes that tier as "name, length of service, credit card information, email address(es), and a recent login/logout IP address(es)." Note the article — a recent login IP, singular, not a history. The transactional tier is described as "message headers and IP addresses, in addition to the basic subscriber records," and the content tier as "messages, photos, videos, timeline posts, and location information." Meta also states that for end-to-end encrypted conversations it "will continue to provide message and call logs, as well as IP data" — the content goes, the pattern of contact does not.
Where the account holder is a client or a cooperating witness, a self-service export is a partial substitute that needs no process: Meta's Download Your Information and its newer data logs export, announced by Meta Engineering on 4 February 2025, and Google Takeout. Two cautions travel with any of them. An export is the account holder's view, produced on demand, with no custodian certification. And its completeness is defined by the platform rather than by the request — Meta's announcement does not state how far back the data logs export reaches, so I assume no period for it. The forensic point is that an export and a production are not the same object.
Frequently Asked Questions
Does a photo posted to a social platform still contain its EXIF data?
Normally not. The one measured study in this research — Soni, Perspectives in Legal and Forensic Sciences, 18 June 2025 — found roughly 16.67% field retention for image and chat mode transfers, which is effectively resolution and nothing else, against 100% retention for USB copies, email attachments and document-mode sends. No major operator publishes a policy stating what it strips, so the defensible way to describe a specific platform's behavior is to upload a file with known EXIF, download it back, and diff the result on a stated date.Can EXIF data be faked?
Yes, in one command. Freely available tools rewrite any field, including the original date and the GPS coordinates, and nothing binds a consumer photograph to its metadata cryptographically. The clock is the device's and is user-settable, and a stated time without a timezone offset is ambiguous by up to a day. That does not make EXIF worthless; it makes it a lead that needs corroboration from something the author did not control, such as file system timestamps on an imaged device, a second copy from a different custodian, or the platform's server-side upload time.What has to be asked for so that metadata survives discovery?
Native files, in native format, with hashes recorded at the point of collection. A Word document printed to PDF loses its core properties. An image re-saved by a viewer is re-encoded and may lose the camera fields while gaining a software field naming the editor. A photograph forwarded through a messaging app in image mode arrives effectively empty. Each of those steps is ordinary rather than deliberate, and all three destroy the record before an examiner sees it. What form of request achieves that in a given matter is a question for counsel.Does the platform know when the file was uploaded?
Yes. The server-side upload timestamp is the platform's own record and survives whatever stripping happened to the file, which makes it one of the more useful anchors available. It is reachable by legal process rather than by inspection, and it dates the arrival of the file at the platform, not the creation of the underlying photograph or document. Pairing a server-side upload time with a device-side artifact is how an examiner tests whether a claimed creation date is plausible without relying on the file's own assertions.Is a Download Your Information export the same as a platform production?
No, and the difference matters. A self-service export is the account holder's view of the account, generated on demand, with no custodian certification attached and no record of what the platform chose not to include. Completeness is defined by the platform, not by the request. Meta's engineering announcement of its data logs export, dated 4 February 2025, does not state how far back the export reaches, so no retention period should be assumed for it. Exports are genuinely useful where the account holder is cooperating; they do not substitute for records produced under process.Can browser history show who posted something?
It can show that a browser profile on a particular machine visited a URL, at a recorded time, and whether the URL was typed or followed from a link. That is one inferential step short of a person, and the step is not free: shared machines, shared profiles and signed-in family members all break it. Private browsing leaves far less behind, history is user-deletable, and timestamps sit in several different epochs, so a mis-converted value can move a timeline by years. Browser artifacts corroborate an attribution; on their own they rarely carry one.Published