A practical source-mapping guide

eDiscovery data sources: devices, cloud and backups

eDiscovery data can sit on computers and phones, in cloud accounts, business systems and backups. Start by mapping where relevant records may exist, who controls them and how long they will remain available.

By Alistair Ewing · Published

At a glance

  • Map people, accounts and shared systems before deciding what to collect.
  • A live account, its local cache and its backup may contain different records.
  • Preserve relevant sources before deletion, account closure or backup expiry changes them.
  • Reduce duplicate review while keeping the record of every source and custodian.

This guide covers potential sources of electronically stored information, often called ESI. It complements our eDiscovery and eDisclosure service. The table is a planning aid: access and collection feasibility are confirmed for each instruction.

Devices · services · recovery sources

Which data sources should an eDiscovery plan consider?

Search by source or provider, or choose a category. The final column explains limits to check before collection.

45 source types shown

Swipe or scroll sideways to compare all five columns.

eDiscovery source map: examples, potential records and collection checks
Source type Category Examples What may be available Collection check
Computers and user profiles Devices and local data Windows PCs, Macs, Linux workstations; desktops and laptops Working documents, downloads, browser records, local logs and user activity. Identify every relevant user profile and disk. Cloud placeholders may contain no local file content.
Phones and tablets Devices and local data iPhone, iPad, Android; work and personal devices Messages, including accessible WhatsApp, Signal or Telegram data; photographs, recordings, contacts and app records. Device state, encryption, app settings and authority determine access. One extraction rarely covers every source.
Removable media Devices and local data USB sticks, external HDDs and SSDs, SD cards Portable working files, transferred documents and camera media. Record which device and person used the media. Preserve it before routine browsing or repair.
Shared storage and file servers Devices and local data SMB/NFS shares, NAS, departmental drives Shared folders, file ownership, permissions and available filesystem records. Map share names to actual storage. Include relevant home folders and departed users; snapshots belong in the backup inventory.
Local email archives Devices and local data PST, MBOX, EML and MSG collections; accessible mail caches Historical messages, attachments, headers and folders outside the live mailbox. A cache can be incomplete. Preserve message and attachment relationships and record the export's origin.
Standalone applications Devices and local data Local accounting, design, CAD and specialist case software Project files, local databases, attachments and application-specific history. A PDF or spreadsheet printout may omit the native relationships. Identify the application and required reader.
On-premises email servers Devices and local data Exchange Server and other locally hosted mail systems Mailboxes, shared folders, journal stores and available message tracking records. Plan a consistent collection of the relevant database or mailboxes. Retain logs and linked attachments within scope.
Microsoft 365 mail Cloud communications and files Exchange Online user, shared and archive mailboxes Email, attachments, calendars and retained mailbox content within the authorised scope. Identify aliases, shared access and leavers. Holds, permissions and licences affect the collection route.
Microsoft 365 documents Cloud communications and files OneDrive and SharePoint sites and libraries Documents, available versions, ownership and collaboration context. Include the right sites and libraries. A downloaded file alone may omit versions, permissions or cloud-only properties.
Microsoft Teams Cloud communications and files Chats, channels and meeting-related records Messages and references to files, recordings or transcripts that may reside in other Microsoft 365 locations. Map associated storage and channel types. Record links between sources instead of counting the same attachment twice.
Google Workspace email Cloud communications and files Gmail accounts and Google Groups Messages, attachments and supported retained content available through authorised exports. Check the actual Vault coverage, account licence and preservation settings. Personal Gmail requires a different access route.
Google Workspace documents Cloud communications and files My Drive and shared drives Cloud documents, uploaded files, available revisions and sharing context. Native Google documents need an agreed export format. Preserve source IDs and relevant context lost during conversion.
Google Chat Cloud communications and files Direct messages, spaces and associated files Retained conversations and references to shared material. History, membership and service settings affect availability. Follow file references to the underlying source.
Slack Cloud communications and files Workspaces, channels and direct messages Available messages, threads, membership context and file links. Export scope depends on plan, permissions and retention. A link in an export is not the attached file itself.
Other hosted mail Cloud communications and files Personal webmail, IMAP providers and hosted business email Messages, folders, attachments and available server-side records. Check local-only folders, forwarding and separate archives. Access to an account does not imply access to provider logs.
Cloud storage and sync Cloud communications and files Dropbox, Box, iCloud Drive and photo libraries Files, photos and available versions, sharing or deletion records. Sync can propagate deletions. Check cloud content separately from the device's downloaded subset.
Meeting platforms Cloud communications and files Zoom, Webex and other hosted meeting services Retained recordings, transcripts, chats and meeting metadata. Separate cloud recordings from local copies. Recording, transcription and retention may never have been enabled.
Customer and sales systems Business applications Salesforce and other CRM platforms Customer records, contacts, activities, attachments and supported change history. Scope related objects and attachments, not just the main report. Audit history depends on configuration.
Finance and resource planning Business applications Accounting, ERP, procurement and expenses systems Transactions, invoices, approvals, supplier changes and audit records. Preserve record IDs and relationships. A current balance or CSV may omit earlier changes and supporting documents.
HR and recruitment Business applications HR platforms, applicant tracking, payroll and learning systems Relevant personnel, recruitment, attendance and approval records. Use a narrow lawful scope for sensitive records. Identify linked documents and separate payroll providers.
Projects, tickets and knowledge Business applications Jira, Confluence, ServiceNow, Asana, Trello and Notion Tickets, pages, comments, attachments and available change history. Include archived projects and linked files where relevant. Flat exports can lose threaded discussion or permissions.
Software development Business applications Git repositories, GitHub, GitLab and CI/CD systems Commits, branches, issues, pull requests, build outputs and deployment records. A repository clone may omit issues, server audit events and external large-file storage. Identify each separately.
Electronic signatures and contracts Business applications DocuSign, Adobe Acrobat Sign and contract management Signed documents, envelopes, completion records and available audit trails. Preserve the signed native document and validation context; a printed copy cannot retain every signature property.
Telephony and contact centres Business applications Hosted VoIP, call recording and customer contact platforms Retained call audio, voicemail, call details and agent notes. Call metadata and audio may have different retention periods. Clock settings and export formats need checking.
AI assistant records Business applications AI chat accounts, workplace assistants and integrated AI tools Accessible prompts, responses, uploaded material and available activity records. Exports vary by product and workspace policy. Do not assume access to provider internals, hidden reasoning or a complete history.
Social media accounts Online and third-party sources X, LinkedIn, Facebook, Instagram and other social platforms Accessible posts, direct messages, uploads and account exports. Public viewing and authorised account export give different coverage. Deletions, edits and unavailable media must be recorded.
Websites and content management Online and third-party sources WordPress and other CMS, hosting control panels Site content, uploads, forms, database records and available access logs. A rendered page omits backend records. Preserve relevant hosting and application data with the site owner's authority.
Public web captures Online and third-party sources Public pages, forums and historical web archives Published material and the context visible at a documented capture time. Coverage is selective. Record URLs, capture time and missing resources; a capture date is not necessarily publication time.
External business portals Online and third-party sources Client portals, marketplaces, payment and supplier platforms Accessible orders, invoices, transactions, correspondence and uploaded records. The account holder may only see a subset. Identify any records that need a provider request or other lawful route.
Cloud object storage Infrastructure and logs Amazon S3, Azure Blob Storage and Google Cloud Storage Objects, metadata and retained versions or audit records where configured. Inventory buckets, containers, regions and lifecycle rules. Versioning and immutability are separate settings.
Databases and data platforms Infrastructure and logs SQL/NoSQL databases, warehouses and managed database services Structured records, relationships and available transaction or query history. Agree a consistent export and preserve its schema. A database snapshot and an application report answer different questions.
Virtual machines and hosted workloads Infrastructure and logs Cloud VM disks, on-premises virtual servers and containers Workload files, virtual disks, configuration and retained application logs. A running snapshot may lack application consistency; container data can be short-lived. Record capture state and dependencies.
Identity, access and security logs Infrastructure and logs Directory services, VPN, firewall, endpoint and SIEM systems Sign-ins, permissions, alerts, access events and recorded data movement. Retention can be short. An account or IP address alone does not identify the person responsible.
Cameras and specialist devices Infrastructure and logs CCTV recorders, body cameras, dashcams and access-control systems Recordings, event logs, configuration and device time settings. Export native footage with required playback material. Overwrite cycles and clock errors can be critical.
Cloud endpoint backups Online backups Carbonite Safe, Backblaze Computer Backup and IDrive Selected backed-up computer files and available historical or deleted versions. Check the product, platform, selected folders and actual restore points. Provider retention rules and exclusions differ.
SaaS backup services Online backups Veeam, Druva, Acronis and other third-party SaaS backup products Backed-up mail, cloud documents or other workloads included in the configured service. Confirm the protected tenant, workload and backup date. A separate subscription does not prove every source was backed up.
Cloud phone backups Online backups iCloud Backup, Android device backup and supported app backups Device or app restore data retained by the relevant backup service. Content already synced separately may be excluded. Encryption, keys, account state and app settings determine access.
Hosted server and disaster-recovery backups Online backups Provider backups, cloud recovery vaults and managed backup services Scheduled recovery points for servers, databases, websites or storage. Ask for job history, retention and application coverage. Restore to a controlled destination without overwriting the source.
Online backup repositories Online backups Network-connected backup appliances, disk repositories and replicated vaults Backup sets, catalogues, job logs and retained recovery points. Internet access is not required for an online repository. Preserve the catalogue, encryption keys and dependent backup chain.
Disconnected backup drives Local and offline backups Rotated USB disks and removable backup cartridges Historical file copies or backup sets held away from the live system. Record labels, custody and dates. A disk's label does not prove when a successful backup last ran.
Tape and optical media Local and offline backups LTO tapes, legacy tape formats, DVDs and archival discs Older backup sets, project archives and exported records. Compatible drives, software, catalogues and complete media sets may be needed before targeted restoration is possible.
Local system and phone backups Local and offline backups Time Machine, File History, system images and Finder/iTunes backups Historical device data retained in local backup stores. Identify backup dates, device association and encryption. A backup is not necessarily a full forensic image.
Storage snapshots Local and offline backups NAS, filesystem and VM snapshot sets Earlier storage states available for the affected volume or workload. Snapshots may still be online and share the original failure risk. Check retention and consistency before use.
Retired systems and retained images Local and offline backups Decommissioned drives, preserved forensic images and legacy servers Historical data no longer present in the active estate. Keep acquisition and custody records. Do not boot a preserved image as a live system before planning the examination.
Long-term records archives Local and offline backups Mail journals, document archives and transferred ZIP/TAR packages Retained business records, historical exports and associated indexes. Archives may be online or offline. Establish completeness, format, retention and how the archive was created.

A category describes the source’s main role. A snapshot or archive can be online; record its actual connection state in the case inventory. Product examples identify places to ask about, rather than a guarantee of recoverability.

What is the difference between sync, backup, archive and a legal hold?

Sync follows the working account

A sync client keeps selected content aligned across locations. Deletions and changes may spread. A laptop may hold only previews or placeholders while the full files remain in the cloud.

A backup offers recovery points

A backup may retain earlier copies of selected data. Its value depends on successful jobs, scope, retention and access. Restoring a file does not establish its complete history.

An archive retains selected records

An archive keeps records for a defined purpose. It may have indexes, retention rules or message journals, but can omit material outside that purpose.

A legal hold preserves defined material

A hold is a preservation measure applied to an agreed scope. It needs the right accounts, systems and monitoring; it cannot bring back data already lost.

For example, Apple explains that iCloud Backup excludes data already synced to iCloud. Both locations may therefore matter. Likewise, Google explains that Vault works with supported Workspace services; it is not a universal backup of everything in an account.

Can Carbonite and other online backups provide evidence?

Potentially. A backup may retain a document version or deleted file absent from the current computer. First establish which computer and folders were protected, which product and platform were used, and which recovery points are still present.

Carbonite’s Windows version-history guidance describes retained file versions and notes that earlier filenames are not retained in that history. Its Mac restore guidance warns that some previous-version restores overwrite files in their original location. The safe collection method therefore depends on the actual product and restore route.

Backblaze version history also depends on the configured retention option. For any provider, record the backup account, protected source, recovery point, export method and exceptions. Preserve the catalogue and logs where relevant. Plan a controlled restore before using an option that could overwrite evidence.

Does a Microsoft 365 export include everything?

No. The selected locations and supported workload matter. Microsoft’s eDiscovery data-source guidance explains how user and group data spans mailboxes, OneDrive and SharePoint locations. Teams can connect several of these sources. A mailbox-only search should not be treated as a complete collection of a person’s collaboration records.

The same distinction matters elsewhere: Slack exports can contain message history and file links. The underlying files, private conversations and older content need separate coverage checks within the authorised scope.

What belongs in an eDiscovery source inventory?

  1. People and owners. List the people who hold or control relevant records, along with shared accounts, teams and external providers.
  2. Locations. Identify devices, mailboxes, sites, apps, archives and backups. Record the relationship between a live source and its copies.
  3. Dates and retention. Note the relevant period, oldest available records, deletion settings, backup expiry and any preservation already applied.
  4. Access and authority. Confirm the account holder, administrator, permissions, encryption and agreed legal basis. Keep credentials out of ordinary email and enquiry forms.
  5. Collection scope. Agree sources, dates, formats, attachments and metadata. Record what is excluded and why.
  6. Delivery and verification. Define the receiving platform, naming, manifests, checksums, exceptions and secure hand-off before collection begins.

For civil matters within its scope, Practice Direction 57AD expressly includes electronic records, metadata and backup sources, with duties to preserve relevant documents. Your lawyers determine the applicable disclosure regime, relevance, privilege and proportionality.

How can duplicate review be reduced without losing context?

Preserve the collected source sets first. Processing can then identify exact duplicate content while retaining all relevant custodians, source paths and locations. Keep attachments linked to their parent messages and retain materially different versions.

A copy in a mailbox, a shared folder and a backup may contain the same document but show different possession or timing. An agreed deduplication rule can reduce repeated review while preserving that context. It should not silently remove the only record of where an item was found.

For interpretation after collection, see the digital evidence artefacts and tools guide. For recordings that need specialist input, see how digital and audiovisual expert roles work together.

Source references and further reading

The linked provider documentation explains specific coverage and restore limits. Check the edition, account settings and current documentation when preparing an actual collection.

Turn the source list into a proportionate collection plan

Describe the systems, approximate volumes, people and deadline. Alistair can assess the technical scope, preservation priorities and hand-off your legal team or review provider needs.