Audit Logging for Data Movement for Compliance
How to Audit Data Movement for Compliance

Your audit log for data movement should be an append-only record of who moved which dataset from where to where, under which credential and configuration, with what row counts and outcome. Most platforms can produce that record for a completed run. The harder task is proving where the data did not go. A residency reviewer or DORA supervisor may need evidence that a regulated table never crossed from the data plane you control into a control plane a vendor runs.
TL;DR
- Your audit log, lineage graph, and provenance record answer different auditor questions; only the audit log names the actor and credential.
- Of SOX, HIPAA, GDPR, PCI DSS v4.0, and DORA, only PCI DSS fixes a log retention period of 12 months; HIPAA's six years and SOX's seven years cover documentation.
- Your boundary proof needs fields that OpenLineage core does not define, and the Open Cybersecurity Schema Framework (OCSF) has no event class for replication jobs.
- You should keep hot storage for detection, warm storage for investigation, and a Write Once Read Many (WORM) cold tier for evidence; chaining detects edits, and WORM storage prevents them.
What Should Your Data Movement Audit Log Record?
Your movement audit record should capture the actor, the credential used, the source and destination datasets, the rows read and written, and the run outcome with error text on failure. Keep it distinct from lineage and provenance, because each answers a question the other two cannot. An audit log names who acted and with what outcome; lineage traces how data flows from input to output; provenance records origin, ownership, and chain of custody.
OpenLineage identifies each job by namespace and name, and each input and output dataset by namespace and name, giving you the path to inspect when a dashboard breaks. Its core specification does not make those references immutable versions; version metadata requires optional facets or implementation-specific extensions. It also carries no core field for who triggered the run or which credential that person or service used. Lineage therefore cannot establish human accountability on its own, and provenance cannot explain the full processing path to a downstream dashboard.
You should also distinguish audit logs kept as regulatory evidence from observability logs. Your audit logs record actions taken on systems and data, including who, what, when, and where, while observability logs focus on system health and troubleshooting. You can rotate health logs, but you need to retain, protect, and produce evidence logs on request. Joining per-event records to lineage across footprints is a separate problem you address through unified logging and lineage. The practical consequence is that you need an audit log for accountability even when you already maintain lineage and provenance.
What Do SOX, HIPAA, GDPR, PCI DSS, and DORA Require of Your Data Movement Logs?
Of the five, PCI DSS v4.0 clearly enumerates per-event fields and fixes a retention period for operational logs. SOX, HIPAA, and GDPR impose control, documentation, and register obligations that compliance blogs often restate as log requirements. DORA is Regulation 2022/2554, with Regulatory Technical Standards (RTS) 2024/1774; it names events but leaves retention to your information and communication technology (ICT) risk assessment. You can find the card-data field mapping under PCI DSS compliance.
HIPAA sets no retention period for your audit logs: the six years in 45 CFR 164.316(b)(2)(i) cover documentation, and 164.312(b) asks only that you record and examine activity in systems holding electronic protected health information (ePHI). The Office for Civil Rights guidance states that the Security Rule does not specify which data the audit controls must capture.
GDPR Article 30 covers a register of processing activities. Per-transfer event logging goes beyond that obligation. The Information Commissioner's Office guidance lists "the name of any third countries" as a category-level entry, with no per-transfer event. You should log each transfer anyway as an internal control beyond Article 30's explicit requirements.
SOX section 802 applies to the external auditor's workpapers. Your pipeline logs fall outside that provision. The SEC implementing rule requires seven years of retention for audit documentation. If you use a movement log as section 404 evidence, treat it as part of your compliance audit controls.
The DORA logging standards published as Delegated Regulation 2024/1774 set no fixed retention period. They require financial entities to set the retention period themselves, weighing their business and information security objectives, the reason each event is recorded, and the results of their ICT risk assessment. You should therefore reject the claim that DORA requires one year, because the text does not specify that period.
PCI DSS is the exception, because it gives you explicit event and retention requirements. The card-data logging rules in Requirement 10.2 name seven event types and six data fields per event, 10.5.1 sets 12 months of retention with three months immediately available; and 10.4.1 requires daily automated review of cardholder data environment (CDE) logs.
These obligations separate actual log requirements from broader documentation obligations.
Base your operational log policy on the applicable primary text. Documentation periods do not automatically govern every logging system.
Which Fields Prove Data Stayed Inside Its Boundary?
Your movement event can prove boundary control when it combines run identity with actor, location, classification, and transfer-integrity evidence. OpenLineage gives you run and dataset identity, OCSF gives you actor and outcome, and the Databricks community audit table gives you src_rec_count and target_rec_count. None combines all of these fields or records where the data physically went.
You Must Reconcile Two Logs in a Split Architecture
In a hybrid control plane, your orchestrator schedules and configures the run while an in-boundary worker reads the source and writes the destination. Each side logs separately. The orchestrator logs who triggered the run and the applicable config_version. The worker logs bytes, rows, checksums, and regions. Correlate them on run.runId, which the OpenLineage spec requires as a universally unique identifier (UUID), so each movement record carries both the actor and the physical evidence. An architecture diagram cannot replace that record.
You Must Treat the Audit Log as Regulated Data
Your execution-side log describes regulated rows, so you should treat it as regulated data. Keep the evidence layer in the same region or realm as the activity it documents. If you ship it to a control plane in another jurisdiction, the audit trail becomes a transfer, so keep the evidence tier where the data plane sits under your data residency compliance program.
You Must Log Metadata Sent to the Orchestrator Separately
You also need a record for metadata that the data plane sends to the orchestrator. Schema information, row counts, and column names can qualify as regulated metadata in some jurisdictions. Where your jurisdiction treats that metadata as regulated, the data plane should emit a separate audit event for each metadata transmission. The event should carry the same region and classification fields.
At minimum, your schema should include the fields needed to prove who moved data and where it went.
With the schema settled, the next question is how long these records survive and whether anyone can alter them.
How Should You Store and Protect Data Movement Audit Logs?
How you store movement logs decides whether they hold up as evidence when a reviewer asks.
Log at Job Level by Default and Drop to Row Level Only When Required
You should log at job level by default and drop to row level only where a regulation or forensic obligation requires every intermediate row state. Job-level records such as status, row counts, and timestamps per run cover operational observability. Row-level capture answers forensic questions that job-level status rows cannot, but it also adds CDC overhead.
You Need WORM Storage to Prevent Tampering
A keyed hash chain makes any edit to one entry invalidate every later integrity check. It detects modification. Someone with storage access can still reconstruct the entire chain, so use separate controls to prevent edits.
For your evidence tier, put the archive on S3 Object Lock in Compliance mode. S3 Object Lock blocks deletion and overwriting for the retention period, including actions by the bucket owner and root account. Keep hash chaining for detection and let Object Lock prevent the edit.
Limit expensive Security Information and Event Management (SIEM) retention to the period your detection and investigation requirements justify. A practical starting point is 30 to 90 days in hot or warm storage, with older logs in a cold archive, but your regulatory and forensic requirements should determine the actual periods. This gives you hot data for detection, warm data for investigation, and cold WORM storage for evidence preservation. The hot tier feeds pipeline compliance monitoring, while the cold WORM tier goes to the auditor.
You Should Not Rely on Native Platform Retention
Your platform's native retention may be shorter than the compliance window you need, so verify it before the first production run. Against the PCI DSS 12-month floor, configure an export to your evidence tier on the day the pipeline goes live.
How Does Airbyte Flex Keep Movement Evidence In-Boundary?
Airbyte Flex is a hybrid deployment in which Airbyte runs the hybrid control plane while your data plane, credentials, and compute stay inside your boundary. The workers that read sources and write destinations run in your environment, so your execution-side movement log originates in-boundary and stays there. Role-Based Access Control (RBAC) scopes who can change a connection, and audit logging records the change. The same 700+ connectors run across cloud, hybrid, and in-boundary deployments, so the event schema you define survives a footprint change.
Airbyte's open-source foundation lets you inspect the platform and move deployments across cloud, hybrid, and in-boundary environments, which reduces dependence on a single deployment model. Airbyte moves 26 billion records daily across customer deployments, and the same governance controls hold whether a pipeline runs in the cloud or inside your walls.
Where Should You Start?
Before your next audit, add boundary fields to every movement event in the schema. Without them, your residency claim is a policy statement rather than evidence a reviewer can sample.
For a split architecture, Airbyte Flex pairs in-boundary audit logging with RBAC, so one record carries both the actor and the physical-transfer evidence a residency reviewer needs.
Get a demo to see how Airbyte Flex keeps movement logs and data inside your boundary.
Frequently Asked Questions
Should You Capture Row-Level Changes With Triggers or Logical Decoding?
Use logical decoding when you need row-level capture. Tools such as wal2json or pgoutput can read the Write-Ahead Log (WAL) without application code changes, though CDC still adds operational overhead that you should test against your workload.
How Do You Reconcile Immutable Audit Logs With GDPR Erasure Requests?
You can use crypto-erasure by encrypting each subject's personal data under a per-subject key, hashing the ciphertext into the chain, and destroying the key on erasure. Your chain still verifies end-to-end, but the plaintext becomes unrecoverable.
Which Platform Audit Events Should You Check Before Deployment?
Check whether your team has enabled Databricks verbose audit logs, because Databricks omits query and command text until you enable verbose logging. You should also configure AWS Database Migration Service (DMS) data-plane events for CloudTrail where needed and verify Google Cloud data access logging at the project level.
Can You Use OpenTelemetry Spans as Pipeline Audit Events?
You should not use OpenTelemetry spans as a replacement for pipeline audit events. OpenTelemetry traces form a tree of services calling downstream services, while pipelines pull from upstream and push downstream, and lineage rarely reaches row level when traces remain at request level. No merged OpenTelemetry-plus-OpenLineage specification exists.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
