Audit Logging for Data Movement for Compliance

How to Audit Data Movement for Compliance

Summarize with AI:

Your audit log for data movement should be an append-only record of who moved which dataset from where to where, under which credential and configuration, with what row counts and outcome. Most platforms can produce that record for a completed run. The harder task is proving where the data did not go. A residency reviewer or DORA supervisor may need evidence that a regulated table never crossed from the data plane you control into a control plane a vendor runs.

TL;DR

  • Your audit log, lineage graph, and provenance record answer different auditor questions; only the audit log names the actor and credential.
  • Of SOX, HIPAA, GDPR, PCI DSS v4.0, and DORA, only PCI DSS fixes a log retention period of 12 months; HIPAA's six years and SOX's seven years cover documentation.
  • Your boundary proof needs fields that OpenLineage core does not define, and the Open Cybersecurity Schema Framework (OCSF) has no event class for replication jobs.
  • You should keep hot storage for detection, warm storage for investigation, and a Write Once Read Many (WORM) cold tier for evidence; chaining detects edits, and WORM storage prevents them.

Try Airbyte Flex

What Should Your Data Movement Audit Log Record?

Your movement audit record should capture the actor, the credential used, the source and destination datasets, the rows read and written, and the run outcome with error text on failure. Keep it distinct from lineage and provenance, because each answers a question the other two cannot. An audit log names who acted and with what outcome; lineage traces how data flows from input to output; provenance records origin, ownership, and chain of custody.

OpenLineage identifies each job by namespace and name, and each input and output dataset by namespace and name, giving you the path to inspect when a dashboard breaks. Its core specification does not make those references immutable versions; version metadata requires optional facets or implementation-specific extensions. It also carries no core field for who triggered the run or which credential that person or service used. Lineage therefore cannot establish human accountability on its own, and provenance cannot explain the full processing path to a downstream dashboard.

You should also distinguish audit logs kept as regulatory evidence from observability logs. Your audit logs record actions taken on systems and data, including who, what, when, and where, while observability logs focus on system health and troubleshooting. You can rotate health logs, but you need to retain, protect, and produce evidence logs on request. Joining per-event records to lineage across footprints is a separate problem you address through unified logging and lineage. The practical consequence is that you need an audit log for accountability even when you already maintain lineage and provenance.

What Do SOX, HIPAA, GDPR, PCI DSS, and DORA Require of Your Data Movement Logs?

Of the five, PCI DSS v4.0 clearly enumerates per-event fields and fixes a retention period for operational logs. SOX, HIPAA, and GDPR impose control, documentation, and register obligations that compliance blogs often restate as log requirements. DORA is Regulation 2022/2554, with Regulatory Technical Standards (RTS) 2024/1774; it names events but leaves retention to your information and communication technology (ICT) risk assessment. You can find the card-data field mapping under PCI DSS compliance.

HIPAA sets no retention period for your audit logs: the six years in 45 CFR 164.316(b)(2)(i) cover documentation, and 164.312(b) asks only that you record and examine activity in systems holding electronic protected health information (ePHI). The Office for Civil Rights guidance states that the Security Rule does not specify which data the audit controls must capture.

GDPR Article 30 covers a register of processing activities. Per-transfer event logging goes beyond that obligation. The Information Commissioner's Office guidance lists "the name of any third countries" as a category-level entry, with no per-transfer event. You should log each transfer anyway as an internal control beyond Article 30's explicit requirements.

SOX section 802 applies to the external auditor's workpapers. Your pipeline logs fall outside that provision. The SEC implementing rule requires seven years of retention for audit documentation. If you use a movement log as section 404 evidence, treat it as part of your compliance audit controls.

The DORA logging standards published as Delegated Regulation 2024/1774 set no fixed retention period. They require financial entities to set the retention period themselves, weighing their business and information security objectives, the reason each event is recorded, and the results of their ICT risk assessment. You should therefore reject the claim that DORA requires one year, because the text does not specify that period.

PCI DSS is the exception, because it gives you explicit event and retention requirements. The card-data logging rules in Requirement 10.2 name seven event types and six data fields per event, 10.5.1 sets 12 months of retention with three months immediately available; and 10.4.1 requires daily automated review of cardholder data environment (CDE) logs.

These obligations separate actual log requirements from broader documentation obligations.

RegulationObligationEnumerated eventsRetentionIntegrityReview
HIPAA 45 CFR 164.312(b)Record and examine ePHI system activityNone specifiedNone for logs; 6 years for documentation under 164.316(b)(2)(i)Protect logs from tampering; no mechanism mandated"Regularly review," frequency undefined
PCI DSS v4.0 Req. 10Implement, protect, review, retain CDE logsYes: 7 event types, 6 data fields per event (Req. 10.2)12 months total, 3 months immediately available (Req. 10.5.1)File integrity monitoring or change detection (Req. 10.3.4)Daily automated review of CDE logs (Req. 10.4.1)
GDPR Art. 30 / Art. 5(2)Register of processing activitiesNot applicableNo fixed periodNot specifiedAvailable to supervisory authority on request
DORA 2022/2554 + RTS 2024/1774Logging framework; major-incident reportingSpecified in RTS Art. 12Risk-based, set by the entityLogging must remain functional; integrity checks on recovery (Art. 12)24/7 anomaly detection for critical functions (RTS Art. 24)
SOX §§ 302/404/802Internal controls; audit documentation retentionNone specified7 years for auditor documentation (SEC rule; Public Company Accounting Oversight Board (PCAOB) AS 3)Not specified for system logsManagement assessment (§404); no IT log frequency

Base your operational log policy on the applicable primary text. Documentation periods do not automatically govern every logging system.

Which Fields Prove Data Stayed Inside Its Boundary?

Your movement event can prove boundary control when it combines run identity with actor, location, classification, and transfer-integrity evidence. OpenLineage gives you run and dataset identity, OCSF gives you actor and outcome, and the Databricks community audit table gives you src_rec_count and target_rec_count. None combines all of these fields or records where the data physically went.

You Must Reconcile Two Logs in a Split Architecture

In a hybrid control plane, your orchestrator schedules and configures the run while an in-boundary worker reads the source and writes the destination. Each side logs separately. The orchestrator logs who triggered the run and the applicable config_version. The worker logs bytes, rows, checksums, and regions. Correlate them on run.runId, which the OpenLineage spec requires as a universally unique identifier (UUID), so each movement record carries both the actor and the physical evidence. An architecture diagram cannot replace that record.

You Must Treat the Audit Log as Regulated Data

Your execution-side log describes regulated rows, so you should treat it as regulated data. Keep the evidence layer in the same region or realm as the activity it documents. If you ship it to a control plane in another jurisdiction, the audit trail becomes a transfer, so keep the evidence tier where the data plane sits under your data residency compliance program.

You Must Log Metadata Sent to the Orchestrator Separately

You also need a record for metadata that the data plane sends to the orchestrator. Schema information, row counts, and column names can qualify as regulated metadata in some jurisdictions. Where your jurisdiction treats that metadata as regulated, the data plane should emit a separate audit event for each metadata transmission. The event should carry the same region and classification fields.

At minimum, your schema should include the fields needed to prove who moved data and where it went.

FieldPurposeSource coverage or missing field
event_id (UUID), eventTime (Coordinated Universal Time (UTC))Stable audit-event identity and event timeOpenLineage defines eventTime as a top-level RunEvent property; add event_id to the audit schema
run.runId, job. namespace, job.nameIdentifies the run and jobOpenLineage core identifiers; immutable job versions require optional facets or implementation-specific extensions
actor_id, actor_type (USER, SERVICE, SYSTEM), credential referenceNames who or what moved the dataNot in OpenLineage core
Source dataset namespace, name, versionDataset the job read fromOpenLineage core identifies inputs by namespace and name; versions require optional facets or implementation-specific extensions
Destination dataset namespace, name, versionDataset the job wrote toOpenLineage core identifies outputs by namespace and name; versions require optional facets or implementation-specific extensions
source_region, destination_region, data_classificationBoundary and residency proofNot in any open schema
src_rec_count, target_rec_count, rejected countSource-to-target reconciliationOpenLineage only as the optional DataQualityMetricsFacet, never required
Checksum of transferred batch, bytes transferredIntegrity of the transfer itselfNot in OpenLineage or OCSF
config_versionConfiguration governing the runSOX reconciliation schemas only
eventType (START, RUNNING, COMPLETE, ABORT, FAIL), errorMessageOutcome and failure detaileventType is a top-level OpenLineage RunEvent property; errorMessage requires a run facet or audit-schema field
Optional hash-chain linkTamper evidenceNot in OpenLineage or OCSF; evidence tier only

With the schema settled, the next question is how long these records survive and whether anyone can alter them.

How Should You Store and Protect Data Movement Audit Logs?

How you store movement logs decides whether they hold up as evidence when a reviewer asks.

Log at Job Level by Default and Drop to Row Level Only When Required

You should log at job level by default and drop to row level only where a regulation or forensic obligation requires every intermediate row state. Job-level records such as status, row counts, and timestamps per run cover operational observability. Row-level capture answers forensic questions that job-level status rows cannot, but it also adds CDC overhead.

You Need WORM Storage to Prevent Tampering

A keyed hash chain makes any edit to one entry invalidate every later integrity check. It detects modification. Someone with storage access can still reconstruct the entire chain, so use separate controls to prevent edits.

For your evidence tier, put the archive on S3 Object Lock in Compliance mode. S3 Object Lock blocks deletion and overwriting for the retention period, including actions by the bucket owner and root account. Keep hash chaining for detection and let Object Lock prevent the edit.

Limit expensive Security Information and Event Management (SIEM) retention to the period your detection and investigation requirements justify. A practical starting point is 30 to 90 days in hot or warm storage, with older logs in a cold archive, but your regulatory and forensic requirements should determine the actual periods. This gives you hot data for detection, warm data for investigation, and cold WORM storage for evidence preservation. The hot tier feeds pipeline compliance monitoring, while the cold WORM tier goes to the auditor.

You Should Not Rely on Native Platform Retention

Your platform's native retention may be shorter than the compliance window you need, so verify it before the first production run. Against the PCI DSS 12-month floor, configure an export to your evidence tier on the day the pipeline goes live.

How Does Airbyte Flex Keep Movement Evidence In-Boundary?

Airbyte Flex is a hybrid deployment in which Airbyte runs the hybrid control plane while your data plane, credentials, and compute stay inside your boundary. The workers that read sources and write destinations run in your environment, so your execution-side movement log originates in-boundary and stays there. Role-Based Access Control (RBAC) scopes who can change a connection, and audit logging records the change. The same 700+ connectors run across cloud, hybrid, and in-boundary deployments, so the event schema you define survives a footprint change.

Airbyte's open-source foundation lets you inspect the platform and move deployments across cloud, hybrid, and in-boundary environments, which reduces dependence on a single deployment model. Airbyte moves 26 billion records daily across customer deployments, and the same governance controls hold whether a pipeline runs in the cloud or inside your walls.

Where Should You Start?

Before your next audit, add boundary fields to every movement event in the schema. Without them, your residency claim is a policy statement rather than evidence a reviewer can sample.

For a split architecture, Airbyte Flex pairs in-boundary audit logging with RBAC, so one record carries both the actor and the physical-transfer evidence a residency reviewer needs.

Get a demo to see how Airbyte Flex keeps movement logs and data inside your boundary.

Frequently Asked Questions

Should You Capture Row-Level Changes With Triggers or Logical Decoding?

Use logical decoding when you need row-level capture. Tools such as wal2json or pgoutput can read the Write-Ahead Log (WAL) without application code changes, though CDC still adds operational overhead that you should test against your workload.

How Do You Reconcile Immutable Audit Logs With GDPR Erasure Requests?

You can use crypto-erasure by encrypting each subject's personal data under a per-subject key, hashing the ciphertext into the chain, and destroying the key on erasure. Your chain still verifies end-to-end, but the plaintext becomes unrecoverable.

Which Platform Audit Events Should You Check Before Deployment?

Check whether your team has enabled Databricks verbose audit logs, because Databricks omits query and command text until you enable verbose logging. You should also configure AWS Database Migration Service (DMS) data-plane events for CloudTrail where needed and verify Google Cloud data access logging at the project level.

Can You Use OpenTelemetry Spans as Pipeline Audit Events?

You should not use OpenTelemetry spans as a replacement for pipeline audit events. OpenTelemetry traces form a tree of services calling downstream services, while pipelines pull from upstream and push downstream, and lineage rarely reaches row level when traces remain at request level. No merged OpenTelemetry-plus-OpenLineage specification exists.

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.