AWS CloudTrail to Databricks: How to Move Your Data

Move AWS CloudTrail into Databricks with Airbyte. Why each account and region needs its own source, and why event shapes vary by AWS service.

Summarize with AI:

Moving AWS CloudTrail into Databricks builds an audit archive that outlives the ninety days CloudTrail's lookup reaches back. Any compliance question about last year is unanswerable from that interface and perfectly answerable from tables you started filling early enough.

This guide covers the managed path with Airbyte. Two things shape the build: the lookup is scoped to one account in one region, and event records vary in shape from one AWS service to the next.

AWS CloudTrail to Databricks at a glance:

CapabilitySupportedWhat it means for this pipeline
ScopeOne account, one regionA wider estate needs a source for each combination
Event typesManagement onlyInsight events are not supported by this connector
Lookup windowAbout 90 daysThere is no backfill, so start collecting early
Rate limit2 per secondPer account per region, so regions do not contend
Event shapeVaries by serviceRequest and response fields differ across AWS services

Why move data from AWS CloudTrail to Databricks?

Two situations account for most of these pipelines.

The first is retention for audit and investigation. Ninety days covers an incident response and not a compliance review, and a lakehouse holds years of activity cheaply while keeping it queryable, which is the combination this data needs.

The second is analysis across accounts and regions together, which CloudTrail's own interface will not do. If what you want is to search events by principal or resource rather than aggregate them, AWS CloudTrail to Elasticsearch suits that better.

What do you need before you start?

Four things, and the first decides how many pipelines you are building:

A list of your accounts and regions. CloudTrail's event history is recorded in the region where the event happened and searched one account at a time, so coverage is a multiplication rather than a single connection. The AWS CloudTrail source documentation covers the fields.

Credentials with CloudTrail read access. An access key and secret, plus the region name. Set the region deliberately, since the connector defaults to us-east-1 and a wrong value quietly produces an empty or unexpected dataset.

Acceptance that you get management events only. Insight events are not supported, and data events are a separate CloudTrail feature, so an investigation needing object-level access will not be served by this pipeline.

Permission to create Volumes in Unity Catalog. Staging goes through Avro files written into a Volume, which is separate from creating tables and worth requesting early.

If your workspace restricts traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.

How do you build an AWS CloudTrail to Databricks pipeline in Airbyte?

Step 1: Count your accounts and regions

Write down every account and every region you need covered, and multiply. That number is how many sources this project involves, and it is usually larger than anybody assumed because AWS estates spread quietly. Doing the arithmetic now turns an open-ended build into a known one, and it stops somebody concluding months later that an audit table covers the whole organisation when it covers one region of one account.

Step 2: Configure the AWS CloudTrail source

Click Sources in the left navigation, then New Source, and select AWS CloudTrail, following adding a source. Supply the access key, secret key, region and start date, then repeat per account and region. Name each source after its account and region, because a list of identically named CloudTrail sources is unmanageable within a week.

Step 3: Configure the Databricks destination

Click Destinations, then New Destination, and select Databricks, following adding a destination. Supply the workspace details, warehouse or cluster, catalogue and schema. Let events land with their nested structures intact, since Spark reads them natively and an audit record should not be trimmed on the way in.

Step 4: Create the connections and test before you rely on them

Click Connections, then New connection for each source, selecting the stream and a sync mode. Use incremental and schedule comfortably inside ninety days. Airbyte publishes support level and reliability signals for each connector, and this one is community maintained rather than certified, so run it for a fortnight before an auditor depends on it.

Then union the accounts and regions in a silver table, because nobody wants to query eleven tables to answer one question.

Why is one pipeline never enough?

Because CloudTrail records events in the region where they happened, and a lookup searches a single account in a single region. That is a property of the service rather than the connector, and it means the scope of one source is one cell in a grid of accounts against regions.

For a small estate that is two or three sources and unremarkable. For an organisation with separate accounts per environment and workloads in several regions, it is dozens, and the work is in keeping them consistent rather than in any one of them. Name them systematically and generate them if your platform allows it, because hand-maintaining thirty near-identical sources is how one quietly ends up misconfigured.

The rate limit is unusually friendly about this. Lookups are capped at two per second per account per region, so separate sources draw on separate budgets rather than competing, and running many in parallel is genuinely faster than running one broad one would be. The constraint that multiplies your configuration also stops it becoming a throughput problem.

Why do events differ from one service to another?

Because a CloudTrail event describes an API call, and AWS has a great many APIs. The outer envelope is consistent, carrying who made the call, when, from where and against which service, while the request parameters and response elements inside it belong to whichever operation was invoked. An instance launch and a bucket policy change share almost nothing internally.

That is exactly the shape a lakehouse handles well and a strictly typed destination handles badly. Spark reads the nested structures natively, so bronze can hold the event as CloudTrail produced it, with nothing discarded because it did not fit a column. For audit data that fidelity is not a nicety: the field nobody modelled is the one an investigation asks about.

So do the shaping in silver, per question rather than per service. A table of identity and access events, one of resource changes, one unioning every account and region with columns naming both: those answer real questions while the raw events stay available underneath. And record which accounts and regions each table covers, because an audit table whose coverage is undocumented is difficult to defend.

Frequently asked questions

Can one source cover several regions?

No. Events are recorded per region and looked up per account, so you need a source for each account and region combination you want covered.

My table is empty or missing events.

Check the region, since the connector defaults to us-east-1 and reads only where events actually happened. Also confirm you expected management events rather than data or Insight events.

Can I backfill older activity?

Not through this connector, since the lookup reaches back about ninety days. For a longer record, start collecting now or use a trail or event data store.

Should I flatten the event structures?

Not on the way in. Spark reads them natively, and for audit data keeping the original record matters more than tidiness. Shape it in a silver layer instead.

Can I do this without writing code?

The pipelines, yes, though there are several of them. The silver tables unioning accounts and regions are the work that makes the archive usable.

Get your AWS CloudTrail data into Databricks

Count your accounts against your regions and accept that the answer is how many sources you are building, since the lookup covers one of each. Set the region explicitly rather than inheriting the default. Start collecting now, because ninety days is all there is to catch. Then let events land nested, union everything in a silver table with columns naming the account and region, and record what that table covers.

Airbyte's connector catalog includes 600+ pre-built connectors, so cloud activity can be retained long after the platform forgets it. For a document database into the same destination, see MongoDB to Databricks, and for the same source into a search engine, AWS CloudTrail to Elasticsearch.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.